Purpose-built audio data at unmatched scale.

We power leading research institutions with speech, conversations, pair-wise preferences, & labels generated by millions of contributors, across languages and geographies.

Contributors worldwide
1M+
Human preferences
20M+/ day
Recording capacity
20,000+hrs/day

Natural multi-speaker conversations

  • Track-separated multi-speaker conversations.
  • Meticulously human-labeled for conversational events.
Loading conversation sample…

Emotional steering

  • Each script and its non-verbal cues delivered to multiple emotional directions.
  • Voice actors rated head-to-head by millions of raters; each speaker carries an ELO.
Loading a take…

Bilingual recordings

  • Native multilingual speakers for dubbing and accent removal.
  • Natural multilingual conversations, code-switching organically.
  • Emotionally steered speech in every language we run.
Loading…

Pairwise judgments at scale

  • Tens of millions of A/B judgments a day.
  • Trained and quality-assured expert raters.
  • Compare model outputs against each other along any judgement axis.
Loading a pair…

Data labeling and transcription

  • Tens of thousands of trained transcribers across more than fifty locales.
  • Conversation labels, laughter, addressee, speaker turns, sentiment — or the taxonomy you bring.
  • Every label and transcript settled by consensus and checked by experts.
Loading labels…

Custom projects, specific to your needs

  • We design, build and launch custom projects within hours.
  • Reaching millions in any locale, through any modality.
  • Meticulous QA on every project, and a free sample before any larger deliverable.