Skip to content

Free for commercial and research use

Basis Conversations 1500

Multilingual, multi-party, full-duplex conversational speech.

Audio hours
1,502
Unique speakers
2,645
Conversations
1,907

Overview

Basis Conversations 1500 is a multi-party, multilingual, full duplex conversational speech dataset. Each conversation includes up to 4 simultaneous speakers, each with channel-separated, 48 kHz audio. The median conversation lasts 33 minutes and 2,645 unique speakers are represented.

Multi-party: a conversation seats between two and four people at a time; participants come and go over the course of the conversation.

Multilingual: Speakers from 33 countries speaking 22 languages (some never before included in public conversational speech data at any scale).

Highly annotated: about 100 hours of conversations are human annotated (113k human judgements from 7,362 labelers) for categories of backchannels, interruptions, speaker addressee moments, pauses, laughter, repairs and self-repairs (disfluencies).

Labels

Annotated Hours
97.6
Annotations
64,772
Human judgments
113,096

Loading labels…

Composition

2,645 speakersAcross 33 countries
Loading map…
1355 speakers
CountrySpeakers
  • 🌍Country unspecified2

Metrics

Mean segment length
37.61 min
Median average simultaneous participants
3.0
Words
13,744,204
Unique words
525,263
Metrics across the full exported release. Expand a metric for its definition.
MetricP5MedianP95Mean
5.2926.05112.5437.61
2.03.03.72.8
25125.71
3.5110.5521.8411.38
5.7911.2316.8411.25
27.8145.8569.9346.86
3.463.874.043.83

About Basis

We power leading research institutions with purpose-built audio data at unmatched scale. We offer emotionally steered speech, conversations, pairwise preferences, and labels generated by millions of contributors across languages and geographies.

Contributors worldwide
1M+
Recording capacity
20,000+hrs / day
Human preferences
20M+/ day

More audio from Basis

Explore more speech, conversations, and expressive audio.

Explore
How was the dataset collected?

All conversations are recorded on our custom consumer applications, where participants are paid for their contributions.

Contributors consent to the recording and licensing of their contributions for AI research and commercial development. Our contributor agreements explicitly cover model training and evaluation, speech recognition, text-to-speech, voice synthesis, voice cloning, and digital replicas, subject to each contributor’s consent and the applicable dataset license.

Why are we open sourcing it?

Progress in real-time human-AI interaction depends heavily on data that captures people listening, responding, and speaking to one another across languages. We're making 1,500 hours of conversational data available so that researchers and builders worldwide can study human interactions and develop better models: understanding nuances of human conversations, turn-taking, emotional understanding, humor, and more.

This release, like all our open source releases, is free for commercial and research use. Its files, documentation, and license are available on Hugging Face.

Go to Hugging Face

What other data does Basis gather?

Basis works with the world's leading research labs to build custom datasets across voice, video, and interaction, shaping each collection around the labs' research goals and data needs. Our voice datasets include emotionally steered speech, multi-speaker conversations with or without video, bilingual recordings, and more. We also provide transcriptions, human annotations, and pairwise human preferences at unmatched scale through our contributor network of over a million users.

Explore Basis audio data

Access & citation

Get the dataset

Free for commercial and research use

Cite this dataset

@misc{basis2026conversations1500,
  title     = {Basis Conversations 1500},
  author    = {{Basis}},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/datasets/basis-ai/basis-conversations-1500}
}