Free for commercial and research use
Basis Conversations 1500
Multilingual, multi-party, full-duplex conversational speech.
- Audio hours
- 1,502
- Unique speakers
- 2,645
- Conversations
- 1,907
Overview
Basis Conversations 1500 is a multi-party, multilingual, full duplex conversational speech dataset. Each conversation includes up to 4 simultaneous speakers, each with channel-separated, 48 kHz audio. The median conversation lasts 33 minutes and 2,645 unique speakers are represented.
Multi-party: a conversation seats between two and four people at a time; participants come and go over the course of the conversation.
Multilingual: Speakers from 33 countries speaking 22 languages (some never before included in public conversational speech data at any scale).
Highly annotated: about 100 hours of conversations are human annotated (113k human judgements from 7,362 labelers) for categories of backchannels, interruptions, speaker addressee moments, pauses, laughter, repairs and self-repairs (disfluencies).
Labels
- Annotated Hours
- 97.6
- Annotations
- 64,772
- Human judgments
- 113,096
Loading labels…
Composition
- 🌍Country unspecified2
Metrics
- Mean segment length
- 37.61 min
- Median average simultaneous participants
- 3.0
- Words
- 13,744,204
- Unique words
- 525,263
| Metric | P5 | Median | P95 | Mean |
|---|---|---|---|---|
| 5.29 | 26.05 | 112.54 | 37.61 | |
| 2.0 | 3.0 | 3.7 | 2.8 | |
| 2 | 5 | 12 | 5.71 | |
| 3.51 | 10.55 | 21.84 | 11.38 | |
| 5.79 | 11.23 | 16.84 | 11.25 | |
| 27.81 | 45.85 | 69.93 | 46.86 | |
| 3.46 | 3.87 | 4.04 | 3.83 |
About Basis
We power leading research institutions with purpose-built audio data at unmatched scale. We offer emotionally steered speech, conversations, pairwise preferences, and labels generated by millions of contributors across languages and geographies.
- Contributors worldwide
- 1M+
- Recording capacity
- 20,000+hrs / day
- Human preferences
- 20M+/ day
More audio from Basis
Explore more speech, conversations, and expressive audio.
How was the dataset collected?
All conversations are recorded on our custom consumer applications, where participants are paid for their contributions.
Contributors consent to the recording and licensing of their contributions for AI research and commercial development. Our contributor agreements explicitly cover model training and evaluation, speech recognition, text-to-speech, voice synthesis, voice cloning, and digital replicas, subject to each contributor’s consent and the applicable dataset license.
Why are we open sourcing it?
Progress in real-time human-AI interaction depends heavily on data that captures people listening, responding, and speaking to one another across languages. We're making 1,500 hours of conversational data available so that researchers and builders worldwide can study human interactions and develop better models: understanding nuances of human conversations, turn-taking, emotional understanding, humor, and more.
This release, like all our open source releases, is free for commercial and research use. Its files, documentation, and license are available on Hugging Face.
What other data does Basis gather?
Basis works with the world's leading research labs to build custom datasets across voice, video, and interaction, shaping each collection around the labs' research goals and data needs. Our voice datasets include emotionally steered speech, multi-speaker conversations with or without video, bilingual recordings, and more. We also provide transcriptions, human annotations, and pairwise human preferences at unmatched scale through our contributor network of over a million users.
Access & citation
Get the dataset
Free for commercial and research use
Cite this dataset
@misc{basis2026conversations1500,
title = {Basis Conversations 1500},
author = {{Basis}},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/basis-ai/basis-conversations-1500}
}