ESPnet releases YODAS v3, a 1.1-million-hour open speech dataset
The ESPnet team published YODAS v3, an open dataset of about 1.1 million hours of real-world speech in more than 100 languages, with 48kHz audio and timestamped transcripts.

The ESPnet team released YODAS v3 on Hugging Face on September 27. It contains roughly 1.1 million hours of real-world speech across more than 100 languages, 34 of which have over 1,000 hours.
Each example includes 48kHz audio, a transcript with word- and utterance-level timestamps, the detected language and, where available, a timestamped English translation. The team says over 70% of the data has at least two distinct channels and estimates each file's effective sampling rate so users can filter by quality.
The authors say the dataset is large enough to train a Whisper-scale speech recognition model from scratch, and suits speech synthesis and spatial audio research.