← All projects

Dataset · Speech · 2025 – present

Neyshekar

A large-scale, open Persian read-speech corpus for speech recognition and text-to-speech, crowdsourced from native speakers.

Neyshekar logo: a pixel-art bundle of sugarcane beside the word Neyshekar and the subtitle A Large-Scale Open Persian Speech Dataset.
  • 62,279recordings
  • 99 hof speech
  • 190speakers
  • 701,621tokens
  • 29,535word vocabulary

Neyshekar is an open, community-driven Persian speech dataset. Volunteers, all native Persian speakers, record sentences through a web-based crowdsourcing platform at ney.shekar.io. The goal is a freely usable corpus for automatic speech recognition, text-to-speech, speech representation learning and other Persian speech applications.

The dataset is released incrementally. Each release is a stable snapshot, so results stay reproducible and benchmarks stay comparable over time.

Release v6.0

v6.0 is the first release with predefined splits. Splits are assigned per speaker, so no recorder appears in more than one set and the evaluation splits are speaker-disjoint from training. About a quarter of the samples (15,222, or 24.4%) are informal speech, identified with Shekar's rule-based informal-language classifier.

SplitSamplesShareHoursSpeakers
Train58,24493.5%91.99134
Validation1,8863.0%3.1426
Test2,1493.5%3.8830

Growth

The corpus has grown roughly sixfold since the first release in December 2025.

ReleaseDateSamplesHoursVocabulary
v6.02026-09-0162,27999.0229,535
v5.02026-07-1650,02679.2227,250
v4.12026-06-1540,00863.0326,758
v32026-03-2330,01945.7123,972
v22026-01-1520,02029.0820,853
v12025-12-2910,04414.4215,224

Average clip length is about 5.7 seconds. The v4 snapshot from May 2026 had misaligned audio–transcript pairs and was replaced by v4.1.

Terms of use

Neyshekar is released under CC0 1.0 and can be used for any purpose. Any attempt to identify the speakers is strictly prohibited.

Citation

@article{amirivojdan2026neyshekar,
  title   = {Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition},
  author  = {Amirivojdan, Ahmad and Nadiri, Farzad and Alizadeh, Abolfazl and Yaraghi, Shaghayegh},
  journal = {arXiv preprint arXiv:2609.14542},
  year    = {2026},
  doi     = {10.48550/arXiv.2609.14542}
}