Neyshekar
A large-scale, open Persian read-speech corpus for speech recognition and text-to-speech, crowdsourced from native speakers.

- 62,279recordings
- 99 hof speech
- 190speakers
- 701,621tokens
- 29,535word vocabulary
Neyshekar is an open, community-driven Persian speech dataset. Volunteers, all native Persian speakers, record sentences through a web-based crowdsourcing platform at ney.shekar.io. The goal is a freely usable corpus for automatic speech recognition, text-to-speech, speech representation learning and other Persian speech applications.
The dataset is released incrementally. Each release is a stable snapshot, so results stay reproducible and benchmarks stay comparable over time.
Release v6.0
v6.0 is the first release with predefined splits. Splits are assigned per speaker, so no recorder appears in more than one set and the evaluation splits are speaker-disjoint from training. About a quarter of the samples (15,222, or 24.4%) are informal speech, identified with Shekar's rule-based informal-language classifier.
| Split | Samples | Share | Hours | Speakers |
|---|---|---|---|---|
| Train | 58,244 | 93.5% | 91.99 | 134 |
| Validation | 1,886 | 3.0% | 3.14 | 26 |
| Test | 2,149 | 3.5% | 3.88 | 30 |
Growth
The corpus has grown roughly sixfold since the first release in December 2025.
| Release | Date | Samples | Hours | Vocabulary |
|---|---|---|---|---|
| v6.0 | 2026-09-01 | 62,279 | 99.02 | 29,535 |
| v5.0 | 2026-07-16 | 50,026 | 79.22 | 27,250 |
| v4.1 | 2026-06-15 | 40,008 | 63.03 | 26,758 |
| v3 | 2026-03-23 | 30,019 | 45.71 | 23,972 |
| v2 | 2026-01-15 | 20,020 | 29.08 | 20,853 |
| v1 | 2025-12-29 | 10,044 | 14.42 | 15,224 |
Average clip length is about 5.7 seconds. The v4 snapshot from May 2026 had misaligned audio–transcript pairs and was replaced by v4.1.
Terms of use
Neyshekar is released under CC0 1.0 and can be used for any purpose. Any attempt to identify the speakers is strictly prohibited.
Citation
@article{amirivojdan2026neyshekar,
title = {Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition},
author = {Amirivojdan, Ahmad and Nadiri, Farzad and Alizadeh, Abolfazl and Yaraghi, Shaghayegh},
journal = {arXiv preprint arXiv:2609.14542},
year = {2026},
doi = {10.48550/arXiv.2609.14542}
}