WenetSpeech-Min: A Large-Scale Minnan Speech Corpus
with Dual Transcriptions for Dialectal Speech Processing

  1. 1Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
  2. 2School of Intelligence Science and Technology, Nanjing University
  3. 3University of New South Wales
  4. 4WeNet Open Source Community
  5. 5Moonstep AI
  6. 6Nexdata
  7. 7School of Informatics, Xiamen University
  8. 8State Key Laboratory of Novel Software Technology, Nanjing University

* Equal contribution โ€  Corresponding author

Abstract

Progress in dialectal speech technology is hindered by the scarcity of large-scale, real-world corpora. For Minnan speech, existing resources remain limited, and few provide paired Minnan and Mandarin transcripts at scale. To address these gaps, we introduce WenetSpeech-Min, an open-source corpus comprising around 10,000 hours of Minnan speech collected from diverse online media, with paired Minnan and Mandarin transcripts for every utterance. We further establish an automatic speech recognition (ASR) benchmark covering both Minnan and Mandarin transcripts and a text-to-speech synthesis (TTS) benchmark using Minnan transcripts, with manually verified evaluation sets for both tasks. To assess the effectiveness of the corpus, we train ASR and TTS models on WenetSpeech-Min and compare them with representative systems on the proposed benchmarks. The resulting models outperform the evaluated open-source models on most metrics and achieve competitive performance against commercial systems. We will release the corpus, benchmarks, and models to facilitate reproducible research on Minnan speech technology.

Data Construction Pipeline

The unified pipeline processes audio from heterogeneous online media and generates paired Minnan and Mandarin transcripts. Audio processing segments recordings into utterances and filters them by speech content and acoustic quality. For transcript generation, multiple Minnan ASR systems are combined to produce a dialect transcript, followed by subtitle extraction or translation to obtain its Mandarin counterpart.

WenetSpeech-Min data construction pipeline: media collection and segmentation, speech quality annotation, Minnan transcription with three ASR systems and ROVER, then Mandarin transcription using Qwen3 or OCR.
Overview of the WenetSpeech-Min data construction pipeline.

WenetSpeech-Min

Dataset Overview

Distribution of sourced hours by content domain. Drama accounts for 54.7%, Audiobook 19.8%, Entertainment 13.9%, Reading 4.2%, Music Program 2.7%, Education 1.5%, Short Video 1.4%, Culture 1.2%, and Others 0.6%.
(a) Content-domain distribution over sourced hours.
WV-MOS score distributions by duration for general sources and archival TV dramas.
(b) WV-MOS distribution by source group.
Duration distribution of speech across signal-to-noise ratio ranges.
(c) Signal-to-noise ratio distribution.
Duration distribution of utterances, with peaks around 6 to 8 seconds and 19 seconds.
(d) Utterance-duration distribution.

Data Samples

Listen to speech samples alongside their Minnan transcripts, Mandarin translations, and speaker information.

ASR Leaderboard

Minnan Transcription Results

Results on WS-Min-Eval-ASR. Lower D-CER is better. Light green rows indicate models trained on WenetSpeech-Min.

ModelD-CER (%) โ†“
Hy-ASR-3.0-Previewโ€ 32.78
CN-MultiDialect-ASR18.59
Qwen3-ASR34.61
Qwen3-ASR-WSM-Min17.59
Qwen3-ASR-WSM-Min + internal data15.21
FireRedASR2-AED43.39
FireRedASR2-AED-WSM18.95

โ€  Results obtained via a commercial API.

Mandarin Transcription Results

Results on WS-Min-Eval-ASR and the external GigaSpeechBench and MinSpeech test sets. Avg. is the unweighted mean across the three test sets. Lower M-CER is better; higher M-BLEU is better. Light green rows indicate models trained on WenetSpeech-Min.

Model WS-Min-Eval-ASR GigaSpeechBench MinSpeech Avg.
M-BLEU โ†‘M-CER (%) โ†“ M-BLEU โ†‘M-CER (%) โ†“ M-BLEU โ†‘M-CER (%) โ†“ M-BLEU โ†‘M-CER (%) โ†“
Mandarin Transcription
FunASR-Realtimeโ€ 39.0945.4153.3732.6228.2263.2340.2347.09
SeedASR2.0โ€ 38.9147.5855.3232.5068.5326.2154.2535.43
Qwen3-ASR-MinSpeech20.3460.5712.3971.2667.9423.9333.5651.92
Qwen3-ASR-WSM-Mandarin42.9342.5651.2034.0159.8229.4251.3235.33
Qwen3-ASR-WSM-Mandarin + internal data49.3936.8147.6536.6665.6925.6854.2433.05
Paired-Transcript Recognition
FireRedASR2-AED18.0564.4036.9149.3513.6085.0222.8566.26
FireRedASR2-AED-WSM46.0639.5357.4028.6157.7331.8553.7333.33

โ€  Results obtained via commercial APIs.

TTS Leaderboard

Results on WS-Min-Eval-TTS. Lower CER is better; higher is better for the other metrics. Light green rows indicate models trained on WenetSpeech-Min.

Model WS-Min-Eval-TTS-Easy WS-Min-Eval-TTS-Hard
CER (%) โ†“SIM โ†‘I-MOS โ†‘S-MOS โ†‘A-MOS โ†‘ CER (%) โ†“SIM โ†‘I-MOS โ†‘S-MOS โ†‘A-MOS โ†‘
QwenAudio-3.0-TTSโ€ 19.900.6723.883.853.6235.090.6853.753.733.70
Qwen3TTS-Flashโ€ ,โ€ก24.44--3.93--3.9338.02--3.75--3.73
FireRedTTS333.760.7473.303.323.1355.330.7313.153.393.04
VoxCPM228.160.7073.453.543.3840.500.6813.173.483.36
MERaLiON-TTS21.250.6223.833.733.6537.280.6363.583.733.63
CosyVoice363.690.6542.132.131.4960.350.6102.002.401.43
CosyVoice3-WSM20.260.6783.853.753.6535.980.6693.653.853.70

โ€  Results obtained via commercial APIs. โ€ก Uses a single fixed speaker; SIM and S-MOS are therefore not evaluated.

TTS Demo

Each row shows the target text, its reference audio, and synthesized speech from each system. Scroll horizontally to compare the models.