Speech recognition result evaluation method, electronic device, and storage medium
By calculating the acoustic consistency between speech and text using an autoregressive TTS model, the problems of existing methods' dependence on labeled data and acoustic neglect are solved, enabling reference-free fine-grained speech recognition evaluation and improving the reliability and accuracy of the evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AISPEECH CO LTD
- Filing Date
- 2026-04-17
- Publication Date
- 2026-07-17
AI Technical Summary
Existing speech recognition result evaluation methods rely on labeled data, lack acoustic basis, cannot perform fine-grained analysis, and ignore the importance of the original audio, resulting in poor evaluation performance in unsupervised and acoustically sensitive scenarios.
An autoregressive text-to-speech (TTS) model is introduced, which calculates conditional probability and negative log-likelihood for each audio frame. The alignment relationship between audio frames and text is extracted by combining the attention mechanism inside the model, so as to achieve a referenceless acoustic consistency assessment of speech and text.
It provides fine-grained speech recognition result evaluation without reference text, enabling the judgment of the quality of recognition results under unsupervised conditions, accurately locating error positions, reducing data annotation costs, and improving the objectivity and interpretability of the evaluation.
Smart Images

Figure CN122416985A_ABST