Speech recognition result evaluation method, electronic device, and storage medium

By calculating the acoustic consistency between speech and text using an autoregressive TTS model, the problems of existing methods' dependence on labeled data and acoustic neglect are solved, enabling reference-free fine-grained speech recognition evaluation and improving the reliability and accuracy of the evaluation.

CN122416985APending Publication Date: 2026-07-17AISPEECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AISPEECH CO LTD
Filing Date
2026-04-17
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing speech recognition result evaluation methods rely on labeled data, lack acoustic basis, cannot perform fine-grained analysis, and ignore the importance of the original audio, resulting in poor evaluation performance in unsupervised and acoustically sensitive scenarios.

Method used

An autoregressive text-to-speech (TTS) model is introduced, which calculates conditional probability and negative log-likelihood for each audio frame. The alignment relationship between audio frames and text is extracted by combining the attention mechanism inside the model, so as to achieve a referenceless acoustic consistency assessment of speech and text.

Benefits of technology

It provides fine-grained speech recognition result evaluation without reference text, enabling the judgment of the quality of recognition results under unsupervised conditions, accurately locating error positions, reducing data annotation costs, and improving the objectivity and interpretability of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416985A_ABST
    Figure CN122416985A_ABST
Patent Text Reader

Abstract

This application discloses a speech recognition result evaluation method, electronic device, and storage medium. The method includes: acquiring an original speech signal and one or more candidate recognition texts generated by an ASR system; inputting the candidate recognition texts as conditions into a pre-trained autoregressive TTS model, calculating the conditional probability of the TTS model generating a real speech sequence frame by frame based on the original speech signal; calculating the negative log-likelihood frame by frame based on the conditional probability to obtain an acoustic difference sequence as a measure of acoustic consistency between speech and text; extracting the alignment relationship between audio frames and text tokens from within the TTS model; performing aggregate scoring on the negative log-likelihood based on the alignment relationship to obtain an aggregate scoring result; for the same original speech signal, comparing the aggregate scoring result between candidate texts and reference texts to achieve a reference-based evaluation, or comparing the aggregate scoring result between multiple candidate texts to achieve a referenceless evaluation.
Need to check novelty before this filing date? Find Prior Art