Long speech recognition model training method, electronic device, and storage medium

By jointly training endpoint detection and speech recognition models, and combining acoustic embedding features and spectral enhancement techniques, the problem of noise interference in long speech recognition is solved, thereby improving recognition accuracy and robustness.

CN115798460BActive Publication Date: 2026-02-06AISPEECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211573275.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2026-02-06
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

Existing technologies suffer from noise interference affecting recognition accuracy in long speech recognition, especially in complex scenarios, where traditional methods struggle to effectively improve the performance of long speech recognition.

Method used

By acquiring the constructed long speech training data, the endpoint detection model and the speech recognition model are jointly trained. The endpoint detection model is optimized by using the gradient backpropagation of the recognition model. The entire recognition link is optimized by combining acoustic embedding features and spectral enhancement technology.

Benefits of technology

It improves the accuracy and robustness of long speech recognition and effectively enhances recognition performance in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115798460B_ABST
    Figure CN115798460B_ABST
Patent Text Reader

Abstract

The application discloses a long speech recognition model training method, an electronic device and a storage medium, and the method comprises the following steps: obtaining long speech training data which is constructed, wherein the long speech training data comprises extracted acoustic input features, frame-level classification labels for training an endpoint detection model, and text labels for training a speech recognition model; and the endpoint detection model and the speech recognition model are jointly trained by using the long speech training data. According to the embodiment of the application, the endpoint detection model and the speech recognition model are jointly trained by obtaining the long speech training data which is constructed, and the related information provided by the recognition model is introduced to assist the training optimization of the endpoint detection model on the basis of optimizing the endpoint detection model, so that a complete joint optimization method is realized, and the recognition performance of the long speech link is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Joint endpointing and automatic speech recognition

    CN113841195A

  • Method and apparatus for detecting voice end point using acoustic and language modeling information for robust voice recognition

    US20220230627A1