Long speech recognition model training method, electronic device, and storage medium
By jointly training endpoint detection and speech recognition models, and combining acoustic embedding features and spectral enhancement techniques, the problem of noise interference in long speech recognition is solved, thereby improving recognition accuracy and robustness.
Patent Information
- Application Number
- CN202211573275.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-08
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-12-08
AI Technical Summary
Existing technologies suffer from noise interference affecting recognition accuracy in long speech recognition, especially in complex scenarios, where traditional methods struggle to effectively improve the performance of long speech recognition.
By acquiring the constructed long speech training data, the endpoint detection model and the speech recognition model are jointly trained. The endpoint detection model is optimized by using the gradient backpropagation of the recognition model. The entire recognition link is optimized by combining acoustic embedding features and spectral enhancement technology.
It improves the accuracy and robustness of long speech recognition and effectively enhances recognition performance in noisy environments.
Smart Images

Figure CN115798460B_ABST
Abstract
Citation Information
Patent Citations
Joint endpointing and automatic speech recognition
CN113841195A
Method and apparatus for detecting voice end point using acoustic and language modeling information for robust voice recognition
US20220230627A1