语音识别模型的训练方法及装置、存储介质及电子设备
By introducing real noise data into the pre-training stage of the speech recognition model, the model's ability to learn audio features in noisy environments is improved. Furthermore, by constructing a target speech recognition model through masking representation and fine-tuning, the problem of poor speech recognition performance in noisy environments is solved, achieving higher robustness and recognition accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2023-03-23
- Publication Date
- 2026-07-17
AI Technical Summary
Existing speech recognition models are not robust in noisy environments, resulting in poor recognition performance.
By introducing real noise data for pre-training, the model's ability to learn audio features in noisy environments is improved. The noise data is used as a mask representation to ensure that the encoding strategy is consistent between the pre-training and subsequent training stages. The initial speech recognition model is fine-tuned by combining labeled speech data to construct the target speech recognition model.
This improved the robustness and recognition accuracy of the speech recognition model in noisy environments, and enhanced the model's performance under noisy conditions.
Smart Images

Figure CN116343781B_ABST