一种语音识别方法、终端设备及存储介质
By transcoding and data augmentation of the raw audio data, acoustic and language models are constructed, solving the problems of insufficient data collection and channel diversity in low-resource language speech recognition, improving recognition performance and robustness, and reducing storage requirements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAMEN KUAISHANGTONG TECH CORP LTD
- Filing Date
- 2021-09-23
- Publication Date
- 2026-07-17
AI Technical Summary
Existing low-resource language speech recognition methods require the collection of data from similar languages for pre-training and do not consider channel diversity, resulting in insufficient recognition performance and robustness.
By collecting raw audio data, transcoding and data augmentation are performed to construct acoustic features, language models, and speaker recognition models. Combined with the TDNN acoustic model, a WFST graph is constructed as a speech recognition model, which is operated directly in the feature extraction stage, increasing the amount of data and channel diversity.
It improves the effectiveness and robustness of speech recognition, reduces the requirements for hard disk storage, and achieves efficient speech recognition.
Smart Images

Figure CN115862602B_ABST