一种语音识别方法、终端设备及存储介质

By transcoding and data augmentation of the raw audio data, acoustic and language models are constructed, solving the problems of insufficient data collection and channel diversity in low-resource language speech recognition, improving recognition performance and robustness, and reducing storage requirements.

CN115862602BActive Publication Date: 2026-07-17XIAMEN KUAISHANGTONG TECH CORP LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIAMEN KUAISHANGTONG TECH CORP LTD
Filing Date
2021-09-23
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing low-resource language speech recognition methods require the collection of data from similar languages ​​for pre-training and do not consider channel diversity, resulting in insufficient recognition performance and robustness.

Method used

By collecting raw audio data, transcoding and data augmentation are performed to construct acoustic features, language models, and speaker recognition models. Combined with the TDNN acoustic model, a WFST graph is constructed as a speech recognition model, which is operated directly in the feature extraction stage, increasing the amount of data and channel diversity.

Benefits of technology

It improves the effectiveness and robustness of speech recognition, reduces the requirements for hard disk storage, and achieves efficient speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862602B_ABST
    Figure CN115862602B_ABST
Patent Text Reader

Abstract

本发明涉及一种语音识别方法、终端设备及存储介质,该方法中包括:S1:采集原始音频数据;S2:对原始音频数据进行转码和数据增强处理后,将三种音频合并组成训练集;S3:提取训练集中各音频的声学特征;S4:构建3‑gram语言模型并进行训练;S5:构建单音素声学模型,并基于单音素构建三音素声学模型,通过训练集中各音频的声学特征模型进行训练;S6:构建说话人识别模型;S7:构建TDNN声学模型,通过说话人识别模型和三音素声学模型对训练集中各音频的声学特征的识别结果对TDNN声学模型进行训练;S8:通过发音词典、声学模型和语言模型共同构建语音识别模型;S9:通过语音识别模型进行语音识别。本发明增加信道的多样性,提升了系统的识别效果及鲁棒性。
Need to check novelty before this filing date? Find Prior Art