语音识别方法、装置、设备以及计算机可读介质

By fusing speech representation vectors and embedding vectors generated by speaker recognition models in a speech recognition system, the problem of feature mismatch is solved, and the accuracy of speech recognition is improved.

CN116994591BActive Publication Date: 2026-07-17IFLYTEK CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2023-09-04
Publication Date
2026-07-17

Smart Images

  • Figure CN116994591B_ABST
    Figure CN116994591B_ABST
Patent Text Reader

Abstract

本发明提供一种语音识别方法、装置、设备以及计算机可读介质,该方法通过获取语音信号;将语音信号输入到初始系统的预训练模型中,由预训练模型将语音信号转化为语音表示向量;将语音表示向量输入至初始系统中的说话人识别模型,得到嵌入向量;将语音表示向量输入至初始系统的语音识别模型中,由语音识别模型对嵌入向量和语音表示向量进行融合处理,得到融合特征向量;根据融合特征向量和实际的语音识别结果对初始系统中的模型进行训练,得到语音识别系统。由于嵌入向量是通过语音表示向量得到的,因此嵌入向量和语音表示向量的融合不存在特征不匹配现象,进而提升了训练出的语音识别系统的准确性。
Need to check novelty before this filing date? Find Prior Art