一种基于质量感知交错标记融合的音视频语音识别方法
By adopting a layered and decoupled audio-visual-speech recognition model, the problems of high noise propagation risk and insufficient robustness in existing technologies are solved, achieving efficient recognition in complex environments and making it suitable for a variety of application scenarios.
CN122201262BActive Publication Date: 2026-07-17SHANDONG UNIV OF SCI & TECH
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG UNIV OF SCI & TECH
- Filing Date
- 2026-05-13
- Publication Date
- 2026-07-17
Smart Images

Figure CN122201262B_ABST
Abstract
本发明属于语音识别技术领域,具体公开了一种基于质量感知交错标记融合的音视频语音识别方法。该方法搭建了分层解耦的音视频语音识别模型,模型中设计了单模态时序稳定化模块,对投影后的音频特征与视频特征进行双向状态空间建模。同时本发明还设计了局部时序增强模块,通过卷积对稳定化音频特征与稳定化视频特征进一步增强。此外,本发明还设计了质量感知交错标记融合模块,通过内容相关性权重与模态可靠性分数联合生成逐帧融合权重,并通过交错标记构造模块构造带有显式帧级结构约束的交错标记序列,然后通过帧对齐交叉门控残差细化模块对同一时间帧内的双模态信息进行双向受控信息注入。本发明能够在复杂环境下实现高精度的音视频语音识别。
Need to check novelty before this filing date? Find Prior Art