流式多模编码融合的CANN架构的深度特征融合方法

By using the CANN architecture with streaming multimodal coding fusion, voice and video data are received and deeply fused in real time, which solves the shortcomings of existing systems in real-time processing and adaptation to domestic hardware. It achieves efficient and robust multimodal biometric recognition, which is suitable for real-time calls and high-security scenarios.

CN121278646BActive Publication Date: 2026-07-17CHANGJIANG TIMES COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHANGJIANG TIMES COMM CO LTD
Filing Date
2025-10-24
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing multimodal biometric recognition systems have shortcomings in real-time processing capabilities, deep semantic association mining, and adaptation to domestically produced hardware, resulting in insufficient recognition accuracy and robustness in complex environments, making it difficult to meet the needs of real-time calls and high-concurrency online services.

Method used

The CANN architecture, which uses streaming multi-modal coding fusion, receives voice and video data streams in real time via the WebSocket protocol. It extracts features using architectures such as ECAPA-TDNN, YOLOv5+ArcFace, and DFDT, and combines Transformer and Cross-Attention mechanisms for deep fusion and decision modeling to output a comprehensive identity vector.

Benefits of technology

It enables real-time processing and efficient deep fusion of voice and video data, improving the system's robustness and recognition accuracy. It supports low-latency, high-throughput real-time authentication and forgery detection, adapts to enterprise-level API integration environments, and supports model self-learning and feedback optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121278646B_ABST
    Figure CN121278646B_ABST
Patent Text Reader

Abstract

本发明公开了一种流式多模编码融合的CANN架构的深度特征融合方法,步骤包括:实时接收语音数据流以及视频数据流;对语音数据流以及视频数据流进行特征提取;将说话人特征、人脸深度特征以及音视频鉴伪特征进行特征对齐;将特征向量输入到基于CANN架构的融合引擎中进行深度融合与判决建模;由上层业务系统对综合身份向量进行精细化决策。该深度特征融合方法利用CANN架构的多模态特征融合机制,利用Transformer风格的自注意力机制提升模态内特征建模能力,通过Cross‑Attention机制增强模态间协同与对抗建模,采用Feature Gating机制动态调整每个模态的权重以提升融合鲁棒性,最终输出一个综合身份向量,用于身份验证、异常检测或伪造判定。
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Multi-modal forged video detection method based on multi-head addition cross attention mechanism

    CN120635786A

  • KR20250019292A