流式多模编码融合的CANN架构的深度特征融合方法
By using the CANN architecture with streaming multimodal coding fusion, voice and video data are received and deeply fused in real time, which solves the shortcomings of existing systems in real-time processing and adaptation to domestic hardware. It achieves efficient and robust multimodal biometric recognition, which is suitable for real-time calls and high-security scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGJIANG TIMES COMM CO LTD
- Filing Date
- 2025-10-24
- Publication Date
- 2026-07-17
AI Technical Summary
Existing multimodal biometric recognition systems have shortcomings in real-time processing capabilities, deep semantic association mining, and adaptation to domestically produced hardware, resulting in insufficient recognition accuracy and robustness in complex environments, making it difficult to meet the needs of real-time calls and high-concurrency online services.
The CANN architecture, which uses streaming multi-modal coding fusion, receives voice and video data streams in real time via the WebSocket protocol. It extracts features using architectures such as ECAPA-TDNN, YOLOv5+ArcFace, and DFDT, and combines Transformer and Cross-Attention mechanisms for deep fusion and decision modeling to output a comprehensive identity vector.
It enables real-time processing and efficient deep fusion of voice and video data, improving the system's robustness and recognition accuracy. It supports low-latency, high-throughput real-time authentication and forgery detection, adapts to enterprise-level API integration environments, and supports model self-learning and feedback optimization.
Smart Images

Figure CN121278646B_ABST
Abstract
Citation Information
Patent Citations
Multi-modal forged video detection method based on multi-head addition cross attention mechanism
CN120635786A
KR20250019292A