一种基于时频空多维特征挖掘和跨模态注意力融合的数据处理方法
By employing a Transformer-based temporal-frequency-spatial multidimensional feature mining and cross-modal attention fusion method, this study addresses the shortcomings in high-order feature modeling and cross-modal feature interaction in existing depression detection technologies. This approach enables highly accurate depression detection in complex scenarios and is suitable for remote diagnosis and treatment as well as smart wearables.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA CRIMINAL POLICE UNIV
- Filing Date
- 2026-04-03
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies struggle to effectively model high-order nonlinear features related to depression, and fail to capture sufficient interactions between cross-modal features, resulting in limited generalization performance and poor robustness in complex real-world scenarios.
We employ a Transformer-based spatiotemporal multidimensional feature mining method, combined with cross-attention modules and cross-modal attention mechanisms, to construct a hierarchical spatiotemporal graph network that integrates video, speech, and text features to improve detection performance.
It improves the accuracy and interpretability of depression detection, can effectively identify depressive states in complex scenarios, and is suitable for remote diagnosis and treatment, smart wearables, and psychological assessment.
Smart Images

Figure CN121971093B_ABST
Abstract
Citation Information
Patent Citations
Multi-modal fusion depression screening method, device and equipment based on end-to-end
CN117854727A