视频理解方法、模型训练方法、装置以及电子设备
By constructing a hypernetwork video understanding model and combining data encoding, memory neurons, and task neurons, the interactive modeling problem in multimodal video understanding was solved, achieving stronger generalization ability and deeper information understanding, thus improving the performance of video understanding tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
- Filing Date
- 2024-07-23
- Publication Date
- 2026-07-17
AI Technical Summary
Existing video understanding methods lack the ability to deeply model multimodal interactions when dealing with multimodal video information, resulting in poor performance in video understanding tasks.
A video understanding model based on hypernetworks is adopted, including data encoding neuron modules, memory neuron modules, and task neuron modules. Through multi-level neuron module interaction and recurrent memory association modules, the encoding, feature extraction, and task understanding of multimodal videos are realized.
It improves the model's generalization ability, enabling a deeper understanding of multimodal information and enhancing the performance of video understanding tasks, especially in tasks such as pixel-level video understanding, high-level video recognition, video-to-text dialogue, and text-to-video content editing.
Smart Images

Figure CN119068384B_ABST