视频理解方法、模型训练方法、装置以及电子设备

By constructing a hypernetwork video understanding model and combining data encoding, memory neurons, and task neurons, the interactive modeling problem in multimodal video understanding was solved, achieving stronger generalization ability and deeper information understanding, thus improving the performance of video understanding tasks.

CN119068384BActive Publication Date: 2026-07-17SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
Filing Date
2024-07-23
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing video understanding methods lack the ability to deeply model multimodal interactions when dealing with multimodal video information, resulting in poor performance in video understanding tasks.

Method used

A video understanding model based on hypernetworks is adopted, including data encoding neuron modules, memory neuron modules, and task neuron modules. Through multi-level neuron module interaction and recurrent memory association modules, the encoding, feature extraction, and task understanding of multimodal videos are realized.

Benefits of technology

It improves the model's generalization ability, enabling a deeper understanding of multimodal information and enhancing the performance of video understanding tasks, especially in tasks such as pixel-level video understanding, high-level video recognition, video-to-text dialogue, and text-to-video content editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119068384B_ABST
    Figure CN119068384B_ABST
Patent Text Reader

Abstract

本申请公开了一种视频理解方法、模型训练方法、装置以及电子设备,该方法用于通过训练好的视频理解模型进行视频任务理解,以输出视频理解结果;其中,视频理解模型是基于超网络构建的,其包括数据编码神经元模块、记忆神经元模块以及任务神经元模块;该方法包括获取待理解视频;通过数据编码神经元模块对待理解视频中的各模态数据进行编码,以获取各模态编码数据;通过记忆神经元模块对各模态编码数据进行特征提取,以获取各模态特征数据;将各模态编码数据以及各模态特征数据共同输入任务神经元模块,以输出视频理解结果。本申请采用了一种基于超网络构建的视频理解模型处理视频理解任务,这种网络结构能够提高模型性能,具有更强的泛化能力。
Need to check novelty before this filing date? Find Prior Art