基于视觉语言对齐差异优化的参数高效视频文本检索方法及设备
By training a video-level semantic alignment loss, an image-level semantic alignment loss, and an image-to-video alignment distillation loss, the video text retrieval model is improved, addressing the issues of insufficient language modeling capabilities and low cross-modal alignment accuracy in existing methods. This enhances the accuracy and performance of video text retrieval.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2025-06-04
- Publication Date
- 2026-07-17
AI Technical Summary
Existing video text retrieval methods struggle to balance parameter efficiency and model performance, exhibiting issues such as insufficient language modeling capabilities and low cross-modal alignment accuracy, leading to reduced accuracy in video text retrieval.
The video text retrieval model is trained using video-level semantic alignment loss, image-level semantic alignment loss, and image-to-video alignment distillation loss to improve the model's fine-grained alignment capability and alleviate key differences in vision, language, and alignment.
It significantly improves video-to-text retrieval performance, enhances the accuracy of video-to-text retrieval, and strengthens the model's language modeling capabilities and cross-modal alignment accuracy.
Smart Images

Figure CN120561609B_ABST