基于视觉语言对齐差异优化的参数高效视频文本检索方法及设备

By training a video-level semantic alignment loss, an image-level semantic alignment loss, and an image-to-video alignment distillation loss, the video text retrieval model is improved, addressing the issues of insufficient language modeling capabilities and low cross-modal alignment accuracy in existing methods. This enhances the accuracy and performance of video text retrieval.

CN120561609BActive Publication Date: 2026-07-17TSINGHUA UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2025-06-04
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing video text retrieval methods struggle to balance parameter efficiency and model performance, exhibiting issues such as insufficient language modeling capabilities and low cross-modal alignment accuracy, leading to reduced accuracy in video text retrieval.

Method used

The video text retrieval model is trained using video-level semantic alignment loss, image-level semantic alignment loss, and image-to-video alignment distillation loss to improve the model's fine-grained alignment capability and alleviate key differences in vision, language, and alignment.

Benefits of technology

It significantly improves video-to-text retrieval performance, enhances the accuracy of video-to-text retrieval, and strengthens the model's language modeling capabilities and cross-modal alignment accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561609B_ABST
    Figure CN120561609B_ABST
Patent Text Reader

Abstract

本发明提供了一种基于视觉语言对齐差异优化的参数高效视频文本检索方法及设备,涉及机器学习领域。包括:获取样本视频和相匹配的样本文本描述;采样多帧样本图像并为每帧样本图像生成对应的伪样本文本描述;根据样本视频特征和样本文本特征确定视频级相似度,基于视频级相似度得到视频级语义对齐损失;根据样本图像特征和伪样本文本特征确定图像级相似度,基于图像级相似度得到图像级语义对齐损失;基于图像级相似度与视频级相似度得到图像到视频对齐蒸馏损失;基于视频级语义对齐损失、图像级语义对齐损失以及图像到视频对齐蒸馏损失,对待训练的视频文本检索模型进行训练得到目标视频文本检索模型,以提高视频文本检索的精度。
Need to check novelty before this filing date? Find Prior Art