序列视频中无对齐文本的弱监督视频表示学习方法

By learning video representations under unaligned text conditions using a multi-granularity contrastive learning loss function, the problem of video and text misalignment is solved, and the performance of downstream tasks in video understanding is improved.

CN116052054BActive Publication Date: 2026-07-17SHANGHAI TECH UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI TECH UNIV
Filing Date
2023-01-31
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively learn video representations in the presence of unaligned text and video, and multimodal video representation models are ill-suited to weakly supervised environments.

Method used

A multi-granularity contrastive learning loss function, including coarse-grained and fine-grained loss functions, is adopted to constrain video frame features and text features through visual and language models, thereby achieving alignment learning between video and text.

Benefits of technology

It enables the learning of powerful video representations under unaligned text conditions, improving the generalization ability of downstream tasks such as step-by-step video sequence verification and text-to-video matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052054B_ABST
    Figure CN116052054B_ABST
Patent Text Reader

Abstract

本发明公开一种序列视频中无对齐文本的弱监督视频表示学习方法,其特征在于,包括以下步骤:获得帧特征、视频整体特征、句子特征和段落特征;使用多粒度对比学习损失函数来限制帧特征、视频整体特征、句子特征和段落特征,多粒度对比学习损失函数包括粗粒度的损失函数以及细粒度的损失函数。本发明提出了针对连续视频提出了一种新的具有未对齐文本的弱监督视频表征学习框架,并引入了多粒度对比损失来约束该网络模型,使得模型充分考虑了帧和句子之间的伪时间对齐,可以学习强大的具有语义的视频文本对表征。本发明提供的模型还展现出对下游任务的强大泛化能力,例如步骤级视频序列验证和文本到视频的匹配。
Need to check novelty before this filing date? Find Prior Art