序列视频中无对齐文本的弱监督视频表示学习方法
By learning video representations under unaligned text conditions using a multi-granularity contrastive learning loss function, the problem of video and text misalignment is solved, and the performance of downstream tasks in video understanding is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI TECH UNIV
- Filing Date
- 2023-01-31
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies struggle to effectively learn video representations in the presence of unaligned text and video, and multimodal video representation models are ill-suited to weakly supervised environments.
A multi-granularity contrastive learning loss function, including coarse-grained and fine-grained loss functions, is adopted to constrain video frame features and text features through visual and language models, thereby achieving alignment learning between video and text.
It enables the learning of powerful video representations under unaligned text conditions, improving the generalization ability of downstream tasks such as step-by-step video sequence verification and text-to-video matching.
Smart Images

Figure CN116052054B_ABST