用于确定文本和视频之间的相似度的方法和装置
By combining text and video feature extraction models with the similarity of image feature sequences and word features, constrained similarity of similar images is generated, solving the accuracy problem of semantic similarity calculation between text and video and achieving more efficient similarity recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
- Filing Date
- 2023-07-21
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies rely on the representational power of semantic features to accurately calculate the semantic similarity between text and video, resulting in inaccurate calculation results.
Text feature extraction model and video feature extraction model are used to extract features from text and video respectively. Similar image features are determined by the similarity in the image feature sequence, and similar image constraint similarity is generated by combining the similarity between word features and image features. Finally, the similarity between text and video is determined.
It improves the accuracy of semantic similarity calculation between text and video, enhances the ability to recognize cross-modal similarity, and can better distinguish between matching and non-matching pairs.
Smart Images

Figure CN116958868B_ABST