用于确定文本和视频之间的相似度的方法和装置

By combining text and video feature extraction models with the similarity of image feature sequences and word features, constrained similarity of similar images is generated, solving the accuracy problem of semantic similarity calculation between text and video and achieving more efficient similarity recognition.

CN116958868BActive Publication Date: 2026-07-17ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
Filing Date
2023-07-21
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies rely on the representational power of semantic features to accurately calculate the semantic similarity between text and video, resulting in inaccurate calculation results.

Method used

Text feature extraction model and video feature extraction model are used to extract features from text and video respectively. Similar image features are determined by the similarity in the image feature sequence, and similar image constraint similarity is generated by combining the similarity between word features and image features. Finally, the similarity between text and video is determined.

Benefits of technology

It improves the accuracy of semantic similarity calculation between text and video, enhances the ability to recognize cross-modal similarity, and can better distinguish between matching and non-matching pairs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958868B_ABST
    Figure CN116958868B_ABST
Patent Text Reader

Abstract

本说明书的实施例提供了一种用于确定文本和视频之间的相似度的方法和装置。在该用于确定文本和视频之间的相似度的方法中,将所获取的文本视频对包括的文本和视频分别提供给文本特征提取模型和视频特征提取模型,得到对应的词符特征序列和图像特征序列;根据各个词符特征与各个图像特征之间的相似度确定相关词符特征‑图像特征对;针对各个相关词符特征‑图像特征对,对该词符特征与该图像特征之间的相似度和所确定的该图像特征对应的相近图像特征与词符特征序列之间的相似度进行聚合,生成相近图像约束相似度;以及基于所得到的相近图像约束相似度,确定文本视频对中的文本和视频之间的相似度。
Need to check novelty before this filing date? Find Prior Art