A method for constructing a text-video pair similarity evaluation model

CN119149772BActive Publication Date: 2026-09-25SOUTHWESTERN UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411176128.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-09-25
Estimated Expiration
2044-08-26

AI Technical Summary

Technical Problem

[0016]综上,当前主流的文本-视频检索技术,以CLIP为代表,并侧重于增强特征表征学习和粒度对齐策略的研究;然而,这些方法常常忽视一个重要问题:由于文本视觉模态之间存在显著的语义差异,在增强表征学习和交互对齐方面存在困难

Benefits of technology

[0095]本发明的有益效果是:本发明的模型,首先,对样本对的图像序列与文本,进行整体的相似性计算获得;然后,对样本对的图像序列所包含图像帧与文本,进行帧级的相似性计算获得;之后,对图像序列所包含视觉实体与文本所包含单词,进行因子级的相似性计算获得。因此,在检索时,能通过多层次的相似度,引入了更多的特征信息,并从多个粒度对文本和视频的相似性进行比较,能够降低文本特征与视觉特征在语义上的不对等所导致的影响,并显著提升了检索性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119149772B_ABST
    Figure CN119149772B_ABST
Patent Text Reader

Abstract

The present application relates to the field of information retrieval, and discloses a text-video similarity evaluation model construction method, first, input positive sample pair and construct its negative sample pair; then, through the visual encoder, the visual features of the sample pair are obtained, and through the text encoder, the text features of the sample pair are obtained; thereafter, the coarse, medium and fine-grained similarities of the sample pair are calculated respectively; wherein, the coarse-grained similarity is the overall similarity calculation of the video and the text; the medium-grained similarity is the frame-level similarity calculation of the image frames contained in the video and the text; the fine-grained similarity is the factor-level similarity calculation of the visual entities contained in the video and the words contained in the text. In retrieval, more features can be introduced, the similarities of the text and the video are compared from multiple granularities, the influence caused by the semantic inequality of the text and the vision can be reduced, and the retrieval performance is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information retrieval, and more specifically to a method for constructing a text-video pair similarity assessment model. Background Technology

[0002] In recent years, with the rise of multimedia platforms such as TikTok, YouTube, and Netflix, users have increasingly higher demands for the accuracy of text-based video content retrieval. Currently, there are roughly three types of text-video retrieval methods:

[0003] I. Modal retrieval based on traditional manual design

[0004] Traditional text-video multimodal retrieval methods include keyword matching, which retrieves video metadata tags; feature extraction, which extracts low-level features such as color histograms and TF-IDF from videos and texts, and matches them based on feature vector similarity; and multimodal fusion, which uses statistical techniques such as canonical correlation analysis to map text and video features to a common representation space.

[0005] However, traditional text-video multimodal retrieval methods rely on manual annotation, which can lead to incomplete or inaccurate labels that affect performance. Furthermore, traditional feature extraction methods struggle to capture complex semantic information, and fixed rules and features lack flexibility and generalization ability.

[0006] Therefore, although it has some practicality in simple and resource-constrained scenarios, it has been gradually replaced by deep learning and CLIP-based methods for complex semantic matching needs.

[0007] II. Modal Retrieval Based on Deep Learning

[0008] Deep learning-based modal retrieval methods have significantly improved retrieval performance. These methods typically employ a two-stream network encoder to extract features from text and video separately, then perform matching in a fusion layer. For example, recurrent neural networks (RNNs) and convolutional neural networks (CNNs) excel at handling sequential and spatial features, respectively; combining them can effectively capture temporal dependencies and spatial features in both video and text. Cross-modal attention mechanisms, such as the Transformer model, capture the contextual relationships between video and text through self-attention. However, these methods require large amounts of data for training and are computationally expensive.

[0009] Nevertheless, deep learning methods perform exceptionally well in handling complex semantic matching tasks, significantly outperforming traditional methods.

[0010] III. Modality Retrieval Based on Pre-trained CLIP

[0011] The CLIP-based approach has achieved significant progress in text-video multimodal retrieval. CLIP stands for Contrastive Language-Image Pre-training. Through contrastive learning, the CLIP model maps text and video frames to the same high-dimensional vector space. By pre-training on large-scale datasets, CLIP learns rich semantic information, enabling precise matching of text and video content within the same space. In text-video retrieval, CLIP can directly vectorize text and video frames and perform retrieval by calculating the cosine similarity of the vectors. Due to its pre-training on large-scale datasets, CLIP possesses strong feature representation and generalization capabilities, achieving remarkable results in multimodal retrieval tasks.

[0012] Although CLIP requires significant computational resources during training, it performs exceptionally well in practical applications, especially in handling complex semantic matching tasks, where it significantly outperforms traditional methods and other deep learning approaches.

[0013] As mentioned above, the current mainstream text-video retrieval technology, represented by CLIP, has also made significant progress and achievements in further research. Further research on CLIP can be broadly divided into two categories: one is enhancing CLIP's representation learning of extracted text and visual features, and the other is the interactive alignment of text and visual features.

[0014] The core of the first type of work lies in the learning of feature representations. For example, EM-Net uses the expectation-maximization algorithm to compactly represent visual and text features, enhancing the semantic representation ability of text and visual features; T-Mass uses random text modeling and text regularization methods to extract effective frames and text, enhancing the semantic similarity between text and video.

[0015] The core of the second type of work lies in the interaction between textual and visual features. For example, MSIA introduces a multi-level semantic interaction model, using adaptive frames and text-guided attention mechanisms to reduce video redundancy and enhance modal interaction. HBI uses a novel multi-cooperative game theory approach for coarse-grained and fine-grained interactions, involving information such as actions, scenes, and entities in visual information.

[0016] In summary, current mainstream text-video retrieval technologies, represented by CLIP, focus on enhancing feature representation learning and granular alignment strategies. However, these methods often overlook a crucial issue: the significant semantic differences between text and visual modalities pose challenges in enhancing representation learning and interaction alignment. Specifically, images possess richer details and content, while manually labeled text information is typically concise. This semantic asymmetry leads to difficulties in content matching between text and visual features. This core difference directly impacts feature enhancement representation learning and subsequent granular alignment, making it difficult to accurately match features from different modalities within the same space, thus affecting retrieval performance. Summary of the Invention

[0017] The technical problem to be solved by this invention is to propose a method for constructing a text-video similarity evaluation model, which can reduce the impact of semantic asymmetry between text features and visual features and significantly improve retrieval performance.

[0018] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:

[0019] A method for constructing a text-video pair similarity evaluation model, wherein the text-video alignment model includes a visual encoder, a text encoder, and an alignment model, wherein the visual encoder and the text encoder are pre-trained models; the training of the text-video alignment model includes:

[0020] A1. Extract text-video pairs from the original dataset, wherein each text-video pair includes a video and its corresponding text; obtain image sequences from the video of each text-video pair, wherein each image sequence consists of a set of image frames; the image sequences obtained by sampling and the corresponding text of the video constitute the training sample pairs of the text-video pair; the training sample pairs of each text-video pair in the original dataset constitute the training set.

[0021] A2. Input training sample pairs as positive sample pairs for this round of training; construct negative sample pairs for the positive sample pairs, wherein the negative sample pairs include text negative sample pairs and / or video negative sample pairs. The text negative sample pairs are sample pairs consisting of the text contained in the positive sample pairs and the image sequences contained in other training sample pairs besides the positive sample pairs. The video negative sample pairs are sample pairs consisting of the image sequences contained in the positive sample pairs and the text contained in other training sample pairs besides the positive sample pairs.

[0022] For each sample pair containing an image sequence, the visual encoder is input to obtain the visual features of each image frame contained therein; for each sample pair containing text, the text encoder is input to obtain its text features.

[0023] A3. For each sample pair, input its text features and visual features into the alignment model to calculate its coarse-grained similarity, medium-grained similarity, and fine-grained similarity.

[0024] The coarse-grained similarity is obtained by performing an overall similarity calculation on the image sequence and text of the sample pair based on the input text features and visual features.

[0025] The medium-granularity similarity is obtained by performing frame-level similarity calculations on the image frames and text contained in the image sequence of the sample pair, based on the input text features and visual features.

[0026] The fine-grained similarity is obtained by performing factor-level similarity calculations on visual entities contained in the image sequence and words contained in the text, based on input text features and visual features.

[0027] A4. Utilize the coarse-grained, medium-grained, and fine-grained similarity of positive and negative sample pairs to calculate the coarse-grained, medium-grained, and fine-grained feature alignment losses, and use this to calculate the total loss for this round of training; based on the total loss, update the parameters of the alignment model and fine-tune the parameters of the visual encoder and text encoder.

[0028] A5. Repeat steps A2 to A4 until the training termination condition is met, and obtain the aligned model that has completed training.

[0029] Furthermore, the visual encoder and text encoder are the visual encoder and text encoder of the pre-trained CLIP model, respectively.

[0030] In step A1, image sequences are obtained from the videos of each text-video pair by means of random sampling or average sampling;

[0031] In step A2, at least two training sample pairs are input as positive sample pairs for this round of training; the negative sample pairs include text negative sample pairs and video negative sample pairs; for each positive sample pair, the video negative sample is formed by combining the image sequence it contains with the text of other input positive sample pairs, and the text negative sample is formed by combining the text it contains with the image sequence of other input positive sample pairs.

[0032] Furthermore, the alignment model also includes a feature compression module; for the textual and visual features input to the alignment model, the alignment model first uses the feature compression module to compress the input visual features, and then calculates coarse-grained similarity, medium-grained similarity, and fine-grained similarity; the compression of the input visual features includes:

[0033] The visual features of the input alignment model are used as the original visual features v. o The redundant parts contained therein are calculated and used as redundant visual features v. r Remove the original visual features v using the following formula. o Redundant visual features v r Obtain the compressed visual features v c This is used as input to calculate coarse-grained similarity, medium-grained similarity, and fine-grained similarity;

[0034] v c =v o -v r

[0035] The redundant visual features v obtained by the calculation r It should satisfy:

[0036]

[0037] Where s(·) represents the similarity function, v o Let t represent the raw visual features of the input alignment model, and v represent the text features of the input alignment model. r To represent the original visual features v o Redundant visual features of the redundant part, v c To remove visual features v o Visual features after redundancy; Min means minimization, Max means maximization.

[0038] Furthermore, a similarity-aware compression factor C is defined, and v is used as the basis for this definition. r =C or v r =C·ε, calculate to obtain redundant visual features v r Where ε is a random factor; the similarity-perceived compression factor C is calculated using any of the following methods:

[0039] Method 1: C = Expand(S)

[0040] Method 2: C = Expand(exp(γS))

[0041] Method 3

[0042] Method 4: C = Expand(exp(MLP(S)))

[0043] Where S = [s1, s2, ..., s i …,s F ], s i= s(t,fi), MLP(S) = LN(Rule(LN(S))); s(·) represents the similarity function, fi represents the visual feature of the i-th image frame in the image sequence, s i F represents the similarity between the i-th image frame in the image sequence and the text feature t; C is the similarity-aware compression factor; MLP is a fully connected network; LN is a linear layer in the fully connected network; Rule is the Rule activation function; exp is the exponential function; Expand represents the similarity between the i-th image frame and the text feature t. r Expand the dimension to make it compatible with v o An expansion function that is consistent with the dimension.

[0044] Furthermore, the coarse-grained similarity is obtained by performing an overall similarity calculation on the image sequence and text of the sample pair based on input text features and visual features, including:

[0045] First, based on the similarity between the visual features and text features of each image frame in the image sequence of the sample pair, the weights of each image frame in the image sequence of the sample pair are constructed. Then, the visual features of each image frame in the sample pair are aggregated by weight summation to obtain the coarse-grained video representation of the sample pair. After that, the similarity between the coarse-grained video representation and the text features of the sample pair is calculated as the coarse-grained similarity of the sample pair.

[0046] The medium-granularity similarity is obtained by performing frame-level similarity calculations on the image frames and text contained in the image sequence of the sample pair, based on input text features and visual features, including:

[0047] First, a query is constructed using the text features of the sample pair, and keys and values ​​are constructed using the visual features of the sample pair. Using a cross-attention mechanism, the embedded representations of each image frame contained in the image sequence of the sample pair are obtained. Then, based on the embedded representations of each image frame, the granular representation of the sample pair in the video is obtained. After that, the similarity between the granular representation of the sample pair in the video and its text features is calculated as the granular similarity of the sample pair.

[0048] The fine-grained similarity is obtained by performing factor-level similarity calculations on visual entities contained in an image sequence and words contained in the text, based on input text features and visual features. This includes:

[0049] First, the text features of the sample pair are decomposed into K text factors, and the coarse-grained video representation of the input sample pair is decomposed into K video factors, forming K video-text factor pairs, where K is the number of words contained in the text of the input sample pair. Then, the similarity between the text factors and video factors contained in each video-text factor pair is calculated. Finally, based on the similarity between the text factors and video factors of each video-text factor pair, the fine-grained similarity of the sample pair is obtained.

[0050] Furthermore, the calculation of the coarse-grained representation of the video includes:

[0051] First, based on the similarity between the visual features and textual features of each image frame in the image sequence of the sample pair, the weight 'a' of each image frame in the image sequence of the sample pair is constructed according to the following formula. i :

[0052]

[0053] Then, using the following formula and a weighted summation method, the visual features of each image frame contained in the sample pair are aggregated to obtain the coarse-grained video representation v of the sample pair. cg :

[0054]

[0055] Then, the coarse-grained video representation v of the sample pair is calculated according to the following formula. cg The similarity between the text feature t and the sample pair is considered as the coarse-grained similarity s. cg :

[0056]

[0057] Among them, f i and f j Let represent the visual features of the i-th and j-th image frames in the image sequence, respectively; τ is a hyperparameter; F represents the number of image frames in the image sequence; the superscript T indicates matrix transpose; exp is the exponential function; and |·| indicates the modulo operation.

[0058] Furthermore, in the calculation of the medium-granularity similarity, the query is constructed using the textual features of the sample pair, and the key and value are constructed using the visual features of the sample pair. A cross-attention mechanism is then used to obtain the embedded representation of each image frame contained in the image sequence of the sample pair:

[0059]

[0060] Q t =LN(t) T W Q

[0061] K v =LN(v)W K

[0062] V v =LN(v c Wv

[0063] Among them, W Q W K and W V Construct query representation Q respectively t K-bond characterization v Sum value characterization V v The transformation matrix; D p For K v The feature dimension; softmax represents the softmax function, the superscript T represents matrix transpose, and LN represents a linear layer;

[0064] Calculate the granularity representation v in the video of the sample pair using the following formula. mg The similarity between the text feature t and the sample pair is considered as the medium-granularity similarity s. mg :

[0065]

[0066] Where |·| represents the modulo operation;

[0067] In the calculation of medium-granularity similarity, the medium-granularity representation v of the sample pair is obtained based on the embedding representation of each image frame according to the following formula. mg :

[0068] v mg =LN(r(v|t))+FC(r(v|t))) T

[0069] r(v|t)=LN(Attention(Q t ,K v V v W1)

[0070] in, This represents the embedding representation of each image frame contained in the image sequence, where W1 is the learnable weight, LN is the linear layer, and FC represents the fully connected layer.

[0071] Furthermore, the calculation of the fine-grained similarity includes:

[0072] First, the text features of the sample pair are decomposed into K text factors fw according to the following formula. k The coarse-grained video representation of the input sample pair is decomposed into K video factors fe. kAnd form K video text factor pairs:

[0073] fw k =W k t

[0074] fe k =W k v cg

[0075] Among them, v cg For coarse-grained representation of the video, t represents the text feature, and W represents the text feature. k Let k be the k-th learnable decomposition factor, where k is the index and k = 1, 2, ..., K, and K is the number of words contained in the text of the sample pair.

[0076] Then, calculate the similarity between the text factors and video factors contained in each video text factor pair according to the following formula;

[0077] s=(FW) T FE

[0078] Where the superscript T denotes matrix transpose, FE = [fe1,fe2,…,fe k …,fe K [FW] is the video factor matrix, where FW = [fw1, fw2, ..., fw] k …,fw K [ ] represents the text factor matrix.

[0079] Furthermore, in the calculation of the fine-grained similarity, the fine-grained similarity s of the sample pairs is obtained based on the similarity between the text factors and video factors of each video text factor, according to the following formula. fg :

[0080] s fg =s·g

[0081] g = MLP(cat[FW,FE])

[0082] MLP(cat[FW,FE])=LN(Rule(LN(cat[FW,FE])))

[0083] Where g is the similarity confidence score, MLP is a fully connected network, LN is a linear layer in a fully connected network, Rule is the Rule activation function, and cat is a concatenation function.

[0084] Furthermore, the negative sample pairs include text negative sample pairs and video negative sample pairs;

[0085] Calculate the text-to-video alignment loss using the following formula.

[0086]

[0087] Calculate the video-to-text alignment loss using the following formula.

[0088]

[0089] Calculate the alignment loss using the following formula:

[0090]

[0091] Where M represents the number of negative video samples plus one, and the plus one represents the positive sample corresponding to the negative video sample; N represents the number of negative text samples plus one, and the plus one represents the positive sample corresponding to the negative text sample; B is the number of positive sample pairs in this round of training; e represents the natural index; λ is the hyperparameter representing the scaling factor; and s(·) represents the similarity function.

[0092] The total loss for:

[0093]

[0094] Where α and β are weight hyperparameters, This represents the coarse-grained feature alignment loss, which is the loss that aligns the coarse-grained representation of the video v. cg The text feature t is taken as input and obtained by alignment loss calculation; This represents the medium-granularity feature alignment loss, which is the loss that aligns the granular representations v in the video. mg The text feature t is taken as input and obtained by alignment loss calculation; This represents the fine-grained feature alignment loss, which is calculated by taking the video factor matrix FE = [fe1,fe2,…,fe2]. k …,fe K ] and the text factor matrix FW = [fw1, fw2, ..., fw k …,fw K ] is used as input and obtained through alignment loss calculation; fe k For the k-th video factor, fw k Let k be the k-th text factor.

[0095] The beneficial effects of this invention are as follows: First, the model of this invention performs an overall similarity calculation on the image sequence and text of the sample pair. Then, it performs a frame-level similarity calculation on the image frames and text contained in the image sequence of the sample pair. Finally, it performs a factor-level similarity calculation on the visual entities contained in the image sequence and the words contained in the text. Therefore, during retrieval, more feature information is introduced through multi-level similarity, and the similarity of text and video is compared at multiple granularities. This reduces the impact of semantic asymmetry between text features and visual features and significantly improves retrieval performance. Attached Figure Description

[0096] Figure 1 This is a schematic diagram illustrating how video semantic information can interfere with modal retrieval.

[0097] Figure 2 This is a schematic diagram of the retrieval process based on the text-video alignment model of the present invention;

[0098] Figure 3 This is a visualization of the search results based on an embodiment of the present invention. Detailed Implementation

[0099] This invention aims to propose a method for constructing a text-video similarity evaluation model. The constructed text-video alignment model is used to compare the similarity between video and text, and can be applied to both text-based video retrieval and video-based text retrieval. First, it achieves coarse-grained information interaction based on the interaction alignment of the entire video with the entire sentence. Second, it directly obtains the video frames with the most representative information using a cross-attention mechanism, and achieves medium-grained information interaction through frame alignment with the entire sentence, supplementing the coarse-grained information. Third, it refines the coarse-grained information by obtaining the alignment of visual entities with text words, achieving fine-grained information interaction. This invention introduces more feature information through hierarchical feature alignment and compares the similarity between text and video from multiple dimensions, reducing the impact of semantic asymmetry between text features and visual features, and significantly improving retrieval performance.

[0100] Secondly, modality retrieval tasks rely on the similarity between modalities; however, the richness and diversity of semantic information in videos can interfere with modality retrieval. For example... Figure 1As shown, visual information such as "tutor" and "microphone" failed to match the manually labeled text information. This weak matching relationship often stems from the text information's inability to fully encompass the complete visual information. Furthermore, visual information itself is diverse, and this diversity often manifests as redundancy in modality retrieval. Therefore, this invention further introduces a feature compression module, using a visual semantic compression method to remove redundant visual information, enabling text to interact with effective visual information. This significantly enhances the representativeness and flexibility of visual information, providing support for hierarchical granular alignment.

[0101] Example

[0102] This embodiment provides a method for constructing a text-video pair similarity evaluation model. The text-video alignment model includes a visual encoder, a text encoder, and an alignment model. The visual encoder and text encoder are pre-trained models, which can adopt any existing video encoding model and text encoding model. In this embodiment, the visual encoder and text encoder are the visual encoder and text encoder of the pre-trained CLIP model, respectively.

[0103] The training of the text-video alignment model includes:

[0104] S1. Data Preparation

[0105] The purpose of this step is to construct training sample pairs that meet the input requirements of subsequent models based on the original text-video pairs. More specifically, this involves extracting text-video pairs from the original dataset, where each text-video pair includes a video and its corresponding text; obtaining image sequences from the videos of each text-video pair, where each image sequence consists of a set of image frames; using the sampled image sequences and their corresponding text to form the training sample pairs for each text-video pair; and finally, using the training sample pairs from each text-video pair in the original dataset to form the training set.

[0106] Existing models struggle to process video directly. Therefore, it's necessary to obtain image sequences from the videos of each text-video pair. These image sequences can include all image frames from the video, or they can be obtained from the videos of each text-video pair through random sampling or average sampling. In this embodiment, average sampling is used.

[0107] In one exemplary approach, the MSRVTT dataset is used. First, text-video pairs are obtained from the MSRVTT dataset, such as the video video9770.mp4 and the corresponding text "Aperson is connecting something to system". From the video, F image frames are sampled using average sampling, and the video IDs are recorded. Thus, paired training sample pairs are obtained, such as data{'video_ID':video_ID,'video':imgs,'text':caption}. For ease of description, the image sequence containing the F image frames (imgs) is represented as V, and the text description (caption) associated with the images is represented as T.

[0108] S2. Construct positive and negative sample pairs and extract features.

[0109] Text-video retrieval ultimately involves calculating the similarity between several text samples and visual images to obtain a similarity matrix. The goal is to maximize the values ​​of the diagonal elements and minimize the values ​​of the remaining elements. Therefore, the commonly used training loss function is the symmetric cross-entropy function, which requires constructing positive and negative sample pairs between text and video.

[0110] Therefore, in this step, the input training sample pairs are used as positive sample pairs for this round of training; negative sample pairs are constructed for the positive sample pairs, which include text negative sample pairs and / or video negative sample pairs. The text negative sample pairs are sample pairs consisting of the text contained in the positive sample pairs and the image sequences contained in other training sample pairs besides the positive sample pairs. The video negative sample pairs are sample pairs consisting of the image sequences contained in the positive sample pairs and the text contained in other training sample pairs besides the positive sample pairs.

[0111] The aforementioned text negative sample pairs and video negative sample pairs can be used to calculate the loss for text-to-video retrieval and video-to-text retrieval, respectively. Therefore, one can be chosen based on actual needs and computing resources; for example, if only text-to-video retrieval is performed, only text negative samples can be used. In this embodiment, to obtain better results and achieve bidirectional retrieval, the negative sample pairs include both text negative sample pairs and video negative sample pairs.

[0112] The construction of negative sample pairs can be achieved by building a text library and a video library from the text and image sequences in the training set, respectively. Then, video negative sample pairs are constructed based on the positive sample videos and text randomly obtained from the text library, and text negative samples are constructed based on the positive sample text and image sequences randomly obtained from the video library.

[0113] However, for ease of calculation, it is recommended that at least two training sample pairs be input in this step as positive sample pairs for this round of training; the negative sample pairs include text negative sample pairs and video negative sample pairs; for each positive sample pair, the video negative sample is formed by combining the image sequence it contains with the text of other input positive sample pairs, and the text negative sample is formed by combining the text it contains with the image sequence of other input positive sample pairs.

[0114] At this point, negative samples are constructed by cross-referencing the text and image sequences of each positive input sample. This eliminates the need to introduce sample resources beyond the text and image sequences contained in the positive samples. Therefore, feature extraction can be performed either before or after the construction of negative samples. In this embodiment, Batch_size = 32.

[0115] Feature extraction aims to extract rich text-visual semantic information from text and image sequences, which serves as input for subsequent feature alignment. For each sample pair, the image sequence is input into a visual encoder to obtain the visual features of each image frame. For each sample pair, the text is encoded using a text encoder to obtain its text features.

[0116] In one exemplary scheme, based on the input image sequence V and text T, V is injected into the CLIP visual encoder φ. v To obtain the visual feature v, here Similarly, inject text T into CLIP's text encoder φ t Obtaining text features Where B is the batch size of the input training sample pairs in this step, F is the number of image frames contained in the image sequence, and D is the feature dimension. Let the number be the real number field. The formula is as follows:

[0117] t=φ t (T), v=φ v (V)

[0118] S3, Hierarchical Feature Alignment

[0119] In this step, for each sample pair, its textual and visual features are input into the alignment model to calculate its coarse-grained, medium-grained, and fine-grained similarity. The coarse-grained similarity is calculated based on the input textual and visual features, performing an overall similarity calculation between the image sequence and the text of the sample pair. The medium-grained similarity is calculated based on the input textual and visual features, performing a frame-level similarity calculation between the image frames and the text contained in the image sequence. The fine-grained similarity is calculated based on the input textual and visual features, performing a factor-level similarity calculation between the visual entities contained in the image sequence and the words contained in the text.

[0120] However, considering that modality retrieval tasks rely on the similarity between modalities, the richness and diversity of video semantic information can interfere with modality retrieval. Therefore, in this step, before performing hierarchical alignment, a visual semantic compression method is introduced to remove redundant visual information, allowing text to interact with effective visual information. This method greatly enhances the representativeness and flexibility of visual information, providing support for subsequent coarse-grained and fine-grained hierarchical alignment. Furthermore, the alignment model also includes a feature compression module; for the text and visual features input to the alignment model, the alignment model first uses the feature compression module to compress the input visual features, and then calculates coarse-grained similarity, medium-grained similarity, and fine-grained similarity.

[0121] In summary, this step includes:

[0122] S31, Compressed Visual Features

[0123] Compression of input visual features includes:

[0124] The visual features of the input alignment model, i.e., the visual features v obtained in step S2, are used as the original visual features v. o The redundant parts contained therein are calculated and used as redundant visual features v. r Remove the original visual features v using the following formula. o Redundant visual features v r Obtain the compressed visual features v c This is used as input to calculate coarse-grained similarity, medium-grained similarity, and fine-grained similarity;

[0125] v c =v o -v r

[0126] The redundant visual features v obtained by the calculation r It should satisfy:

[0127]

[0128] Where s(·) represents the similarity function, v o Let t represent the raw visual features of the input alignment model, and v represent the text features of the input alignment model. r To represent the original visual features v o Redundant visual features of the redundant part, v c To remove visual features v o Visual features after redundancy; Min means minimization, Max means maximization.

[0129] The above calculation indicates that for redundant visual features v r Representation learning aims to reduce redundant visual features v r Similarity with text feature t, while preserving the original visual feature v o and compressed visual features v c It has a greater similarity to the text feature t, that is, s(v) r ,t) should be as small as possible, s(v) o ,t) and s(v c ,t) should be as large as possible, while s(v) c ,t) is greater than s(v) o Redundant visual features v r The calculation can be performed using any scheme based on the above conditions, for example: by defining a similarity-aware compression factor C, and using v r =C or v r =C·ε, calculate to obtain redundant visual features v r , where ε is a random factor.

[0130] The similarity-aware compression factor C can also be calculated using any of the following methods:

[0131] Method 1: C = Expand(S)

[0132] Method 2: C = Expand(exp(γS))

[0133] Method 3

[0134] Method 4: C = Expand(exp(MLP(S)))

[0135] Where S = [s1, s2, ..., s i …,s F ], s i =s(t,f i MLP(S) = LN(Rule(LN(S))); s(·) represents the similarity function, f i s represents the visual features of the i-th image frame in the image sequence. i F represents the similarity between the i-th image frame in the image sequence and the text feature t; C is the similarity-aware compression factor; MLP is a fully connected network; LN is a linear layer in the fully connected network; Rule is the Rule activation function; exp is the exponential function; Expand represents the similarity between the i-th image frame and the text feature t. r Expand the dimension to make it compatible with v o An expansion function that is consistent with the dimension.

[0136] In this embodiment, the redundant visual feature v is calculated using the following formula. r :

[0137] v r =C·ε

[0138] C = Expand(exp(MLP(S)))

[0139] MLP(S) = LN(Rule(LN(S)))

[0140] S = [s1, s2, ..., s i …,s F ]

[0141] s i =s(t,f i )

[0142] Where s(·) represents the cosine similarity function, f i s represents the visual features of the i-th image frame in the image sequence. i Let F represent the similarity between the i-th image frame in the image sequence and the text feature t, where F is the number of image frames in the image sequence; C is the similarity-aware compression factor, and ε is a random factor that follows a normal distribution; MLP is a fully connected network, LN is a linear layer in the fully connected network, Rule is the Rule activation function, exp is the exponential function, and Expand represents the similarity of the i-th image frame with respect to v. r Expand the dimension to make it compatible with v o An expansion function that is consistent with the dimension.

[0143] In this embodiment, firstly, the cosine similarity between text features and visual features is calculated to obtain a similarity matrix. Next, a fully connected network is used to map redundant visual features onto the similarity matrix, obtaining a similarity-perceived compression factor C; simultaneously, a random distribution satisfying a normal distribution is introduced. Redundant visual features v r It is characterized as C·ε.

[0144] S32. Calculate coarse-grained similarity.

[0145] The coarse-grained similarity refers to the coarse-grained alignment between the entire video and the sentence. That is, it is obtained by calculating the overall similarity between the image sequence and the text of the sample pair based on the input text features and visual features. Therefore, it is necessary to first aggregate the image frames contained in the image sequence to obtain the overall visual features of the video. The aggregation method can be any existing method, such as summation, averaging, weighted summation, etc.

[0146] The recommended approach, which uses a weighted sum method and aggregates data based on text conditions, includes:

[0147] First, based on the similarity between the visual features and textual features of each image frame in the image sequence of the sample pair, the weights of each image frame in the image sequence of the sample pair are constructed. Then, the visual features of each image frame in the sample pair are aggregated by weight summation to obtain the coarse-grained video representation of the sample pair. After that, the similarity between the coarse-grained video representation and the textual features of the sample pair is calculated as the coarse-grained similarity of the sample pair.

[0148] Specifically, in this embodiment, the calculation of the video coarse-grained representation includes:

[0149] First, based on the similarity between the visual features and textual features of each image frame in the image sequence of the sample pair, the weight 'a' of each image frame in the image sequence of the sample pair is constructed according to the following formula. i :

[0150]

[0151] Then, using the following formula and a weighted summation method, the visual features of each image frame contained in the sample pair are aggregated to obtain a coarse-grained video representation of the sample pair.

[0152]

[0153] Then, the coarse-grained video representation v of the sample pair is calculated according to the following formula. cg The similarity between the sample pair and its textual feature t, as a coarse-grained similarity sc. g :

[0154]

[0155] Among them, f i and f j Let represent the visual features of the i-th and j-th image frames in the image sequence, respectively; τ is a hyperparameter; F represents the number of image frames in the image sequence; the superscript T indicates matrix transpose; exp is the exponential function; and |·| indicates the modulo operation.

[0156] S33, Calculate the similarity of medium granularity.

[0157] Coarse-grained alignment often lacks learnable parameters, includes invalid frames, and misses valid frames. Therefore, medium-grained similarity is introduced as a supplement. This medium-grained similarity is calculated by performing frame-level similarity calculations on the image frames and text within the image sequence of the sample pair, based on input text and visual features. Essentially, it selects the video frames with the most representative information through frame-level similarity comparison, reducing the impact of invalid frames and increasing the contribution of valid frames. This can be achieved by using the attention between each frame and the text as a weight or selection metric, or by directly using keyframes from the selected video.

[0158] Among them, the recommended approach using cross-attention mechanisms includes:

[0159] First, a query is constructed using the text features of the sample pair, and keys and values ​​are constructed using the visual features of the sample pair. Using a cross-attention mechanism, the embedded representations of each image frame contained in the image sequence of the sample pair are obtained. Then, based on the embedded representations of each image frame, the granular representation of the sample pair in the video is obtained. After that, the similarity between the granular representation of the sample pair in the video and its text features is calculated as the granular similarity of the sample pair.

[0160] Specifically, in this embodiment, the calculation of medium-granularity similarity is performed using the following formula: a query is constructed using the textual features of the sample pair, and keys and values ​​are constructed using the visual features of the sample pair. A cross-attention mechanism is then used to obtain the embedded representations of each image frame contained in the image sequence of the sample pair:

[0161]

[0162] Q t =LN(t) T W Q

[0163] K v =LN(v)W K

[0164] V v =LN(v c W V

[0165] Among them, W Q W K and W V Construct query representation Q respectively t K-bond characterization v Sum value characterization V v The transformation matrix; D p For K v Feature dimensions, is the scaling factor for attention; softmax represents the softmax function, the superscript T represents the matrix transpose, and LN represents a linear layer.

[0166] Based on the embedded representations of each image frame, the granular representation v of the sample pair in the video can be obtained simply by summing. mg In this embodiment, to obtain the final effective frame for alignment with the text, firstly, the embedding representation of the image frame is weighted according to the following formula. We obtain the aggregated video embedding r(v|t) based on text features t, and embed the image frame embedding representation into the joint space related to the text features:

[0167] r(v|t)=LN(Attention(Q t ,K v V v W1)

[0168] in, W1 represents the embedding representation of each image frame contained in the image sequence, and LN is the learnable weight.

[0169] Then, following the formula below, a fully connected layer and a residual connection method were used to obtain the final granular representation v in the video. mg :

[0170] v mg =LN(r(v|t))+FC(r(v|t))) T

[0171] Where LN represents a linear layer and FC represents a fully connected layer.

[0172] Finally, calculate the granularity representation v in the video of the sample pair according to the following formula. mg The similarity between the text feature t and the sample pair is considered as the medium-granularity similarity s. mg :

[0173]

[0174] Here, |·| represents the modulo operation.

[0175] S34. Calculate fine-grained similarity

[0176] To fully extract richer semantic information, such as detailed actions of characters and scene information, fine-grained similarity is further introduced. This fine-grained similarity is obtained by calculating factor-level similarity between visual entities contained in an image sequence and words contained in the text, based on input text features and visual features.

[0177] The features of words contained in text can be obtained by decomposing the text into text features, or by obtaining them from words through a pre-trained language model; the features of visual entities contained in image sequences can also be obtained by extracting them through existing visual models, such as obtaining fine-grained features from image patches, and then obtaining the features of visual entities through clustering.

[0178] To simplify computational complexity and gain an advantage in processing time while maintaining granularity, the following methods are recommended:

[0179] First, the text features of the sample pair are decomposed into K text factors, and the coarse-grained video representation of the input sample pair is decomposed into K video factors, forming K video-text factor pairs, where K is the number of words contained in the text of the input sample pair. Then, the similarity between the text factors and video factors contained in each video-text factor pair is calculated. Finally, based on the similarity between the text factors and video factors of each video-text factor pair, the fine-grained similarity of the sample pair is obtained.

[0180] Specifically, in this embodiment, the calculation of fine-grained similarity includes:

[0181] First, the text features of the sample pair are decomposed into K text factors fw according to the following formula. k The coarse-grained video representation of the input sample pair is decomposed into K video factors fe. k And form K video text factor pairs:

[0182] fw k =W k t

[0183] fe k =W k v cg

[0184] Among them, v cg For coarse-grained representation of the video, t represents the text feature, and W represents the text feature. k Let k be the k-th learnable decomposition factor, where k is the index and k = 1, 2, ..., K, and K is the number of words contained in the text of the sample pair.

[0185] Then, calculate the similarity between the text factors and video factors contained in each video text factor pair according to the following formula;

[0186] s=(FW) T FE

[0187] Where the superscript T denotes matrix transpose, FE = [fe1,fe2,…,fe k …,fe K[FW] is the video factor matrix, where FW = [fw1, fw2, ..., fw] k …,fw K [ ] represents the text factor matrix.

[0188] Based on the similarity between text factors and video factors of each video text factor, fine-grained similarity of sample pairs can be obtained by averaging, summing, etc. In this embodiment, a confidence level g is introduced, and the similarity is calculated using the following formula. With confidence level The dot product, through an adaptive pooling operation, yields the final fine-grained similarity based on the similarity between the text factors and video factors of each video text factor.

[0189] s fg =s·g

[0190] g = MLP(cat[FW,FE])

[0191] MLP(cat[FW,FE])=LN(Rule(LN(cat[FW,FE])))

[0192] Where g is the similarity confidence score, MLP is a fully connected network, LN is a linear layer in a fully connected network, Rule is the Rule activation function, and cat is a concatenation function.

[0193] S4, Parameter Update

[0194] For the design of training objectives, alignment losses at three granularities can be directly incorporated. That is, by utilizing the coarse-grained, medium-grained, and fine-grained similarities of positive and negative sample pairs, coarse-grained feature alignment losses, medium-grained feature alignment losses, and fine-grained feature alignment losses are calculated, and the total loss for this round of training is obtained accordingly. Based on the total loss, the parameters of the alignment model are updated, and the parameters of the visual encoder and text encoder are fine-tuned.

[0195] Specifically, the negative sample pairs include text negative sample pairs and video negative sample pairs;

[0196] Calculate the text-to-video alignment loss using the following formula.

[0197]

[0198] Calculate the video-to-text alignment loss using the following formula.

[0199]

[0200] Calculate the alignment loss using the following formula:

[0201]

[0202] Where M represents the number of negative video samples plus one, and the plus one represents the positive sample corresponding to the negative video sample; N represents the number of negative text samples plus one, and the plus one represents the positive sample corresponding to the negative text sample; B is the number of positive sample pairs in this training round; e represents the natural exponent; λ is a hyperparameter representing the scaling factor; and s(·) represents the similarity function. In this implementation example, the similarity function s(·) is cosine similarity.

[0203] The total loss for:

[0204]

[0205] Where α and β are weight hyperparameters, This represents the coarse-grained feature alignment loss, which is the loss that aligns the coarse-grained representation of the video v. cg The text feature t is taken as input and obtained by alignment loss calculation; This represents the medium-granularity feature alignment loss, which is the loss that aligns the granular representations v in the video. mg The text feature t is taken as input and obtained by alignment loss calculation; This represents the fine-grained feature alignment loss, which is calculated by taking the video factor matrix FE = [fe1,fe2,…,fe2]. k …,fe K ] and the text factor matrix FW = [fw1, fw2, ..., fw k …,fw K ] is used as input and obtained through alignment loss calculation; fe k For the k-th video factor, fw k Let k be the k-th text factor.

[0206] In this embodiment, negative samples are constructed by cross-referencing the text and image sequences of each positive input sample, therefore, M = N = B.

[0207] S5. Repeat steps S2 to S4 until the training termination condition is met, and obtain the aligned model that has completed training.

[0208] After training the text-video alignment model, it can be used to perform retrieval tasks, such as... Figure 1 As shown, the text-to-video retrieval method includes the following steps:

[0209] B1. Input the search text, and construct search sample pairs based on the search text and each video contained in the video library;

[0210] B2. Using the visual encoder and text encoder of the CLIP model, obtain the visual and text features of each retrieved sample pair.

[0211] B3. Using the alignment model obtained from training, obtain the coarse-grained similarity, medium-grained similarity, and fine-grained similarity of each retrieved sample pair;

[0212] B4. Sort the similarity of each search sample pair by summing the coarse-grained similarity, medium-grained similarity, and fine-grained similarity, and output the L videos with the highest similarity as the search results for the search text.

[0213] The video-to-text retrieval method includes the following steps:

[0214] B1. Input the search video, and construct search sample pairs based on the search video and the text contained in the text library;

[0215] B2. Using the visual encoder and text encoder of the CLIP model, obtain the visual and text features of each retrieved sample pair.

[0216] B3. Using the alignment model obtained from training, obtain the coarse-grained similarity, medium-grained similarity, and fine-grained similarity of each retrieved sample pair;

[0217] B4. Sort the similarity of each search sample pair by summing the coarse-grained similarity, medium-grained similarity, and fine-grained similarity, and output the L texts with the highest similarity as the search results for the search video.

[0218] Experimental verification:

[0219] To verify the effectiveness of the present invention, the inventors compared the solutions of the above embodiments with existing methods on six benchmark text-video retrieval datasets. The six benchmark text-video retrieval datasets include MSRVTT, MSVD, ActivityNet, Charades, DiDeMo, and VATEX.

[0220] Specifically, the MSRVTT dataset contains 10,000 videos, each with 20 text descriptions, and videos are approximately 10 to 30 seconds long. Following the MSRVTT-test-1k split, 9,000 videos are used for training and 1,000 for testing. The MSVD dataset contains 1,970 videos, each with approximately 40 text descriptions, and videos are approximately 10 to 25 seconds long. Following the official split, 1,200 videos are used for training and 670 for testing. The ActivityNet dataset contains 20,000 videos, each approximately 180 seconds long. All sentence descriptions from the videos are concatenated into a single query. In this split, 10,009 videos are used for training and 4,917 for testing. The Charades dataset contains 9,848 video clips, totaling approximately 38.8 hours, with each clip averaging 30 seconds in length and accompanied by a text description. The DiDeMo dataset contains 10,000 unedited videos of 25 to 30 seconds randomly selected from YFCC100M, labeled with 40,000 text descriptions. The VATEX dataset is a new large-scale multilingual video description dataset containing over 41,250 videos and 825,000 English and Chinese subtitles.

[0221] Feature extraction was performed using CLIP's backbone models ViT-B / 32 and ViT-B / 16, with a latent space dimension of D = 512, weight decay of 0.2, and a dropout rate of 0.3. The batch size was set to Batch_size = 32, and the training epochs were 5. An initial learning rate of 1e-5 was maintained, and the module for feature alignment was fine-tuned at a rate of 1e-6 for the CLIP model. The training process used the AdamW optimizer with cosine scheduling and a coefficient τ of 0.1. 12 frames were uniformly sampled from the video and resized to 224×224 to suit all datasets. Experiments were conducted on an A800 GPU with parameters set to α = 1e-3, β = 1e-1, and K = 8 to decompose the number of visual entities and text words.

[0222] The test metrics include R@k, MdR, and MnR.

[0223] R@k is an evaluation metric in retrieval systems, used to measure the proportion of correct targets found in the first k search results. It reflects the model's retrieval ability to varying degrees. Specifically, R@1 represents the proportion of correct targets found in the first 1 search result, R@5 represents the proportion found in the first 5 search results, and R@10 represents the proportion found in the first 10 search results. A higher R@k value indicates that the model has better accuracy and recall within the corresponding k search results range.

[0224] Median Rank (MdR) is a metric for evaluating the performance of a retrieval system, representing the median rank of a correct result in the search results list. Specifically, it measures where the correct result typically appears in the search results for a set of queries. A lower MdR indicates better retrieval performance, as it means the correct result generally appears higher up in the search results. Compared to the mean rank (MnR), MdR is less sensitive to extreme values ​​of abnormally high or low rank, more accurately reflecting the typical performance of the retrieval system, and is an effective indicator for evaluating system stability and reliability.

[0225] MnR (Mean Responsibility) is a metric for measuring the performance of a retrieval system, representing the average rank of correct results across all queries. It reflects the overall retrieval effectiveness by calculating the rank of correct results for each query and averaging the results. A lower MnR value indicates better system performance because it means correct results typically appear higher in the results. MnR averages all ranks, thus providing a comprehensive reflection of the system's performance across various queries, but it is also susceptible to extreme rank values. Overall, MnR provides a comprehensive assessment of the accuracy and effectiveness of a retrieval system.

[0226] The experimental results are shown below. Table 1 shows the performance of various schemes on the MSRVTT dataset, Table 2 shows the performance of various schemes on the DiDeMo dataset, Table 3 shows the performance of various schemes on the ActivityNet and MSVD datasets, and Table 4 shows the performance of various schemes on the Charades and VATEX datasets. In the table, Ours represents the schemes of this invention, and Publications represents the journals or conferences in which the comparative scheme models were published.

[0227] Table 1 shows the performance of each scheme on the MSRVTT dataset.

[0228]

[0229] Table 2 shows the performance of various schemes on the DiDeMo dataset.

[0230]

[0231] Table 3 shows the performance of various schemes on the ActivityNet and MSVD datasets.

[0232]

[0233] Table 4 shows the performance of various schemes on the Charades and VATEX datasets.

[0234]

[0235] Table 5 shows the performance of the scheme of the present invention using different compression factors C on the MSRVTT dataset, wherein:

[0236] The first group uses v directly without compression. o The behavior of subsequent operations, where w / o C means without C;

[0237] The second group directly uses the S method, that is: v r =C,C=Expand(S);

[0238] The third group uses the method of adding coefficient γ to S, that is: v r =C,C=Expand(exp(γS));

[0239] The fourth group uses the method of averaging, that is:

[0240] Groups 5 and 6 both use the MLP method, that is: v r =C·ε,C=Expand(exp(MLP(S))); where, the ε of the fifth group follows a normal distribution and is represented as N, and the ε of the sixth group follows a uniform distribution and is represented as U.

[0241] Table 5 shows the performance of the scheme of the present invention using different compression factors C on the MSVTT dataset.

[0242]

[0243] Figure 3 The image shows the results of the retrieval. In the image, Query represents the given retrieval text, the image represents the top five query results ranked by similarity, and the numbers represent the similarity values. From the results, we can see that the Rank 1 query result, which has the highest similarity, is the correct retrieval result.

[0244] Finally, it should be noted that the above embodiments are merely preferred embodiments and are not intended to limit the present invention. It should be pointed out that those skilled in the art can make various modifications, equivalent substitutions, and improvements without departing from the spirit and scope of the claims, and all such modifications, substitutions, and improvements should be included within the scope of protection of the present invention.

Claims

1. A method for constructing a text-video pair similarity evaluation model, characterized in that: The text-video alignment model includes a visual encoder, a text encoder, and an alignment model, wherein the visual encoder and the text encoder are pre-trained models. The training of the text-video alignment model includes: A1. Extract text-video pairs from the original dataset, wherein each text-video pair includes a video and its corresponding text; obtain image sequences from the video of each text-video pair, wherein each image sequence consists of a set of image frames; the image sequences obtained by sampling and the corresponding text of the video constitute the training sample pairs of the text-video pair; the training sample pairs of each text-video pair in the original dataset constitute the training set. A2. Input training sample pairs as positive sample pairs for this round of training; construct negative sample pairs for the positive sample pairs, wherein the negative sample pairs include text negative sample pairs and / or video negative sample pairs. The text negative sample pairs are sample pairs consisting of the text contained in the positive sample pairs and the image sequences contained in other training sample pairs besides the positive sample pairs. The video negative sample pairs are sample pairs consisting of the image sequences contained in the positive sample pairs and the text contained in other training sample pairs besides the positive sample pairs. For each sample pair containing an image sequence, the visual encoder is input to obtain the visual features of each image frame contained therein; for each sample pair containing text, the text encoder is input to obtain its text features. A3. For each sample pair, input its text features and visual features into the alignment model to calculate its coarse-grained similarity, medium-grained similarity, and fine-grained similarity. The coarse-grained similarity is obtained by performing an overall similarity calculation on the image sequence and text of the sample pair based on the input text features and visual features. The medium-granularity similarity is obtained by performing frame-level similarity calculations on the image frames and text contained in the image sequence of the sample pair, based on the input text features and visual features. The fine-grained similarity is obtained by performing factor-level similarity calculations on visual entities contained in the image sequence and words contained in the text, based on input text features and visual features. A4. Utilize the coarse-grained, medium-grained, and fine-grained similarity of positive and negative sample pairs to calculate the coarse-grained, medium-grained, and fine-grained feature alignment losses, and use this to calculate the total loss for this round of training; based on the total loss, update the parameters of the alignment model and fine-tune the parameters of the visual encoder and text encoder. A5. Repeat steps A2 to A4 until the training termination condition is met to obtain the aligned model that has been trained. The text-video alignment model also includes a feature compression module. For the text and visual features input to the alignment model, the alignment model first uses the feature compression module to compress the input visual features, and then calculates coarse-grained similarity, medium-grained similarity, and fine-grained similarity. The compression of the input visual features includes: The visual features of the input alignment model are used as the raw visual features. The redundant parts contained therein are calculated and used as redundant visual features. Remove the original visual features using the following formula. Redundant visual features Obtain compressed visual features This is used as input to calculate coarse-grained similarity, medium-grained similarity, and fine-grained similarity; ; The redundant visual features obtained by calculation It should satisfy: ; in, Represents the similarity function. This represents the original visual features of the input alignment model. The text features are used as input for the alignment model. To represent the original visual features Redundant visual features of the redundant parts in the middle. To remove visual features Visual features after the redundant parts; Indicates minimization. Maximize the table; Define similarity-aware compression factor and with or Calculate and obtain redundant visual features ,in, The random factor; the similarity-aware compression factor The calculation can be performed using any of the following methods: Method 1 ; Method 2 ; Method 3 ; Method 4 ; in, , , ; Represents the similarity function. Represents the first image in the image sequence. Visual features of an image frame Represents the first image in the image sequence. Image frames with text features Similarity between them The number of image frames contained in the image sequence; A similarity-aware compression factor; For a fully connected network, For a linear layer in a fully connected network, for Activation function It is an exponential function. Indicates to Expand the dimensions to make it compatible with An expansion function that is consistent with the dimension.

2. The method for constructing a text-video pair similarity evaluation model as described in claim 1, characterized in that, The visual encoder and text encoder are the visual encoder and text encoder of the pre-trained CLIP model, respectively. In step A1, image sequences are obtained from the videos of each text-video pair by means of random sampling or average sampling; In step A2, at least two training sample pairs are input as positive sample pairs for this round of training; the negative sample pairs include text negative sample pairs and video negative sample pairs; for each positive sample pair, the video negative sample is formed by combining the image sequence it contains with the text of other input positive sample pairs, and the text negative sample is formed by combining the text it contains with the image sequence of other input positive sample pairs.

3. A method for constructing a text-video pair similarity evaluation model as described in claim 1 or 2, characterized in that, The coarse-grained similarity is obtained by performing an overall similarity calculation on the image sequence and text of the sample pair based on input text features and visual features, including: First, based on the similarity between the visual features and text features of each image frame in the image sequence of the sample pair, the weights of each image frame in the image sequence of the sample pair are constructed. Then, the visual features of each image frame in the sample pair are aggregated by weight summation to obtain the coarse-grained video representation of the sample pair. After that, the similarity between the coarse-grained video representation and the text features of the sample pair is calculated as the coarse-grained similarity of the sample pair. The medium-granularity similarity is obtained by performing frame-level similarity calculations on the image frames and text contained in the image sequence of the sample pair, based on input text features and visual features, including: First, a query is constructed using the text features of the sample pair, and keys and values ​​are constructed using the visual features of the sample pair. Using a cross-attention mechanism, the embedded representations of each image frame contained in the image sequence of the sample pair are obtained. Then, based on the embedded representations of each image frame, the granular representation of the sample pair in the video is obtained. After that, the similarity between the granular representation of the sample pair in the video and its text features is calculated as the granular similarity of the sample pair. The fine-grained similarity is obtained by performing factor-level similarity calculations on visual entities contained in an image sequence and words contained in the text, based on input text features and visual features. This includes: First, the text features of the sample pairs are decomposed into Each text factor decomposes the coarse-grained video representation of the input sample pair into... Each video factor constitutes a video element. A video text factor pair, where... The number of words contained in the text of the input sample pair is determined; then, the similarity between the text factors and video factors contained in each video text factor pair is calculated; subsequently, based on the similarity between the text factors and video factors of each video text factor, fine-grained similarity of the sample pair is obtained.

4. The method for constructing a text-video pair similarity evaluation model as described in claim 3, characterized in that, The calculation of the coarse-grained representation of the video includes: First, based on the similarity between the visual features and textual features of each image frame in the image sequence of the sample pair, the weights of each image frame in the image sequence of the sample pair are constructed according to the following formula. : ; Then, using the following formula and a weighted summation method, the visual features of each image frame contained in the sample pair are aggregated to obtain a coarse-grained video representation of the sample pair. : ; Then, the coarse-grained video representation of the sample pairs is calculated using the following formula. Its textual features The similarity between them, as coarse-grained similarity of sample pairs. : ; in, and These represent the number of images contained in the image sequence. The and the first Visual features of an image frame For hyperparameters, Indicates the number of image frames contained in an image sequence; superscript Indicates matrix transpose. It is an exponential function. This indicates the modulo operation.

5. The method for constructing a text-video pair similarity evaluation model as described in claim 3, characterized in that, In the calculation of medium-granularity similarity, the query is constructed using the text features of the sample pair, and the key and value are constructed using the visual features of the sample pair. A cross-attention mechanism is then used to obtain the embedded representation of each image frame contained in the image sequence of the sample pair: ; ; ; ; in, , and Construct query representations respectively Key characterization Sum value representation The transformation matrix; for Feature dimensions; express Function, superscript Indicates matrix transpose. Indicates a linear layer; Calculate the granularity representation in the video of the sample pair using the following formula. Its textual features The similarity between them, as a medium-granularity similarity of sample pairs. : ; in, This represents the modulo operation; In the calculation of medium-granularity similarity, the medium-granularity representation of the sample pair is obtained based on the embedding representation of each image frame according to the following formula. : ; ; in, This represents the embedding representation of each image frame contained in an image sequence. For learnable weights, For linear layers, This indicates a fully connected layer.

6. The method for constructing a text-video pair similarity evaluation model as described in claim 3, characterized in that, The calculation of the fine-grained similarity includes: First, the text features of the sample pair are decomposed into the following formula: Text factors The video coarse-grained representation of the input sample pair is decomposed into Video Factors and constitute One video text factor pair: ; ; in, For coarse-grained characterization of video, For text features, For the first A learnable factorization factor, subscript The serial number and , The number of words contained in the text of a sample pair; Then, calculate the similarity between the text factors and video factors contained in each video text factor pair according to the following formula; ; Among them, superscript Indicates matrix transpose. For video factor matrix, This is a text factor matrix.

7. The method for constructing a text-video pair similarity evaluation model as described in claim 6, characterized in that, In the calculation of fine-grained similarity, the fine-grained similarity of sample pairs is obtained based on the similarity between the text factors and video factors of each video text factor, according to the following formula. : ; ; ; in, For similarity confidence, For a fully connected network, For a linear layer in a fully connected network, for Activation function It is a series function.

8. The method for constructing a text-video pair similarity evaluation model as described in claim 6, characterized in that, The negative sample pairs include text negative sample pairs and video negative sample pairs; Calculate the text-to-video alignment loss using the following formula. : ; Calculate the video-to-text alignment loss using the following formula. : ; Calculate the alignment loss using the following formula: ; in, This indicates that the number of negative samples in the video has been incremented by one, and the increment represents the positive sample corresponding to the negative sample in the video. This indicates that the number of negative samples in the text has been incremented by one, and the increment represents the positive sample corresponding to the negative sample in the text. This represents the number of positive sample pairs in this training round. Represents the natural index. The hyperparameter representing the scaling factor, Represents the similarity function; The total loss for: ; in, and For weight hyperparameters, This represents the coarse-grained feature alignment loss, which is used to represent the video at a coarse-grained level. and text features It is obtained as input after alignment loss calculation; This represents the medium-granularity feature alignment loss, which is used to granularize the features in the video. and text features It is obtained as input after alignment loss calculation; This represents the fine-grained feature alignment loss, which is the loss calculated from the video factor matrix. and text factor matrix It is obtained as input after alignment loss calculation; For the first Each video factor, For the first Each text factor.