Video-text cross-modal retrieval method based on cross-granularity self-distillation

By using the cross-particle size self-distillation method and lightweight label screening network in video-text cross-modal retrieval, the problems of information loss and binary label absolute are solved, and the search performance and accuracy are significantly improved.

CN114548293BActive Publication Date: 2025-05-13COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210174138.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2025-05-13
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

The existing video-text cross-modal search method loses information during the encoding process, resulting in a lack of fine-grained interaction during retrieval, affecting the matching accuracy; at the same time, because the binary label is too absolute, it is difficult to conform to the real situation, which affects network training.

Method used

The video-text cross-modal retrieval method based on cross-particle size self-distillation is adopted to provide the self-distillation loss of soft labels through fine-grained interaction, and the deep learning network is screened in combination with lightweight labels, and representative marks are selected for soft label calculations to build a network structure suitable for self-distillation.

Benefits of technology

Retrieval performance is significantly improved, and relatively gentle labels are provided through fine-grained interaction, helping the network better learn and match video-text data, improving the accuracy and efficiency of retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548293B_ABST
    Figure CN114548293B_ABST
Patent Text Reader

Abstract

The present invention discloses a video-text cross-modal retrieval method based on cross-granularity self-distillation. The method aims to provide pseudo-labels through fine-grained interactive similarity to solve the problem that binary labels in cross-modal contrastive learning are not smooth enough and do not conform to the actual situation. The method first designs a screening module to screen a part of key tokens for each modality for calculating token-level fine-grained similarity. Then, this fine-grained similarity is used as a soft label, combined with contrast loss, to jointly optimize the encoders of each modality. This method introduces cross-granularity self-distillation during the training stage to improve the natural defects of contrastive learning labels, but there is no additional computational consumption during retrieval, so it is an efficient method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video-text cross-modal retrieval method based on cross-granularity self-distillation, and belongs to the technical field of artificial intelligence. Background Art

[0002] The rapid development of information technology in recent years, especially the high-speed mobile Internet and the diversification of social media, has made the forms of media communication more diverse. Media content is transmitted in the form of text, pictures, audio and video, or even in a combination of various forms. In the era of explosive growth of media information, how to efficiently process massive amounts of information to better serve the audience has become an increasingly urgent issue. Video and text are two important information carriers, and video-text cross-modal retrieval is also an important field of multimedia information processing. This task aims to query the database for text descriptions (or videos) related to a video (or text description). In order to achieve efficient retrieval, existing methods usually use video and text encoders to vectorize the data of the two modalities respectively to embed them into a joint space, so that the distance between the corresponding video-text vectors is less than the distance between the unpaired video-text vectors. These methods focus on the representation learning of video and text and the cross-modal alignment of video and text. Recent studies have shown the superiority of Transformer or BERT in video-text cross-modal retrieval tasks, and more and more methods choose Transformer-based structures as encoders.

[0003] Existing video-text cross-modal retrieval methods can be roughly divided into two categories based on the interaction between the modalities: the first category is to first use two independent encoders to encode the content of different modalities into feature vectors, and then calculate the similarity between the two through simple functions or operations; the second category is to input the content of the two modalities into the encoder together, let them fully interact within the model through self-attention, and finally directly output the matching score of the two. Since the first method can calculate the vector of the content of the resource library to be retrieved offline in advance, it is more efficient and has a wide range of applications in real life. However, this type of method faces some problems:

[0004] 1. For the sake of efficiency, the video encoder and text encoder encode the video and text into a representation vector respectively, and then calculate the matching degree between a video and a text based on the vector distance. The advantage of this is that the vectors of all database resources can be calculated offline in advance, and the retrieval efficiency is relatively high. However, due to the loss of a large amount of information during the encoding process, there is a lack of fine-grained interaction (such as frames in a video and words in a text) during retrieval, which affects the matching accuracy.

[0005] 2. Since Transformer-based structures usually have strong expressions, if the amount of data is not large enough, it is easy to cause overfitting. In addition, existing methods mainly optimize the network by pulling positive samples closer and pushing negative samples farther in the joint space. InfoNCE is a common contrast loss based on this idea. This loss distributes a label "1" to positive sample pairs and a label "0" to negative sample pairs, and optimizes it through cross entropy. However, such binary discrete labels are too absolute and not conducive to network learning. In fact, the label distribution should be smoother.

[0006] In view of the above problems, the present invention designs a video-text cross-modal retrieval method based on cross-granularity self-distillation, which uses the fine-grained interaction between cross-modalities to provide self-distillation loss of soft labels, and realizes video-text cross-modal retrieval by constructing a lightweight label screening deep learning network suitable for self-distillation. Compared with the previous two methods, this method can solve the problem that binary labels in cross-modal contrastive learning are not smooth enough and do not conform to the actual situation, and can significantly improve the retrieval performance. Summary of the invention

[0007] The main purpose of the present invention is to design a video-text cross-modal retrieval method based on cross-granularity self-distillation, the core of which is to use the fine-grained interaction between cross-modalities to provide self-distillation loss of soft labels. In addition, considering that the role of each token is not equivalent, in order to ensure the reliability of soft labels as much as possible, the present invention also proposes a lightweight label screening network adapted to self-distillation, selects representative labels for each modality, and then performs subsequent soft label calculations.

[0008] To achieve the above purpose, the training and retrieval process of the present invention is as follows: Figure 1 , as follows:

[0009] Step 1: Given a mini-batch of size b Where {V k , T k} is the kth video-text pair. i (i=1, 2, .., b) and a video V with a frame length of M j (j=1, 2, ..., b) input their respective encoders f and g, and obtain and Where d is the dimension of the representation vector.

[0010] Step 2: Calculate f(T i ) and g(V j ) is averaged, and we get and have Calculate their similarity by vector inner product Therefore, the InfoNCE loss can be calculated according to formula (1):

[0011]

[0012] Step 3: Figure 2 As shown, f(T i ) and g(V j ) are input into two tokens filtering modules with the same structure, and n (n≤N) "key" text tokens and m (m≤M) "key" video tokens are obtained respectively. Taking video as an example, in the tokens filtering module, g(V j )First, a weight generator consisting of "Linear-ReLU-Linear-Softmax" is used to obtain a vector w = (w0, w1, ..., w M-1 ), each digit is a number between 0 and 1, indicating the "importance" of the corresponding token. Then, based on the importance score of each token, the top m "most important" tokens are selected; similarly, the same operation is performed on text tokens.

[0013] Step 4: After filtering out the “key” tokens of the video and text according to step 3, perform pairwise interactions (vector inner product) between cross-modal tokens to obtain a fine-grained interaction matrix Each element in I represents the similarity between a text token and another video token. Then we can get the text T by taking the maximum value and then averaging each row of I. i For video V j The fine-grained similarity score s t2v (T i , V j ); Similarly, first take the maximum value of each column of I and then find the average value to get the video V j For text T i The fine-grained similarity score s v2t (V j , T i ).

[0014] Step 5: For a mini-batch of samples, obtain any two T i and V j The fine-grained interaction similarity can form a text-video soft label matrix The jth element in the i-th row represents the text T i For video V j The fine-grained similarity score oft2v (T i , V j ); Similarly, the video-text soft label matrix can be obtained At the same time, the coarse-grained similarity matrix can be obtained from step 2 The jth element in the i-th row represents the text T i and video V j The cross-granularity self-distillation loss is the KL divergence of the coarse-grained similarity distribution and the fine-grained similarity soft label, that is:

[0015]

[0016]

[0017] Among them, S i is the i-th row of matrix S, S T is the transpose of S, P and Q are probability distributions of dimension b respectively.

[0018] Step 6. Calculate the final loss based on InfoNCE loss and cross-granularity self-distillation loss:

[0019]

[0020] Among them, λ is an adjustable hyperparameter.

[0021] Step 7: Train according to the loss obtained in the previous step and optimize the network parameters. After several iterations, the final encoder is obtained. During reasoning and retrieval, fine-grained interaction is no longer required, but the average pooling result of the last layer of the encoder is directly used as the representation vector.

[0022] Compared with the prior art, the present invention proposes a video-text cross-modal retrieval method based on cross-granularity self-distillation. The core of this method is to design a method that uses token-level fine-grained interactions to supplement soft labels for contrast loss, obtain relatively flat labels, which are more in line with the actual situation and help network training. In addition, in order to improve the reliability of soft labels, the present invention also designs a tokens screening module, which aims to select tokens that are more important to the entire sequence for fine-grained interactions. The module and loss proposed by the present invention can help the network significantly improve retrieval performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 The following is a specific training and retrieval flow chart of the present invention. (a) is the network training process; (b) is the retrieval process.

[0024] Figure 2 This is a diagram of the fine-grained soft label calculation method proposed in the present invention. DETAILED DESCRIPTION

[0025] The core of the present invention is mainly the generation of soft labels and the calculation of self-distillation loss. The specific training process and encoder selection are relatively flexible. The implementation process of the present invention is described below through a specific example, which specifically includes the following steps:

[0026] 1. Collect training data: It needs to include several videos and corresponding text descriptions. Each video needs at least one corresponding text description. Divide the training set into several mini-batches of size b, each mini-batch contains b video-text pairs.

[0027] 2. Select the pre-trained BERT-base-uncased as the text encoder, initialize a 4-layer 4-attention-head BERT as the video encoder, and set the hidden layer size to 512. Add a linear layer after the text encoder and video encoder to transform the dimensions of both modalities to 1024.

[0028] 3. Select a mini-batch and use the pre-trained expert model to extract features for each video. The expert models used are: S3D pre-trained on the Kinetics action recognition dataset, VGGish pre-trained on the Youtube8M dataset, and SENet-154 pre-trained on ImageNet. The extracted features of different videos are concatenated together to form a feature sequence T, which is input into the video encoder to obtain f(T). At the same time, according to the configuration of bert-base-uncased, the text is embedded into the dense space to obtain V, and then input into the text encoder to obtain g(V).

[0029] 4. Calculate the loss of the entire mini-batch according to equations (1), (2), (3), and (4), and calculate the gradient through back propagation. -5 The encoder parameters are updated with a learning rate of .

[0030] 5. Repeat steps 2, 3, and 4 several times to complete the training.

[0031] 6. During inference, select the average pooling result of the last layer of the encoder as the representation vector of the resource. When retrieving videos (texts), calculate the representation vectors of all texts (videos) in the database offline in advance. Calculate the representation vector of the query video (text) online, match it with the representation vectors of the database resources calculated in advance, and recall the ones with the highest similarity as the retrieval results.

Claims

1. A video-text cross-modal retrieval method based on cross-granularity self-distillation includes the following steps: Step 1: Pass the input of different modes through each modality encoder based on Transformer to obtain a series of feature vectors; Step 2: Average pool the feature sequences generated by the encoders of each modality to obtain the coarse-grained representation features of each modality; during inference, directly use the representation features for vector retrieval; during training, use this representation feature in a mini-batch of size b Calculate the cross-modal similarity s(T i ,V j ), thus forming a coarse-grained similarity matrix And calculate the InfoNCE loss by cross entropy Step 3: This step is only performed during training. Through the tokens screening module, a portion of "more important" tokens are selected for each modality, and then the cross-modal similarity pseudo-labels s of any two samples are generated through the fine-grained interaction between these key tokens. t2v (T i ,V j ) and s v2t (V j ,T i ), and calculate the cross-granularity self-distillation loss through KL divergence Step 4: This step is only performed during training; the final loss is calculated using the hyperparameter λ The network is iteratively optimized based on this loss.

2. The video-text cross-modal retrieval method based on cross-granularity self-distillation according to claim 1 is characterized in that: The cross-granularity self-distillation loss in the training phase is calculated as follows: The feature sequence f(T i ) and g(V j ) are sent to their respective tokens screening modules to select the "key" tokens, and then the inner products between the cross-modal tokens are calculated to obtain the fine-grained interaction matrix I; Then, by taking the maximum value of each row of I and then finding the average, we get the text T i For video V j The fine-grained similarity score s t2v (T i ,V j ); Similarly, we can first take the maximum value of each column of I and then find the average value to get the video V j For text T i The fine-grained similarity score s v2t (V j ,T i ); Furthermore, t2v (T i ,V j ) and s v2t (V j ,T i ) can form a text-video soft label matrix and video-text soft label matrix Using KL divergence to calculate cross-granularity self-distillation loss The formula is as follows: Among them, S i is the i-th row of matrix S, and Similarly; S T is the transpose of S; P and Q represent the probability distribution of dimension b respectively; Finally, the final loss is calculated by the following formula: Among them, λ is an adjustable hyperparameter.

3. The cross-granularity self-distillation loss in the training phase according to claim 2, characterized in that The structure of the tokens screening module is as follows: The tokens filtering module consists of a weight generator and a sorting selection operation; the weight generator is a lightweight network composed of "Linear-ReLU-Linear-Softmax", which is responsible for generating a weight between 0 and 1 for each token; The sorting selection operation is responsible for sorting each token from large to small according to its corresponding weight, and selecting the first k as the "key" tokens. At the same time, these "key" tokens are placed in their original order as the output of this module and provided to subsequent fine-grained interactions.

Citation Information

Patent Citations

  • Cross-modal pedestrian re-identification method based on self-imitation mutual distillation

    CN112926451A

  • Cross-modal retrieval method and device, electronic device and storage medium

    CN113157739A