Video Text Retrieval Method Based on Differential Multi-Scale and Multi-Granularity Feature Fusion

Through the differential multi-scale multi-grained feature fusion method, the problem of failing to effectively utilize video timing and fine-grained features in the prior art is solved, and higher-precision video text matching is achieved, cross-modal retrieval ability is enhanced, and the accuracy and efficiency of search results are improved.

CN116226449BActive Publication Date: 2025-07-29XIDIAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310050175.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-01
Publication Date
2025-07-29
Estimated Expiration
2043-02-01

AI Technical Summary

Technical Problem

The prior art fails to effectively utilize the timing and fine-grained features of videos, resulting in insufficient video text matching accuracy, especially when dealing with natural language text queries with complex semantics.

Method used

By constructing a differential multi-scale multi-grained feature fusion method, the global and local features of video and text are extracted using visual feature encoder and text feature encoder, the differential features of video frames are calculated, and the timing feature extraction module is combined to construct cross-similarity and polynomial loss functions for network training, enhancing feature matching capabilities.

Benefits of technology

It improves the accuracy of video text retrieval, reduces the semantic gap between modals, enhances the ability of cross-modal retrieval, and improves the accuracy and efficiency of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116226449B_ABST
    Figure CN116226449B_ABST
Patent Text Reader

Abstract

The present invention discloses a video text retrieval method based on differential multi-scale and multi-granularity feature fusion, which mainly solves the problem of low video text matching accuracy caused by the failure to fully utilize video temporal features and fine-grained information text annotation in the prior art. The implementation solution is as follows: obtaining a video frame sequence and a text annotation sequence; constructing a feature extraction network and extracting global and local features of the text annotation; differentiating the video frame features according to the time series and combining them with the frame features through a sequence feature extraction network to obtain local and global features of the video; calculating the global similarity and local similarity between the video and the text annotation, and calculating a loss function; training the network using the loss function; calculating the similarity between the video and the text annotation using the trained network and sorting to obtain the retrieval result. The present invention can reduce the semantic gap between different modalities, mine the temporal information in video modality data, improve the cross-modal retrieval accuracy, and can be used for video topic detection and content recommendation in video applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and further relates to a video text retrieval method, which can be used for video theme detection and content recommendation of video applications. Background Art

[0002] With the development of big data and 5G technologies, multimedia data on the Internet has shown explosive growth, and at the same time, many new retrieval requirements have emerged. Traditional retrieval methods mainly support concept-based retrieval for simple natural language text queries, which are ineffective for complex long natural language text queries with complex semantics, that is, they cannot meet diverse retrieval needs.

[0003] In recent years, cross-modal retrieval methods based on shared subspaces have emerged as a hot topic in the current multimedia research field. It maps video and natural language text modalities to a joint visual semantic shared space to calculate cross-modal semantic similarity as the basis for retrieval work, which can well meet the search needs of users between different media data.

[0004] Hunan University disclosed in its patent document with the application number CN202111312233.1 "A Cross-modal Text-Video Retrieval Method Based on Enhancement of Spatiotemporal Relationships". It uses a variety of pre-trained models to extract video global and local object features respectively, and finally realizes text-video retrieval through technical means of affine transformation mapping. Although the data processing methods for different modalities of this method improve the retrieval accuracy to a certain extent, the heterogeneity brought about by separately extracting features will lead to the inability to accurately find the common embedding subspace between video and text even through mapping, thus causing poor matching between video and text due to modality differences and affecting the retrieval accuracy.

[0005] Xidian University disclosed in its patent document with the application number: CN202110968279 a "Video Natural Language Text Retrieval Method Based on Spatial-Temporal Features". It uses three different types of neural networks to perform hierarchical fine-grained and comprehensive video unified representation of the spatial-temporal semantic information of the video, constructs a video text common semantic embedding network to fit the semantic gap of cross-modal data, and trains the network using a contrastive ranking loss function. The disadvantage of this method is that video and text features are extracted separately and not well fitted. At the same time, the method of using a common subspace mapping cannot effectively remove feature redundancy, which not only increases a large amount of computational cost during the training process of the model, but also has poor retrieval effects.

[0006] Yang X, Dong J, Cao Y et al. proposed a tree-structured augmented video natural language text retrieval method for complex natural language text queries in their published paper "Tree-Augmented Cross-Modal Encoding for Complex-Query Video Retrieval" (International ACM SIGIR Conference on Research and Development in Information Retrieval, 2020: 1339-1348). It performs fine-grained encoding by jointly learning the language structure of the query natural language text and the temporal representation of the video. Since the video spatial entity objects correspond to the "noun" part of the natural language text, which is the key information for retrieval, and this method only focuses on the model of temporal modeling, it is difficult to capture the spatial object information at the video region level, thus affecting the retrieval accuracy.

[0007] Huaishao Luo, Lei Ji, Ming Zhong et al. proposed a cross-modal matching method for end-to-end video clip retrieval in their published paper "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval" (arXiv: Computer Vision and Pattern Recognition, 2104.08860(2021)). It extracts the information of videos and texts through a pre-trained model, transfers the knowledge of the image-text retrieval model to video-language retrieval, and achieves SOTA results on multiple public datasets. However, the drawback of this method is that it aggregates the entire video without directly considering the text. Whether it is mean pooling or self-attention on the frames, without combining with the text expression, it cannot achieve thorough collaborative feature extraction, resulting in a lot of redundant information being fused into the video features, increasing the training burden, and generating a lot of un-described visual information for text encoding, affecting the retrieval results.

[0008] In summary, most of the existing technologies only use the static feature information of the video and do not effectively utilize the temporal features of the video, resulting in the inability of the model to effectively understand some content manifested by the temporal information and causing ambiguity in retrieval. At the same time, most of the existing technologies only use the global features of the video to match the global features of the text. Although semantic information can be basically paired macroscopically, due to the neglect of some fine-grained information of the video frames and the fine-grained information at the text word level, the retrieval results are often inaccurate. Therefore, how to find a better way of feature extraction and aggregation for the above problems is an urgent problem to be solved at present. Summary of the Invention

[0009] The purpose of the present invention is to overcome the deficiencies of the above-mentioned existing technologies and propose a video text retrieval method based on differential multi-scale multi-granularity feature fusion to effectively utilize the temporal features and fine-grained features of the video and improve the retrieval accuracy.

[0010] To achieve the above purpose, the technical solution of the present invention includes the following:

[0011] (1) Process the video dataset:

[0012] (1a) Select the video dataset to be trained and its corresponding text annotations, and use a video image generation tool to extract key frames from the video dataset according to the amount of information to obtain a video sequence set composed of pictures after sampling: V = {V i}, where: V i represents the i-th video sequence of the video dataset, and each video sequence is composed of n picture frames, i = 1, 2, 3,..., N, and N is the size of the video dataset;

[0013] (1b) Split the text annotations corresponding to the video by spaces to obtain the split text annotations;

[0014] (2) Construct a feature extraction network, that is, use a visual feature encoder and a text feature encoder as the feature extraction network, and initialize the feature network with the parameters in the existing CLIP pre-trained model;

[0015] (3) Obtain the global feature S i and the local feature T i , and obtain the visual feature sequence F i of the video sequence V i :

[0016] (3a) For a video sequence V i , extract its RGB pixel information, that is, red, green, and blue color features, to obtain 3 groups of feature matrices;

[0017] (3b) Construct a fully connected layer with the number of neuron nodes being the same as the dimension of each group of feature matrices obtained in (3a), and the parameters can be randomly initialized;

[0018] (3c) Slice each frame in the video sequence V i according to a given step size, then flatten the sliced features by group, and input them into this fully connected layer to be mapped into one-dimensional vectors;

[0019] (3d) Input the sliced text annotation obtained in (1b) into the text feature encoder, and output the global feature S i of the text annotation and the local feature Input the one-dimensional vectors of the video obtained in (3c) into the video feature encoder, and output the visual feature sequence i of the video sequence V where m represents the number of words in the current text annotation, n is the length of the video frames in this sequence, represents the feature of the p-th word in the i-th text annotation, represents the k-th frame visual feature of the video sequence V i ;

[0020] (4) Calculate the local feature and global feature of the video sequence V i :

[0021] (4a) Differentiate the visual feature sequence F i with different step sizes to obtain the differential features of the video frames:

[0022]

[0023] where represents the differential feature between the j-th frame and the k-th frame of the video sequence V i , represents the j-th frame visual feature of the video sequence V i , and k represents the differential step size;

[0024] (4b) Calculate all the differential features of a video frame, form them into a sequence, and insert the visual feature sequence of the current frame at the head, that is, for the j-th frame i in the video sequence V its differential feature sequence is: Similarly, calculate the differential feature sequences of other frames to obtain the multi-scale differential feature sequence i of the video sequence V

[0025] (4c) Construct a temporal feature extraction module, take the differential feature sequence Δ i obtained in (4b) as the input of this module, and extract the video sequence Vi The timing information of i local features of wherein represents the k-th local feature of the video sequence V i ;

[0026] (4d) According to the global feature S of the text annotation i and the corresponding video local feature L vi , calculate the global feature A of the video sequence V i ; i

[0027] (5) Calculate the final similarity between the video and the text annotation:

[0028] (5a) Calculate the cross-similarity Sim i between the global feature S of the text annotation vi and the local feature L of the video sequence S-f ;

[0029] (5b) According to the global feature A of the video sequence V i and the local feature T of the text annotation i , calculate the cross-similarity Sim i from the video global feature to the text annotation local feature: V-w :

[0030] (5c) According to the full-local feature A of the video V i and the global feature S of the text annotation i , calculate the global feature similarity Sim i from the video to the text annotation; S-A

[0031] (5d) According to the results of (5a), (5b), and (5c), obtain the following final similarity between the video and the text annotation:

[0032] Sim(S, V) = (Sim S-A + Sim V-w + Sim S-f ) / 3

[0033] where S represents the text annotation and V represents the video;

[0034] (6) Train the feature extraction network:

[0035] (6a) According to the final similarity between the video and the text annotation, construct the total loss function L of the feature extraction network:

[0036] (6a1) Calculate the prior probability of video features to text annotation features and the prior probability of text annotation features to video features based on the final similarity Sim(S, V) between the video and text annotation obtained from (5d). and the prior probability of text annotation features to video features

[0037] (6a2) Calculate the matching loss from video to text annotation and the matching loss from text annotation to video respectively using the cross - entropy function based on the prior probabilities obtained from (6a1). and the matching loss from text annotation to video

[0038] (6a3) Calculate the polynomial loss from video to text annotation and the polynomial loss from text annotation to video based on the final similarity Sim(S, V) between the video and text annotation obtained from (5d). and the polynomial loss from text annotation to video

[0039] (6a4) Obtain the total loss function L of the feature extraction network according to the results of (6a1), (6a2), and (6a3):

[0040]

[0041] where λ1 and λ2 represent loss weights;

[0042] (6b) Update the parameters of the feature extraction network:

[0043] (6b1) Set the initial value of the learning rate of the feature extraction network to 1e - 7, the initial value of the learning rate of the temporal feature extraction module to 1e - 4, and the initial value of the neuron dropout rate to 0.1;

[0044] (6b2) Train the model using the Adam optimizer, set the batch size to 64, calculate the total loss L according to the current network parameter values, and iteratively update the learning rate, neuron dropout rate, and parameter values of the loss function through the backpropagation of L, recalculate to obtain a new round of L, and repeat this process until the total loss function L converges to the minimum to obtain the trained video - text matching model;

[0045] (7) Input the video - text to be retrieved into the trained video - text matching model and sort the model output results to obtain the final retrieval results.

[0046] Compared with the prior art, the present invention has the following effects:

[0047] First, since the present invention calculates the multi - scale difference sequence of the video during the calculation of video features, it adds temporal information to the frame features of the video, not only enhancing the temporal representation ability of the video but also eliminating the ambiguous representation of text annotations.

[0048] Second, in the process of constructing the loss function, the present invention adds the calculation of cross similarity, enabling the feature extraction network to not only focus on the matching between global features but also on the matching relationship between the frame-level features of the video and the word-level features of the text annotation, enhancing the ability of the feature extraction network to align fine-grained information, and at the same time providing guidance for the training of the global features of the video and text annotation, enhancing the cross-modal retrieval ability.

[0049] Third, in the process of constructing the loss function, the present invention adds the calculation of the polynomial loss function, fully excavating the difficult sample information, solving the problem that most video text annotations contribute insufficiently to the training of the feature extraction network in the later stage of training, further enhancing the cross-modal retrieval ability, and improving the retrieval accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is the implementation flowchart of the present invention;

[0051] Figure 2 is the structural diagram of the temporal feature extraction module in the present invention;

[0052] Figure 3 is the enlarged structural diagram of the aggregation module in the present invention;

[0053] Figure 4 is the schematic diagram of the example of video text annotation matching in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The embodiments and effects of the present invention will be further described in detail below with reference to the drawings.

[0055] Refer to Figure 1 , the implementation steps of this example are as follows:

[0056] Step 1, process the data set.

[0057] 1.1) Select the video data set to be trained and its corresponding text annotation, and extract the key frames from the video data set according to the information volume through a video image generation tool to obtain a video sequence set composed of pictures after sampling: V = {V i}, where: V i represents the i-th video sequence of the video data set, and each video sequence is composed of n frame pictures, i = 1, 2, 3,..., N, N is the size of the video data set, and existing video image generation tools include ffmpeg and opencv. This example uses but is not limited to ffmpeg;

[0058] 1.2) Split the text annotation corresponding to the video by spaces to obtain the split text annotation.

[0059] Step 2: Construct a feature extraction network.

[0060] The specific implementation of this step is to select a visual feature encoder and a text feature encoder composed of 12 layers of Transformer respectively as the feature extraction network, and use the parameters in the existing contrastive language-image pre-training model, that is, the CLIP model, to initialize this feature network.

[0061] The structures of the visual feature encoder and the text feature encoder are as follows:

[0062] The first 6 layers of the visual feature encoder are encoders, and the last 6 layers are decoders. Each layer of the encoder is composed of 8-head self-attention layers and a feedback network layer. Each layer of the decoder is composed of 8-head self-attention layers, masked self-attention layers and a feedback network layer, and its output dimension is 512 dimensions; use the parameters in the CLIP pre-training model as the initialization parameters of the visual feature encoder;

[0063] The structure and output dimension of the text feature encoder are the same as those of the video feature encoder, and use the CLS position feature output by the last layer of Transformer as the global feature of the text annotation; use the parameters in the CLIP pre-training model as the initialization parameters of the text feature encoder.

[0064] Step 3: Obtain the global feature S i and local feature T i , and obtain the visual feature sequence F i of the video sequence V i .

[0065] 3.1) For the video sequence V i , extract its RGB pixel information, that is, red, green, and blue color features, to obtain 3 groups of feature matrices. The dimension of each group of feature matrices is the width × height of the image pixels, and the elements in the matrix take values from 0 to 255;

[0066] 3.2) Construct a fully connected layer, the number of its neuron nodes is the same as the dimension of each group of feature matrices obtained in (3.1), and the parameters can be randomly initialized;

[0067] 3.3) Slice each frame in the video sequence V i at a given step size, then flatten the sliced features by group, and input them into this fully connected layer to map them into one-dimensional vectors;

[0068] 3.4) Input the sliced text annotation obtained in step 1.2) into the text feature encoder, and output the global feature S i and local feature where m represents the number of words in the current text annotation, Represents the feature of the p-th word in the i-th text annotation;

[0069] 3.5) Input the one-dimensional vector of the video obtained in step 3.3) into the video feature encoder to output the video sequence V i of the visual feature sequence where n is the length of the video frames in this sequence, represents the video sequence V i of the k-th frame visual feature.

[0070] Step 4, calculate the local and global features of the video sequence V i of.

[0071] 4.1) Differentiate the visual feature sequence F i with different step sizes to obtain the differential features of the video frames:

[0072]

[0073] where represents the differential feature between the j-th frame and the k-th frame of the video sequence V i of, represents the j-th frame visual feature of the video sequence V i of, and k represents the differential step size;

[0074] 4.2) Calculate all the differential features of a video frame, form them into a sequence, and insert the visual feature sequence of the current frame at the head, that is, for the j-th frame i in the video sequence V its differential feature sequence is:

[0075]

[0076] Similarly, calculate the differential feature sequences of other frames to obtain the multi-scale differential feature sequence i of the video sequence V

[0077] 4.3) Refer to Figure 2 , construct a temporal feature extraction module composed of a cascade of a time series extraction part and an aggregation part. The time series extraction part uses a one-layer LSTM or a four-layer transformer with randomly initialized parameters as the network structure, and the input is the multi-scale differential feature sequence Δ i of the video V i , which is used to extract the temporal information of the video sequence V i ; the aggregation part is to aggregate the output result of the last layer of the time series extraction part with the frame features in a superimposed manner to output the local feature i of the video sequence V where Denote the video sequence as V i as the k-th local feature;

[0078] 4.4) According to the global feature S i annotated by the text and the corresponding video local feature L vi , calculate the global feature A i of the video sequence V i :

[0079] 4.4.1) Calculate the cross-similarity Sim i from the globally text-annotated feature S vi to the video local feature L

[0080]

[0081] where denotes the j-th local feature of video V i , T represents the transpose, and ||·|| represents the modulus operation;

[0082] 4.4.2) Calculate the global feature A i of the video sequence V i :

[0083]

[0084] where denotes the j-th local feature of video V i .

[0085] Step 5, calculate the final similarity between the video and the text annotation:

[0086] Referring to Figure 3 , the specific implementation of this step is as follows:

[0087] 5.1) Calculate the cross-similarity Sim i between the globally text-annotated feature S vi and the local feature L S-f of the video sequence:

[0088]

[0089] where ω is a hyperparameter set to 0.1 and exp is the exponential operation;

[0090] 5.2) According to the global feature A i of the video sequence V i and the local feature T i annotated by the text, calculate the cross-similarity Sim V-w from the video global feature to the text-annotated local feature:

[0091]

[0092] Among them represents the video V i corresponding to the feature of the j-th word in the text annotation, and m represents the number of words in the text annotation;

[0093] 5.3) According to the video V j and its global and local features A j and the global feature S of the text annotation i calculate the global feature similarity Sim S-A from the video to the text annotation:

[0094]

[0095] where T represents transpose, and ||·|| represents the modulus operation;

[0096] 5.4) According to the results of 5.1), 5.2), and 5.3), obtain the following final similarity between the video and the text annotation:

[0097] Sim(S, V) = (Sim S-A + Sim V-w + Sim S-f ) / 3

[0098] where S represents the text annotation and V represents the video.

[0099] Step 6: Train the feature extraction network.

[0100] 6.1) According to the final similarity between the video and the text annotation, construct the total loss function L of the feature extraction network:

[0101] 6.1.1) According to the final similarity Sim(S, V) between the video and the text annotation obtained in step 5.4), calculate the prior probability of the video feature to the text annotation feature

[0102]

[0103]

[0104] where respectively represent the prior probability of the video to the similarities of all text annotations in a batch and the prior probability of the text annotation to the similarities of all videos in a batch; η represents a hyperparameter, and its value is set to 100;

[0105] 6.1.2) According to the prior probabilities obtained in step 6.1.1), use the cross-entropy function to calculate the matching loss The matching loss between the video and the text annotation

[0106]

[0107]

[0108] where respectively represent the matching loss from the video to the text annotation and the matching loss from the text annotation to the video, κ is a hyperparameter set to 0.5, and B is the batch size;

[0109] 6.1.3) Calculate the polynomial loss from the video to the text annotation based on the final similarity Sim(S, V) between the video and the text annotation obtained in step 5.4) and the polynomial loss from the text annotation to the video

[0110]

[0111]

[0112] where represents the polynomial loss from the video to the text annotation, represents the polynomial loss from the text annotation to the video, B represents the training batch size, α, β represent adjustment parameters, p represents the polynomial order, and [·] + = max(0, ·);

[0113] Using the polynomial function gives a larger weight to the sample pairs that contribute more to the model and reduces the weight of the sample pairs that contribute less to the model, which can make the training more focused and make better use of the training data. Under the constraints of the above loss function, the feature extraction network highly retains the cross-modal similarity semantics of the video and text annotation features, making the positions of the semantically similar video and text annotation features closer and the positions of the semantically dissimilar video and text annotation features farther in the common semantic space.

[0114] 6.1.4) According to the results of steps 6.1.1), 6.1.2), and 6.1.3), obtain the total loss function L of the feature extraction network:

[0115]

[0116] where λ1 and λ2 represent loss weights;

[0117] 6.2) Update the parameters of the feature extraction network:

[0118] 6.2.1) Set the initial learning rate of the feature extraction network to 1e-7, the initial learning rate of the temporal feature extraction module to 1e-4, and the initial neuron dropout rate to 0.1;

[0119] 6.2.2) Use the Adam optimizer to train the model, set the batch size to 64, calculate the total loss L according to the current network parameter values, and iteratively update the learning rate of the network, the neuron dropout rate, and the parameter values of the loss function through the backpropagation of L. Recalculate to obtain a new round of L, and repeat this process until the total loss function L converges to the minimum to obtain the trained video-text matching model;

[0120] Step 7, input the video text to be retrieved into the trained video-text matching model, and sort the model output results to obtain the final retrieval results.

[0121] Refer to Figure 4 , in this example, a video and a text annotation are selected and input into the video feature encoder and text feature encoder in the feature extraction network respectively to obtain the frame features of the video and the global and local features of the text annotation; the frame features of the video are obtained through differential with different strides to obtain multi-scale differential features, which are input into the temporal feature extraction module to obtain the local and global features of the video; calculate the cross-similarity from the video global to the text annotation local and the cross-similarity from the text annotation global to the video local, and then calculate the similarity from the video global to the text annotation global. Aggregate these three similarities together to obtain the total similarity. The sorting result of the total similarities of multiple video-text annotations is the retrieval result.

[0122] The effects of the present invention will be further described below in conjunction with simulation experiments.

[0123] 1. Simulation conditions:

[0124] The hardware platform for the simulation experiment of the present invention is: NVIDIA GEFORCE RTX 3090 GPU.

[0125] The software platform for the simulation experiment of the present invention is: Ubuntu22 operating system and PyTorch 1.7.1.

[0126] The present invention has conducted experimental tests on two benchmark datasets: MSR-VTT and MSVD to evaluate the performance of the model proposed by the present invention.

[0127] The MSR-VTT dataset contains 10,000 video clips, and each video clip has 20 different descriptive text matches. In this paper, the 1K Data Split segmentation method is used for training. The training set and the test set contain 9,000 and 1,000 video clips respectively.

[0128] The MSVD dataset contains 1970 videos, with each video ranging from 1 to 62 seconds in length. The training set, validation set, and test set contain 1200, 100, and 670 videos respectively. Each video has approximately 40 related English sentences, with a total of 1970 video data and 78800 text data.

[0129] 2. Simulation experiment content and analysis of simulation results:

[0130] Simulation 1: The present invention and the existing four retrieval methods, namely HGR, Frozen, CF-GNN, and CLIP4Clip, are respectively used to perform retrieval on the MSR-VTT dataset, and the recall rate, median rank, and mean rank of each method are calculated. The results are shown in Table 1.

[0131] Table 1 Comparison table of retrieval accuracies between the present invention and the prior art in the MSR-VTT dataset

[0132]

[0133] In the table, the recall rate Recall at K (R@K) represents the probability of correctly predicting the item to be retrieved among the top K retrieval results for the ordered retrieval results; the median rank Median Rank (MedR) represents the median of the positions where the item to be retrieved is correctly predicted for the ordered retrieval results; the mean rank Mean Rank (MnR) represents the average of the positions where the item to be retrieved is correctly predicted for the ordered retrieval results. The larger the evaluation criterion R@K, the higher the retrieval accuracy; the smaller the evaluation criteria MedR and MnR, the higher the retrieval accuracy.

[0134] It can be seen from Table 1 that for the method of the present invention, when using text annotation to retrieve videos, the probabilities of correctly predicting the item to be retrieved among the top 1, 5, and 10 retrieval results, namely R@1, R@5, and R@10, are 43.8, 72.6, and 84.3 respectively; when using videos to retrieve text annotations, the probabilities of correctly predicting the item to be retrieved among the top 1, 5, and 10 retrieval results, namely R@1, R@5, and R@10, are 45.8, 77.6, and 83.5 respectively, all far higher than the 2020 SOTA model, the fine-grained hierarchical graph reasoning HGR method, and higher than the 2021 SOTA model CLIP4Clip proposed in this field. In particular, the R@1 index for retrieving videos with text annotation has increased by 376.1% compared to HGR and also increased by 1.3% compared to CLIP2Video.

[0135] Simulation 2: The present invention and the existing four retrieval methods, namely HGR, Frozen, CF-GNN, and CLIP4Clip, are respectively used to perform retrieval on the MSVD dataset, and the recall rate, median rank, and mean rank of each method are calculated. The results are shown in Table 2.

[0136] Table 2 Comparison Table of Retrieval Precision between the Present Invention and the Prior Art in the MSVD Dataset

[0137]

[0138] As can be seen from Table 2, for the method of the present invention to retrieve videos with text annotations, the probabilities R@1, R@5, and R@10 of correctly predicting the item to be retrieved in the top 1, 5, and 10 retrieval results are 46.8, 76.7, and 85.5 respectively. The R@1 value has increased by 403.2% compared to HGR, by 38.9% compared to Frozen, by 105.2% compared to CF-GNN, and by 1.3% compared to CLIP4Clip.

[0139] In the above simulation experiment, the following 4 prior arts were adopted:

[0140] The prior art HGR refers to the video natural language text retrieval method proposed by Chen S et al. in "Fine-grained video-text retrieval with hierarchical graph reasoning." (Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 10638-10647, 2020), abbreviated as the fine-grained hierarchical graph reasoning HGR video natural language text retrieval method.

[0141] The prior art Frozen refers to the cross-modal information retrieval method proposed by Max Bain et al. in "Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, international conference on computer vision, 2021", abbreviated as Frozen.

[0142] The prior art CF-GNN refers to the cross-modal information retrieval method proposed by Wei Wang et al. in "Learning Coarse-to-Fine Graph Neural Networks for Video-Text Retrieval, IEEE TRANSACTIONS ON MULTIMEDIA 2021", abbreviated as CF-GNN.

[0143] The prior art CLIP4Clip refers to the cross-modal information retrieval method proposed by Huaishao Luo et al. in "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval. arXiv: Computer Vision and Pattern Recognition 2021", abbreviated as CLIP4Clip.

[0144] The above experimental results prove that the present invention can more accurately realize the mutual retrieval of videos and text annotations, indicating that the method of extracting features through a pre-trained model, enhancing the temporal features of videos by using multi-scale differential frame features, and calculating cross-similarity to train the feature extraction network can improve the retrieval accuracy of video text annotations, proving the correctness and effectiveness of the method proposed by the present invention.

[0145] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these modifications and changes based on the idea of the present invention are still within the scope of protection of the claims of the present invention.

Claims

1. A video text retrieval method based on differential multi-scale and multi-granularity feature fusion, characterized in that, It includes the following: (1) Process the video dataset: (1a) Select the video dataset to be trained and its corresponding text annotations, and extract key frames from the video dataset according to the amount of information through a video image generation tool to obtain a video sequence set composed of pictures after sampling: V = {V i}, where: V i represents the i-th video sequence of the video dataset, and each video sequence is composed of n frames of pictures, i = 1, 2, 3,..., N, and N is the size of the video dataset; (1b) Split the text annotation corresponding to the video by spaces to obtain the split text annotation; (2) Construct a feature extraction network, that is, use a visual feature encoder and a text feature encoder as the feature extraction network, and initialize the feature network with the parameters in an existing pre-trained CLIP model; (3) Obtain the global feature S of the text annotation i and the local feature T i , and obtain the video sequence V i 's visual feature sequence F i : (3a) For a video sequence V i , extract its RGB pixel information, that is, the red, green, and blue color features, to obtain three groups of feature matrices; (3b) Construct a fully connected layer, the number of its neuron nodes is the same as the dimension of each group of feature matrices obtained in (3a), and the parameters can be randomly initialized; (3c) Segment each frame in the video sequence V according to the given step size, then flatten the segmented features by group and input them into this fully connected layer to map them into a one-dimensional vector; i ​ (3d) Input the segmented text annotation obtained in (1b) into the text feature encoder to output the global feature S of the text annotation i and local features Input the one-dimensional vector of the video obtained in (3c) into the video feature encoder to output the video sequence V i of the visual feature sequence F i ={f i 1 , f i 2 ,..., f i k ,...., f i n}, where m represents the number of words in the current text annotation, n is the length of the video frames in the sequence, w i p represents the feature of the p-th word in the i-th text annotation, and f i k represents the visual feature of the k-th frame of the video sequence V i ; (4) Calculate the local and global features of video sequence V i : (4a) Differentiate the visual feature sequence F i with different step sizes to obtain the differential features of the video frames: d i jk = f i k -f i j , k = 1, 2,... i - 1, i + 1,..., n where d i jk represents the differential feature between the j-th frame and the k-th frame of the video sequence V i and f i j represents the visual feature of the j-th frame of the video sequence V i and k represents the differential step size; (4b) Calculate all the differential features of a video frame, form them into a sequence, and insert the visual feature sequence of the current frame at the head, that is, for the video sequence V i for the j-th frame f i j in it, its differential feature sequence is: Similarly, calculate the differential feature sequences of other frames to obtain the multi-scale differential feature sequence of the video sequence V i (4c) Construct a temporal feature extraction module, and use the differential feature sequence Δ i obtained in (4b) as the input of this module to extract the temporal information of the video sequence V i and output the local features of the video sequence V i where represents the k-th local feature of the video sequence V i ;​​ (4d) Calculate the global feature A of the video sequence V according to the global feature S marked by the text i and the corresponding local video feature L vi , calculate the global feature A of the video sequence V i ; i ; (5) Calculate the final similarity between the video and the text annotation: (5a) Calculate the global feature S of the text annotation i and the local feature L of the video sequence vi to obtain the cross similarity Sim S-f ; (5b) According to the global feature A of video sequence V i and the local feature T of text annotation i , calculate the cross similarity Sim from the video global feature to the text annotation local feature i : V-w ​ (5c)According to the full local feature A of video V i and the global feature S of the text annotation i , calculate the global feature similarity Sim from the video to the text annotation i ; S-A ​ (5d) Obtain the following final similarity between the video and the text annotation according to the results of (5a), (5b), and (5c): Sim(S,V) = (Sim S-A + Sim V-w + Sim S-f ) / 3 where S represents the text annotation and V represents the video; (6) Train the feature extraction network: (6a) Construct the total loss function L of the feature extraction network according to the final similarity between the video and the text annotation: (6a1) Calculate the prior probability of video features to text annotation features and the prior probability of text annotation features to video features based on the final similarity Sim(S, V) between the video and text annotations obtained from (5d). and the prior probability of text annotation features to video features Based on the prior probability obtained in (6a1), using the cross-entropy function, calculate the matching loss from video to text annotation and the matching loss from text annotation to video respectively and the matching loss from text annotation to video (6a3) The final similarity Sim(S, V) between the video and text annotations obtained according to (5d), calculate the polynomial loss from the video to the text annotation and the polynomial loss from the text annotation to the video (6a4) Obtain the total loss function L of the feature extraction network according to the results of (6a1), (6a2), and (6a3): where λ1 and λ2 represent loss weights; (6b) Update the parameters of the feature extraction network: (6b1) Set the initial value of the learning rate of the feature extraction network to 1e-7, the initial value of the learning rate of the temporal feature extraction module to 1e-4, and the initial value of the neuron dropout rate to 0.1; (6b2) Use the Adam optimizer to train the model, set the batch size to 64, calculate the total loss L according to the current network parameter values, and iteratively update the learning rate, neuron dropout rate, and parameter values of the loss function through the backpropagation of L, recalculate to obtain a new round of L, and repeat this process until the total loss function L converges to the minimum to obtain a trained video-text matching model; (7) Input the video text to be retrieved into the trained video-text matching model, and sort the model output results to obtain the final retrieval results.

2. The method according to claim 1, characterized in that: Step (2) constitutes the visual feature encoder and the text feature encoder in the feature extraction network, and the structure is as follows: The visual feature encoder is composed of 12 layers of Transformers. The first 6 layers are encoders, and the last 6 layers are decoders. Each layer of the encoder is composed of a multi-head self-attention layer and a feedback network layer. Each layer of the decoder is composed of a multi-head self-attention layer, a masked self-attention layer, and a feedback network layer, and its output dimension is 512 dimensions; The text feature encoder has the same structure and output dimension as the video feature encoder, and uses the CLS position feature output by the last layer of the Transformer as the global feature of the text annotation.

3. The method according to claim 1, wherein: Step (3a) obtains 3 groups of feature matrices, which respectively represent the RGB of the color feature images of the red, green, and blue channels of the image. The dimension of each group of feature matrices is the width × height of the image pixels, and the elements in the matrix take values from 0 to 255, forming a total of three groups of feature matrices.

4. The method according to claim 1, wherein: The temporal feature extraction module constructed in step (4c) includes a time series extraction part and an aggregation part. The time series extraction part uses a one-layer LSTM or a four-layer Transformer with randomly initialized parameters as the network structure, and the input is the multi-scale differential feature sequence Δ of the video V i ; The aggregation part aggregates the output result of the last layer of the time series extraction part with the frame features in a superimposed manner, and finally outputs the local feature L of the video i vi .​ 5. The method according to claim 1, wherein: The global feature S annotated according to the text in step (4d) i and the corresponding local video feature L vi , calculate the global feature A i of the video sequence V i , which is implemented as follows: (4d1) Calculate the cross-similarity Sim between the global feature S of text annotation i and the local feature L of the video vi as follows: Among them represents the j-th local feature of video V i , T represents transpose, and ||·|| represents the modulus operation; (4d2) Calculate the global feature A of video sequence V i i :​ where a i j represents the j-th local feature of video V i ​ 6. The method according to claim 1, wherein: Calculate the global feature S of the text annotation in step (5a) i and the local feature L of the video sequence vi The cross similarity Sim S-f is as follows: where ω is a hyperparameter set to 0.1, and exp is the exponential operation.

7. The method according to claim 1, wherein: In step (5b), according to the global feature A i of the video sequence V i and the local feature T i of the text annotation, calculate the cross similarity Sim V-w from the video global feature to the text annotation local feature. The formula is as follows: Among them represents video V i corresponds to the feature of the j-th word in the text annotation, and m represents the number of words in the text annotation.

8. The method according to claim 1, wherein: In step (5c), according to the global and local features A j of video V j and the global feature S of the text annotation i , calculate the global feature similarity Sim S-A from the video to the text annotation, with the formula as follows: where T represents the transpose and ||·|| represents the modulus operation.

9. The method according to claim 1, wherein: In step (6a1), based on the final similarity Sim(S, V) between the video and the text annotation, calculate the prior probability of the video feature with respect to the text annotation feature and the prior probability of the text annotation feature with respect to the video feature The formula is as follows: where respectively represent the prior probability of the video for the similarity of all text annotations in a batch and the prior probability of the text annotation for the similarity of all videos in a batch. η represents a hyperparameter, which is set to 100.

10. The method according to claim 1, wherein: The prior probability of text annotation features based on video features in step (6a2) and the prior probability of video features based on text annotation features are used to calculate the matching loss from video to text annotation and the matching loss from text annotation to video The formulas are as follows: Among them, respectively represent the matching loss from video to text annotation and the matching loss from text annotation to video. κ is a hyperparameter set to 0.5, and B is the batch size.

11. The method according to claim 1, wherein: Step (6a3) calculates the polynomial loss from the video to the text annotation and the polynomial loss from the text annotation to the video according to the final similarity Sim(S, V) between the video and the text annotation. and the polynomial loss from the text annotation to the video The formulas are as follows: Among them, represents the polynomial loss from video to text annotation, represents the polynomial loss from text annotation to video, B represents the training batch size, α, β represent adjustment parameters, p represents the polynomial order, and [·] + = max(0, ·).

Citation Information

Patent Citations

  • Cross-modal text-video retrieval method based on space-time relationship enhancement

    CN114048351A

  • Video natural language text retrieval method based on space time sequence characteristics

    CN113704546A

  • Cross-modal image-text retrieval method based on multi-granularity feature fusion

    CN115033670A