A Cross-modal Video Retrieval Method Inspired by Reading Strategies
By introducing a cross-modal video retrieval method inspired by reading strategies in video representation learning, using preview branches and intensive reading branches to learn video representation together, the problem of incomplete video information capture and lack of information interaction in the existing technology is solved, and more efficient video information capture and robustness is achieved.
Patent Information
- Application Number
- CN202111084182.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-09-15
AI Technical Summary
The existing video representation learning model has a coarse grain size from a global perspective in the capture of video information, and cannot fully capture video information. Moreover, the multi-branch model has a lower performance due to its independence between branches and lacks information interaction and transmission.
The cross-modal video retrieval method inspired by reading strategies is adopted to extract the initial video features through a pre-trained convolutional neural network, and the preview branch and intensive reading branch are used to learn video representation together. The preview branch captures video overview information through bidirectional GRU network and average pooling operations, and the intensive reading branch extracts multi-grained fragment features through convolutional neural network and perceptual preview attention operations, encodes the text with the BERT model, and finally calculates the video text similarity in the mixed space.
The best balance in performance and model complexity is achieved, and the video information can be captured more comprehensively, with good robustness and the potential to process complex videos.
Smart Images

Figure CN114003770B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video cross-modal retrieval, and in particular to a cross-modal video retrieval method inspired by reading strategies. Background Art
[0002] With the increasing popularity of video streaming platforms such as YouTube and TikTok, video data has exploded in growth. The goal of the present invention is to achieve language-based video retrieval. Given a query in the form of a natural language sentence, it is required to retrieve videos semantically related to the given query from a large number of unlabeled videos.
[0003] To build such a video retrieval model, it is crucial to calculate the semantic similarity between two modalities, namely video and text. Early language-based video retrieval was concept-based methods that represented videos and text queries into a predefined concept space and calculated similarity through concept matching. Due to the limited performance of concept-based methods, cross-modal representation learning-based methods are more favored, which learn a joint embedding space in a concept-free manner for cross-modal similarity measurement and show better performance.
[0004] Based on the cross-modal representation learning method, the present invention focuses on video representation learning, which is an important part of language-based video retrieval. A typical method of video representation learning is to first extract visual features from video frames through a pre-trained CNN model, and then aggregate the frame-level features into video-level features through average pooling or max pooling operations. Subsequently, a fully connected layer is used to further map the video-level features into the joint embedding space. Current video representation learning models can be roughly divided into two categories according to their structures: single-branch models and multi-branch models. Single-branch models mainly replace the above simple pooling strategy and further explore sequence-aware deep neural networks, but they usually perform coarse-grained video representation learning from a global perspective, so they may not be able to comprehensively capture video information. Multi-branch models encode videos by using multiple branches. Although they have better performance improvement, in this structure, the branches are independent of each other and there is no further information interaction and transmission, so this method is considered sub-optimal. Summary of the Invention
[0005] In order to solve the above technical problems existing in the prior art, the present invention provides a cross-modal video retrieval method inspired by reading strategies, and its specific technical solutions are as follows:
[0006] A cross-modal video retrieval method inspired by reading strategies, comprising the following steps:
[0007] (1) Extract the initial features of the video modality using a pre-trained convolutional neural network to obtain the initial feature sequence of the video;
[0008] (2) Input the initial feature sequence and encode it through the preview branch to obtain the preview features in the video;
[0009] (3) Input the initial feature sequence and encode it through the intensive reading branch to obtain multi-granularity segment features, then perceive and integrate the preview features to extract the intensive reading features;
[0010] (4) Use a pre-trained BERT model to encode the text modality to obtain the text multi-level encoded features;
[0011] (5) Map the video preview features and intensive reading features to the corresponding hybrid spaces respectively with the text multi-level encoded features, and calculate the similarity between the video modality and the text modality through the hybrid spaces for cross-modal matching;
[0012] (6) Optimize and train the retrieval model established through steps (1) to (5), and finally input the video and text into the trained retrieval model to achieve cross-modal retrieval from text to video.
[0013] Further, the specific operation of step (2) is as follows: input the video frame feature sequence into the bidirectional GRU network of the preview branch. The bidirectional GRU consists of a forward GRU and a backward GRU. Concatenate the hidden states at all specific time steps {t = 1, …, m} in the forward GRU and the backward GRU as the output of the bidirectional GRU to obtain a feature vector sequence H = {h 1 , h 2 , …, h m}, with a size of m × 1024 dimensions; then apply average pooling operation to the feature vector sequence H along the time dimension to obtain the preview feature vector, that is
[0014]
[0015] Further, step (3) specifically includes the following steps:
[0016] (3-1) First, use the fully connected layer of the intensive reading branch to reduce the dimension of the visual feature sequence to obtain the reduced visual feature sequence V′;
[0017] (3-2) Then input V′ into a convolutional neural network CNN with a convolutional kernel size of n, a stride of s, and a number of convolutional kernels of r to extract segment features of different lengths. The specific formula is expressed as:
[0018] C n = δ(Conv1D r,n,s (V′))
[0019] where δ represents the Relu activation function;
[0020] Put the piecewise features generated by convolutional kernels of different sizes together to obtain multi-granularity segment features, that is:
[0021]
[0022] where φ represents the size of the convolutional kernel, m n represents the number of segments of length n, r is the dimension of the segment feature vector, after vectorizing the segment features it is C′, and the visual feature sequence V′ is used as the segment feature of length 1;
[0023] (3-3) Perform a perceptual preview attention operation on the multi-granularity segment features to obtain a refined feature vector.
[0024] Furthermore, the specific steps of step (3-3) are as follows:
[0025] First, map the preview feature vector p to a query feature vector Q of dimension d k , map the segment feature vector C′ to a key feature vector K of dimension d k and a value feature vector V of dimension d v respectively, then use the query and value to calculate the attention weights through dot product, and then perform a weighted sum of the obtained attention weights and the value feature vector to obtain an attention feature vector O, that is:
[0026] O = W 4 Attention(pW 1 , C′W 2 , C′W 3 )
[0027]
[0028] where W 1 , W 2 , W 3 and W 4 are learnable mapping matrix parameters;
[0029] Then use a residual operation to enhance the input to obtain an updated attention feature vector O′, that is:
[0030] O′ = LN(O + maxpool(C′))
[0031] where LN represents layer normalization and max pooling operations, and a pooling operation is performed on the segment feature vector along the time dimension;
[0032] After obtaining the attention feature vector, a feed-forward network with residual and layer normalization is used to further enhance the above feature vector. The feed-forward network is implemented by using a multi-layer perceptron MLP, that is:
[0033] PaA(p, C′) = LN(o′ + MLP(o′))
[0034] where MLP consists of two fully connected layers and a Relu activation function;
[0035] Finally, for the multi-granularity segment features, the above-mentioned perceptual preview attention operation is performed on each granularity in parallel, and the outputs of each granularity are concatenated as the final output of the intensive reading branch to obtain the intensive reading feature vector g, which is specifically expressed as follows:
[0036]
[0037] where ConCat represents the concatenation operation.
[0038] Furthermore, the hybrid space includes a preview hybrid space and an intensive reading hybrid space. In the hybrid space, a fully connected layer is used to map the video-text pair to a concept space and a latent space respectively, and the cosine similarity is used to calculate the similarity of the video-text pair.
[0039] Furthermore, the concept space uses a binary cross-entropy loss and a triplet ranking loss to jointly constrain the learning of this space; the latent space uses a triplet ranking loss to constrain the learning of this space.
[0040] Furthermore, step (6) is specifically: adding the loss of the concept space and the loss of the latent space as the joint loss of the hybrid space, optimizing the retrieval model by minimizing the sum of the losses of the preview hybrid space and the intensive reading hybrid space, training the retrieval model using the Adam optimizer, mapping the text and the video to the preview hybrid space and the intensive reading hybrid space respectively, and in these two spaces, sorting the candidate videos by calculating the cosine similarity between the video-text pairs, and taking the result with better similarity as the final return result, and implementing the cross-modal matching task through this process.
[0041] Furthermore, when training the retrieval model using the Adam optimizer, during the training process, an early stopping training strategy is adopted. If the validation loss does not decrease in three consecutive epochs, the learning rate is divided by 2. If the validation performance does not improve in 10 consecutive epochs, early stopping occurs.
[0042] Advantages of the present invention:
[0043] A cross-modal video retrieval method inspired by reading strategies of the present invention uses a preview branch and a intensive reading branch to jointly learn to represent videos and establish a retrieval model. Compared with models based on heavyweight Transformers, it achieves the best balance in terms of performance and model complexity. Moreover, the video representation architecture of the two branches is insensitive to the implementation of attention, has good robustness, and has the potential to process complex videos. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a schematic flowchart of the cross-modal video retrieval method inspired by reading strategies according to an embodiment of the present invention;
[0045] Figure 2 It is a schematic flowchart of the perceptual preview attention operation of the retrieval model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] In order to make the objectives, technical solutions, and technical effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings of the specification.
[0047] As Figure 1 shown, the present invention proposes a cross-modal video retrieval method inspired by reading strategies, which specifically includes the following steps:
[0048] (1) Different feature extraction methods are used for the video to obtain the initial features of the video modal data. Specifically, given a video, first, a video frame sequence with a pre-specified interval of 0.5 seconds is uniformly extracted, and the frame features are extracted using a 2D CNN convolutional neural network pre-trained on ImageNet. In addition, a 3D CNN convolutional neural network can also be used for feature extraction, and the frame segments are processed as separate items.
[0049] Through the above steps, the initial feature sequence of the video is obtained, and then it is input into the preview branch and the intensive reading branch to capture more accurate and in-depth video information.
[0050] (2) The feature sequence obtained in step (1) is input into the preview branch for encoding to obtain the visual overview information in the video. The specific steps are as follows:
[0051] The video frame feature sequence obtained in step (1) is input into a bidirectional GRU (bi-GRU) network for encoding to extract the overview information of the video. The bidirectional GRU is composed of a forward GRU and a backward GRU. The hidden states at all specific time steps {t = 1,..., m} in the forward GRU and the backward GRU are concatenated as the output of the bidirectional GRU to obtain a feature vector sequence H = {h 1 , h 2 , …, h m}, with a size of m×1024 dimensions; then, an average pooling operation is applied to the feature vector sequence H along the time dimension to obtain a preview feature vector, that is
[0052]
[0053] In addition, the preview branch can be any lightweight model that models the video as a sequence of feature vectors.
[0054] (3) Input the video feature sequence obtained in step (1) into the intensive reading branch for encoding. Since a video usually contains multiple objects and complex scenes, the present invention introduces an intensive reading branch to encode multi-granularity feature information and extract deeper information under the guidance of the visual overview information learned by the preview branch. The specific steps are as follows:
[0055] (3-1) First, use a fully connected layer to map the visual feature sequence obtained in step (1) into a low-dimensional feature space for dimensionality reduction to reduce the computational complexity, obtaining a visual feature sequence V′ of 2048×m dimensions.
[0056] (3-2) Then, use a convolutional neural network CNN with a convolutional kernel size of n, a stride of s, and a number of convolutional kernels of r to extract segment features of different lengths. The specific formula can be expressed as:
[0057] C n =δ(Conv1D r,n,s (V′))
[0058] where δ represents the Relu activation function.
[0059] Put the segment features generated by convolutional kernels of different sizes together to obtain multi-granularity segment features, that is:
[0060]
[0061] where φ represents the size of the convolutional kernel, m n represents the number of segments of length n, r is the dimension of the segment feature vector, and after vectorizing the segment features, it is C′. To reduce the amount of computation, the visual feature sequence V′ is directly used as the segment feature of length 1.
[0062] (3-3) For the multi-granularity segment features obtained through step (3-1), in order to adaptively enhance the segment features more relevant to the video semantics, a perceptual preview attention operation is performed, as Figure 2 shown. The perceptual preview attention operation draws on the idea of the multi-head attention mechanism in Transformer. The specific steps are as follows:
[0063] First, map the preview feature vector p to a dk -dimensional query feature vector Q of the query, and maps the segment feature vector C′ into a d k -dimensional key feature vector K and a d v -dimensional value feature vector V respectively. Then, the query and the value are used to calculate the attention weights through dot product, and then the obtained attention weights are weighted and summed with the value feature vector to obtain an attention feature vector O, that is:
[0064] O = W 4 Attention(pW 1 , C′W 2 , C′W 3 )
[0065]
[0066] where W 1 , W 2 , W 3 and W 4 are learnable mapping matrix parameters.
[0067] Similar to the Transformer, the present invention also uses a residual operation to enhance the input to obtain an updated attention feature vector O′, that is:
[0068] O′ = LN(O + maxpool(C′))
[0069] where LN represents layer normalization and max pooling operations. Since the sizes of the segment feature and the attention feature vector do not match, a pooling operation is performed on the segment feature vector along the time dimension.
[0070] After obtaining the attention feature vector, a feed-forward network with residual and layer normalization is used to further enhance the above-mentioned feature vector, and the feed-forward network is implemented by using a multi-layer perceptron MLP, that is:
[0071] PaA(p, C′) = LN(o′ + MLP(o′))
[0072] where MLP is composed of two fully connected layers and a Relu activation function.
[0073] Finally, for the multi-granularity segment features, the above-mentioned perceptual preview attention operation is performed on each granularity in parallel, and the outputs of each granularity are concatenated as the final output of the intensive reading branch to obtain an intensive reading feature vector g, which is specifically expressed as follows:
[0074]
[0075] where ConCat represents the concatenation operation.
[0076] (4) Since the BERT model has made great progress in the field of natural language processing, the present invention uses a pre-trained BERT model to encode the text modality to obtain a multi-level encoded feature vector s of the text.
[0077] (5) Through the video and text encoding in step (3) and step (4), the preview feature vector p, the intensive reading feature vector g, and the text multi-level encoding feature vector s of the video are obtained, and then the preview feature vector p and the intensive reading feature vector g are respectively input into two different hybrid spaces for learning with the text multi-level encoding feature vector s, namely, the preview hybrid space and the intensive reading hybrid space.
[0078] In the hybrid space, a fully connected layer is used to map the video-text pairs into a concept space and a latent space, and cosine similarity is used to calculate the similarity of the video-text pairs. Since the concept space is expected to be used for interpretability and video-text matching, a binary cross-entropy loss and a marginal ranking loss are used for joint constraints.
[0079] For the latent space, in order to shorten the distance between related video-text pairs and push away the distance between unrelated video-text pairs, a ternary ranking loss marginal ranking loss is used to constrain the learning of the space.
[0080] (6) The above three losses are added together as the joint loss of the mixed space, that is:
[0081] L(v,s)=L l +L c +L ce
[0082] L(v,s) represents the joint loss of video v and text s in the mixed space;
[0083] L l Represents the triple ranking loss of the latent space;
[0084] L c A triple ranking loss representing the concept space;
[0085] L ce represents the binary cross entropy loss.
[0086] The combined loss of the above-mentioned hybrid space is used to optimize the preview hybrid space and the intensive reading hybrid space, and finally optimize by minimizing the sum of the losses of the preview hybrid space and the intensive reading hybrid space, that is:
[0087] L = L(p, s) + L(g, s)
[0088] Among them, L(p, s) represents the loss of the preview hybrid space, that is, the combined loss of the hybrid space calculated when the video is represented by the preview feature vector p; L(g, s) represents the loss of the intensive reading hybrid space, that is, the combined loss of the hybrid space calculated when the video is represented by the intensive reading feature vector g.
[0089] Use an Adam optimizer with a mini-batch size of 128, set the initial learning rate to 0.0001, and train the retrieval model established by the above steps (1) to (5). During the training process, adopt the early stop training strategy. Once the validation loss does not decrease in three consecutive epochs, divide the learning rate by 2. If the validation performance does not improve in 10 consecutive epochs, early stopping will occur.
[0090] Through the above steps, the retrieval model can capture the overview information of the video through the preview branch, and then further guide the intensive reading branch, so that the intensive reading branch is further strengthened, so that the fine-grained feature perception to the coarse-grained video overview information. Next, the corresponding video can be retrieved by matching the given text with all candidate videos, and the steps are as follows:
[0091] Map the text and the video into the preview hybrid space and the intensive reading hybrid space respectively, and in these two spaces, calculate the cosine similarity between the video-text pairs to rank the candidate videos, and use the better similarity result as the final return result to implement the cross-modal matching task through this process.
[0092] The above is only the preferred embodiment of the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention, or modify it into an equivalent embodiment with equivalent changes, without departing from the scope of the technical solution of the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still belong to the scope of protection of the technical solution of the present invention.
Claims
1. A cross-modal video retrieval method inspired by reading strategies, characterized in that, it includes the following steps: (1) Use a pre-trained convolutional neural network to extract the initial features of the video modality to obtain the initial feature sequence of the video; (2) Input the initial feature sequence and encode it through the preview branch to obtain the preview features in the video; (3) Input the initial feature sequence and encode it through the intensive reading branch to obtain multi-granularity segment features, then perceive and integrate the preview features to extract the intensive reading features; (4) Use the pre-trained BERT model to encode the text modality to obtain the text multi-level encoded features; (5) Map the visual preview features and intensive reading features to the corresponding hybrid spaces respectively with the text multi-level encoded features, and calculate the similarity between the video modality and the text modality through the hybrid space to perform cross-modal matching; (6) Optimize and train the retrieval model established through steps (1) to (5), and finally input the video and text into the trained retrieval model to achieve cross-modal retrieval from text to video.
2. A cross-modal video retrieval method inspired by reading strategies according to claim 1, characterized in that, The specific content of the step (2) is as follows: input the video frame feature sequence into the bidirectional GRU network of the preview branch. The bidirectional GRU consists of a forward GRU and a backward GRU. Concatenate the hidden states at all specific time steps {t = 1,..., m} in the forward GRU and the backward GRU as the output of the bidirectional GRU, and obtain a feature vector sequence H = {h 1 , h 2 ,..., h m}, with a size of m × 1024 dimensions; then apply average pooling operation to the feature vector sequence H along the time dimension to obtain the preview feature vector, that is 3. A cross-modal video retrieval method inspired by reading strategies according to claim 2, characterized in that, The step (3) specifically includes the following steps: (3-1) First, use the fully connected layer of the intensive reading branch to reduce the dimension of the visual feature sequence to obtain the reduced visual feature sequence V'; (3-2) Then input V' into a convolutional neural network CNN with a convolutional kernel size of n, a stride of s, and a number of convolutional kernels of r to extract segment features of different lengths. The specific formula is expressed as: C n = δ(Conv1D r,n,s (V')) where δ represents the Relu activation function; Put the segment features generated by convolutional kernels of different sizes together to obtain multi-granularity segment features, that is: where φ represents the size of the convolutional kernel, m n represents the number of segments of length n, r is the dimension of the segment feature vector, after vectorizing the segment features it is C′, and the visual feature sequence V′ is used as the segment feature of length 1; (3-3) Perform a perceptual preview attention operation on the multi-granularity segment features to obtain the intensive reading feature vector.
4. A cross-modal video retrieval method inspired by reading strategies according to claim 3, characterized in that, The step (3-3) is specifically: First, map the preview feature vector p to a query feature vector Q of dimension d k Map the segment feature vector C' to a key feature vector K of dimension d k and a value feature vector V of dimension d v respectively. Then, use the query and value to calculate the attention weights through dot product, and then perform a weighted sum of the obtained attention weights and the value feature vector to obtain an attention feature vector O, that is: O = W 4 Attention (pW 1 , C′W 2 , C′W 3 ) Among them, W 1 , W 2 , W 3 and W 4 are learnable mapping matrix parameters; Then use the residual operation to enhance the input to obtain the updated attention feature vector O', that is: O' = LN(O + maxpool(C')) where LN represents layer normalization and max pooling operation, and the pooling operation is performed on the segment feature vector along the time dimension; After obtaining the attention feature vector, use a feed-forward network with residual and layer normalization to further enhance the above feature vector, and implement the feed-forward network through a multi-layer perceptron MLP, that is: PaA(p, C') = LN(o' + MLP(o')) where MLP consists of two fully connected layers and the Relu activation function; Finally, for the multi-granularity segment features, perform the above perceptual preview attention operation on each granularity in parallel, and splice the outputs of each granularity together as the final output of the intensive reading branch to obtain the intensive reading feature vector g, which is specifically expressed as follows: where ConCat represents the splicing operation.
5. A cross-modal video retrieval method inspired by reading strategies according to claim 1, characterized in that, The mixed space includes a preview mixed space and a intensive reading mixed space. In the mixed space, a fully connected layer is used to map the video-text pair into a concept space and a latent space respectively, and the cosine similarity is used to calculate the similarity of the video-text pair.
6. A cross-modal video retrieval method inspired by reading strategies as claimed in claim 5, wherein, the concept space jointly constrains the learning of this space using a binary cross-entropy loss and a triplet ranking loss; the latent space constrains the learning of this space using a triplet ranking loss.
7. A cross-modal video retrieval method inspired by reading strategies as claimed in claim 6, wherein, step (6) is specifically as follows: adding the loss of the concept space and the loss of the latent space as the joint loss of the mixed space, optimizing the retrieval model by minimizing the sum of the losses of the preview mixed space and the intensive reading mixed space, training the retrieval model using the Adam optimizer, mapping the text and the video into the preview mixed space and the intensive reading mixed space respectively, and in these two spaces, sorting the candidate videos by calculating the cosine similarity between the video-text pairs, and taking the result with better similarity as the final returned result, and realizing the cross-modal matching task through this process.
8. A cross-modal video retrieval method inspired by reading strategies as claimed in claim 7, wherein, when training the retrieval model using the Adam optimizer, during the training process, an early stopping training strategy is adopted. If the validation loss does not decrease in three consecutive epochs, the learning rate is divided by 2. If the validation performance does not improve in 10 consecutive epochs, early stopping occurs.
Citation Information
Patent Citations
Text-to-video cross-modal retrieval method based on multistage coding
CN111309971A
Cross-modal video retrieval method and system based on multi-head self-attention mechanism and storage medium
CN112241468A