Semantic enhancement and multi-level alignment video text retrieval method
Through semantic enhancement and multi-level alignment methods, and by utilizing external knowledge and cross-modal information fusion modules, the problem of ignoring weak semantic descriptions in existing technologies is solved, precise matching and information interaction of video text retrieval are achieved, and the accuracy of retrieval is improved.
Patent Information
- Application Number
- CN202510848794.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-03
AI Technical Summary
Existing video text retrieval methods ignore the existence of weak semantic descriptions and their potential information, fail to deeply explore the information interaction between text details and complex visual semantics, and limit the retrieval performance of the model.
The method of semantic enhancement and multi-level alignment is adopted. The latent semantic information is mined through the external knowledge retrieval module. The cross-modal information fusion module is used for feature fusion. The inter-modal and intra-modal similarity loss functions are combined to achieve accurate matching between video and text by using global, action and entity alignment.
The precision and practicality of video text retrieval are improved. The information interaction between text details and complex visual semantics is enhanced through a multi-level alignment strategy, thereby improving the accuracy of retrieval.
Smart Images

Figure CN120744178A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data retrieval, and in particular to a semantically enhanced and multi-level aligned video text retrieval method. Background Art
[0002] The nonlinear learning capabilities of deep learning have driven the development of cross-modal video-text retrieval methods, primarily including global alignment and local alignment strategies. Global alignment directly calculates similarity by mapping the global features of video and text into a joint embedding space, while local alignment focuses on measuring the local similarity between each frame of the video and each word of the text. Both methods strive to address the semantic gap between different modalities. In recent years, CLIP has been widely used in this field due to its powerful visual and language feature extraction capabilities. For example, CLIP4Clip and X-Pool improve retrieval performance by fine-tuning pre-trained models and introducing cross-modal attention mechanisms, respectively. However, most of these methods assume a strong semantic relationship between text descriptions and videos, ignoring the existence of weak semantic descriptions and their potential information. They also fail to deeply explore the information interaction between text details and complex visual semantics, thereby limiting the model's retrieval performance. Therefore, it is necessary to develop a new method to overcome these limitations and improve the accuracy and practicality of retrieval. Summary of the Invention
[0003] The purpose of the present invention is to provide a video text retrieval method with semantic enhancement and multi-level alignment to address the deficiencies of the above-mentioned prior art and to solve the problems in the prior art.
[0004] The present invention specifically provides the following technical solution, a semantically enhanced and multi-level aligned video text retrieval method, comprising:
[0005] Get the original text-video data pair dataset;
[0006] Use the external knowledge retrieval module to retrieve external texts and videos that are similar to the original videos and texts;
[0007] Use the cross-modal information fusion module to fuse complementary information and extract enhanced feature representations of video and text;
[0008] Using inter-modal and intra-modal similarity loss functions to eliminate semantic gaps and achieve accurate retrieval;
[0009] The query text is decomposed and encoded by part of speech, and the video frames are encoded and clustered to obtain the global, action, and entity encoding features of the text and video respectively;
[0010] Similarity measurement between video and text is achieved using global alignment, action alignment and entity alignment.
[0011] Video encoders and text encoders can be used to encode external video sets and external text sets respectively, to mine the potential semantic information between video and text, and to enhance the corresponding feature representation and extraction.
[0012] Based on a cross-modal information fusion module, an adaptive embedding matrix is designed to give different weights to video frames and text words, and the text and data are fused using improved cross-attention. The output of the module serves as the final fused video and text features.
[0013] The symmetric cross entropy loss calculation formula can be used to maximize the similarity between matching video text pairs and perform inter-modal similarity constraints.
[0014] The Euclidean distance formula is used to calculate the data similarity in the video feature space, and the Tanimoto coefficient is used to calculate the data similarity in the text feature space. Finally, the contrast loss is used to calculate the same-modality similarity loss to optimize the model parameters and perform intra-modality similarity constraints.
[0015] The query text is decomposed and encoded according to parts of speech to obtain a global code containing overall information, a verb code with temporal sequence, and a noun code containing specific entity information. The video frames are encoded and clustered to obtain a global code describing the overall information of the video and sub-action codes of different sequence frames.
[0016] The query text is decomposed and encoded according to parts of speech to obtain a global code containing overall information, a verb code with temporal sequence, and a noun code containing specific entity information. The video frames are encoded and clustered to obtain a global code describing the overall information of the video and sub-action codes of different sequence frames.
[0017] The contrast loss is used for the similarity of the same modality. The specific expression is:
[0018]
[0019] Among them, J2 is the video modality similarity loss, B is the batch size, Q j is the feature vector of the jth video, V i is the feature vector of the i-th query text, j is the sample video index, i is the index of the query text, and τ is the temperature parameter;
[0020]
[0021] Among them, J3 is the similarity loss within the text modality, t i For i query texts, P j is the feature vector of all video samples, P i is the feature vector of the sample video corresponding to the i-th query text;
[0022] J = J1 + α (J2 + J3);
[0023] Among them, J1 measures the similarity between samples within the same modality, α is a hyperparameter, and J represents the final loss.
[0024] The text encoder and the video encoder are used to encode the query text and video data in the historical data set to obtain video block features and word features, specifically:
[0025] A visual encoder is used to encode the video, and video features are obtained by temporally aggregating video frame features;
[0026] The text encoder is used to encode the query text and obtain word features.
[0027] The clustering operation can use K-means clustering, the specific expression is:
[0028]
[0029] in, is the action feature of the lth cluster of the vector, W v is the learnable transformation matrix, b v is the bias term, is the feature vector of the original video frame of the lth cluster, and l is the cluster action index.
[0030] Compared with the prior art, the present invention has the following significant advantages:
[0031] This paper proposes a video text retrieval method with semantic enhancement and multi-level alignment. This method uses external knowledge to achieve semantic enhancement of video text, while using a multi-level alignment strategy to achieve information interaction between text details and complex visual semantics, thereby improving the accuracy of video text retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 Flowchart of a video text retrieval method with semantic enhancement and multi-level alignment according to the present invention
[0033] Figure 2 This is a model architecture diagram of a semantically enhanced and multi-level aligned video text retrieval method of the present invention;
[0034] Figure 3 It is the cross-modal information fusion module of the method proposed in the present invention;
[0035] Figure 4 This is an example display of the video retrieval text task of the present invention. DETAILED DESCRIPTION
[0036] The following is a clear and complete description of the technical solutions of the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0037] The overall structure of the model is as follows Figure 1 As shown, the video text retrieval method with semantic enhancement and multi-level alignment mainly includes two modules: a semantic enhancement module and a multi-level alignment module. The semantic enhancement module designs two external knowledge retrieval modules to mine the potential semantic information in the video and text, and constructs an adaptive cross-attention cross-modal information fusion module to perform feature fusion on complementary information to strengthen the feature representation of the video text pair. At the same time, it introduces inter-modal and intra-modal similarity loss functions to achieve consistency in cross-modal and intra-modal data representation, eliminate the semantic gap and achieve accurate retrieval; the multi-level alignment module decomposes and encodes the query text according to the part of speech to obtain a global encoding containing overall information, a verb encoding with temporal sequence and a noun encoding containing specific entity information; the video frame is encoded and a K-means clustering operation is performed to obtain a global encoding describing the overall information of the video and a sub-action encoding of different sequence frames. The video and text interact semantically from three levels: global, action and entity, so as to capture important correlation information, weaken redundant information and enhance the relevance of the video text. The following is a further detailed description through a specific implementation method: In this implementation method, a video text retrieval method with semantic enhancement and multi-level alignment includes the following steps:
[0038] Step S1: Obtain the original text-video data pair dataset.
[0039] This paper conducts experiments on four public datasets: MSR-VTT, MSVD, LSMDC, and DiDeMo. The MSR-VTT dataset contains more than 10,000 video clips of 10 to 30 seconds, each with approximately 20 natural language descriptions. The MSVD dataset contains approximately 1,970 videos, each with an average of approximately 40 text descriptions, for a total of approximately 85,550 descriptions. The LSMDC dataset contains 128,000 movie clips, each lasting between 2 and 30 seconds and described by text titles or subtitles extracted from the movies, and contains 118,081 video-text pairs. The DiDeMo dataset consists of 10,464 video clips, each with 3 to 5 text descriptions. All sentence descriptions of each video are concatenated as the final description of the entire video.
[0040] Furthermore, in the original text-video data pair dataset, each video corresponds to one or more text descriptions, which aim to capture the main content of the video, including but not limited to the scene, actions, and people involved.
[0041] Step S2: Use the external knowledge retrieval module to retrieve external texts and videos that are similar to the original videos and texts.
[0042] definition is an external knowledge source, which contains m video-text pairs. These video-text pairs come from videos on YouTube and their corresponding titles, covering a variety of fields. When performing information fusion, external knowledge information that is opposite to the input modality is selected. The subscript t or v is used to represent the modality retrieved from the external knowledge source, that is, P t ={T1,...,T j ,...,T k} represents the text set corresponding to the top k external videos most relevant to video v; Q v ={V1,...,V j ,...,V k} represents the video set corresponding to the top k external texts most relevant to text t.
[0043] The video encoder (ViT-B / 16) of CLIP is used to decode the external video set Q v ={V1,...,V j ,...,V k} to encode and obtain the encoding features of the external video set right Perform the averaging operation to obtain the final external video coding feature f Q .
[0044] The CLIP text encoder (ViT-B / 16) is used to encode the external text set P t ={T1,...,T j ,...,T k} to encode and obtain the encoding features of the external text set right Perform the averaging operation to obtain the final external text encoding feature f P .
[0045] Going a step further, based on the acquired dataset of original text-video data pairs, external texts and videos similar to existing video clips and their natural language descriptions can be found by querying search engines or using pre-trained models such as CLIP encoders. For example, the external video collection can be encoded using the CLIP video encoder, and the external text collection can be encoded using the CLIP text encoder to mine the potential semantic information between the video and text and strengthen the corresponding feature representation and extraction. In this process, for each original video and text pair, the closest match will be searched in a larger database, with the aim of discovering additional resources that can supplement and enhance the original content. The result of this is an expanded dataset that contains external texts and videos similar to the original videos and texts, providing richer materials for subsequent steps.
[0046] Step S3: Use the cross-modal information fusion module to perform feature fusion on the complementary information to extract the enhanced feature representation of the video and text.
[0047] The data to be retrieved is input into the final cross-modal retrieval model, and the retrieval results are obtained by sorting the cosine similarity of the output feature vectors. The steps include:
[0048] The data to be retrieved is input into the semantic enhancement module, and the model generates three feature spaces: a text feature space, which contains text encoding features and external text encoding features; a video feature space, which contains video encoding features and external video encoding features; and a fusion feature space, which contains video encoding features fused from video and external knowledge, and text encoding features fused from text and external knowledge.
[0049] The cosine similarity formula is used to calculate the similarity between and , the Euclidean distance is used to calculate the similarity between and , and the Tanimoto coefficient is used to calculate the similarity between and . The three similarities are added together to obtain the final similarity score of the semantic enhancement module.
[0050] The data to be retrieved is input into the multi-level alignment module to obtain the global encoding of text and video, text action encoding, video sub-action encoding, text first noun encoding and video key frame encoding in the video key frame segment.
[0051] The cosine similarity formula is used to calculate the global similarity of and , the action similarity of and , and the entity similarity of and . The above three similarities are added together to obtain the final similarity score of the multi-level alignment module.
[0052] Finally, the similarity scores of the semantic enhancement module and the multi-level alignment module are added together, and the similarity results are sorted to retrieve the video most relevant to the descriptive text or the text description that best matches the query video.
[0053] Cross-modal information fusion module such as Figure 2 As shown, the external text encoding feature f corresponding to the video v is P Perform global pooling operation to obtain the global encoding features of the text Combined with the video global encoding features The two methods interact with each other to dynamically assign different weights to video frames and remove redundant information in the video.
[0054]
[0055] De-redundant video coding features As a query, the external text encoding feature f P As the key and value, we get the video encoding feature F after interaction with external knowledge. v-an , use the residual structure to encode the interactive video features F v-an and video coding after redundancy removal Connect and enhance feature representation to obtain the final video coding feature F v .
[0056]
[0057] Furthermore, a global pooling operation is performed on the external text encoding features corresponding to the video to obtain the global encoding features of the text. This global encoding feature is interacted with the global encoding features of the video. Through this process, different weights are dynamically assigned to video frames, thereby removing redundant information in the video. Next, the video encoding features after de-redundancy are used as queries and the external text encoding features as keys and values. The video encoding features after interaction with external knowledge are obtained by calculating similarity. In order to enhance feature representation, the video encoding features after interaction are connected to the video encoding features after de-redundancy using a residual structure. Specifically, this connection method helps to retain the original features while introducing new information, thereby generating the final video encoding features. This process not only improves the expressiveness of video features, but also promotes more accurate matching between video and text.
[0058] Step S4: Use inter-modal and intra-modal similarity loss functions to eliminate the semantic gap and achieve accurate retrieval.
[0059] Following the common practice of video-text retrieval models, we perform bidirectional objective learning and define J1 as the inter-modal similarity loss, which is calculated using the symmetric cross-entropy loss formula. It aims to maximize the similarity between matching video-text pairs and minimize the similarity between other pairs.
[0060] For video data, because it is rotation invariant, the Euclidean distance is used to calculate the data similarity in the video feature space. The calculation formula is as follows:
[0061]
[0062] For text data, when semantically related text pairs differ significantly in length, the value obtained by using Euclidean distance to measure similarity is often too large. Therefore, this chapter uses the Tanimoto coefficient, which is unaffected by text length and better handles text similarity measurement issues, to calculate similarity in text feature space data.
[0063] Furthermore, an inter-modal similarity loss is defined, employing a symmetric cross-entropy loss formula to maximize the similarity between matching video-text pairs and minimize the similarity between non-matching pairs. This bidirectional objective learning approach helps ensure that videos and text are close to each other in the feature space, while mismatched samples are further away. For video data, due to its rotational invariance, Euclidean distance is used to calculate data similarity in the video feature space. For example, a smaller Euclidean distance between two video feature vectors indicates a higher similarity between the two videos in the feature space. However, for text data, when semantically related text pairs differ significantly in length, Euclidean distance may not be the optimal choice, as it may result in large numerical discrepancies. Therefore, the Tanimoto coefficient is used to calculate data similarity in the text feature space. This method is unaffected by text length and can more effectively handle the text similarity measurement problem. In this way, the similarity within video and text (intra-modality) and between video and text (inter-modality) can be simultaneously optimized, thereby improving the accuracy and efficiency of cross-modal retrieval.
[0064] Step S5: Decompose and encode the query text according to part of speech, and encode the video frames and perform Kmeas clustering operations to obtain the global, action, and entity encoding features of the text and video respectively.
[0065] In order to avoid the influence of the model on the retrieval performance caused by the copula being misclassified as verbs, the query text needs to be Preprocessing is done by first using the NLTK toolbox to remove stop words from the query text, and then decomposing the processed query text by part of speech to filter out verbs m t and noun sets Among them, m t Indicates the verb decomposed from t, represents the jth noun decomposed from t, and n is the number of nouns decomposed from the text.
[0066] After text decomposition, the text encoding module contains three inputs: query text t, decomposed verb mt and noun set n t To obtain the global features of the query text t, [CLS] and [EOS] are added to the start and end positions of t respectively, and t is encoded using the pre-trained CLIP text encoder. The CLIP text encoder is a Transformer structure consisting of 12 layers and 8 attention heads. The query, key, and value feature dimensions are all 512 dimensions. The encoded word token sequence is Among them, f t j is a 512-dimensional vector representing the j-th word t of t j The semantic features of f t CLS and f t EOS are the semantic features of [CLS] and [EOS] respectively. t Perform average pooling operation to obtain the global encoding f of t t all . t and n t Encode using CLIP text encoder, denoted as and in, and are all 512-dimensional vectors, and Represent the action characteristics of t and the jth noun respectively Entity features.
[0067] Furthermore, the query text is preprocessed using the NLTK toolkit to remove stop words. The processed text is then decomposed according to part of speech to filter out verb and noun sets. Here, represents the verbs extracted from the query text, and represents the total number of nouns extracted from the query text. After the text decomposition is complete, the text encoding phase begins, which includes three input components: the original query text, the decomposed verbs, and the noun set.
[0068] To capture the global features of the query text, the [CLS] and [EOS] tokens are added to the start and end of the original query text, respectively. The pre-trained CLIP text encoder is then used to encode the text with the tokens. The CLIP text encoder uses a Transformer architecture, comprising a 12-layer network and 8 attention heads. Its query, key, and value feature dimensions are all 512. After encoding, the output word token sequence is , where each represents the semantic features of the word at the corresponding position, while and represent the semantic features of the [CLS] and [EOS] tokens, respectively. Average pooling is then performed on to obtain the global encoding of the query text.
[0069] In addition, the verb and noun sets are encoded separately using the CLIP text encoder, and the results are recorded as and, where is a 512-dimensional vector representing the action features corresponding to the verb, and is a 512-dimensional vector representing the entity features corresponding to the word.
[0070] At the same time, the video frames are encoded and Kmeans clustering is performed to extract the video's global, action, and entity features. Through this series of operations, the global, action, and entity encoding features of the text and video are ultimately obtained, providing more fine-grained semantic support for subsequent cross-modal interactions.
[0071] Step S6: Use global alignment, action alignment, and entity alignment to measure the similarity between video and text. Text-video alignment is broken down into three levels: global alignment, action alignment, and entity alignment. Each level uses cosine similarity, which excels at processing high-dimensional data, to measure the similarity between video and text. The losses generated by these three alignments are then weighted and summed for retrieval.
[0072] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. For those skilled in the art to which the present invention belongs, several simple deductions or replacements can be made without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.
Claims
1. A video text retrieval method with semantic enhancement and multi-level alignment, characterized in that: The steps include: Get the original text-video data pair dataset; Use the external knowledge retrieval module to retrieve external texts and videos that are similar to the original videos and texts; Use the cross-modal information fusion module to fuse complementary information and extract enhanced feature representations of video and text; Using inter-modal and intra-modal similarity loss functions to eliminate semantic gaps and achieve accurate retrieval; The query text is decomposed and encoded by part of speech, and the video frames are encoded and clustered to obtain the global, action, and entity encoding features of the text and video respectively; Similarity measurement between video and text is achieved using global alignment, action alignment and entity alignment.
2. The video text retrieval method with semantic enhancement and multi-level alignment according to claim 1, characterized in that: Video encoders and text encoders can be used to encode external video sets and external text sets respectively, to mine the potential semantic information between video and text, and to enhance the corresponding feature representation and extraction.
3. The video text retrieval method with semantic enhancement and multi-level alignment according to claim 1, characterized in that: Based on a cross-modal information fusion module, an adaptive embedding matrix is designed to give different weights to video frames and text words, and the text and data are fused using improved cross-attention. The output of the module serves as the final fused video and text features.
4. The video text retrieval method with semantic enhancement and multi-level alignment as claimed in claim 3, characterized in that: The symmetric cross entropy loss calculation formula can be used to maximize the similarity between matching video text pairs and perform inter-modal similarity constraints.
5. The video text retrieval method with semantic enhancement and multi-level alignment according to claim 1, characterized in that: The Euclidean distance formula is used to calculate the data similarity in the video feature space, and the Tanimoto coefficient is used to calculate the data similarity in the text feature space. Finally, the contrast loss is used to calculate the same-modality similarity loss to optimize the model parameters and perform intra-modality similarity constraints.
6. The video text retrieval method with semantic enhancement and multi-level alignment according to claim 1, characterized in that: The query text is decomposed and encoded according to parts of speech to obtain a global code containing overall information, a verb code with temporal sequence, and a noun code containing specific entity information. The video frames are encoded and clustered to obtain a global code describing the overall information of the video and sub-action codes of different sequence frames.
7. The video text retrieval method with semantic enhancement and multi-level alignment according to claim 1, characterized in that: Text-video alignment can be decomposed into three levels: global alignment, action alignment, and entity alignment. The cosine similarity, which is good at processing high-dimensional data, can be used to measure the similarity between video and text. Then, the losses generated by the above three alignments are added according to the weights to obtain the retrieval results.
8. The video text retrieval method with semantic enhancement and multi-level alignment according to claim 4, characterized in that: The contrast loss is used for the similarity of the same modality. The specific expression is: Among them, J2 is the video modality similarity loss, B is the batch size, Q j is the feature vector of the jth video, V i is the feature vector of the i-th query text, j is the sample video index, i is the index of the query text, and τ is the temperature parameter; Among them, J3 is the similarity loss within the text modality, t i For i query texts, P j is the feature vector of all video samples, P i is the feature vector of the sample video corresponding to the i-th query text; J = J1 + α (J2 + J3); Among them, J1 measures the similarity between samples within the same modality, α is a hyperparameter, and J represents the final loss.
9. The video text retrieval method with semantic enhancement and multi-level alignment according to claim 1, characterized in that: The text encoder and the video encoder are used to encode the query text and video data in the historical data set to obtain video block features and word features, specifically: A visual encoder is used to encode the video, and video features are obtained by temporally aggregating video frame features; Use the text encoder to encode the query text and obtain word features X.
10. The video text retrieval method with semantic enhancement and multi-level alignment according to claim 1, characterized in that: The clustering operation can use K-means clustering, the specific expression is: in, is the action feature of the lth cluster of the vector, W v is the learnable transformation matrix, b v is the bias term, is the feature vector of the original video frame of the lth cluster, and l is the cluster action index.