Text-video retrieval-oriented hierarchical understanding method

By using a visual transformer and Transformer architecture model to perform hierarchical interaction between video frames and descriptions, the problem of gap between video domain and text description and insufficient information matching in text-video retrieval is solved, achieving high-precision video and text alignment and retrieval.

CN120929622APending Publication Date: 2025-11-11UNIV OF CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511027759.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

现有文本-视频检索方法在处理视频域的文本描述时存在显著差距,尤其在查询文本抽象性时,缺乏细粒度与全局信息的匹配,且未能充分利用辅助信息,导致匹配准确性较低。

Method used

A visual transformer is used to encode video frames, and video embeddings are extracted through a cross-frame attention mechanism. Combined with video descriptions, hierarchical interaction is performed to generate hybrid video embeddings. The Transformer architecture model is used for multi-stage matching, and video-text contrast loss and positive sample perceptual embedding are set for training to improve matching accuracy.

Benefits of technology

It significantly improves the retrieval performance of complex queries, enhances the alignment between videos and text, and improves matching accuracy and retrieval precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929622A_ABST
    Figure CN120929622A_ABST
Patent Text Reader

Abstract

The invention discloses a text-video retrieval-oriented hierarchical understanding method, which comprises the following steps of: coding a video frame to extract video frame features to obtain video embedding, and coding a query text to generate text embedding; setting video description; fusing the video embedding and the video description to generate mixed video embedding; and obtaining a query text, predicting whether the mixed video embedding is matched with the query text, and taking a video corresponding to a matching result as a query result. According to the method disclosed by the invention, the problem of alignment in the process of processing complex queries in the prior art is solved, the retrieval performance is remarkably improved, and the challenge of the complex queries is effectively coped with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a hierarchical understanding method for text-video retrieval, belonging to the field of artificial intelligence technology. Background Technology

[0002] Multimodal information matching is a challenge in text-to-video retrieval (TVR). Most existing techniques rely on pre-trained image-text models (such as CLIP) for subsequent cross-modal retrieval. Among these methods, CLIP4Clip, X-CLIP, and Cap4Video are typical examples. These methods utilize large-scale image-text pairs for pre-training and fine-tuning on video tasks. Cap4Video is a method that enhances video representation by generating video titles or descriptions. They have achieved some success in aligning text and video, but still have significant drawbacks:

[0003] (1) Differences between text and video description domains: Existing image-text retrieval models show significant gaps when processing text descriptions in the video domain, especially when the query text is abstract;

[0004] (2) Lack of matching fine-grained and global information: Existing methods cannot effectively handle the details and overall semantics in video segments, resulting in low matching accuracy;

[0005] (3) Lack of full utilization of auxiliary information: Existing methods fail to effectively utilize auxiliary information (such as video description and generated video description) to enhance the alignment between text and video.

[0006] Therefore, it is necessary to conduct more in-depth research on existing text-video retrieval methods to solve the above problems. Summary of the Invention

[0007] To overcome the above problems, in-depth research was conducted, and a hierarchical understanding method for text-video retrieval was proposed, including the following steps:

[0008] S1. Encode video frames to extract video frame features and obtain video embeddings; encode query text to generate text embeddings.

[0009] S2. Set video description;

[0010] S3. Merge the video embedding and video description to generate a hybrid video embedding;

[0011] S4. Obtain the query text, predict whether the mixed video embedding matches the query text, and take the video corresponding to the matching result as the query result.

[0012] In a preferred embodiment, in S1, the video encoder is performed using a visual transformer.

[0013] In a preferred embodiment, a cross-frame attention mechanism is set in the visual transformer.

[0014] In a preferred embodiment, in S2, the video is divided into multiple segments, each segment containing a series of adjacent frames, an independent video description is set for each segment, and an overall video description is set for the entire video.

[0015] In a preferred embodiment, in S3, fusion is performed through a hierarchical interaction module, which is a Transformer architecture model that enables the embedding of video frames and video descriptions to interact with each other through an attention mechanism.

[0016] In a preferred embodiment, in S4, the query text sentence is used as the "query" and the flattened hybrid video embedding vector is used as the "key" and "value" to predict whether the video-text pair is a positive match or a non-match through a linear layer.

[0017] In a preferred embodiment, a loss for hybrid video embeddings and query text is set during training, referred to as video-text contrast loss. This includes text-to-video contrast loss and video-to-text contrast loss. The text-to-video contrast loss is achieved by maximizing the similarity between the correct text and its corresponding video, while minimizing its similarity to other videos. The video-to-text contrast loss is achieved by maximizing the similarity between the correct video and its corresponding text, while minimizing its similarity to other texts. Here, the video refers to the hybrid video embedding, and the text refers to the query text.

[0018] In a preferred embodiment, S4 includes the following sub-steps:

[0019] S41. Obtain the similarity between each hybrid video embedding and the query text, and select the top K most similar candidate videos as the query results, wherein the hybrid video embedding is obtained based on the overall video description of the video;

[0020] S42. For any candidate video, obtain the similarity between the hybrid video embedding corresponding to each video description in the candidate video and the query text, combine the similarity of different video segments in the candidate video to obtain a comprehensive matching result, sort the candidate videos according to the comprehensive matching result and display them to the user.

[0021] The present invention also provides an electronic device, comprising:

[0022] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the methods described above.

[0023] The present invention also provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in any of the preceding claims.

[0024] The beneficial effects of this invention include:

[0025] (1) It solves the alignment problem of existing technologies when dealing with complex queries, significantly improves retrieval performance, and effectively addresses the challenges of complex queries;

[0026] (2) Improve matching accuracy: By segmented perception video description generation and hierarchical interaction, the matching effect of complex queries is significantly improved;

[0027] (3) By using the positive sample perception embedding training mechanism, the gap between the generated video description and the training query is narrowed, thereby improving the alignment effect between video and text. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of a hierarchical understanding method for text-video retrieval according to a preferred embodiment of the present invention.

[0029] Figure 2 This is a comparison of Example 1 and Comparative Example 1. Detailed Implementation

[0030] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Through these descriptions, the features and advantages of the present invention will become clearer and more apparent.

[0031] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments. Although various aspects of embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless specifically indicated otherwise.

[0032] According to the present invention, a hierarchical understanding method for text-video retrieval is provided, such as... Figure 1 As shown, it includes the following steps:

[0033] S1. Encode video frames to extract video frame features and obtain video embeddings; encode query text to generate text embeddings.

[0034] S2. Set video description;

[0035] S3. Merge the video embedding and video description to generate a hybrid video embedding;

[0036] S4. Obtain the query text, predict whether the mixed video embedding matches the query text, and take the video corresponding to the matching result as the query result.

[0037] In S1, the main function of encoding video frames is to decompose the input video into frames and extract features from each frame.

[0038] Preferably, the video encoder employs a visual transformer (ViT), which is a transformer specifically designed for computer vision and whose powerful representational capabilities are particularly suitable for capturing complex spatial and temporal features in video.

[0039] According to the present invention, each video frame generates an embedded representation by a visual transformer, which contains the visual features of the frame and is then used to align with query text.

[0040] Preferably, the video is first segmented into multiple frames. The visual transformer then segments the first frame into multiple fixed-size image patches and extracts features from each patch. In this way, ViT can capture important information in each video frame while maintaining efficient computation.

[0041] In a preferred embodiment, a cross-frame attention mechanism is incorporated into the visual transformer to enhance the temporal modeling capability of the video representation.

[0042] Specifically, the visual transformer has L layers, and the first K layers of the visual transformer generate frame-level embeddings.

[0043]

[0044] Where 1 represents the [CLS] flag, P is the number of patches per frame, d is the dimension of the latent variable layer, and F represents the number of video frames.

[0045] The CLS tags of video frames are processed through (LK) temporal layers after K layers of the visual transformer. In each layer, the [CLS] tags generate message embeddings.

[0046]

[0047] Self-attention (SA) is used to capture temporal dependencies:

[0048]

[0049] Where, mk This indicates a message marker, and SA indicates self-attention.

[0050] Message tag m k and frame-level embedding The video is fed back to the visual transformer, and after passing through (LK) time layers, the encoder generates the final video embedding.

[0051]

[0052] According to the present invention, a text encoder is used to encode the query text to generate a text embedding.

[0053] Preferably, the text encoder uses the BERT model, which encodes the input text through a self-attention mechanism, captures the contextual relationships in the sentence, and generates an embedding that can represent the semantic information of the entire query text.

[0054] BERT (Bidirectional Encoder Representations from Transformers), as a pre-trained language model, is widely used in semantic recognition. It can effectively capture semantic information in text and generate high-quality text embeddings. This text encoder transforms query text into vector-represented text embeddings. Furthermore, this encoder can also effectively align these text embeddings with the embeddings of video frames.

[0055] In S2, the video description is used to provide auxiliary information for the video to help better align the query text.

[0056] Preferably, the video is divided into multiple segments, each containing a series of adjacent frames, and an independent video description is set for each segment to capture fine-grained semantic information.

[0057] Preferably, an overall video description is also set for the entire video to provide a global semantic description.

[0058] In this invention, the method for obtaining video descriptions is not limited. Those skilled in the art can use any existing method based on experience, such as using open-source video description generation models like FreeVA.

[0059] According to the present invention, for each video segment, the generated video description not only provides specific details, but also works in conjunction with the global video description to provide richer contextual information.

[0060] In S3, fusion can enhance the richness and expressiveness of video descriptions, thereby improving the alignment accuracy between video and query text in subsequent processes.

[0061] Preferably, the fusion is achieved through a hierarchical interaction module, which is a Transformer architecture model. This module uses a multi-layered attention mechanism to enable the embeddings of video frames and video descriptions to interact with each other.

[0062] Specifically, the hierarchical interaction module processes video descriptions and video embeddings through a joint attention mechanism, generating a fused multimodal representation of the hybrid video embedding. This representation incorporates joint information from both video and text, enabling better matching of subsequent query text. Through the joint attention mechanism, the hierarchical interaction module effectively fuses video and video description information, enhancing the semantic representation of the video. This allows for high matching accuracy even in complex query scenarios, especially when dealing with abstract queries spanning multiple video segments.

[0063] In S4, a multi-stage matching model is used to perform matching video queries based on the query text.

[0064] The multi-stage matching model uses the query text sentence as the "query" and the flattened hybrid video embedding vector as the "key" and "value". Through a linear layer, it predicts whether the video-text pair is a positive match (positive example) or a non-match (negative example).

[0065] In a preferred embodiment, during training, a loss for the hybrid video embedding and query text needs to be set, called the video-text contrast loss. This includes text-to-video contrast loss and video-to-text contrast loss. The text-to-video contrast loss is achieved by maximizing the similarity between the correct text and its corresponding video, while minimizing its similarity to other videos. The video-to-text contrast loss is achieved by maximizing the similarity between the correct video and its corresponding text, while minimizing its similarity to other texts. Here, the video refers to the hybrid video embedding, and the text refers to the query text.

[0066] Furthermore, the video-text contrast loss function is expressed as:

[0067]

[0068] in, Indicates video-text contrast loss. This represents the contrast loss between text and video. Let B represent the video-to-text contrast loss, and let B represent the number of samples. Represents the i-th text embedding, e vi Let represent the i-th hybrid embedding, and τ represent the temperature coefficient.

[0069] In a preferred embodiment, during training, the text encoder is further configured with a video-text matching loss. For each pair of video-text samples, its positive matching score is first predicted; subsequently, negative samples are generated by randomly replacing the video or text, and negative matching scores are calculated for each. The loss function effectively distinguishes between positive and negative samples by minimizing the logarithm of the positive matching score and maximizing the logarithm of the negative matching score, and is expressed as:

[0070]

[0071] in, This represents the video-text matching loss. This represents a correctly matched video-text positive sample pair. Negative sample pairs representing video label mismatches. This represents a negative sample pair where the text labels do not match.

[0072] In a preferred embodiment, S4 includes the following sub-steps:

[0073] S41. Obtain the similarity between each hybrid video embedding and the query text, and select the top K most similar candidate videos as the query results.

[0074] Preferably, in step S41, the hybrid video embedding is obtained based on the overall video description of the video. This process can quickly filter out videos relevant to the query, thereby narrowing the search scope.

[0075] S42. For any candidate video, obtain the similarity between the hybrid video embedding corresponding to each video description in the candidate video and the query text. Combine the similarity of different video segments in the candidate video to obtain a comprehensive matching result. Sort the candidate videos with the comprehensive matching result and display them to the user so that the video most relevant to the query text is ranked first.

[0076] In a preferred embodiment, the training process further includes the step of generating a positive sample perceptual embedding based on the generated video description and text embedding, and generating a hybrid video embedding by replacing the original video description with the obtained positive sample perceptual embedding, so that the video description can be better aligned with the query text, thereby improving the accuracy of retrieval.

[0077] Traditional text-to-video retrieval methods typically use video descriptions directly as positive examples for training. However, there may be domain differences between the video descriptions and the original training query text, which can lead to performance degradation during the testing phase.

[0078] In this invention, positive sample-aware embedding is used as a supplementary adjustment to abstract queries, which refer to queries with a similarity score to the video that is lower than a certain threshold or significantly lower than the average similarity score.

[0079] Preferably, the positive sample-aware embedding is generated through one of the following methods:

[0080] Concatenation: The text embedding and video description are directly concatenated to generate a positive sample perceptual embedding. This method is simple and direct, preserving all the information from both.

[0081] Mean: Positive sample perceptual embeddings are generated by averaging the video description and text embeddings. This method provides a balanced representation of both.

[0082] Fermat points: Using video descriptions, text embeddings, and video embeddings as vertex vectors, the vectors at their Fermat points are obtained as positive sample perceptual embeddings. That is, the vector of the point with the smallest sum of distances to the video description, text embedding, and video embedding in the embedding space is the positive sample perceptual embedding. This method ensures that the generated positive sample perceptual embeddings are representative across modal spaces.

[0083] Query description cross-attention: This method uses text embeddings as query embedding vectors and video descriptions as keys and values, applying a cross-attention mechanism to generate positive sample perceptual embeddings. This approach selectively focuses on relevant descriptions, thereby improving query alignment.

[0084] In a preferred embodiment, to make the positive sample-aware embedding closer to the query embedding, the loss of the positive sample-aware embedding is set as follows:

[0085]

[0086]

[0087] in, This represents the loss for positive sample-aware embedding. This represents the comparison loss when a positive example is found. Let represent the contrastive loss describing the positive examples, B represent the training batch sample size, i represent the i-th sample in the training batch, and j represent all other samples in the batch except the i-th sample. Represents the query embedding vector, r pi Let r represent the positive sample-aware embedding generated for the i-th sample. pj This represents the positive sample perceptual embedding generated for the j-th sample, where τ represents the temperature coefficient. This represents the text embedding of the description of the i-th video.

[0088] Among them, the contrast loss for finding positive examples. The method maximizes the similarity between the query embedding and the positive example-aware description embedding, while minimizing the similarity with other descriptions; the contrastive loss from description to positive examples. It emphasizes the alignment of the description embedding with the positive example-aware description embedding, while distinguishing it from other descriptions.

[0089] Various embodiments of the methods described above in this invention can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0090] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0091] Example

[0092] Example 1

[0093] Text-video retrieval experiments were conducted on the publicly available MSVD dataset, which contains approximately 120,000 captions describing 1,970 videos ranging in length from 1 second to 62 seconds. The dataset was divided into training, validation, and test sets, containing 1,200, 100, and 670 videos, respectively.

[0094] The text-video retrieval experiment includes the following steps:

[0095] S1. Encode video frames to extract video frame features and obtain video embeddings; encode query text to generate text embeddings.

[0096] S2. Set video description;

[0097] S3. Merge the video embedding and video description to generate a hybrid video embedding;

[0098] S4. Obtain the query text, predict whether the mixed video embedding matches the query text, and take the video corresponding to the matching result as the query result.

[0099] In S1, the video encoder uses a visual transformer, and the text encoder uses the BERT model. The text encoder has a video-text matching loss.

[0100]

[0101] In S2, the video is divided into multiple segments, each containing a series of adjacent frames. Each segment is given an independent video description, and an overall video description is also given for the entire video, providing a global semantic description. The video description generation model uses FreeVA.

[0102] The training process also includes the following steps: generating positive sample perceptual embeddings based on the generated video descriptions and text embeddings, and then using the obtained positive sample perceptual embeddings to replace the original video descriptions in the generation of hybrid video embeddings, so that the video descriptions can be better aligned with the query text, thereby improving the accuracy of retrieval.

[0103] In S3, fusion is achieved through a hierarchical interaction module, which is a Transformer architecture model.

[0104] In S4, the query text sentence is used as the "query," and the flattened hybrid video embedding vector is used as the "key" and "value." A linear layer predicts whether the video-text pair is a positive match or a non-match. During training, the loss function is set as follows:

[0105]

[0106] Comparative Example

[0107] Comparative Example 1

[0108] Text-video retrieval experiments were conducted using the same dataset as in Example 1. The difference was that the CE, SUPPORT, CLIP4Clip, X-Pool, X-CLIP, Unmask, and Cap4Video methods were used on the MSVD dataset, respectively.

[0109] For the CE method, see the literature Yang Liu, Samuel Albanie, Arsha Nagrani, and Andrew Zisserman. Use what you have: Video retrieval using representations from collaborative experts. arXiv preprint arXiv:1907.13487,2019.2,7;

[0110] The SUPPORT method can be found in the literature Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander Hauptmann, Joao Henriques, and Andrea Vedaldi. Support-set bottlenecks for video-text representation learning. arXiv preprint arXiv:2010.02824, 2020.7;

[0111] The CLIP4Clip and X-CLIP methods can be found in the literature Hongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu, Ruihua Song, Houqiang Li, and Jiebo Luo. Clip-vip: Adapting pretrained image-text model to video-language representation alignment. arXiv preprint arXiv:2209.06430, 2022.2, 3, 6, 7;

[0112] The X-Pool method can be found in the literature Satya Krishna Gorti, No¨el Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, and Guangwei Yu. X-pool: Cross-modal language-video attention for text-video retrieval. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pages 5006–5015, 2022.1, 2, 7, 8;

[0113] The Unmask method can be found in the literature Kunchang Li, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wang, and Yu Qiao. Unmasked teacher: Towards training-efficient video foundation models, 2023.2, 6, 7;

[0114] For the Cap4Video method, please refer to the literature Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang, and Wanli Ouyang. Cap4video: What can auxiliary captions do for text-video retrieval? In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 10704–10713, 2023.2, 5, 6, 7.

[0115] Comparing the results of Example 1 with those of Comparative Example 1, such as Figure 2 As shown, R@1, R@5, R@10, MdR, and MnR are commonly used evaluation metrics for text-video retrieval.

[0116] from Figure 2 As can be seen from the results, the method in Example 1 has better performance in all indicators, with R@1, R@10 and MdR achieving the best results.

[0117] The present invention has been described above with reference to preferred embodiments; however, these embodiments are merely exemplary and illustrative. Various substitutions and modifications can be made to the present invention based on these embodiments, all of which fall within the scope of protection of the present invention.

Claims

1. A hierarchical understanding method for text-video retrieval, characterized in that, Includes the following steps: S1. Encode video frames to extract video frame features and obtain video embeddings; encode query text to generate text embeddings. S2. Set video description; S3. Merge the video embedding and video description to generate a hybrid video embedding; S4. Obtain the query text, predict whether the mixed video embedding matches the query text, and take the video corresponding to the matching result as the query result.

2. The hierarchical understanding method for text-video retrieval according to claim 1, characterized in that, In S1, the video encoder uses a visual transformer.

3. The hierarchical understanding method for text-video retrieval according to claim 2, characterized in that, Set up a cross-frame attention mechanism in the visual transformer.

4. The hierarchical understanding method for text-video retrieval according to claim 1, characterized in that, In S2, the video is divided into multiple segments, each containing a series of adjacent frames. Each segment is given an independent video description, and the entire video is also given an overall video description.

5. The hierarchical understanding method for text-video retrieval according to claim 1, characterized in that, In S3, fusion is achieved through a hierarchical interaction module, which is a Transformer architecture model. It enables the embedding of video frames and video descriptions to interact with each other through an attention mechanism.

6. The hierarchical understanding method for text-video retrieval according to claim 1, characterized in that, In S4, the query text sentence is used as the "query" and the flattened hybrid video embedding vector is used as the "key" and "value". A linear layer is used to predict whether the video-text pair is a positive match or a non-match.

7. The hierarchical understanding method for text-video retrieval according to claim 1, characterized in that, During training, a loss is set for the hybrid video embedding and query text, called the video-text contrast loss. This includes text-to-video contrast loss and video-to-text contrast loss. The text-to-video contrast loss is achieved by maximizing the similarity between the correct text and the corresponding video while minimizing its similarity to other videos. The video-to-text contrast loss is achieved by maximizing the similarity between the correct video and the corresponding text while minimizing its similarity to other texts. Here, the video refers to the hybrid video embedding, and the text refers to the query text.

8. The hierarchical understanding method for text-video retrieval according to claim 1, characterized in that, S4 includes the following sub-steps: S41. Obtain the similarity between each hybrid video embedding and the query text, and select the top K most similar candidate videos as the query results, wherein the hybrid video embedding is obtained based on the overall video description of the video; S42. For any candidate video, obtain the similarity between the hybrid video embedding corresponding to each video description in the candidate video and the query text, combine the similarity of different video segments in the candidate video to obtain a comprehensive matching result, sort the candidate videos according to the comprehensive matching result and display them to the user.

9. An electronic device, comprising: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

10. A computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.