A Video-Text Retrieval Method Based on Multimodal Large Models

By generating supplementary text and enhancing videos for videos and using multimodal big models for information alignment, the problem of limited text semantic space and cross-modal alignment in the video-text search method in the prior art is solved, and a more efficient and accurate video-text search effect is achieved.

CN119336923BActive Publication Date: 2025-06-13GUANGDONG UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411271756.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2025-06-13
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

In the prior art, the video-text retrieval method ignores intramodal alignment due to the limited text semantic space and excessive focus on cross-modal alignment, resulting in poor retrieval effect.

Method used

The video-text search method based on multimodal large model is adopted. By generating supplementary text and enhanced videos for the original video, the fused text features and video global features are calculated, and cross-attention encoding and contrast learning constraints are used to promote information alignment and improve retrieval effect.

Benefits of technology

By enriching the semantic expression of text modality and enhancing the expression ability of video modality, the accuracy and generalization ability of video-text retrieval are improved, and the dependence on manual labeled data is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119336923B_ABST
    Figure CN119336923B_ABST
Patent Text Reader

Abstract

The present invention discloses a video-text retrieval method based on a multimodal large model. The method includes the following steps: obtaining a dataset of original text-original video pairs; generating supplementary text and enhanced videos for the original videos by using a preset method; calculating the fused text features of the original text and the supplementary text, and respectively extracting the video global features of the original videos and the enhanced videos by using cross-attention encoding; calculating a contrastive learning constraint by using the text features and the video global features; promoting the model to achieve information alignment by using the contrastive learning constraint; generating and saving the feature information of the video to be retrieved by using the trained model, and obtaining the retrieval result by calculating the similarity between the features of the input video or text and the saved feature information. Through data augmentation based on a multimodal large model and quadruple video-text contrastive learning, the present invention fully aligns videos and texts, and obtains high recall accuracy for video-text retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of video retrieval technology, and more specifically, relates to a video-text retrieval method based on a multimodal large model. Background Art

[0002] Currently, video-text retrieval is mainly based on the image-text contrast learning framework and its derivative frameworks. Directly applying such a framework to the video-text field does not take into account the large information deviation between the video modality and the text modality, which is mainly manifested in that the text often only describes a part of the content in the video, or the text is a highly abstract summary of the video content. To make up for the information deviation, previous methods often rely on a large amount of data for secondary training, which requires a lot of time for manual annotation and additional training costs.

[0003] The prior invention with the publication number CN114817627A discloses a cross-modal retrieval method from text to video based on multi-faceted video representation learning, including: obtaining preliminary features of video and text; after using a video shot segmentation tool to group the initial frames of the video according to different scenes and inputting them into the explicit encoding branch for explicit encoding, obtaining explicit multi-faceted representations of different scenes of the video; inputting the initial video features into the implicit encoding branch, and performing implicit encoding on the initial video features through the leading feature multi-attention network to obtain implicit multi-faceted representations expressing different semantic contents of the video; fusing the multi-faceted encodings of the two branches to obtain a multi-faceted video feature representation; mapping the multi-faceted video feature representation and the text features into a common space respectively, and using the common space learning algorithm to learn the correlation between the two modalities, and training the model in an end-to-end manner to achieve cross-modal retrieval from text to video. This solution only utilizes the original text, resulting in a relatively limited semantic expression space, prone to semantic differences, and making the text unable to comprehensively cover the video information. In addition, the solution overly focuses on cross-modal alignment while ignoring intra-modal alignment. Summary of the Invention

[0004] The present invention aims to solve the problems in the prior art of limited text semantic space, over-focus on cross-modal alignment while ignoring intra-modal alignment, and provides a video-text retrieval method based on a multimodal large model. The method includes the following steps:

[0005] Obtain a dataset of original text-original video pairs;

[0006] Generate supplementary text and enhanced video for the original video using a preset method;

[0007] Calculate the fused text features of the original text and the supplementary text, and respectively extract the video global features of the original video and the enhanced video using cross-attention encoding;

[0008] Calculate the contrastive learning constraint using the text features and video global features;

[0009] Use the contrastive learning constraint to promote the model to achieve information alignment;

[0010] Use the trained model to generate the feature information of the video to be retrieved and save it. During retrieval, calculate the similarity between the features of the input video or text and the saved feature information to obtain the retrieval result.

[0011] Furthermore, the method for generating supplementary text is to use a multi-modal large language model to generate a detailed text description for the original video.

[0012] Furthermore, the method for generating an enhanced video includes the following steps: performing frame sampling on the original video; applying image enhancement techniques to the sampled video frames; and recombining the processed frame sequences in sequence to generate an enhanced video.

[0013] Furthermore, the image enhancement technique includes at least one of the following: rotating, flipping, color transformation, spatial transformation, style transformation, and geometric transformation of the original video.

[0014] Furthermore, the specific method for calculating the fused text features is as follows:

[0015] Use a tokenizer to separately segment the original text and the supplementary text into discrete words;

[0016] Map the discrete words to continuous low-dimensional vectors through an embedding layer;

[0017] Input the low-dimensional vectors into a deep learning model to extract the original text features and supplementary text features;

[0018] Obtain the fused text features by summing the original text features and the supplementary text features.

[0019] Furthermore, the specific method for extracting video global features is as follows:

[0020] Sample video frames from the video;

[0021] Divide each video frame into non-overlapping image patches;

[0022] Perform convolution operations on each image patch to generate low-dimensional vectors, and add position encoding and class labels;

[0023] Input the low-dimensional vectors with position encoding and class labels into a video deep learning model, and use the obtained class labels as the feature vectors for each video frame;

[0024] The global features of the video are calculated by using cross-attention encoding for the fused text features and the feature vectors of each video frame.

[0025] Furthermore, the contrastive learning constraint includes the contrastive learning constraint between cross-modal text and video, and the global feature constraint between the original video and the enhanced video within the modality.

[0026] Furthermore, the contrastive learning constraint between the text and the video includes the original text-original video contrastive learning constraint and the original text-enhanced video contrastive learning constraint.

[0027] Furthermore, the contrastive learning constraint between the text and the video has the following calculation formula:

[0028]

[0029] where represents calculating the similarity between features, exp() represents the exponential function, is a learnable temperature coefficient, represents the fused text features of video sample i in the batch training videos, and represent the video global features of two different video samples i and j.

[0030] Furthermore, the contrastive learning constraint between the original video and the enhanced video has the following calculation formula:

[0031]

[0032] where is a learnable temperature coefficient, and respectively represent the global features of the original video and the enhanced video of video sample i in the batch training videos, represents the global feature of the enhanced video of video sample j in the batch training videos.

[0033] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:

[0034] The present invention utilizes the multimodal capabilities of the multimodal large language model to generate supplementary text with rich details for each video, and then fuses the supplementary text and the original text features in an additive manner, enriching the semantic information expressed in the text modality. This not only reduces the dependence on manually labeled data but also expands the semantic space of text expression. By generating an enhanced video for each video and using the self-supervised contrastive learning method, the expression ability of the video modality is further enhanced.

[0035] By constructing a "quadruple" consisting of the original text, supplementary text, original video, and enhanced video, and through cross-attention encoding, video features are extracted based on the association between the text and the video, promoting effective and meaningful information interaction and fusion between the video and the text, enhancing the representational ability of the model, and enabling the generated representation to more comprehensively and deeply reflect the internal connections of multimodal data.

[0036] Design two types of analogy learning constraints: inter-modal contrastive learning constraint and intra-modal contrastive learning constraint. The inter-modal contrastive learning constraint includes original text-original video and original text-enhanced video contrastive learning, which promotes effective alignment of modalities while taking into account the generalization ability of the model; intra-modal contrastive learning, through self-supervised contrastive learning between the original video and the enhanced video, forces the model to learn the internal structure and meaningful feature representations of the data, further improving the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] To make the objectives and technical solutions of the present invention clearer, the present invention provides the following drawings and descriptions:

[0038] Figure 1 It is a flowchart of the method provided in an embodiment of the present invention;

[0039] Figure 2 It is a schematic diagram of functional modules provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] In order to more clearly understand the above objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.

[0041] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0042] The present invention provides a video-text retrieval method based on a multimodal large model, as Figure 1 shown is a flowchart of a video-text retrieval method based on a multimodal large model, Figure 2 shown is a schematic diagram of functional modules of a video-text retrieval method based on a multimodal large model provided by the present invention. The specific steps are as follows:

[0043] S1: Obtain a dataset of original text-original video pairs;

[0044] In a specific embodiment, an original text - original video pair dataset is obtained from the MSRVTT dataset. This dataset contains 10,000 videos, and each video contains 20 manually annotated captions.

[0045] S2: Use a preset method to generate supplementary text with rich details and enhanced videos for the original videos;

[0046] More specifically, the method for generating supplementary text in step S2 is to use a multimodal large language model to generate detailed text descriptions for the original videos. The method for generating enhanced videos includes the following steps: frame sampling of the original videos; applying image enhancement techniques to the sampled video frames; recombining the processed frame sequences in order to generate enhanced videos. The image enhancement techniques include at least one of the following: rotating the original video, flipping it, color transformation, spatial transformation, style transformation, geometric transformation.

[0047] In a specific embodiment, samples are extracted from the dataset as the original text and the original video respectively. Taking one sample as an example, the original text corresponding to the original video is "People are singing on the beach". The original video is input into the multimodal large language model VideoLlava, and the text prompt is used: "Please generate a new caption for the video that not only rephrases the scene but also adds layers of interpretation or context that might not be explicit in the original caption. Think about elements that could be present in the scene in the video". Through this prompt, the model is guided to generate supplementary text with rich details: "As the sun sets on the beach, a group of people gather to sing......The group is a mix of individuals,......".

[0048] At the same time, operations such as rotating and flipping the original video frames are performed to obtain enhanced videos. Finally, a quadruple for model input is constructed: original text, supplementary text, original video, enhanced video.

[0049] S3: Calculate the fused text features of the original text and the supplementary text, and use cross-attention encoding to extract the global video features of the original video and the enhanced video respectively;

[0050] More specifically, the specific method for calculating the fused text features in step S3 is as follows:

[0051] Use a tokenizer to split the original text and the supplementary text into discrete words respectively;

[0052] Map the discrete words to continuous low-dimensional vectors through an embedding layer;

[0053] Input the low-dimensional vectors into a Transformer model to extract the original text features and the supplementary text features;

[0054] Obtain the fused text features by summing the original text features and the supplementary text features.

[0055] The specific method for extracting the global video features in step S3 is as follows:

[0056] Sample video frames from the video;

[0057] Divide each video frame into non-overlapping image patches;

[0058] Perform convolution operations on each image patch to generate low-dimensional vectors, and add position encoding and class labels;

[0059] Input the low-dimensional vectors with position encoding and class labels into a video deep learning model, and use the obtained class labels as the feature vectors of each video frame;

[0060] Calculate the global features of the video by using cross-attention encoding for the fused text features and the feature vectors of each video frame.

[0061] In a specific embodiment, for the original text and the supplementary text , first use a tokenizer to split the text into a set of discrete words ["People", "are", "singing", "on", "the", "beach"], then map the set of words to a set of numbers [123, 456, 987, 565, 158, 681] through a predefined vocabulary, and then map the discrete information to continuous low-dimensional vectors through a feed-forward neural network embedding layer , where bs represents the batch size, set to 64, seq represents the sequence length, i.e., the number of tokens, set to 32, and dim represents the feature dimension, set to 512. This vector is only the vector representation of the text and does not include feature information. After adding positional encoding and class labels to the obtained vector, it is input into the Transformer model to extract the original text features and supplementary text features , and the fused text feature is obtained by summing the text features .

[0062] For the original video or enhanced video, first sample video frames from the video and split each video frame into non-overlapping image patches. In this embodiment, the resolution of the video is 224*224, and the size of the image patch is 32*32. Therefore, each video frame will be split into 7*7 non-overlapping image patches. These image patches are converted into low-dimensional vectors through convolutional operations, and positional encoding is added to each image patch. In addition, a CLS token class label is added to each video frame. The positional encoding is randomly initialized by a normal distribution at the start of training and remains fixed throughout the training process, while the CLS token is randomly initialized by a normal distribution each time it is processed. Then, these vectors are input into the Vision Transformer model, and finally, the output of the CLS token is used as the feature vector of each video frame .

[0063] Apply the cross-attention encoding method, using the fused text feature as the query , the video frame feature as the value , after matrix transposition as the key . Through the cross-attention mechanism, the similarity of each video frame is dynamically calculated and used as the dynamic weight. Based on the dynamic weight, the video frame features are fused to obtain the global feature of the video. In this embodiment, 12 frames are sampled for each video. The fused text feature is calculated for similarity with 12 frame features, and a set of 12 similarity weights is obtained, representing the similarity between the fused text feature and each video frame. Multiply the 12 similarities with the 12 frame features correspondingly and then add them up to obtain the global feature of the video. The calculation formula is:

[0064]

[0065] where, represents the th sampled frame, is the The video frame features of a sampling frame is the similarity between the video frame features of the nth sampling frame and the fused text features. Calculated according to this method, the global features of the original video can be calculated and the global features of the enhanced video .

[0066] S4: Using the text features and the global video features, calculate the contrastive learning constraints;

[0067] More specifically, the contrastive learning constraints in step S4 include cross-modal original text-original video contrastive learning constraints and original text-enhanced video contrastive learning constraints, as well as global feature constraints between the original video and the enhanced video within the modality.

[0068] For the contrastive learning constraints between the original text and the original video across modalities , the calculation formula is:

[0069]

[0070] For the contrastive learning constraints between the original text and the enhanced video across modalities , the calculation formula is:

[0071]

[0072] For the contrastive learning constraints between the original video and the enhanced video within the modality , the calculation formula is:

[0073]

[0074] Among them, represents calculating the similarity between features, exp() represents the exponential function, and are learnable temperature coefficients, represents the fused text features of video sample i in the batch training video, and respectively represent the global features of the original video of video samples i and j in the batch training video, respectively represent the global features of the enhanced video of video samples i and j in the batch training video.

[0075] S5: Using the contrastive learning constraints to promote the model to achieve information alignment;

[0076] In contrastive learning, the optimization objective usually aims to minimize the constraints. Specifically, when using three contrastive learning constraint formulas to promote the model to achieve information alignment, a key point is to balance the relationship between the numerator and the denominator. The denominator measures the dissimilarity between samples, that is, the similarity of negative sample pairs. Ideally, it is expected that the smaller the denominator, the higher the discrimination between negative sample pairs; while the numerator represents the similarity between the target sample and the positive sample pair, and it is expected that the larger the numerator, the higher the matching degree between positive sample pairs. Therefore, the goal is to adjust the model to increase the numerator and decrease the denominator to optimize the contrastive learning effect.

[0077] S6: Use the trained model to generate the feature information of the video to be retrieved and save it. During retrieval, calculate the similarity between the features of the input video or text and the saved feature information to obtain the retrieval result.

[0078] The present invention shows higher recall accuracy than other methods in existing video question answering datasets such as MSRVTT and MSVD. The following experimental model code is implemented based on Pytorch and trained on a computer with 8 NVIDIA Tesla A100 GPUs. The data batch size is 64, the maximum number of sampled frames for videos is 10, and the maximum text length limit is 70. The weights of the text encoder and video encoder are initialized with the weights of the pre-trained CLIP model ViT-B / 16. In addition, the Adamax optimizer is used with a learning rate of 4e-6. The experimental results on the MSRVTT and MSVD datasets are shown in Table 1 and Table 2. Table 1 shows the test results on the MSRVTT data, and Table 2 shows the test results on the MSVD data. In the tables, CLIP2TV, Clip4Clip, X-CLIP, and CLIP-VIP are existing algorithm models. The additional data column indicates whether additional data volume is required for training in the model training stage. R@1, R@5, and R@10 are recall rates, representing the accuracy of the retrieval results among the top 1, 5, and 10 results. These metrics directly reflect the accuracy of the retrieval system, and the higher the value, the better the retrieval effect. The MeanR column represents the average recall rank, which measures the average position of the correct answer in the returned results. The smaller the value, the more forward the correct answer is. For data points that cannot be effectively measured, '-' is used for marking.

[0079] Table 1

[0080]

[0081] Table 2

[0082]

[0083] The present invention demonstrates excellent performance in the field of video retrieval, especially in the two-way tasks of retrieving videos from text and retrieving text from videos, where its performance is particularly prominent. As can be seen from the tabular data, the present invention has reached a very high level in multiple evaluation metrics such as R@1, R@5, and R@10, far exceeding other comparative models. At the same time, from the perspective of the MeanR average recall ranking metric, the value of the present invention is lower, which means that the average ranking of the correct answers is more forward, further verifying its excellent retrieval performance. It is worth mentioning that without using additional manually annotated data, the performance of the present invention brings a 15.4% accuracy improvement in the R@1 metric of the MSRVTT data compared to the CLIP-VIP model trained with an additional 10 million video-text pairs of data, proving that using a multimodal large language model to generate additional supplementary text effectively reflects video details and expands the expression space of the text. In addition, the results on the MSVD dataset show that compared with the Clip4Clip method that uses mean aggregation to obtain video global features, the method of combining quadruple contrast learning framework with cross-attention encoding brings a 22.1% R@1 accuracy improvement in performance, indicating that extracting video features according to the association between text and video effectively promotes the effective and meaningful information interaction and fusion between video and text, enhances the representation ability of the model, and enables the generated representation to more comprehensively and deeply reflect the internal connection of multimodal data.

[0084] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A video-text retrieval method based on a multimodal large model, characterized in that: The steps include: Get the original text-original video pair dataset; Use a multimodal large language model to generate detailed text descriptions for the original video as supplementary text; use a preset method to generate enhanced videos for the original video; Calculate the fused text features of the original text and the supplemented text, and use cross-attention encoding to extract the global video features of the original video and the enhanced video respectively; Utilizing the text features and the video global features, calculating contrastive learning constraints, wherein the contrastive learning constraints include contrastive learning constraints between text and video across modalities, and global feature constraints between original video and enhanced video within a modality; Using the contrastive learning constraint to promote the model to achieve information alignment; The trained model is used to generate feature information of the video to be retrieved and saved. During retrieval, the similarity between the features of the input video or text and the saved feature information is calculated to obtain the retrieval results.

2. A video-text retrieval method based on a multimodal large model according to claim 1, characterized in that: The method for generating enhanced video comprises the following steps: sampling frames of original video; applying image enhancement technology to the sampled video frames; and recombining the processed frame sequences in order to generate enhanced video.

3. The video-text retrieval method based on a multimodal large model according to claim 2, characterized in that: The image enhancement technology includes at least one of the following: rotation, flipping, color transformation, space transformation, style transformation, and geometric transformation of the original video.

4. The video-text retrieval method based on a multimodal large model according to claim 1, characterized in that: The specific method for calculating the fusion text features is: Use a word segmenter to segment the original text and the supplementary text into discrete words respectively; Mapping discrete words into continuous low-dimensional vectors through an embedding layer; Inputting the low-dimensional vector into a deep learning model to extract original text features and supplementary text features; The fused text features are obtained by summing the original text features and the supplementary text features.

5. The video-text retrieval method based on a multimodal large model according to claim 1, characterized in that: The specific method for extracting the global features of the video is: Sampling video frames from the original video; Split each video frame into non-overlapping image blocks; Perform convolution operations on each image block to generate a low-dimensional vector and add position encoding and category labels; The low-dimensional vector with position encoding and category label is input into the video deep learning model, and the obtained category label is used as the feature vector of each video frame; The global features of the video are calculated by fusion text features and feature vectors of each video frame using cross attention encoding.

6. The video-text retrieval method based on a multimodal large model according to claim 1, characterized in that: The contrast learning constraints between text and video include original text-original video contrast learning constraints and original text-enhanced video contrast learning constraints.

7. The video-text retrieval method based on a multimodal large model according to claim 1, characterized in that: The contrastive learning constraint L between text and video tv The calculation formula is: Among them, sim() represents the similarity between calculated features, exp() represents the exponential function, τ tv is the learnable temperature coefficient, represents the fused text features of video sample i in the batch training video, v i and v j Represents the global video features of two different video samples i and j.

8. The video-text retrieval method based on a multimodal large model according to claim 1, characterized in that: The contrast learning constraint L between the original video and the enhanced video va The calculation formula is: Among them, τ va is the learnable temperature coefficient, and They represent the global features of the original video and the global features of the enhanced video of video sample i in the batch training video, Represents the global features of the enhanced video of video sample j in the batch training video.

Citation Information

Patent Citations

  • Text-to-video cross-modal retrieval method based on multi-face video representation learning

    CN114817627A

  • Video text cross-modal retrieval method based on double contrast learning and related equipment

    CN117235305A

  • Video text retrieval method based on BEiT-3 multi-mode large model

    CN118377930A