Multi-label cross-modal video-text retrieval method based on Mamba and storage medium

By using the Mamba model and the video-text interaction module VAT, the correlation problem of long video sequence data was solved, the high accuracy and reliability of video-text retrieval was achieved, the defects of traditional models were overcome, and the performance of multi-label cross-modal retrieval was improved.

CN120780871APending Publication Date: 2025-10-14GUILIN UNIV OF ELECTRONIC TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510860267.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively learning long-distance dependencies when processing long video sequence data, resulting in insufficient accuracy and reliability of video-text retrieval results and an inability to fully consider the overall semantics of the video.

Method used

The Mamba model is used to encode text and video frames, and a video-text interaction module VAT is constructed. The bimodal similarity function and bidirectional retrieval loss function are used to achieve efficient interaction and multimodal information fusion between video and text.

Benefits of technology

It significantly improves the accuracy and reliability of video-text retrieval, solves the modal gap problem of video and text data, and can more accurately grasp the correlation between the previous and next content in the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780871A_ABST
    Figure CN120780871A_ABST
Patent Text Reader

Abstract

The invention discloses a Mama-based multi-label cross-modal video-text retrieval method, which is characterized in that text and video frames are coded by using Mama and Vision Mama, so that a model can effectively learn a long-distance dependency relationship, the defects of a traditional model during processing of long video sequence data are overcome, and the retrieval efficiency is improved. The relevance between the front content and the back content in the video can be more accurately grasped, and the accuracy and the reliability of video-text retrieval are remarkably improved; the multiple tags are input into the Mamba model in sequence for feature extraction, rich information contained in the multi-level tags can be more fully utilized, and the performance of the video text retrieval model is further improved; a video-text interaction module is utilized to enable the model to realize modal interaction in low-level feature and high-level semantic levels, a Mama model is utilized to construct multi-modal information interaction and association between a video and a text, and a video-text bidirectional retrieval loss function is combined to maximize the similarity between video features and text features. And the modal gap problem of video and text data is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer application, and particularly relates to a multi-label cross-modal video-text retrieval method based on Mamba and a storage medium. BACKGROUND

[0002] Cross-modal video-text retrieval refers to establishing association and mapping between video and text data, so as to realize retrieving semantically related text according to given video, or retrieving semantically related video according to given text. Video is composed of continuous image frames and contains rich visual information, while text is composed of discrete words and expresses semantics through language rules. It is a challenging task to effectively associate and match these two completely different forms of representation.

[0003] At present, the related technologies mainly model and construct association between video and text based on recurrent neural network, convolutional neural network and Transformer architecture, including patent CN114282060A discloses a fine-grained video-text retrieval method based on context Transformer network, patent CN117112838A discloses a video text retrieval method based on CLIP contrast learning, patent CN119166851A discloses a text retrieval video method, system and device based on instruction guided GPT, patent CN119066222A discloses a video text retrieval method based on time token merging, patent CN118377930B discloses a video text retrieval method based on BEiT-3 multi-modal large model, etc.

[0004] However, the above-mentioned prior art has difficulty in effectively learning long-distance dependency when processing long video sequence data, which further affects the accurate grasp of the correlation between the front and rear content in the video, resulting in that the overall video semantics cannot be comprehensively considered during retrieval, and the retrieval result may be biased. SUMMARY

[0005] In view of the deficiencies of the prior art, the present application provides a multi-label cross-modal video-text retrieval method based on Mamba. The present application overcomes the defects of traditional models in processing long video sequence data, can more accurately grasp the correlation between the front and rear content in the video, significantly improves the accuracy and reliability of video-text retrieval, and also solves the problem of modal gap between video and text data.

[0006] The technical scheme of the present application specifically includes:

[0007] (1) Obtain a video-text pair dataset Where v i is the video modality, t i is the text modality, and yi is a multi-label text set, y i ={α1,α2,...,α m} represents that the video-text pair (v i , t i ) has m labels, and n represents the number of video-text pairs in the video-text pair dataset.

[0008] (2) Randomly divide the video-text pair dataset O into training set, validation set and test set according to the ratio of 8:1:1.

[0009] (3) Train and verify the multi-label cross-modal video-text retrieval model using the training set and the validation set, and obtain the trained video-text retrieval model file.

[0010] (4) Input the test set into the trained multi-label cross-modal video-text retrieval model to obtain the retrieval result.

[0011] The construction of the multi-label cross-modal video-text retrieval model includes the following steps:

[0012] (1) Use the first Mamba (i.e. Mamba1) model to encode the text t i to obtain the initial feature of the text

[0013] (2) Input the multi-label text set y i to the second Mamba (i.e. Mamba2) model in sequence for encoding to obtain the multi-label text feature

[0014] (3) Extract frames from the video v i at a constant speed to obtain j frames v i ={z1,z2,...,z j}, and input the frame sequence to the Vision Mamba model for encoding to obtain the initial feature of the video frame sequence

[0015] (4) Construct a video-text interaction module VAT and insert it into the middle layer of the Mamba1 and Vision Mamba models to realize efficient interaction between video and text modalities. The interaction formula is:

[0016]

[0017] where Layer is the number of layers of the feature extraction network, and is the cross-modal context vector. After Layer layer feature extraction, the final text feature and video frame sequence features

[0018] (5) combine the text features and video frame sequence features to obtain first fusion features

[0019] (6) input the first fusion features C i 1 , the position index ps of the frame sequence vector and the multi-label text features into a third Mamba (i.e., Mamba3) model to perform multi-modal information interaction between the video and the text, to obtain second fusion features of the fusion of the text, the video, the position information and the multi-label

[0020] (7) further learn the second fusion features C i 2 using two fully connected layers, and calculate the similarity value of the video v i and the text t i using a bi-modal similarity function P(v v2t , t t2v ) = F(Sigmoid(F(C i i )), wherein F(·) represents a fully connected layer, and Sigmoid(·) is an activation function.

[0021] (8) construct a video-text bidirectional retrieval loss function to train the constructed multi-label cross-modal video-text retrieval model, and the formula of the video-text bidirectional retrieval loss function is as follows:

[0022]

[0023] wherein e() represents an exponential function with base e, and τ is a model learnable parameter.

[0024] Preferably, the construction of the video-text interaction module VAT comprises the following steps:

[0025] (1) generate query vectors and value vectors and from the text features and the video frame sequence features respectively. wherein and are network parameter matrices.

[0026] (2) calculate the attention matrix of the text to the video and video-to-text attention matrix

[0027]

[0028] where d is the feature dimension, T denotes the transpose, and the Softmax function is used to convert scores into probability distributions.

[0029] (3) Calculate the cross-modal context vector F v2t and F t2v :

[0030]

[0031] where, and is the output network parameter matrix.

[0032] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the multi-label cross-modal video-text retrieval method.

[0033] The technical features and beneficial effects of the application are as follows:

[0034] (1) By using Mamba and Vision Mamba to encode text and video frames, the model can effectively learn long-distance dependencies, overcoming the shortcomings of traditional models in processing long video sequence data, and more accurately grasping the relevance of the content before and after the video, thereby comprehensively considering the overall semantics of the video, significantly improving the accuracy and reliability of video-text retrieval.(2) The multi-label is sequentially input into the Mamba model for feature extraction, which can more fully utilize the rich information contained in the multi-level labels, further improving the performance of the video-text retrieval model.(3) The video-text interaction module is constructed to enable the model to realize modal interaction at both low-level features and high-level semantics, use the Mamba model to construct multi-modal information interaction and association between video and text, and combine the video-text bidirectional retrieval loss function to maximize the similarity of video features and text features, solving the modal gap problem of video and text data. BRIEF DESCRIPTION OF DRAWINGS

[0035] Figure 1 is a flowchart of the application;

[0036] Figure 2 is a flowchart of model training, model verification and model testing in the embodiment. DETAILED DESCRIPTION

[0037] The application will be further described in detail below with reference to the accompanying drawings and embodiments, so as to better understand the technical solutions of the application.

[0038] As Figure 1 shown, the present application mainly includes the following steps:

[0039] (1) Obtain a video-text pair dataset where v i is a video modality, t i is a text modality, y i is a label set, and y i ={α1,α2,...,α m} indicates that the video-text pair (v i , t i ) has m labels.

[0040] (2) Randomly divide the video-text pair dataset O into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0041] (3) Train and validate the Mamba-based multi-label cross-modal video-text retrieval model using the training set and the validation set, and obtain a trained video-text retrieval model file.

[0042] (4) Input the test set into the trained video-text retrieval model to obtain a retrieval result.

[0043] In this embodiment, the model training process of the Mamba-based multi-label cross-modal video-text retrieval method is as follows:

[0044] Obtain a video-text pair dataset where v i is a video modality, t i is a text modality, y i is a label set, and y i ={α1,α2,...,α m} indicates that the video-text pair (v i , t i ) has m labels.

[0045] Randomly divide the video-text pair dataset O into a training set, a validation set, and a test set in a ratio of 8:1:1.

[0046] Encode the text t i using a first Mamba (Mamba1) model to obtain text initial features

[0047] Input the multi-label text y i ={α1,α2,...,α m} into a second Mamba (Mamba2) model in sequence for encoding to obtain multi-label text features Syi .

[0048] Uniform speed from video v i Extract the frame and get j frames v i ={z1,z2,...,z j}, and input the frame sequence into the Vision Mamba (Vim) model for encoding to obtain the initial features of the video frame sequence

[0049] A video-text interaction module (VAT) is constructed and inserted into the middle layer of the Mamba1 and Vim models to achieve efficient interaction between video and text modalities. The interaction formula is:

[0050]

[0051] Among them, Layer is the number of layers of the feature extraction network, and Cross-modal context vector. After layer-by-layer feature extraction, the final text feature is obtained. and video frame sequence features

[0052] The construction of the video-text interaction module VAT includes the following steps:

[0053] Text features and video frame sequence features Generate query vectors separately and Value vector and

[0054]

[0055] in, and is the network parameter matrix.

[0056] Calculating the attention matrix for text-to-video and the video-to-text attention matrix

[0057]

[0058] Among them, d is the feature dimension, represents the transpose, and the Softmax function is used to convert the score into a probability distribution.

[0059] Calculate the cross-modal context vector F v2t and F t2v :

[0060]

[0061] wherein, and is an output network parameter matrix.

[0062] The text features and the video frame sequence features are connected to obtain first fusion features

[0063] The first fusion features C i 1 , the position index ps of the frame sequence vector and the multi-label text features are merged to input a third Mamba (Mamba3) model to perform multi-modal information interaction between the video and the text, and to obtain second fusion features of the fusion of the text, the video, the position information and the multi-label

[0064] The second fusion features C i 2 are further learned by using two full connection layers, and the similarity value of the video v i and the text t i is calculated by using a dual-modal similarity function P(v ,t

[0001] )=F(Sigmoid(F(C )), wherein F(·) represents a full connection layer, and Sigmoid(·) is an activation function.

[0065] A video-text bidirectional retrieval loss function is constructed to train the video-text retrieval model, and the formula of the video-text bidirectional retrieval loss function is as follows:

[0066]

[0067] wherein e() represents an exponential function with e as the base, and τ is a model learnable parameter.

[0068] In the embodiment, the model verification process of the Mamba-based multi-label cross-modal video-text retrieval method is as follows:

[0069] The performance of the model in two retrieval tasks, including the video retrieval text task and the text retrieval video task, is evaluated by using the verification set.

[0070] The video set and the text set in the verification set are used to verify the model, and the model file with the best performance is selected as the trained complete video-text retrieval model.

[0071] In the present embodiment, the model test flow of the Mamba-based multi-label cross-modal video-text retrieval method is as follows:

[0072] The test set is input into the trained video-text retrieval model, and the model returns the retrieval result.

Claims

1. A multi-label cross-modal video-text retrieval method based on Mamba, characterized by: The method comprises: Get the video-text pair dataset where v i is the video mode, t i Is text mode, y i is a multi-label text set, y i ={α1,α2,...,α m } represents a video-text pair (v i ,t i ) has m labels, and n represents the number of video-text pairs in the video-text pair dataset; The video-text pair dataset O is randomly divided into training set, validation set and test set in the ratio of 8:1:1; Use the training set and validation set to train and validate the multi-label cross-modal video-text retrieval model to obtain a fully trained video-text retrieval model file; The construction of the multi-label cross-modal video-text retrieval model includes the following steps: (1) Use the first Mamba (i.e. Mamba1) model to analyze the text t i Encode and obtain the initial features of the text (2) The multi-label text set y i Input the second Mamba (i.e. Mamba2) model in sequence for encoding to obtain multi-label text features (3) Constant speed from video v i Extract the frame and get j frames v i ={z1,z2,...,z j }, and input the frame sequence into the Vision Mamba model for encoding to obtain the initial features of the video frame sequence (4) Construct the video-text interaction module VAT and insert it into the middle layer of Mamba1 and Vision Mamba models to achieve efficient interaction between video and text modalities; after Layer-layer feature extraction, the final text feature is obtained and video frame sequence features (5) Text features and video frame sequence features Connect them to get the first fusion feature C i 1 ; (6) The first fusion feature C i 1 , the position index ps of the frame sequence vector and the multi-label text feature Merge and input the third Mamba (i.e. Mamba3) model to perform multimodal information interaction between video and text, and obtain the second fusion feature that integrates text, video, location information and multiple labels (7) Use two fully connected layers to fusion the second feature C i 2 Further learning and using the bimodal similarity function P(v i ,t i )=F(Sigmoid(F(C i 2 )))Calculate video v i and text t i , where F(·) represents the fully connected layer and Sigmoid(·) is the activation function; (8) Construct a video-text bidirectional retrieval loss function to train the constructed multi-label cross-modal video-text retrieval model.

2. The method according to claim 1, characterized in that The construction of the video-text interaction module VAT includes the following steps: (1) Text features and video frame sequence features Generate query vectors separately and Value vector and Among them, W t q 、 W t p and is the network parameter matrix; (2) Calculate the attention matrix from text to video and the video-to-text attention matrix Where,d is the feature dimension, represents the transpose, and the Softmax function is used to convert the score into a probability distribution; (3) Calculate the cross-modal context vector F v2t and F t2v : Among them, W t out and is the output network parameter matrix.

3. A non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-label cross-modal video-text retrieval method according to claim 1 or 2.

Citation Information

Patent Citations

  • Fine-grained video-text retrieval method based on context Transform network

    CN114282060A

  • Video text retrieval method based on CLIP comparative learning

    CN117112838A

  • A video text retrieval method based on BEiT-3 multimodal large model

    CN118377930B

  • Video text retrieval method based on time sequence token combination

    CN119066222A

  • Text retrieval video method, system and equipment based on instruction guide GPT

    CN119166851A