Video text retrieval method and system based on improved fine-grained alignment
By introducing text-guided objects-text alignment module and similarity-driven frame aggregation module in video text retrieval, the problem of insignificant improvement of fine-grained alignment accuracy in the prior art is solved, and more efficient video text retrieval performance is achieved.
Patent Information
- Application Number
- CN202510164149.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
AI Technical Summary
Existing video text retrieval methods have problems with insignificant improvement in precision in fine-grained alignment, especially when frame embedding ignores visual details, while block embedding is limited by the size and position of Vision Transformer, resulting in low accuracy.
A video text retrieval method based on improved fine-grained alignment is proposed. By introducing text-guided object-text alignment module and similarity-driven frame aggregation module, combining video encoder and text encoder, the fine-grained alignment of objects in video frames and text is realized, and the overall video-text similarity is improved through frame aggregation.
The performance of video text retrieval has been significantly improved, especially in text-to-video and video-to-text retrieval tasks, which perform better on the MSR-VTT dataset compared to the most recent state-of-the-art models.
Smart Images

Figure CN119988676A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video text retrieval, and in particular relates to a video text retrieval method and system based on improved fine-grained alignment. Background Art
[0002] Video-text retrieval is a two-way task, which is divided into two tasks: text-to-video (T2V) and video-to-text (V2T). It takes a text as input and retrieves the corresponding video file. It has received more and more attention in recent years, and it meets the universal needs of life.
[0003] The current video text retrieval methods, whether using fine-tuning or zero-shot learning, have obvious deficiencies. The paper CLIP "Learning transferable visual models from natural language supervision. In International conference on machine learning" has surpassed previous work in terms of accuracy and speed in image-text retrieval, which has led to a paradigm shift in the field of video text retrieval today. Since video text retrieval evolved from image text retrieval, the paper "Clip4clip: An empirical study of clip for end-to-end video clip retrieval and captioning" CLIP4Clip transfers the knowledge of image text retrieval to video text retrieval.
[0004] In video-text retrieval, CLIP4CLIP initially transferred the knowledge of CLIP to the global embedding alignment of video and text, which is a coarse-grained alignment strategy that surpassed previous state-of-the-art models based on non-CLIP. However, this strategy ignores the fine-grained details between the modalities. Therefore, researchers began to adopt a fine-grained alignment strategy, using fine-grained video representations, mainly divided into frames and patches. The results show that DRL in the paper "Disentangled representation learning" uses frame embedding as fine-grained video representation, showing higher accuracy than coarse-grained methods, but frame embedding inevitably ignores many visual details in the frame. At the same time, other studies use patch embedding of CLIP as a finer-grained video representation. Although it provides richer visual details than frame embedding, it is limited by the size and position limitations of the VisionTransformer (ViT) proposed in the paper "An image is worth 16x16words: Transformers for image recognition at scale", resulting in its accuracy being less than that of using frames as fine-grained feature alignment.
[0005] In summary, fine-grained alignment strategies and fine-grained video representations (divided into frames and blocks) are used in video text retrieval methods. The results show that fine-grained video representation using frame embedding is more accurate than coarse-grained methods, but it will ignore many visual details in the frames. Other studies use CLIP's block embedding as a finer-grained video representation. Although it can provide richer visual details, it is limited by the size and position of the Vision Transformer, and its accuracy is lower than that of fine-grained feature alignment using frames, making the accuracy improvement brought by fine-grained alignment of video text retrieval not significant. Summary of the invention
[0006] In view of the deficiencies in the prior art, the present invention proposes a video text retrieval method and system based on improved fine-grained alignment to solve the defects of fine-grained alignment in the existing video text retrieval methods.
[0007] In a first aspect, the present invention provides a video text retrieval method based on improved fine-grained alignment, comprising the following steps:
[0008] Step 1: Obtain a video-text pair dataset; the video-text pair dataset includes a number of samples, each sample is a video-text pair, including a video and its corresponding text, and the text is used to describe its corresponding video;
[0009] Step 2: preprocessing the video-text pair data set to obtain a preprocessed video-text pair data set; the preprocessing comprises: selecting a set number of video frames from each video in the video-text pair data set;
[0010] Step 3: Build a video text retrieval model based on object-based improved fine-grained alignment;
[0011] The video text retrieval model based on object-improved fine-grained alignment includes a video encoder, a text encoder, an object-text alignment module and a similarity-driven frame aggregation module;
[0012] The video encoder is used to encode a plurality of video frames selected from a certain video to obtain a fine-grained representation V of the video = {F 1 ,F 2 ,...,F k ,...,F Nf}, where F k is the kth video frame after encoding, k is the number of the video frame, N f is the number of selected video frames;
[0013] The text encoder is used to encode the text to obtain a fine-grained representation of the text T = {t 1 ,t 2 ,...,t i ...,t Nt}, where t i is the word vector obtained after encoding the i-th word in the text, i is the number of the word in the text, N t is the number of words;
[0014] The object-text alignment module is used to perform fine-grained alignment of objects and texts in the input video frame to obtain the alignment result of objects and texts in the video frame, that is, the local frame-level similarity matrix, including the local frame-level similarity matrix S from video to text direction. v2t-local and the local frame-level similarity matrix S in the text-to-video direction t2v-local ;
[0015] The similarity-driven frame aggregation module is used to perform frame aggregation based on a local frame-level similarity matrix, that is, to assign weights to video frames according to the local frame-level similarity matrix to obtain a final overall video-text similarity matrix S(v,p), where v represents video and p represents text;
[0016] Further, the object-text alignment module includes a target extraction module and an object-text alignment module;
[0017] The object extraction module is used to identify objects in the video frame according to the input text, so as to encode and obtain the object embedding of the extracted object;
[0018] The object and text alignment module is used to align objects and texts in a video frame according to the object embedding and the fine-grained representation of the text, thereby obtaining a local frame-level similarity matrix;
[0019] Further, the object extraction module includes a text-guided object detector and an object encoder;
[0020] The text-guided object detector adopts GroundingDINO, which is used to identify objects in video frames according to the input text, thereby obtaining the bounding box B of the object = {b 1 ,b 2 ,…,b j …,b No}, b j is the bounding box of the jth object identified in the video frame, j represents the object number, and No represents the number of objects;
[0021] The target encoder is used to encode the bounding box B of the input object to obtain the object embedding O = {o 1 ,o 2 ,…,o j ,…,o No}, where o j is the object embedding of the jth object identified in the video frame;
[0022] The target encoder includes a SAM model and several identical cross-attention layers;
[0023] The SAM model uses the hint encoder to calculate the bounding box B of the object according to 1 ,b 2 ,…,b j …,b No}Generate a box token; Encode a video frame f selected from a video through an image encoder to obtain an image embedding f'; After receiving the box token and the image embedding f', the mask decoder decodes the box token to obtain a box embedding L = {l 1 ,l 2 ,…,l j ,…,l No}, where l j is the box embedding of the jth object identified in the video frame;
[0024] The cross-attention layer uses the box embedding as the query and the encoded video frame F as the key and value to generate the object embedding O = {o 1 ,o 2 ,…,o j ,…,o No};
[0025] Furthermore, the object and text alignment modules respectively use a text multi-layer perceptron MLP t And a softmax activation function, object multi-layer perceptron MLP o And a softmax activation function to calculate the weight vector w of the i-th word in the text t i and the weight vector w of the jth object in the video frame o j ; Then use the weighted token-wise maixmum idea to align the objects and texts in the video frame. Specifically, according to the weight vector w of the i-th word in the text t i and the weight vector w of the jth object in the video frame o j , calculate the local frame-level similarity matrix S from video to text in the video frame v2t-local The local frame-level similarity matrix S in the text-to-video direction t2v-local ; The text multi-layer perceptron MLP t and object multilayer perceptron MLP o Both include two linear layers and a sigmoid activation function;
[0026] The formula for the weight vector of a word and the weight vector of an object is:
[0027] w o =softmax(MLP o (O))(3)
[0028] w t =softmax(MLP t (T))(4)
[0029] Among them, w o is the weight vector of the object, w t is the weight vector of the word, softmax represents the softmax activation function, MLP o Represents an object multi-layer perceptron, MLP t Represents a text multilayer perceptron;
[0030] The local frame-level similarity matrix S in the video-to-text direction v2t-local The local frame-level similarity matrix S in the text-to-video direction t2v-local for:
[0031]
[0032] Among them, max j N oAmong the objects, the jth object with the highest similarity to the i-th word, max i N t Among the words, the i-th word has the highest similarity with the j-th object, and T represents transposition.
[0033] Furthermore, the similarity-driven frame aggregation module first uses a multi-layer perceptron MLP in the text-to-video direction t2v and a softmax activation function, based on the local frame-level similarity matrix S in the text-to-video direction t2v-local Calculate the similarity driving weight w from text to video t2v ; Using multi-layer perceptron MLP in the video-to-text direction v2t and a softmax activation function, based on the local frame-level similarity matrix S in the video-to-text direction v2t-local Calculate the similarity driving weight w from video to text v2t ; Then, the local frame-level similarity matrix is aggregated into the overall video-text similarity matrix S(v,p) through the dot product operation;
[0034] w t2v =softmax(MLP t2v (S t2v-local ))(7)
[0035] w v2t =softmax(MLP v2t (S v2t-local ))(8)
[0036]
[0037] in, represents the similarity-driven weight of the text-to-video direction of the kth video frame, represents the similarity-driven weight of the video-to-text direction of the kth video frame, represents the local frame-level similarity matrix of the text-to-video direction of the kth video frame, S v2t-local The local frame-level similarity matrix representing the video-to-text direction of the k-th video frame;
[0038] Step 4: Use the preprocessed video text pair dataset to train the video text retrieval model based on object improved fine-grained alignment, and obtain a trained video text retrieval model based on object improved fine-grained alignment;
[0039] Step 5: Use the trained object-based improved fine-grained alignment video text retrieval model to perform text-to-video and video-to-text retrieval tasks to obtain retrieval results;
[0040] When performing a text-to-video retrieval task, first obtain the text input by the user, and use the trained video text retrieval model based on object-improved fine-grained alignment to obtain the overall video-text similarity matrix between the text input by the user and all videos. The video with the largest overall video-text similarity matrix with the text input by the user is taken as the retrieval result.
[0041] When performing a video-to-text retrieval task, obtain the video input by the user, use the trained video-text retrieval model to obtain the overall video-text similarity matrix between the video input by the user and all texts, and take the text with the largest overall video-text similarity matrix with the video input by the user as the retrieval result;
[0042] In a second aspect, the present invention provides a video text retrieval system based on improved fine-grained alignment, which is used to implement the video text retrieval method based on improved fine-grained alignment, including a data acquisition module and a video text retrieval model based on object improved fine-grained alignment;
[0043] The data acquisition module is used to acquire text or video input by the user;
[0044] The object-based improved fine-grained alignment video text retrieval model is used to retrieve text or video input by a user to obtain retrieval results;
[0045] Specifically, when performing a text-to-video retrieval task, the video text retrieval model based on object-improved fine-grained alignment is used to obtain the overall video-text similarity matrix between the text input by the user and all videos, and the video with the largest overall video-text similarity matrix with the text input by the user is used as the retrieval result;
[0046] When performing a video-to-text retrieval task, the video-text retrieval model is used to obtain the overall video-text similarity matrix between the video input by the user and all texts, and the text with the largest overall video-text similarity matrix with the video input by the user is used as the retrieval result;
[0047] In a third aspect, the present invention provides an electronic device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the video text retrieval method based on improved fine-grained alignment are performed;
[0048] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, executes the steps of a video text retrieval method based on improved fine-grained alignment.
[0049] Compared with the prior art, the beneficial effects of the present invention are:
[0050] (1) We introduce a text-guided object-text alignment (TOTAL) module, which innovatively aligns text with objects extracted from video frames, significantly improving the performance.
[0051] (2) In order to solve the problem of frames with low contribution in the video, a similarity frame aggregation (SIFA) module is proposed to improve the overall effect of the video by assigning weights to the object-text aligned frames in the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A flowchart of a method for improving fine-grained alignment of object-based video text retrieval in an embodiment of the present invention;
[0053] Figure 2 It is a structural diagram of a video text retrieval model based on object-improved fine-grained alignment in an embodiment of the present invention;
[0054] Figure 3 It is an internal structure diagram of the target encoder in an embodiment of the present invention;
[0055] Figure 4 4 is an internal structure diagram of a similarity driven frame aggregation module in an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The present invention is based on CLIP and adopts target detection algorithm to improve fine-grained alignment in video text retrieval. A video text retrieval method based on improved fine-grained alignment is provided. Figure 1 The steps shown include:
[0057] Step 1: Obtain a video-text pair dataset; the video-text pair dataset includes a number of samples, each sample is a video-text pair, including a video and its corresponding text, and the text is used to describe its corresponding video;
[0058] Step 2: Preprocess the video text pair dataset to obtain a preprocessed video text pair dataset;
[0059] The preprocessing is: selecting a set number of video frames from each video in the video text pair data set;
[0060] Step 3: Build a video text retrieval model based on object-improved fine-grained alignment (OFIA);
[0061] like Figure 2 As shown, the object-improved fine-grained alignment-based video text retrieval model (OFIA) includes a video encoder, a text encoder, an object-text alignment (TOTAL) module, and a similarity-wise Frame Aggregation (SIFA) module;
[0062] The video encoder is used to encode a plurality of video frames selected from a certain video to obtain a fine-grained representation V of the video = {F 1 ,F 2 ,...,F k ,...,F Nf}, where F k is the kth video frame after encoding, k is the number of the video frame, N f is the number of selected video frames;
[0063] The text encoder is used to encode the text to obtain a fine-grained representation of the text T = {t 1 ,t 2 ,...,t i ...,t Nt}, where t i is the word vector obtained after encoding the i-th word in the text, i is the number of the word in the text, N t is the number of words;
[0064] In this embodiment, the video encoder and the text encoder are respectively ViT-B / 32 and Transformer in CLIP;
[0065] The object-text alignment (TOTAL) module is used to perform fine-grained alignment of objects and texts in the input video frame to obtain the alignment result of objects and texts in the video frame, that is, the local frame-level similarity matrix, including the local frame-level similarity matrix S from video to text direction. v2t-local and the local frame-level similarity matrix S in the text-to-video direction t2v-local ;
[0066] Furthermore, the object-text alignment (TOTAL) module includes an object extraction unit (OEU) and an object-text alignment module;
[0067] The object extraction module is used to identify objects in the video frame according to the input text, so as to encode and obtain the object embedding of the extracted object;
[0068] Furthermore, the object extraction module includes a text-guided object detector and an object encoder;
[0069] The text-guided object detector adopts GroundingDINO, which is used to identify relevant objects in the video frame according to the input text, so as to obtain the bounding box B of the object = {b 1 ,b 2 ,…,b j …,b No}, b j is the bounding box of the jth object identified in the video frame, j represents the object number, and No represents the number of objects;
[0070] The target encoder is used to encode the bounding box B of the input object to obtain the object embedding O = {o 1 ,o 2 ,…,o j ,…,o No}, where o j is the object embedding of the jth object identified in the video frame;
[0071] Furthermore, if Figure 3 As shown, the target encoder includes a SAM model and several identical cross-attention layers;
[0072] The SAM model uses the hint encoder to calculate the bounding box B of the object according to 1 ,b 2 ,…,b j …,b No Generate box tokens; Encode a video frame f selected from a video through an image encoder to obtain an image embedding f'; After receiving the box token and the image embedding f', the mask decoder decodes the box token to obtain a box embedding L = {l 1 ,l 2 ,…,l j ,…,l No}, where l j is the box embedding of the jth object identified in the video frame;
[0073] The box embedding is represented as:
[0074] L=SAM(B,f)(1)
[0075] Among them, SAM represents the SAM model;
[0076] To generate object embeddings, several identical cross-attention layers are introduced, which use box embeddings as queries and the encoded video frames F (encoded by the video encoder) as keys and values to generate object embeddings O = {o 1 ,o 2 ,…,o j ,…,o No}:
[0077] O=CMA(L,F,F)(2)
[0078] Among them, CMA represents the cross-attention layer;
[0079] The object and text alignment module is used to align objects and texts in a video frame according to the object embedding and the fine-grained representation of the text, thereby obtaining a local frame-level similarity matrix;
[0080] Furthermore, the object and text alignment modules respectively use a text multi-layer perceptron MLP t And a softmax activation function, object multi-layer perceptron MLP o And a softmax activation function to calculate the weight vector w of the i-th word in the text t i and the weight vector w of the jth object in the video frame o j ; Then, the weighted-token wise maximum proposed in DRL and the token-wise interaction proposed based on FILIP and ColBERT are used to align the objects and texts in the video frame. Specifically, according to the weight vector w of the i-th word in the text t i and the weight vector w of the jth object in the video frame o j , calculate the local frame-level similarity matrix in the video-to-text (V2T) and text-to-video (T2V) directions within the video frame; the text multi-layer perceptron MLP t and object multilayer perceptron MLP o Both include two linear layers and a sigmoid activation function;
[0081] The formula for the weight vector of a word and the weight vector of an object is:
[0082] w o =softmax(MLP o(O))(3)
[0083] w t =softmax(MLP t (T))(4)
[0084] Among them, w o is the weight vector of the object, w t is the weight vector of the word, softmax represents the softmax activation function, MLP o Represents an object multi-layer perceptron, MLP t Represents a text multilayer perceptron;
[0085] The local frame-level similarity matrix S in the video-to-text direction v2t-local The local frame-level similarity matrix S in the text-to-video direction t2v-local , the formula is:
[0086]
[0087] Among them, max j N o Among the objects, the jth object with the highest similarity to the i-th word, max i N t Among the words, the i-th word has the highest similarity with the j-th object, and T represents transposition;
[0088] In order to alleviate the negative impact of inefficient frames, that is, there is inefficient alignment in the selected video frames, the similarity-driven frame aggregation module is used to perform frame aggregation based on the local frame-level similarity matrix, that is, weights are assigned to video frames according to the local frame-level similarity matrix, aiming to improve the final similarity score of the correct video-text pair, and obtain the final overall video-text similarity matrix S(v,p), where v represents video and p represents text;
[0089] Furthermore, if Figure 4 As shown, the similarity-driven frame aggregation module first uses a multi-layer perceptron MLP in the text-to-video direction t2v and a softmax activation function, based on the local frame-level similarity matrix S in the text-to-video direction t2v-local Calculate the similarity driving weight w from text to video t2v ; Using multi-layer perceptron MLP in the video-to-text direction v2t and a softmax activation function, based on the local frame-level similarity matrix S in the video-to-text direction v2t-local Calculate the similarity driving weight w from video to text v2t ; Then, the local frame-level similarity matrix is aggregated into the overall video-text similarity matrix S(v,p) through the dot product operation;
[0090] w t2v =softmax(MLP t2v (S t2v-local ))(7)
[0091] w v2t =softmax(MLP v2t (S v2t-local ))(8)
[0092]
[0093] in, represents the similarity-driven weight of the text-to-video direction of the kth video frame, represents the similarity-driven weight of the video-to-text direction of the kth video frame, represents the local frame-level similarity matrix of the text-to-video direction of the kth video frame, S v2t-local The local frame-level similarity matrix representing the video-to-text direction of the k-th video frame;
[0094] Step 4: Use the preprocessed video text pair dataset to train the video text retrieval model based on object improved fine-grained alignment (OFIA), and obtain the trained video text retrieval model based on object improved fine-grained alignment (OFIA);
[0095] Step 5: Use the trained object-improved fine-grained alignment-based video text retrieval model (OFIA) to perform text-to-video (T2V) and video-to-text (V2T) retrieval tasks to obtain retrieval results;
[0096] Specifically, when performing a text-to-video (T2V) retrieval task, the text input by the user is first obtained, and the overall video-text similarity matrix between the text input by the user and all videos is obtained using the trained object-improved fine-grained alignment-based video text retrieval model (OFIA). The video with the largest overall video-text similarity matrix with the text input by the user is taken as the retrieval result.
[0097] When performing a video-to-text (V2T) retrieval task, obtain the video input by the user, and use the trained video text retrieval model (OFIA) to obtain the overall video-text similarity matrix between the video input by the user and all texts. The text with the largest overall video-text similarity matrix with the video input by the user is taken as the retrieval result.
[0098] The performance of the Object Improved Fine-grained Alignment (OFIA) based video text retrieval model on MSR-VTT is evaluated by comparing with recent state-of-the-art models.
[0099] The performance of the video text retrieval model based on object-improved fine-grained alignment (OFIA) on the MSR-VTT dataset significantly surpasses all baseline models. Specifically, in text-to-video retrieval, OFIA improves the state-of-the-art DRL model by 2.3% on R@1 using the ViT-B / 32 video encoder, and reaches 54.1% in the ViT-B / 16 setting. In video-to-text retrieval, the video text retrieval model based on object-improved fine-grained alignment improves DiffusionRet and HBI by 0.5% and 1.4% respectively using the ViT-B / 32 backbone network. When using ViT-B / 16 as the backbone network, the video text retrieval model based on object-improved fine-grained alignment achieves excellent results of 91.1% and 96.9% on R@5 and R@10, respectively.
[0100] Table 1 Comparison of OFIA and previous work on the MSRVTT dataset
[0101]
[0102]
[0103] In this embodiment, a video text retrieval system based on improved fine-grained alignment is used to implement a video text retrieval method based on improved fine-grained alignment, including a data acquisition module and a video text retrieval model based on object improved fine-grained alignment;
[0104] The data acquisition module is used to acquire text or video input by the user;
[0105] The object-based improved fine-grained alignment video text retrieval model is used to retrieve text or video input by a user to obtain retrieval results;
[0106] Specifically, when performing a text-to-video (T2V) retrieval task, the object-improved fine-grained alignment-based video text retrieval model (OFIA) is used to obtain the overall video-text similarity matrix between the text input by the user and all videos, and the video with the largest overall video-text similarity matrix with the text input by the user is taken as the retrieval result;
[0107] When performing video-to-text (V2T) retrieval tasks, the video text retrieval model (OFIA) is used to obtain the overall video-text similarity matrix between the video input by the user and all texts, and the text with the largest overall video-text similarity matrix with the video input by the user is used as the retrieval result.
[0108] In this embodiment, an electronic device includes: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the video text retrieval method based on improved fine-grained alignment are performed;
[0109] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the video text retrieval method based on improved fine-grained alignment are performed.
Claims
1. A video text retrieval method based on improved fine-grained alignment, characterized in that: The steps include: Step 1: Obtain a video-text pair dataset; the video-text pair dataset includes a number of samples, each sample is a video-text pair, including a video and its corresponding text, and the text is used to describe its corresponding video; Step 2: preprocessing the video-text pair data set to obtain a preprocessed video-text pair data set; the preprocessing comprises: selecting a set number of video frames from each video in the video-text pair data set; Step 3: Build a video text retrieval model based on object-based improved fine-grained alignment; Step 4: Use the preprocessed video text pair dataset to train the video text retrieval model based on object improved fine-grained alignment, and obtain a trained video text retrieval model based on object improved fine-grained alignment; Step 5: Use the trained object-based improved fine-grained alignment video text retrieval model to perform text-to-video and video-to-text retrieval tasks to obtain retrieval results.
2. The video text retrieval method based on improved fine-grained alignment according to claim 1, characterized in that: The video text retrieval model based on object-improved fine-grained alignment in step 3 includes a video encoder, a text encoder, an object-text alignment module, and a similarity-driven frame aggregation module; The video encoder is used to encode a number of video frames selected from a video to obtain a fine-grained representation of the video V = {F1, F2, ..., F k ,...,F Nf }, where F k is the kth video frame after encoding, k is the number of the video frame, N f is the number of selected video frames; The text encoder is used to encode the text to obtain a fine-grained representation of the text T = {t1, t2, ..., t i ...,t Nt }, where t i is the word vector obtained after encoding the i-th word in the text, i is the number of the word in the text, N t is the number of words; The object-text alignment module is used to perform fine-grained alignment of objects and texts in the input video frame to obtain the alignment result of objects and texts in the video frame, that is, the local frame-level similarity matrix, including the local frame-level similarity matrix S from video to text direction. v2t-local and the local frame-level similarity matrix S in the text-to-video direction t2v-local ; The similarity-driven frame aggregation module is used to perform frame aggregation based on the local frame-level similarity matrix, that is, to assign weights to video frames according to the local frame-level similarity matrix to obtain the final overall video-text similarity matrix S(v,p), where v represents video and p represents text.
3. The video text retrieval method based on improved fine-grained alignment according to claim 2, characterized in that: The object-text alignment module includes a target extraction module and an object-text alignment module; The object extraction module is used to identify objects in the video frame according to the input text, thereby obtaining the object embedding of the extracted object through encoding; The object and text alignment module is used to achieve object and text alignment within a video frame based on object embedding and fine-grained representation of text, thereby obtaining a local frame-level similarity matrix.
4. The video text retrieval method based on improved fine-grained alignment according to claim 3 is characterized in that: The object extraction module includes a text-guided object detector and an object encoder; The text-guided object detector adopts GroundingDINO, which is used to identify objects in video frames according to the input text, so as to obtain the bounding box B of the object = {b1, b2, ..., b j …,b No }, b j is the bounding box of the jth object identified in the video frame, j represents the object number, and No represents the number of objects; The target encoder is used to encode the bounding box B of the input object to obtain the object embedding O = {o1, o2, ..., o j ,…,o No }, where o j is the object embedding of the jth object identified in the video frame; The target encoder includes a SAM model and several identical cross-attention layers; The SAM model uses the hint encoder to calculate the bounding box B = {b1, b2, ..., b j …,b No }Generate a box token; encode a video frame f selected from a video through the image encoder to obtain the image embedding f'; after receiving the box token and the image embedding f', the mask decoder decodes the box token to obtain the box embedding L = {l1,l2,…,l j ,…,l No }, where l j is the box embedding of the jth object identified in the video frame; The cross-attention layer uses the box embedding as the query and the encoded video frame F as the key and value to generate the object embedding O = {o1,o2,…,o j ,…,o No }.
5. The video text retrieval method based on improved fine-grained alignment according to claim 3 is characterized in that: The object and text alignment modules use text multi-layer perceptron MLP t And a softmax activation function, object multi-layer perceptron MLP o And a softmax activation function to calculate the weight vector w of the i-th word in the text t i and the weight vector w of the jth object in the video frame o j ; Then use the weighted token-wise maixmum idea to align the objects and texts in the video frame. Specifically, according to the weight vector w of the i-th word in the text t i and the weight vector w of the jth object in the video frame o j , calculate the local frame-level similarity matrix S from video to text in the video frame v2t-local The local frame-level similarity matrix S in the text-to-video direction t2v-local ; The text multi-layer perceptron MLP t and object multilayer perceptron MLP o Both include two linear layers and a sigmoid activation function; The formula for the weight vector of a word and the weight vector of an object is: w o = softmax(MLP o (O)) (3) w t = softmax(MLP t (T)) (4) Among them, w o is the weight vector of the object, w t is the weight vector of the word, softmax represents the softmax activation function, MLP o Represents an object multi-layer perceptron, MLP t Represents a text multilayer perceptron; The local frame-level similarity matrix S in the video-to-text direction v2t-local The local frame-level similarity matrix S in the text-to-video direction t2v-local for: Among them, max j N o Among the objects, the jth object with the highest similarity to the i-th word, max i N t Among the words, the i-th word has the highest similarity with the j-th object, and T represents transposition.
6. The video text retrieval method based on improved fine-grained alignment according to claim 2, characterized in that: The similarity-driven frame aggregation module first uses a multi-layer perceptron MLP in the text-to-video direction t2v and a softmax activation function, based on the local frame-level similarity matrix S in the text-to-video direction t2v-local Calculate the similarity driving weight w from text to video t2v ; Using multi-layer perceptron MLP in the video-to-text direction v2t and a softmax activation function, based on the local frame-level similarity matrix S in the video-to-text direction v2t-local Calculate the similarity driving weight w from video to text v2t ; Then, the local frame-level similarity matrix is aggregated into the overall video-text similarity matrix S(v,p) through the dot product operation; w t2v =softmax(MLP t2v (S t2v-local )) (7) w v2t =softmax(MLP v2t (S v2t-local )) (8) in, represents the similarity-driven weight of the text-to-video direction of the kth video frame, represents the similarity-driven weight of the video-to-text direction of the kth video frame, represents the local frame-level similarity matrix of the text-to-video direction of the kth video frame, S v2t-local Represents the local frame-level similarity matrix in the video-to-text direction for the k-th video frame.
7. The video text retrieval method based on improved fine-grained alignment according to claim 1, characterized in that: The step 5 is specifically as follows: when performing a text-to-video retrieval task, firstly obtain the text input by the user, use the trained video text retrieval model based on object-improved fine-grained alignment to obtain the overall video-text similarity matrix between the text input by the user and all videos, and take the video with the largest overall video-text similarity matrix with the text input by the user as the retrieval result; When performing a video-to-text retrieval task, obtain the video input by the user, and use the trained video-text retrieval model to obtain the overall video-text similarity matrix between the video input by the user and all texts. The text with the largest overall video-text similarity matrix with the video input by the user is used as the retrieval result.
8. A video text retrieval system based on improved fine-grained alignment, used to implement a video text retrieval method based on improved fine-grained alignment as claimed in any one of claims 1 to 7, characterized in that: It includes a data acquisition module and a video text retrieval model based on object-based improved fine-grained alignment; The data acquisition module is used to acquire the text or video input by the user; The object-based improved fine-grained alignment video text retrieval model is used to retrieve text or video input by a user to obtain retrieval results; Specifically, when performing a text-to-video retrieval task, the video text retrieval model based on object-improved fine-grained alignment is used to obtain the overall video-text similarity matrix between the text input by the user and all videos, and the video with the largest overall video-text similarity matrix with the text input by the user is used as the retrieval result; When performing a video-to-text retrieval task, the video-text retrieval model is used to obtain the overall video-text similarity matrix between the video input by the user and all texts, and the text with the largest overall video-text similarity matrix with the video input by the user is used as the retrieval result.
9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of a video text retrieval method based on improved fine-grained alignment are performed as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of a video text retrieval method based on improved fine-grained alignment as described in any one of claims 1 to 7.