A video retrieval method and system driven by semantic of large language model

By adopting a large language model-driven video feature extraction and cross-attention mechanism in text video retrieval, the problem of lack of semantic reasoning in cross-modal matching in the prior art is solved, and a more efficient and interpretable text video retrieval effect is achieved.

CN119397057BActive Publication Date: 2025-06-20JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411990362.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-06-20
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The existing text video retrieval technology ignores the semantic reasoning process in the cross-modal matching process, resulting in low retrieval accuracy and lack of interpretability.

Method used

The video retrieval method based on the large language model is adopted, and the frame-level features and patch-level features of candidate videos are extracted through the video feature extractor, and time-sequence modeling and spatial modeling are performed. Multi-grained interaction is performed by combining the dual-stream cross attention mechanism, and finally the video features and query text are embedded in the large language model for semantic reasoning and cross-modal matching.

Benefits of technology

It improves the accuracy and efficiency of text video retrieval, enhances the interpretability of the model, and can more accurately understand and match cross-modal semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119397057B_ABST
    Figure CN119397057B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of text-video retrieval, and in particular to a video retrieval method and system driven by the semantics of a large language model, including: obtaining a query text and a candidate video set; constructing a text-video retrieval model; inputting the candidate video set into the text-video retrieval model, and obtaining the video features of each candidate video through a video feature extractor; after embedding the query text into a preset prompt statement, inputting the video features of all candidate videos and the query text into the large language model, and outputting the similarity between the query text and each candidate video; and outputting a video retrieval result based on the similarity between the query text and each candidate video. The present invention constructs video features including dynamic changes and spatial details, and uses the powerful semantic reasoning ability of the large language model to obtain cross-modal semantic relationships, which conforms to the cognitive behavior of human retrieval, enhances the interpretability of the model, and improves the accuracy and efficiency of cross-modal text-video retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text-video retrieval, and in particular to a video retrieval method and system driven by large language model semantics. Background Art

[0002] With the popularization of the Internet and the development of digital media technology, videos have become an important carrier for information dissemination, and the number of video contents has increased explosively. How to meet the demand for efficient information retrieval has become an urgent problem to be solved. Text to Video Retrieval is a cross-modal matching task that combines natural language processing and computer vision, aiming to quickly locate and retrieve relevant video segments according to text queries in a large-scale video dataset.

[0003] Benefiting from the excellent performance of the pre-trained image-text model (Contrastive Language Image Pre-training, CLIP) in image retrieval tasks, some technologies have migrated CLIP to video retrieval tasks and achieved good retrieval results. As a representative technology, the CLIP4Clip model migrates the image encoder and text encoder of CLIP to video retrieval tasks, realizing end-to-end video retrieval. First, this technology uniformly samples video frames and inputs them into the image encoder to extract frame-level features, encodes the text description using the text encoder to obtain text features; then uses a network structure containing several layers of Transformer modules to perform temporal modeling on the frame-level features and performs average pooling to obtain video features; finally, calculates the cosine similarity between the video features and the text features for cross-modal matching, and trains the model using contrastive learning loss. Many subsequent excellent technologies, such as Cap4Video, HBI, and DRL, etc., have also been improved on this basis, such as auxiliary information enhancement, auxiliary signal supervision, and fine-grained interaction, etc.

[0004] However, the cross-modal matching process in existing text-video retrieval technologies is in the form of "tokens to tokens", that is, the text is segmented into text tokens such as words, sub-words, or characters, and the video is segmented into video frames or video blocks as video tokens, and then feature extraction is performed on the text tokens and video tokens respectively, and they are converted into feature vectors of a fixed dimension for association and alignment, such as video frame-level feature alignment with text word-level features, video patch-level feature alignment with text word-level features, and video-level feature alignment with text global features. No matter which alignment method is used, it measures the semantic relevance between modalities, while ignoring the cross-modal semantic reasoning process, making text-video retrieval lack interpretability and difficult to complete some complex retrieval tasks that require in-depth understanding of multi-modal semantic information, thus limiting the accuracy of text-video retrieval. Summary of the Invention

[0005] To this end, the technical problem to be solved by the present invention is to overcome the prior art, which only measures the semantic relevance between modalities while ignoring the cross-modal semantic reasoning process, thereby reducing the accuracy of text-video retrieval.

[0006] To solve the above technical problem, the present invention provides a video retrieval method driven by semantic of large language model, including:

[0007] Constructing a text-video retrieval model; the text-video retrieval model includes a video feature extractor and a large language model; the video feature extractor includes a video encoder, a temporal modeling module, a spatial modeling module, a patch-level clustering module, a two-stream cross-attention module, and a multi-layer perceptron;

[0008] Inputting a candidate video set into the text-video retrieval model, and obtaining the video features of each candidate video through the video feature extractor, including:

[0009] The candidate video passes through the video encoder to output a first frame-level feature and a first patch-level feature;

[0010] Inputting the first frame-level feature into the temporal modeling module to output a second frame-level feature; inputting the first patch-level feature into the spatial modeling module to output a second patch-level feature;

[0011] Using the patch-level clustering module to cluster the second patch-level features to obtain third patch-level features;

[0012] Inputting the second frame-level feature and the third patch-level feature into the two-stream cross-attention module for interaction to output a target frame-level feature and a target patch-level feature;

[0013] Inputting the target frame-level feature and the target patch-level feature into the multi-layer perceptron to output the video features of the candidate video;

[0014] After embedding the query text into a preset prompt statement, inputting it together with the video features of all candidate videos into the large language model to output the similarity between the query text and each candidate video;

[0015] Outputting a video retrieval result based on the similarity between the query text and each candidate video.

[0016] Preferably, the video encoder is constructed based on the image encoder of the CLIP model.

[0017] Preferably, both the temporal modeling module and the spatial modeling module are composed of a number of stacked Transformer network layers.

[0018] Preferably, inputting the second frame-level feature and the third patch-level feature into the dual-stream cross-attention module for interaction and outputting the target frame-level feature and the target patch-level feature includes:

[0019] The dual-stream cross-attention module includes a patch-level cross-attention mechanism and a frame-level cross-attention mechanism;

[0020] The patch-level cross-attention mechanism takes the third patch-level feature as the input query, takes the second frame-level feature as the input key and value, and outputs the target patch-level feature;

[0021] The frame-level cross-attention mechanism takes the second frame-level feature as the input query, takes the third patch-level feature as the input key and value, and outputs the target frame-level feature.

[0022] Preferably, the large language model is the LLaMA model.

[0023] Preferably, when training the text-video retrieval model, freeze the parameters of the video encoder of the video feature extractor and the large language model.

[0024] Preferably, when training the text-video retrieval model, it further includes a masking and reconstruction module;

[0025] The masking and reconstruction module includes an importance prediction layer, a conditional masking layer, and a feature reconstruction layer;

[0026] The importance prediction layer is used to generate the importance score of each patch-level feature in the second patch-level feature, including a first linear layer, a RELU activation layer, a second linear layer, and a Softmax layer; wherein, the first linear layer is used to map the second patch-level feature from d dimensions to 2*d dimensions for projection transformation; the RELU activation layer is used to maintain the non-negativity of the feature and add non-linear factors; the second linear layer is used to map the activated feature from 2*d dimensions to 1 dimension; the Softmax layer is used to map the 1-dimensional feature to the interval (0,1) to generate the importance score of each patch-level feature;

[0027] The formula representation of the importance prediction layer is:

[0028] ;

[0029] wherein, and respectively represent the projection parameters of the first linear layer and the second linear layer, is the activation function, is the Softmax function; represents the second patch-level feature, represents the importance score of the second patch-level feature;

[0030] The conditional masking layer is used to mask the K features with the highest importance scores in each frame of the second patch-level features using random noise, obtaining the masked patch-level features;

[0031] The feature reconstruction layer is composed of a stack of encoding layers of several Transformer network layers; in the multi-head attention layer of the encoding layer of each Transformer network layer, the masked patch-level features are used as the input query, and the second frame-level features are used as the input key and value to reconstruct the masked patch-level features, and the reconstructed patch-level features are output, and the formula is expressed as:

[0032] ;

[0033] Among them, represents the feature reconstruction module, , and are the linear transformation parameters of the query, key, and value respectively; represents the reconstructed patch-level features, represents the masked patch-level features, represents the second frame-level features.

[0034] Preferably, the loss function of the text video retrieval model includes a contrastive learning loss function and a reconstruction loss function; the reconstruction loss function is used to calculate the difference between the reconstructed patch-level features and the second patch-level features.

[0035] Preferably, the patch-level clustering module clusters the second patch-level features based on the Kmeans clustering algorithm.

[0036] The present invention also provides a video retrieval system driven by the semantics of a large language model, including:

[0037] A retrieval model construction module for constructing a text video retrieval model; the text video retrieval model includes a video feature extractor and a large language model; the video feature extractor includes a video encoder, a temporal modeling module, a spatial modeling module, a patch-level clustering module, a two-stream cross-attention module, and a multi-layer perceptron;

[0038] A feature extraction module for inputting a candidate video set into the text video retrieval model and obtaining the video features of each candidate video through the video feature extractor, including:

[0039] The candidate video passes through the video encoder to output the first frame-level features and the first patch-level features;

[0040] Input the first frame-level feature into the temporal modeling module to output the second frame-level feature; input the first patch-level feature into the spatial modeling module to output the second patch-level feature;

[0041] Use the patch-level clustering module to cluster the second patch-level feature to obtain the third patch-level feature;

[0042] Input the second frame-level feature and the third patch-level feature into the two-stream cross-attention module for interaction to output the target frame-level feature and the target patch-level feature;

[0043] Input the target frame-level feature and the target patch-level feature into a multi-layer perceptron to output the video feature of the candidate video;

[0044] The similarity acquisition module is used to embed the query text into a preset prompt statement, and then input the video features of all candidate videos into a large language model to output the similarity between the query text and each candidate video;

[0045] The retrieval output module is used to output the video retrieval result according to the similarity between the query text and each candidate video.

[0046] The above technical solution of the present invention has the following beneficial effects compared with the prior art:

[0047] In the video retrieval method based on semantic driving of a large language model described in the present invention, the frame-level feature and the patch-level feature of the candidate video are extracted, and the temporal modeling is performed on the frame-level feature to capture the changes of the video content on the time axis, which helps to extract the temporal dynamics of the video and understand the dynamic content of the video. The spatial modeling and clustering are performed on the patch-level feature to capture the spatial relationship between the parts of the video picture, which helps to extract the spatial structure distribution of the video picture; further, based on the cross-attention mechanism, the frame-level feature and the patch-level feature are interacted at multiple granularities, which can integrate the information of the video in the time and space dimensions, construct a video feature containing dynamic changes and spatial details, retain the key information of the video to the greatest extent, and provide rich and accurate video information for the text-video retrieval task to improve the accuracy and efficiency of the text-video retrieval. Moreover, the present invention inputs the extracted video feature and the prompt statement embedding the query text into the large language model for semantic reasoning and cross-modal semantic matching, and uses the powerful semantic reasoning ability of the large language model to obtain more reliable and more interpretable cross-modal semantic relationships, which conforms to the cognitive behavior of human retrieval, enhances the interpretability of the model, and improves the accuracy and efficiency of cross-modal text-video retrieval. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to the specific embodiments of the present invention and in combination with the accompanying drawings, wherein:

[0049] Figure 1 This is a flowchart of a video retrieval method based on semantic drive of a large language model according to the present invention. Detailed implementation manners

[0050] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited do not limit the present invention. Embodiment 1

[0051] Refer to Figure 1 As shown, the present invention provides a video retrieval method based on semantic drive of a large language model, including:

[0052] S1. Obtain a query text and a candidate video set.

[0053] S2. Construct a text-video retrieval model.

[0054] The text-video retrieval model includes a video feature extractor and a large language model.

[0055] The video feature extractor includes a video encoder, a temporal modeling module, a spatial modeling module, a patch-level clustering module, a two-stream cross-attention module, and a multi-layer perceptron.

[0056] S3. Input the candidate video set into the text-video retrieval model, and obtain the video features of each candidate video through the video feature extractor, including:

[0057] S3-1. The candidate video passes through the video encoder to output a first frame-level feature and a first patch-level feature.

[0058] Preferably, in this embodiment, the video encoder is constructed based on the image encoder of the CLIP model, and the video encoder is stacked by 12 Transformer network layers.

[0059] Specifically, the input of the video encoder is the video frame images uniformly extracted from the candidate video , where represents the number of sampled frames. Each frame image is divided into non-overlapping patches of the same size, and after simple encoding, a classification marker [CLS] token is added to each frame, and then input into the video encoder.

[0060] The video encoder performs attention operations on each frame image and outputs a first frame-level feature ; performs attention operations on the patches after dividing each frame image, and outputs a first patch-level feature , where, represents the m patch-level features of the first frame image, represent the m patch-level features of the nth frame image; m represents the number of non-overlapping patches into which each frame image is divided, represent the dimension of the features.

[0061] S3-2. Input the first frame-level feature into the temporal modeling module to output the second frame-level feature; input the first patch-level feature into the spatial modeling module to output the second patch-level feature.

[0062] Preferably, the temporal modeling module is composed of a stack of several Transformer network layers. In this embodiment, the number of Transformer network layers in the temporal modeling module is 4.

[0063] The temporal modeling module outputs the corresponding second frame-level feature after temporal modeling according to the first frame-level feature .

[0064] Preferably, the spatial modeling module is composed of a stack of several Transformer network layers. In this embodiment, the number of Transformer network layers in the spatial modeling module is 4.

[0065] The spatial modeling module outputs the corresponding second patch-level feature after spatial modeling according to the first patch-level feature .

[0066] S3-3. Use the patch-level clustering module to cluster the second patch-level features to obtain the third patch-level features.

[0067] Preferably, the patch-level clustering module clusters the second patch-level features based on the Kmeans clustering algorithm.

[0068] The patch-level clustering module clusters the second patch-level features After clustering, the m patches in each frame image generate o cluster centers, and after splicing, the third patch-level features are obtained , where represents the patch-level features of o clusters of the first frame image, represents the patch-level features of o clusters of the nth frame image; o represents the number of cluster centers.

[0069] S3-4. Input the second frame-level feature and the third patch-level feature into the dual-stream cross-attention module for multi-granularity interaction to output the target frame-level feature and the target patch-level feature.

[0070] Specifically, the dual-stream cross-attention module includes a patch-level cross-attention mechanism and a frame-level cross-attention mechanism, and is used for the second frame-level feature and the third patch-level feature ​​Interact to generate target frame-level features and target patch-level features .

[0071] Both the patch-level cross-attention mechanism and the frame-level cross-attention mechanism include a feed-forward linear layer and multi-head attention. The feed-forward linear layer is used to adjust the dimension and distribution of the input data, convert the input data into a vector representation, and provide a suitable input for the subsequent multi-head attention mechanism. The multi-head attention is used to enhance the feature representation of the input data.

[0072] Specifically, in the patch-level cross-attention mechanism, the third patch-level feature is used as the input query , and the second frame-level feature is used as the input key and value , and the target patch-level feature is output.

[0073] In the frame-level cross-attention mechanism, the second frame-level feature is used as the input query , and the third patch-level feature is used as the input key and value , and the target frame-level feature is output.

[0074] S3-5. Input the target frame-level feature and the target patch-level feature into a multi-layer perceptron to output the video feature of the candidate video.

[0075] Specifically, this embodiment constructs a multi-layer perceptron based on a fully connected network. The multi-layer perceptron outputs the video feature by fusing the target frame-level feature and the target patch-level feature .

[0076] S4. After embedding the query text into a preset prompt statement, input it together with the video features of all candidate videos into a large language model to output the similarity between the query text and each candidate video.

[0077] The prompt statement is used to obtain the input of the large language model according to the query text, guide the large language model to calculate the similarity of each video in the candidate video set in descending order according to the relevance between the candidate video and the query text, and limit the range of the similarity to between 0 and 1, and output it in the form of a list.

[0078] In this embodiment, the preset prompt statement is set as: Please compute the similarities (ranging from 0 to 1) of the videos in the video collection in descending order according to their relevance to the sentence " " and output them as a list;

[0079] After obtaining the query text, it is embedded into the " " position of the preset prompt statement to obtain the input of the large language model.

[0080] Preferably, the large language model adopted in this embodiment is the LLaMA model.

[0081] Specifically, the large language model performs semantic-driven reasoning based on the prompt statement embedded with the query text and the video features of all candidate videos and outputs the similarity between the query text and each candidate video.

[0082] S5. Output the video retrieval result according to the similarity between the query text and each candidate video.

[0083] Preferably, when training the text-video retrieval model, a masking and reconstruction module is further included.

[0084] The masking and reconstruction module is used in the model training stage and includes an importance prediction layer, a conditional masking layer, and a feature reconstruction layer.

[0085] Specifically, the importance prediction layer includes a first linear layer, a RELU activation layer, a second linear layer, and a Softmax layer, and is used to generate a non-negative importance score for each patch-level feature in the second patch-level feature . Among them, the first linear layer is used to projectively transform the second patch-level feature from d dimensions to 2*d dimensions; the RELU activation layer is used to maintain the non-negativity of the feature and add non-linear factors to increase the non-linear expression ability of the importance prediction layer; the second linear layer is used to map the activated feature from 2*d dimensions to 1 dimension; the Softmax layer is used to map the 1-dimensional feature to the interval, indicating the importance degree of the feature. The formula representation of the importance prediction layer is:

[0086]

[0087] ;

[0088] where and respectively represent the projection parameters of the first linear layer and the second linear layer, is the activation function, and is the Softmax function; represents the second patch-level feature, represents the importance score of the second patch-level feature.

[0089] The conditional masking layer performs conditional masking on the second patch-level feature based on the non-negative importance score specifically: using random noise to mask the K features with the highest importance scores in each frame of the second patch-level feature to obtain the masked patch-level feature , where, represents the feature after masking the j-th patch in the first frame of the image, represents the feature after masking the k-th patch in the i-th frame of the image.

[0090] The feature reconstruction layer reconstructs the masked patch-level feature according to the second frame-level feature Specifically, the feature reconstruction layer is composed of the encoding layers of several Transformer network layers stacked, and the core of the encoding layer of each Transformer network layer is the multi-head attention layer. In the multi-head attention layer, the masked patch-level feature is used as the input query

[0091] , the second frame-level feature is used as the input key and value . The output of the feature reconstruction module is the reconstructed patch-level feature . .

[0092] The formula of the feature reconstruction layer is expressed as:

[0093] ;

[0094] where, represents the feature reconstruction module, , and are the linear transformation parameters of the query, key, and value respectively; represents the reconstructed patch-level feature, represents the masked patch-level feature, represents the second frame-level feature.

[0095] The masking and reconstruction module reconstructs the masked patch-level feature using the second frame-level feature to enhance the representation ability of the second frame-level feature during training.

[0096] The process of training the text-video retrieval model in this embodiment includes:

[0097] (1) Obtain the video-text pair data in the training set, initialize the model parameters, and set the learning rate, decay strategy, and number of iterations.

[0098] Specifically, this embodiment uses the official data division to obtain the training set; uses Initialize the video encoder, Initialize the large language model, The method initializes the parameters of the remaining modules based on the normal distribution; the learning rate is set to ; Select the cosine decay strategy; the number of iterations is set to 5.

[0099] (2) Batch input n video-text pairs into the pre-constructed text-video retrieval model. For each video, fixed-length frame-level features and patch-level features are extracted. The video features extracted by the video feature extractor and the prompt statement after embedding the query text are jointly input into the large language model, and a similarity matrix with an output size of is output.

[0100] (3) Use the contrastive learning loss and the reconstruction loss function to construct the total loss function to supervise the training of the text-video retrieval model. The model is iteratively optimized through backpropagation according to the pre-set learning rate to obtain the trained text-video retrieval model.

[0101] Specifically, the contrastive learning loss function obtains the loss value based on the similarity matrix output by the large language model. Through backpropagation, the model parameters are updated to increase the similarity between positive samples and decrease the similarity between negative samples. Its formula is expressed as:

[0102] ;

[0103] Among them, represents the contrastive learning loss function, represents the batch size, represents the similarity output by the large language model, and respectively represent the th text and the th video in the batch, and respectively represent the th text and the th video in the batch; represents the temperature coefficient.

[0104] The reconstruction loss function calculates the reconstructed patch-level feature and the second patch-level feature The difference between them generates a loss value. Through backpropagation, the model parameters are updated to reduce the difference between the two sets of patch-level features and maintain the consistency between the reconstructed patch-level features and the second patch-level features. Its formula is expressed as:

[0105] ;

[0106] Among them, represents the reconstruction loss function.

[0107] Therefore, the total loss function of the training process is expressed as: , where represents the balance parameter.

[0108] Preferably, when training the text-video retrieval model, the video encoder of the video feature extractor and the parameters of the large language model are frozen, and only the parameters of the video feature extractor are iteratively updated, which reduces the training cost to a certain extent.

[0109] In summary, for the video retrieval method based on large language model semantic drive described in the present invention, the frame-level features and patch-level features of the candidate video are extracted, and the frame-level features are temporally modeled to capture the changes in the video content on the time axis, which helps to extract the temporal dynamics of the video and understand the dynamic content of the video. The patch-level features are spatially modeled and clustered to capture the spatial relationship between different parts of the video frame, which helps to extract the spatial structure distribution of the video frame; further, based on the cross-attention mechanism, the frame-level features and patch-level features are interacted at multiple granularities, which can integrate the information of the video in both the temporal and spatial dimensions, construct video features containing dynamic changes and spatial details, retain the key information of the video to the greatest extent, and provide rich and accurate video information for the text-video retrieval task to improve the accuracy and efficiency of text-video retrieval. Moreover, the video features extracted by the present invention and the prompt statement embedded with the query text are input into the large language model for semantic reasoning and cross-modal semantic matching. By using the powerful semantic reasoning ability of the large language model, more reliable and interpretable cross-modal semantic relationships are obtained, which conforms to the cognitive behavior of human retrieval, enhances the interpretability of the model, and improves the accuracy and efficiency of cross-modal text-video retrieval. Embodiment 2

[0110] Based on the video retrieval method based on large language model semantic drive described in Embodiment 1, this embodiment provides a video retrieval system based on large language model semantic drive, including:

[0111] A retrieval model construction module for constructing a text-video retrieval model; the text-video retrieval model includes a video feature extractor and a large language model; the video feature extractor includes a video encoder, a temporal modeling module, a spatial modeling module, a patch-level clustering module, a two-stream cross-attention module, and a multi-layer perceptron;

[0112] A feature extraction module for inputting a candidate video set into the text-video retrieval model and obtaining video features of each candidate video through the video feature extractor, including:

[0113] The candidate video passes through the video encoder to output a first frame-level feature and a first patch-level feature;

[0114] Input the first frame-level feature into the temporal modeling module to output a second frame-level feature; input the first patch-level feature into the spatial modeling module to output a second patch-level feature;

[0115] Use the patch-level clustering module to cluster the second patch-level feature to obtain a third patch-level feature;

[0116] Input the second frame-level feature and the third patch-level feature into the two-stream cross-attention module for interaction to output a target frame-level feature and a target patch-level feature;

[0117] Input the target frame-level feature and the target patch-level feature into the multi-layer perceptron to output the video features of the candidate video;

[0118] A similarity acquisition module for embedding the query text into a preset prompt statement and then inputting it together with the video features of all candidate videos into the large language model to output the similarity between the query text and each candidate video;

[0119] A retrieval output module for outputting a video retrieval result based on the similarity between the query text and each candidate video.

[0120] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0121] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0122] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0124] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to exhaustively list all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A video retrieval method based on semantics driven by a large language model, characterized in that: include: Constructing a text video retrieval model; the text video retrieval model includes a video feature extractor and a large language model; the video feature extractor includes a video encoder, a temporal modeling module, a spatial modeling module, a patch-level clustering module, a dual-stream cross attention module and a multi-layer perceptron; The candidate video set is input into the text video retrieval model, and the video features of each candidate video are obtained through the video feature extractor, including: Uniformly extract video frame images from the candidate video, and divide each video frame image into patches of the same size and without overlap to obtain video frame image patches; pass the video frame image and the video frame image patches through a video encoder to output a first frame-level feature and a first patch-level feature respectively; Input the first frame-level feature into the temporal modeling module and output the second frame-level feature; input the first patch-level feature into the spatial modeling module and output the second patch-level feature; Clustering the second patch-level features using a patch-level clustering module to obtain a third patch-level feature; Inputting the second frame-level feature and the third patch-level feature into a two-stream cross-attention module, wherein the two-stream cross-attention module includes a patch-level cross-attention mechanism and a frame-level cross-attention mechanism; The patch-level cross-attention mechanism takes the third patch-level feature as input query, the second frame-level feature as input key and value, and outputs the target patch-level feature; The frame-level cross-attention mechanism takes the second frame-level features as input query, the third patch-level features as input keys and values, and outputs the target frame-level features; Input the target frame-level features and the target patch-level features into a multi-layer perceptron, and output the video features of the candidate video; After embedding the query text into the preset prompt sentence, it is input into the large language model together with the video features of all candidate videos, and the similarity between the query text and each candidate video is output; Output video retrieval results based on the similarity between the query text and each candidate video.

2. The video retrieval method based on semantics driven by a large language model according to claim 1, characterized in that: The video encoder is built based on the image encoder of the CLIP model.

3. The video retrieval method based on semantics driven by a large language model according to claim 1, characterized in that: The temporal modeling module and the spatial modeling module are both composed of a stack of several Transformer network layers.

4. The video retrieval method based on semantics driven by a large language model according to claim 1, characterized in that: The large language model is an LLaMA model.

5. The video retrieval method based on semantics driven by a large language model according to claim 1, characterized in that: When training the text video retrieval model, the parameters of the video encoder and the large language model of the video feature extractor are frozen.

6. The video retrieval method based on semantics driven by a large language model according to claim 1, characterized in that: When training the text video retrieval model, a masking and reconstruction module is also included; The mask reconstruction module includes an importance prediction layer, a conditional mask layer and a feature reconstruction layer; The importance prediction layer is used to generate the importance score of each patch-level feature in the second patch-level feature, and includes a first linear layer, a RELU activation layer, a second linear layer and a Softmax layer; wherein the first linear layer is used to map the second patch-level feature from d dimensions to 2*d dimensions for projection transformation; the RELU activation layer is used to maintain the non-negativity of the feature and add nonlinear factors; the second linear layer is used to map the activated feature from 2*d dimensions to 1 dimension; the Softmax layer is used to map the 1-dimensional feature to the (0,1) interval to generate the importance score of each patch-level feature; The formula of the importance prediction layer is expressed as: ; in, and denote the projection parameters of the first and second linear layers respectively, for Activation function, is the Softmax function; represents the second patch-level feature, represents the importance score of the second patch-level feature; The conditional masking layer is used to mask the K features with the highest importance scores of each frame image in the second patch-level features using random noise to obtain the masked patch-level features; The feature reconstruction layer is composed of a stack of encoding layers of several Transformer network layers; in the multi-head attention layer of the encoding layer of each Transformer network layer, the masked patch-level features are used as input queries, the second frame-level features are used as input keys and values, the masked patch-level features are reconstructed, and the reconstructed patch-level features are output, which is expressed as follows: ; in, represents the feature reconstruction module, , and are the linear transformation parameters for query, key, and value, respectively; represents the reconstructed patch-level features, represents the patch-level features after masking, Represents the second frame level feature.

7. The video retrieval method based on semantics driven by a large language model according to claim 6, characterized in that: The loss function of the text video retrieval model includes a contrastive learning loss function and a reconstruction loss function; the reconstruction loss function is used to calculate the difference between the reconstructed patch-level feature and the second patch-level feature.

8. The video retrieval method based on semantics driven by a large language model according to claim 1, characterized in that: The patch-level clustering module clusters the second patch-level features based on a Kmeans clustering algorithm.

9. A video retrieval system based on semantics driven by a large language model, characterized in that: include: A retrieval model building module, used to build a text video retrieval model; the text video retrieval model includes a video feature extractor and a large language model; the video feature extractor includes a video encoder, a temporal modeling module, a spatial modeling module, a patch-level clustering module, a dual-stream cross attention module and a multi-layer perceptron; The feature extraction module is used to input the candidate video set into the text video retrieval model and obtain the video features of each candidate video through the video feature extractor, including: Uniformly extract video frame images from the candidate video, and divide each video frame image into patches of the same size and without overlap to obtain video frame image patches; pass the video frame image and the video frame image patches through a video encoder to output a first frame-level feature and a first patch-level feature respectively; Input the first frame-level feature into the temporal modeling module and output the second frame-level feature; input the first patch-level feature into the spatial modeling module and output the second patch-level feature; Clustering the second patch-level features using a patch-level clustering module to obtain a third patch-level feature; Inputting the second frame-level feature and the third patch-level feature into a two-stream cross-attention module, wherein the two-stream cross-attention module includes a patch-level cross-attention mechanism and a frame-level cross-attention mechanism; The patch-level cross-attention mechanism takes the third patch-level feature as input query, the second frame-level feature as input key and value, and outputs the target patch-level feature; The frame-level cross-attention mechanism takes the second frame-level features as input query, the third patch-level features as input keys and values, and outputs the target frame-level features; Input the target frame-level features and the target patch-level features into a multi-layer perceptron, and output the video features of the candidate video; A similarity acquisition module is used to embed the query text into a preset prompt sentence, input the query text and the video features of all candidate videos into the large language model, and output the similarity between the query text and each candidate video; The retrieval output module is used to output the video retrieval results based on the similarity between the query text and each candidate video.

Citation Information

Patent Citations

  • Performing multiple queries within a robust video search and retrieval mechanism

    CN108780457A

  • Multi-granularity comparative learning collaborative generation method for long-span video questions and answers

    CN118170885A