Multi-modal video retrieval method and system based on low-rank fine-grained prompt
By introducing a low-rank fine-grained prompt in multimodal video retrieval, the problem of unscalable modal quantity and type in the prior art is solved, and efficient and flexible multimodal video retrieval effect is achieved, and the calculation cost is reduced.
Patent Information
- Application Number
- CN202510542738.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
Existing multimodal learning technologies are difficult to adapt to different tasks, have high computational costs, and are difficult to achieve scalability of the number and types of modals, especially in the field of multimodal video retrieval.
A multimodal video retrieval method based on low-rank fine-grained hints is adopted. By introducing a prompt update module in the video characterization generation module, each modal prompt is generated in the fine-grained hints, conducting deep interactions, and fixing the parameters of the multimodal model in pre-training and fine-tuning training, and independently fine-tuning the prompt update module.
It realizes efficient and flexible expansion of the number and types of modals, improves the effect of multimodal video retrieval, reduces calculation costs, and improves the adaptability of the model.
Smart Images

Figure CN120067390A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimodal video retrieval, and in particular, to a multimodal video retrieval method and system based on low-rank fine-grained prompts. Background Art
[0002] Multimodal tasks refer to tasks involving multiple modalities of information (such as text, images, and speech), which require the use of technologies from various fields such as computer vision, natural language processing, and audio processing. To enable artificial intelligence to better solve real-world problems, it is necessary to enhance its capabilities in multimodal tasks, that is, to understand, process, and correlate information from multiple modalities. In recent years, multimodal learning has gradually shown its great potential and importance, playing an indispensable role in fields such as gaming, robotics, and education.
[0003] Existing research on multimodal learning technologies can be divided into the construction of multimodal large language models and the research on lightweight methods.
[0004] The emergence of large language models and the limitations of traditional large language models in processing non-text data have made the construction of multimodal large language models a research hotspot. Such methods usually solve certain specific multimodal tasks by designing elaborate but large and complex model structures, but they are often difficult to adapt to different tasks, have high computational costs, and are difficult to fine-tune. The research on lightweight methods is represented by prompt learning. Prompts adjust a small number of learnable parameters to make downstream tasks adapt to pre-trained models. The lightweight and efficient advantages of prompt learning have made it popularized in the multimodal field. However, current multimodal prompt learning methods basically only perform shallow interactions between two specific modalities and are limited by the model structure, making it difficult to be extended to other numbers and types of modalities. Summary of the Invention
[0005] To overcome the limitations of existing technologies that are restricted to fixed numbers and types of modalities and do not perform deep interactions on multimodal information, the present invention provides a multimodal video retrieval method and system based on low-rank fine-grained prompts to achieve multimodal prompt learning with expandable numbers and types of modalities.
[0006] The specific technical solutions adopted by the present invention are as follows:
[0007] In a first aspect, the present invention proposes a multimodal video retrieval method based on low-rank fine-grained prompts for matching videos and captions. The multimodal video retrieval method includes:
[0008] (1) Pre-train a multimodal model including a video representation generation module and a caption representation generation module;
[0009] (2)In each of the first N - 1 layers of the encoder in the video representation generation module, a prompt update module is introduced to generate modality - specific prompts for each layer in a fine - grained prompt manner; where N represents the total number of encoder layers;
[0010] During the process of fine - tuning and training the prompt update module, the modality - specific video features are concatenated with the corresponding modality - specific prompts to obtain a multi - modality input. The multi - modality input enters the current encoder layer for processing. At the same time, the modality - specific prompts input to the current encoder layer are concatenated with each other and enter the prompt update module of the same layer to update the modality - specific prompts. The updated modality - specific prompts replace the corresponding modality - specific prompts in the output of the current encoder layer, and are concatenated with the modality - specific video feature parts in the output result of the current encoder layer to form the multi - modality input for the next encoder layer. The modality - specific prompts entering the last encoder layer do not need to be updated anymore. The last token in the output result of the last encoder layer is converted into a video representation for matching the caption representation;
[0011] (3)Fix the modality - specific prompts obtained through fine - tuning. In the video representation generation module, the modality - specific video feature parts in the output of the previous encoder are concatenated with the modality - specific prompts of the current layer as the multi - modality input of the current layer. The video representation finally obtained by the video representation generation module is used to match the caption representation to achieve the multi - modality video retrieval task.
[0012] Furthermore, in step (1), the pre - training of the multi - modality model includes:
[0013] Obtain the video representation of the candidate video using the video representation generation module;
[0014] Obtain the caption representation of the candidate caption using the caption representation generation module;
[0015] Based on the similarity between the caption representation and the video representation, use the contrastive learning method to pre - train the video representation generation module and the caption representation generation module.
[0016] Furthermore, the video representation generation module includes a multi - modality feature extraction network, a dimension adjustment layer, a backbone network composed of N layers of encoders, and a fully - connected layer. During the pre - training process of the multi - modality model, the multi - modality initial video features generated by the multi - modality feature extraction network are unified in dimension through the dimension adjustment layer, and then concatenated and input into the backbone network. The last token in the output result of the backbone network passes through the fully - connected layer to obtain the video representation.
[0017] Furthermore, the multi - modality initial video features include at least two modality features among the audio modality, the visual modality, and the optical character recognition modality.
[0018] Furthermore, during the process of fine - tuning and training the prompt update module, the modality - specific prompts input to the first encoder layer are randomly initialized.
[0019] Furthermore, the multi-modal input is represented as:
[0020] ;
[0021] ;
[0022] ;
[0023] wherein, represents the multi-modal input of the (i + 1)-th encoder layer, represents the modality of the i-th encoder layer hint, is the total number of modalities, represents the modality of the i-th encoder layer hint, respectively represent the video features of the modality input to the i-th encoder layer and the (i + 1)-th encoder layer, represents the i-th encoder layer, represents the hint update module.
[0024] Furthermore, the hint update module generates the hints of each modality of the corresponding layer in the form of fine-grained hints, represented as:
[0025] ;
[0026] ;
[0027] ;
[0028] wherein, represents the token update weight of the in the i-th encoder layer, represents the -th token, represents element-wise multiplication, represents the relationship between the -th token in and different tokens in the remaining n - 1 modality hints, represents the -order matrix representing the importance of, , represents the tensor multiplication of multi-matrices, represents the hint of the modality in the input of the -th encoder layer, represents summing over dimensions of the high-order matrix other than e.
[0029] Further, decompose into tensors of rank 1 to approximate the fine-grained prompt, and update the weights which is simplified to:
[0030] ;
[0031] wherein, represents the number of ranks, represents the th tensor of dimension for the corresponding modality, represents the th tensor of dimension and represent dimensions, and the superscript T represents transpose, represents summation over dimension L, represents the prompt for the modality of the i-th encoder layer.
[0032] Furthermore, the subtitle representation generation module adopts a pre-trained BERT-based encoder structure.
[0033] In a second aspect, the present invention proposes a multi-modal video retrieval system based on low-rank fine-grained prompts for implementing the above-mentioned multi-modal video retrieval method based on low-rank fine-grained prompts.
[0034] Compared with the prior art, the beneficial effects of the present invention are:
[0035] The multi-modal video retrieval method based on low-rank fine-grained prompts proposed by the present invention introduces fine-grained prompts in the encoder layer to achieve deep interaction of modal prompts. The prompt update module is independent of the model backbone. First, a multi-modal model is pre-trained, and then the prompt update module is fine-tuned under the condition of fixing the parameters of the pre-trained multi-modal model. It does not limit the number and type of modalities. The modal prompts obtained after fine-tuning training are directly used for the multi-modal video retrieval task, realizing efficient and flexible expansion of multi-modal learning in terms of the number and type of modalities, and improving the effect of multi-modal video retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a framework diagram of the multi-modal video retrieval method based on low-rank fine-grained prompts of the present invention;
[0037] Figure 2 is a schematic diagram of the fine-grained prompt generation process of the present invention;
[0038] Figure 3 is a schematic diagram of the low-rank fine-grained prompt generation process of the present invention. Detailed Implementation Modes
[0039] The present invention will be further described and explained below in conjunction with the detailed implementation modes. The embodiments are only demonstrations of the disclosed content and do not delimit the scope of limitation. Without conflict, the technical features of each implementation mode in the present invention can be combined accordingly.
[0040] The accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0041] As Figure 1 shown, a multi-modal video retrieval method based on low-rank fine-grained prompts proposed by the present invention mainly includes the following steps:
[0042] S1. Pre-training stage
[0043] Pre-train a multi-modal model including a video representation generation module and a caption representation generation module;
[0044] S2. Fine-tuning training stage
[0045] Introduce a prompt update module into each layer of the first N - 1 layer encoders in the video representation generation module to generate various modal prompts of the corresponding layer in a fine-grained prompt manner; where N represents the total number of encoder layers;
[0046] S3. Inference stage
[0047] Fix the various modal prompts obtained by fine-tuning, and splice the various modal video feature parts in the output result of the previous encoder with the various modal prompts of the current layer as the multi-modal input of the current layer in the video representation generation module; the video representation finally obtained by the video representation generation module is used to match the caption representation to achieve the multi-modal video retrieval task.
[0048] In a specific implementation of the present invention, an optional manner of the multi-modal model is as follows:
[0049] The video representation generation module in the multi-modal model includes a multi-modal feature extraction network, a dimension adjustment layer, a backbone network composed of N layers of encoders, and a fully connected layer. Preferably, the backbone network composed of N layers of encoders adopts an N-layer Transformer encoder structure. The backbone model composed of Transformer encoders is a general model with powerful understanding ability and generalization performance, and can extract rich semantic information from the input. During the pre-training process of the multi-modal model, the multi-modal initial video features generated by the multi-modal feature extraction network are unified in dimension through the dimension adjustment layer, and then concatenated and input into the backbone network. The last token of the output result of the backbone network passes through the fully connected layer to obtain the video representation.
[0050] For example, the multi-modal feature extraction network is a feature extractor for extracting the initial features of different modalities of the video. Taking a video containing a person as an example, the ResNet50 model pre-trained on the VGGFace2 dataset is used as the face feature extractor to extract the face features in the original video sample as the initial visual modality features; the VGGish model pre-trained on the YouTube-8m dataset is used as the audio feature extractor to extract the audio features in the original video sample to obtain the initial audio modality features; the I3D model pre-trained on the Kinetics dataset is used as the action feature extractor to extract the action features in the original video sample as the initial action modality features; however, it is not limited to the above pre-trained models and modality features. Common modalities include optical character recognition modality, visual (such as action, scene, appearance, face, etc.) modality, audio modality, etc. The definitions and extraction processes of each modality are well-known common knowledge in the art and will not be elaborated here.
[0051] The multi-modal video features obtained by the above feature extractors also need to unify the feature dimensions for convenient subsequent processing. An optional acquisition method is as follows:
[0052] Input the initial features of each modality into a one-dimensional convolutional layer, where and respectively represent the sequence length and feature dimension of the initial features of modality Initial features In the one-dimensional convolutional layer, its feature dimension is converted from to a unified target dimension . In this embodiment, the size of the convolutional kernel is set to 1 to ensure that only the feature dimension is transformed without changing the sequence length. The expression is The size of the feature matrix finally output by the one-dimensional convolutional layer is , thus completing the dimension alignment of the features of each modality. The structure for unifying the feature dimensions is not limited to the above one-dimensional convolutional layer, and can also be a fully connected layer or other linear projection methods.
[0053] The subtitle representation generation module in the multi-modal model can adopt a pre-trained BERT-based encoder structure. For example, first encode the subtitle through BERT, and then map the encoded features through a fully connected layer to obtain a subtitle representation with the same dimension as the video representation for subsequent similarity calculation.
[0054] In the above S1, the original video sample and text annotation are obtained, and the text annotation refers to the subtitle of the original video sample; in this embodiment, the pre-training is carried out in a contrastive learning manner, and the calculation formula of the loss function is:
[0055]
[0056] where is the bidirectional maximum margin ranking loss function; is the number of samples; is the video representation and the subtitle representation (negative sample pair) between the similarity scores; is the video representation and its corresponding correct subtitle representation (positive sample pair) between the similarity scores; represents the margin, which is a hyperparameter used to control the minimum gap between the similarity of the positive sample pair and the similarity of the negative sample pair. The above video representation refers to the result generated by the above video representation generation module, and the subtitle representation refers to the result generated by the subtitle text through the above subtitle representation generation module, which has been introduced above. The principle of contrastive learning belongs to the common knowledge in the art and will not be elaborated here. Those skilled in the art can also introduce other training methods and loss functions suitable for multi-modal video retrieval tasks.
[0057] In a specific implementation of the present invention, a prompt update module is introduced into each of the first N - 1 layers of the encoder in the video representation generation module. The prompt update module is independent of the backbone model to enhance flexibility. During the process of fine-tuning and training the prompt update module, the parameters of the multi-modal feature extraction network pre-trained in the first stage, the backbone network composed of N layers of encoders, and the caption representation generation module are fixed and not updated. The prompt update module, as well as the dimension adjustment layer and the fully connected layer in the video representation generation module, are fine-tuned and trained. The multi-modal video features of each modality are concatenated with the corresponding modality prompts to obtain a multi-modal input. The multi-modal input enters the current encoder layer for processing. At the same time, the prompts of each modality input into the current encoder layer are concatenated with each other and enter the prompt update module of the same layer to update the prompts of each modality. The updated prompts of each modality replace the corresponding modality prompts in the output of the current encoder layer and are concatenated with the multi-modal video feature parts in the output result of the current encoder layer to form the multi-modal input of the next encoder layer. The prompts of each modality entering the last encoder layer do not need to be updated anymore. The last token of the output result of the last encoder layer is converted into a video representation for matching the caption representation.
[0058] When fine-tuning and training the prompt update module, the same loss function as in the first stage is adopted, and the contrastive learning method is used to fine-tune the prompt update module to generate the prompts of each modality of the corresponding layer in the form of fine-grained prompts.
[0059] An optional way to obtain the multi-modal input is as follows:
[0060] The prompts of each modality are concatenated in front of the respective modality features, and then concatenated with each other to obtain the multi-modal input, which is expressed as:
[0061]
[0062] Among them, represents the multi-modal input; is the total number of modalities of the original video sample; represents the concatenation operation, represents the prompt of modality n, and respectively represent The sequence length and feature dimension of. The prompts of each modality input into the first encoder layer are randomly initialized.
[0063] Furthermore, the present invention generates the prompts of each modality of the corresponding layer in the form of fine-grained prompts. An optional way for the multi-modal input to enter the encoder layer of the backbone network for processing is as follows:
[0064] The calculation formula is:
[0065]
[0066]
[0067]
[0068] Among them, represents the th modality in the multimodal input of the th encoder layer; is the total number of modalities of the original video sample; represents the hint update module; represents the th layer of the Transformer encoder layer; represents the multimodal input of the
[0069] An optional way of the hint update module is as follows:
[0070] Capture the interaction between modality hints through the fine-grained hint method, supporting flexible expansion of the modality type and quantity; as Figure 2 shown, an optional way of the fine-grained hint method is as follows:
[0071] Each modality hint is updated through the update weight within the hint update module. Each update weight is consistent with the hint size, ensuring that each element in the hint is updated by the corresponding update weight value to achieve fine-grained hint update.
[0072] The calculation formula is:
[0073]
[0074]
[0075] Among them, represents 's th token; is 's target update weight, is 's token update weight; represents element-wise multiplication; is 's sequence length.
[0076] An optional way of the process of obtaining the update weight is as follows:
[0077] First, construct a high-order matrix that integrates the modality hint information, and the calculation formula is:
[0078]
[0079]
[0080] Among them, is defined as the element-wise multiplication of a sequence of tensors; and is defined as the "tensor multiplication" of multiple matrices.
[0081] Next, use to represent , there is:
[0082]
[0083] Among them, represents the element-wise multiplication of each token in ; represents that it only performs operations with when , represents the relationship between the -th token in
[0084] Finally, the target update weight of
[0085]
[0086] Among them, represents the -order matrix representing the importance of ; represents the sum of the dimensions of the high-order matrix except for e.
[0087] In a specific implementation of the present invention, through the low-rank fine-grained prompting method, low-rank decomposition is used to approximate the fine-grained prompting, reducing the complexity of the fine-grained prompting that grows exponentially with the number of modalities to a linear level.
[0088] As Figure 3 shown, an optional way of the low-rank fine-grained prompting method is as follows:
[0089] Apply low-rank decomposition to reduce the -induced complexity in the fine-grained prompting method. can be decomposed into tensors of rank 1:
[0090]
[0091] Among them, represents the tensor; denote tensor denote the tensor after transposition denote the number of ranks
[0092] Therefore, the update weight formula in the fine-grained prompting method can be rewritten as follows
[0093]
[0094]
[0095] The present invention pre-trains the overall framework using a high-resource dataset, then freezes most of the parameters of the framework, and fine-tunes the prompt update module using a low-resource dataset. During the fine-tuning training, using the corresponding loss function and adopting the gradient descent learning method, the trainable parameters in the framework (including the dimension adjustment layer, the prompt update module, and the fully connected layer) are fine-tuned so that the model can complete the multi-modal video retrieval task
[0096] The above method will be applied to the following embodiments to demonstrate the technical effects of the present invention. The specific steps in the embodiments will not be elaborated
[0097] The present invention conducts experiments on two multi-modal datasets, MSR-VTT and MSVD. MSR-VTT (a high-resource dataset for pre-training) and MSVD are used to evaluate the multi-modal video retrieval task
[0098] To objectively evaluate the performance of the present invention, in the selected dataset, the present invention uses evaluation criteria such as Rank@K (R@K), Mean Rank (MeanR), and Median Rank (MedR) to evaluate the video retrieval effect. Among them, except for R@K, the lower the metric value, the better the performance of the model
[0099] To verify the effectiveness of the present invention, this implementation conducts a comparative experiment with the MMT model in 2 to 5 modalities. The tasks are specifically divided into two types: (1) Given a text query, retrieve the most relevant video from the video collection (Text to Video); (2) According to the given video content, retrieve the most relevant text description (Video to Text)
[0100] To verify the effectiveness of the present invention in reducing complexity, this implementation statistically analyzes the actual floating-point operations (FLOPs) of the prompt update module and the memory space requirements of the entire model of the fine-grained prompting method and the low-rank fine-grained prompting method in the present invention in 4 to 5 modalities on the MSVD dataset
[0101] According to the steps described in the specific implementation manner, the experimental results obtained are shown in Tables 1 to 3. In the present invention, the fine-grained prompt method is denoted as Fine-grained Prompt (FP), and the low-rank fine-grained prompt method is denoted as Low-rankFine-grained Prompt (LFP).
[0102] Table 1: Pretraining results of multimodal video retrieval obtained by the present invention for the MSR-VTT dataset
[0103]
[0104] Table 2: Evaluation results of multimodal video retrieval obtained by the present invention for the MSVD dataset
[0105]
[0106] Table 3: Statistical results of computing resources for the MSVD dataset according to the present invention
[0107]
[0108] As can be seen from Tables 1 and 2, the method of the present invention outperforms the baseline model in all metrics under all modal configurations, highlighting the effectiveness of the method of the present invention in the multimodal video retrieval task and its adaptability to different numbers of modalities.
[0109] As can be seen from Table 3, the actual floating-point operation count and storage space requirements under the fine-grained prompt method are significantly higher than those of the low-rank fine-grained prompt method. The low-rank fine-grained prompt method in the present invention has an obvious advantage in terms of computational cost.
[0110] Combining Tables 2 and 3, it can be found that although the low-rank fine-grained prompt method in the present invention is slightly worse than the fine-grained prompt method in the present invention in some metrics, it has an obvious advantage in terms of computational cost, which demonstrates the truly implementable modal scalability of the low-rank fine-grained prompt method in the present invention in practical applications.
[0111] In this embodiment, a modal-scalable low-rank fine-grained prompt system is also provided, which is used to implement the above embodiment. The following terms "module", "unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the system described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible.
[0112] The above system includes:
[0113] A multimodal model, which includes a video representation generation module and a caption representation generation module;
[0114] A hint update module, which is introduced into each layer of the first N-1 encoder layers in the video representation generation module;
[0115] A fine-tuning training module, which is used to fine-tune the training prompt update module based on the pre-trained multimodal model, and generate prompts for each modality of the corresponding layer in a fine-grained prompt manner;
[0116] The multimodal video retrieval module fixes and fine-tunes the modal prompts obtained, and in the video representation generation module, concatenates the modal video feature parts of the previous encoder output results with the modal prompts of the current layer as the multimodal input of the current layer; the video representation finally obtained by the video representation generation module is used to match the subtitle representation to achieve the multimodal video retrieval task.
[0117] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be repeated here. The system embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0118] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, the corresponding computer program instructions in the non-volatile memory are read into the memory by the processor of any device with data processing capabilities and run.
[0119] The above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or associated with the contents disclosed by a person skilled in the art should be considered as the protection scope of the present invention.
Claims
1. A multimodal video retrieval method based on low-rank fine-grained cues for matching videos and subtitles, characterized in that The multimodal video retrieval method comprises: (1) Pre-train a multimodal model that includes a video representation generation module and a subtitle representation generation module; (2) In the video representation generation module, a prompt update module is introduced into each layer of the first N-1 layers of the encoder to generate the prompts of each modality of the corresponding layer in the form of fine-grained prompts; where N represents the total number of encoder layers; During the fine-tuning training prompt update module, each modal video feature is spliced with the corresponding modal prompt to obtain a multimodal input, and the multimodal input enters the current encoder layer for processing. At the same time, the modal prompts input to the current encoder layer are spliced with each other and enter the prompt update module of the same layer to update each modal prompt. The updated modal prompts replace the corresponding modal prompts in the output of the current encoder layer, and are spliced with the modal video feature parts in the output result of the current encoder layer to form the multimodal input of the next encoder layer; the modal prompts entering the last encoder layer do not need to be updated, and the last token of the output result of the last encoder layer is converted into a video representation for matching the subtitle representation; (3) Fix the fine-tuned modal prompts and concatenate the modal video feature parts of the previous encoder output results with the modal prompts of the current layer in the video representation generation module as the multimodal input of the current layer; the video representation finally obtained by the video representation generation module is used to match the subtitle representation to achieve the multimodal video retrieval task.
2. The multimodal video retrieval method based on low-rank fine-grained hints according to claim 1, characterized in that: In step (1), the pre-training of the multimodal model includes: Obtaining video representations of candidate videos using a video representation generation module; Obtaining subtitle representations of candidate subtitles using a subtitle representation generation module; Based on the similarity between the subtitle representation and the video representation, the contrastive learning method is used to pre-train the video representation generation module and the subtitle representation generation module.
3. The multimodal video retrieval method based on low-rank fine-grained hints according to claim 1 or 2, characterized in that: The video representation generation module includes a multimodal feature extraction network, a dimension adjustment layer, a backbone network composed of N layers of encoders, and a fully connected layer. During the pre-training process of the multimodal model, the multimodal initial video features generated by the multimodal feature extraction network are unified in dimension by the dimension adjustment layer, and then spliced and input into the backbone network. The last token of the backbone network output result is obtained by the fully connected layer to obtain the video representation.
4. The multimodal video retrieval method based on low-rank fine-grained hints according to claim 3, characterized in that: The multimodal initial video features include features of at least two modalities of an audio modality, a visual modality, and an optical character recognition modality.
5. The multimodal video retrieval method based on low-rank fine-grained hints according to claim 1, characterized in that: During fine-tuning of the prompt update module, the prompts for each modality input to the first encoder layer are randomly initialized.
6. The multimodal video retrieval method based on low-rank fine-grained hints according to claim 1, characterized in that: The multimodal input is represented as: ; ; ; in, represents the multimodal input of the i+1th encoder layer, represents the mode of the i-th encoder layer Tips, is the total number of modes, represents the mode of the i-th encoder layer Tips, Respectively represent the modes of the input i-th encoder layer and i+1-th encoder layer The video features, represents the i-th encoder layer, Indicates a prompt to update the module.
7. The multimodal video retrieval method based on low-rank fine-grained hints according to claim 6, characterized in that: The prompt update module generates each modal prompt of the corresponding layer in the form of fine-grained prompts, which can be expressed as: ; ; ; in, represents the i-th encoder layer The token update weight, express No. Tokens, represents element-wise multiplication, express Middle The relationship between tokens and the different tokens in the remaining n-1 modal prompts, Representation Importance The matrix of order, represents tensor multiplication of multiple matrices, Indicates The modality in the input of the encoder layer Tips, It means to sum the dimensions of the high-order matrix except e.
8. The multimodal video retrieval method based on low-rank fine-grained hints according to claim 7, characterized in that: Will Decompose into A rank-1 tensor implements an approximation of fine-grained cues and updates weights Simplified to: ; in, represents the number of ranks, Indicates the corresponding mode No. The dimensions are The tensor of Indicates The dimensions are The tensor of and Indicates dimension, superscript T indicates transposition, represents the sum of dimension L, represents the mode of the i-th encoder layer Tips.
9. The multimodal video retrieval method based on low-rank fine-grained hints according to claim 1, characterized in that: The subtitle representation generation module adopts a pre-trained BERT-based encoder structure.
10. A multimodal video retrieval system based on low-rank fine-grained cues, used to implement the method of claim 1, characterized in that: The system comprises: A multimodal model, which includes a video representation generation module and a subtitle representation generation module; A hint update module, which is introduced into each layer of the first N-1 encoder layers in the video representation generation module; A fine-tuning training module, which is used to fine-tune the training prompt update module based on the pre-trained multimodal model, and generate prompts for each modality of the corresponding layer in a fine-grained prompt manner; The multimodal video retrieval module fixes and fine-tunes the modal prompts obtained, and in the video representation generation module, concatenates the modal video feature parts of the previous encoder output results with the modal prompts of the current layer as the multimodal input of the current layer; the video representation finally obtained by the video representation generation module is used to match the subtitle representation to achieve the multimodal video retrieval task.
Citation Information
Patent Citations
Visual inspection multitask learning method based on multimodal prompt cooperation
CN118918447A
Missing perception prompting method and system based on modal specific and general information
CN119763010A
Systems and methods for video and language pre-training
US20230154188A1
Machine-learned multi-modal artificial intelligence (AI) models for understanding and interacting with video content
US20240362272A1