Cross-modal retrieval model construction method based on denoising and momentum distillation
By constructing a cross-modal retrieval model based on denoising and momentum distillation, the problems of insufficient generalization ability and insufficient information interaction of cross-modal retrieval models on noisy datasets are solved, more efficient retrieval accuracy and robustness are achieved, computational complexity is reduced, and application scenarios are expanded.
Patent Information
- Application Number
- CN202310750571.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-21
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2043-06-21
AI Technical Summary
Existing cross-modal retrieval models have problems with image-text datasets, such as insufficient generalization due to noise, slow reasoning in single-stream models, and low accuracy due to the lack of information interaction in two-stream models.
A cross-modal retrieval model construction method based on denoising and momentum distillation is adopted. By constructing an encoding unit, a self-supervised denoising unit, a momentum distillation unit and a similarity calculator, modal interaction loss, fusion denoising loss and cross-modal contrastive learning loss are used for training. Combined with natural language processing and momentum distillation models, modal feature interaction and denoising training are performed.
It improves the model's retrieval accuracy and inference speed, enhances its robustness to sparse data, reduces the encoder's computational complexity and storage requirements, and expands its practical application space.
Smart Images

Figure CN116861021B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cross-modal retrieval, and more specifically, relates to a method for constructing a cross-modal retrieval model based on denoising and momentum distillation. Background Art
[0002] With the gradual maturity of the Internet, the popularization of mobile terminals, and the rapid development of industries such as self-media, data generators are no longer limited to a few large companies and enterprises, and more and more users are participating in the creation of multimedia data. Driven by the data age, users' demand for data utilization is increasing and becoming more complex. Cross-modal retrieval algorithms can query one modality (such as text) through another modality (such as images), and therefore have extremely high practical application requirements. The pre-training + fine-tuning paradigm has demonstrated tremendous capabilities in the field of cross-modal retrieval. First, the model is pre-trained through a large-scale cross-modal dataset, and then fine-tuned for different downstream tasks. This training method significantly improves the performance of multimodal tasks (such as image text retrieval, video text retrieval, visual question answering, etc.), surpassing previous training methods.
[0003] Despite this, current technology still has limitations in practical application scenarios. Current mainstream cross-modal pre-training models are generally categorized as single-stream and two-stream models. Single-stream models achieve good performance through fine-grained interaction, but the high computational cost leads to low retrieval efficiency. While two-stream models significantly improve retrieval efficiency, they often only achieve suboptimal retrieval accuracy due to the lack of fine-grained interaction between features from different modalities.
[0004] Cross-modal retrieval tasks often suffer from data sparsity. Taking image-text cross-modal retrieval as an example, many samples in image and text data contain only partial information, or very little image and text data is available for learning. Furthermore, as image-text pre-training datasets grow in size, large-scale manual annotation becomes nearly impossible. Consequently, these large-scale datasets inevitably contain noise. This type of noisy data can significantly negatively impact model training. How to use noisy datasets to learn visual and textual representations, how to mitigate the impact of noise, and even how to leverage it to improve model robustness are very relevant and pressing issues that need to be addressed. Summary of the Invention
[0005] In response to the defects of the existing technology and the need for improvement, the present invention provides a method for constructing a cross-modal retrieval model based on denoising and momentum distillation, which aims to solve the problem of insufficient model generalization ability caused by the presence of noise in the image-text dataset itself, as well as the problem of slow reasoning of single-stream visual language pre-training models and low accuracy caused by the lack of information interaction in two-stream models.
[0006] To achieve the above objectives, according to one aspect of the present invention, a method for constructing a cross-modal retrieval model based on denoising and momentum distillation is provided, comprising: constructing an encoding unit, wherein the encoding unit comprises N cascaded first modality data encoders and N cascaded second modality data encoders, where N>1; setting an i-th self-supervised denoising unit between the output ends of the i-th first modality data encoder and the i-th second modality data encoder, for sequentially performing denoising and decoding reconstruction on the original joint feature labels of the i-th layer to obtain a reconstructed joint feature label of the i-th layer, i∈(1,N-1); constructing a cross-modal retrieval model, wherein the cross-modal retrieval model comprises: a fusion denoising unit, a momentum distillation unit and a similarity calculator arranged at the output end of the encoding unit, as well as the encoding unit and the self-supervised denoising unit; constructing a modality interaction loss with the goal of minimizing the KL divergence between the reconstructed joint feature label and the original joint feature label, and constructing a total loss function comprising the modality interaction loss, the fusion denoising loss and the cross-modal contrastive learning loss; and training the cross-modal retrieval model with the goal of convergence of the total loss function.
[0007] Furthermore, the first modality is text and the second modality is image, and the method also includes: expanding the text sample set by using one or more methods including synonym replacement, random insertion, random exchange, and random deletion; using random masking to filter out some images from the image samples to form new image samples to expand the image sample set; and training the cross-modal retrieval model includes: training the cross-modal retrieval model using the expanded text sample set and image sample set.
[0008] Furthermore, the i-th self-supervised denoising unit is specifically used to: connect the output of the i-th first modal data encoder and the output of the i-th second modal data encoder to obtain the original joint feature marker of the i-th layer; add noise to the original joint feature marker of the i-th layer in a masking manner, and decode and reconstruct the joint feature marker containing the noise through a lightweight cross-modal decoder to obtain the reconstructed joint feature marker of the i-th layer.
[0009] Furthermore, the modal interaction loss is:
[0010]
[0011] Among them, L msd is the modal interaction loss, KL[] is the KL divergence calculation function, cat() is the connection vector function, trans() is the transformer function, w i is the first modal data feature vector of the i-th layer, v i is the feature vector of the second modal data of the i-th layer, is the feature vector of the first modal data of the i-th layer after noise addition, is the feature vector of the second modal data of the i-th layer after noise addition.
[0012] Furthermore, the fusion denoising unit is used to: sequentially connect and decode the data feature labels of the two modalities output by the last layer of the encoding unit to obtain the data features of the two modalities; use the visible features in the data features of the two modalities to predict the mask features in the data features of the two modalities to obtain the final data features of the two modalities.
[0013] Furthermore, the similarity calculator is used to calculate the similarity between the final data features of the two modalities obtained by the fusion denoising unit, so as to output the second modal data having the highest similarity with the first modal data.
[0014] Furthermore, the momentum distillation unit is used to accumulate the corresponding queues according to the final data features of the two modalities, and correct the similarity obtained by the similarity calculator.
[0015] According to another aspect of the present invention, a cross-modal retrieval method based on denoising and momentum distillation is provided, which uses the trained cross-modal retrieval model obtained by the cross-modal retrieval model construction method based on denoising and momentum distillation as described above to retrieve second modal data that matches the first modal data.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements the cross-modal retrieval model construction method based on denoising and momentum distillation as described above, or implements the cross-modal retrieval method based on denoising and momentum distillation as described above.
[0017] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects:
[0018] (1) A cross-modal retrieval model construction method based on denoising and momentum distillation is provided. By designing denoising training tasks, modal interactions are performed in the middle and late stages of the feature encoder, thereby strengthening the learning of inter-modal associations, improving the accuracy of model retrieval, and achieving more efficient inference speed.
[0019] (2) Combining the text data enhancement algorithm in natural language processing with the momentum distillation model can solve the problem of insufficient data in practical applications and alleviate the impact of model overfitting caused by noise samples in the data itself. Experiments show that the combination of the two can effectively improve the robustness of the model to sparse retrieval data and enhance the model's generalization ability;
[0020] (3) Perform a large-scale random masking on the image blocks, and then use the visual encoder to encode the unmasked image blocks, so as to reduce the time complexity and storage requirements of the encoder calculation, so that the batch size can be made larger, further improving the robustness and generalization ability of the model; at the same time, it also supports the pre-coding and storage of retrieval database features, greatly expanding the application space in real life. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A flowchart of a method for constructing a cross-modal retrieval model based on denoising and momentum distillation provided by an embodiment of the present invention;
[0022] Figure 2 A schematic diagram of the structure of a cross-modal retrieval model based on denoising and momentum distillation provided in an embodiment of the present invention;
[0023] Figure 3 for Figure 2 Schematic diagram of the structure of the self-supervised denoising unit in the shown model;
[0024] Figure 4 for Figure 2 Schematic diagram of the structure of the fusion denoising unit in the shown model;
[0025] Figure 5 for Figure 2 Schematic diagram of the structure of the momentum distillation unit in the shown model. DETAILED DESCRIPTION
[0026] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0027] In the present invention, the terms "first", "second", etc. (if any) in the present invention and the drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0028] Figure 1 Flowchart of the method for constructing a cross-modal retrieval model based on denoising and momentum distillation provided by an embodiment of the present invention. Figure 1 , combined with Figure 2-Figure 5 , a method for constructing a cross-modal retrieval model based on denoising and momentum distillation in this embodiment is described in detail, and the method includes operations S1 to S5.
[0029] In operation S1 , a coding unit is constructed, where the coding unit includes N cascaded first modality data encoders and N cascaded second modality data encoders, where N>1.
[0030] In this embodiment, the constructed encoding unit includes a first modality data encoding unit and a second modality data encoding unit. The first modality data encoding unit includes N cascaded first modality data encoders for performing cascade encoding on first modality data, such as image data; and the second modality data encoding unit includes N cascaded second modality data encoders for performing cascade encoding on second modality data, such as text data.
[0031] See Figure 3 and Figure 4 For example, the first modal data encoder and the second modal data encoder each include a normalization layer 1, a multi-head self-attention layer, a normalization layer 2, and a feed-forward fully connected layer connected in sequence. The first modal data encoder and the second modal data encoder can also be other structures that can achieve encoding functions.
[0032] Operation S2, setting the i-th self-supervised denoising unit between the output ends of the i-th first modal data encoder and the i-th second modal data encoder, for sequentially adding noise and decoding and reconstructing the original joint feature markers of the i-th layer to obtain the reconstructed joint feature markers of the i-th layer, i∈(1,N-1).
[0033] The N first modality data encoders and the N second modality data encoders correspond one to one. In this embodiment, among the first N-1 first modality data encoders and the first N-1 second modality data encoders, a self-supervised denoising unit is connected between the corresponding first modality data encoders and the second modality data encoders, such as Figure 3 shown.
[0034] See Figure 3 , for any i-th self-supervised denoising unit, it is specifically used to perform sub-operation S21-sub-operation S23.
[0035] In sub-operation S21, the output of the i-th first modality data encoder and the output of the i-th second modality data encoder are connected to obtain the original joint feature label of the i-th layer.
[0036] In sub-operation S22, noise is added to the original joint feature signature of the i-th layer by masking to obtain the joint feature signature of the i-th layer containing noise.
[0037] Based on joint feature tagging, a masked visual language modeling task is constructed, using a masking method to add noise. Specifically, taking image text retrieval as an example, each word tag or image region tag is masked with a probability of 15%. The masked word tag has an 80% probability of being blocked, a 10% probability of being replaced with a random tag, and the rest of the time, it remains unchanged. The masked image region tag has a 90% probability of being blocked, and a 10% probability of remaining unchanged.
[0038] In sub-operation S23, the joint feature signature containing noise in the i-th layer is decoded and reconstructed through a lightweight cross-modal decoder Transformer to obtain a reconstructed joint feature signature in the i-th layer.
[0039] Operation S3 constructs a cross-modal retrieval model, which includes: a fusion denoising unit, a momentum distillation unit, and a similarity calculator set at the output end of the encoding unit, as well as the encoding unit and the self-supervised denoising unit.
[0040] Furthermore, a cross-modal retrieval model is constructed, which includes the above-mentioned encoding unit, the above-mentioned self-supervised denoising unit, the fusion denoising unit located at the output end of the encoding unit, the momentum distillation unit located at the output end of the encoding unit, and the similarity calculator located at the output end of the encoding unit. The structure of the constructed cross-modal retrieval model is as follows: Figure 2 shown.
[0041] See Figure 4 The fusion denoising unit is used to: sequentially connect and decode the data feature labels of the two modalities output by the last layer of the encoding unit to obtain the data features of the two modalities; use the visible features in the data features of the two modalities to predict the mask features in the data features of the two modalities to obtain the final data features of the two modalities.
[0042] Taking image-to-text retrieval as an example, during the initial noise addition process, masked noise is added to the text tokens using BERT's masking method. To reconstruct these obscured text tokens, after the visual encoder and text encoder complete the feature representation of the image and sentence, the image feature representation and the remaining visible text feature representation are used to predict the masked words. The underlined words represent the words randomly masked during the initial noise addition process. Prediction is performed through the subsequent mask reconstruction denoising training task, and a cross-modal decoder is used to interactively learn the joint feature representation of image and text.
[0043] The similarity calculator is used to calculate the similarity between the final data features of the two modalities obtained by the fusion denoising unit, so as to output the second modal data with the highest similarity to the first modal data.
[0044] See Figure 5,The momentum distillation unit is used to accumulate the corresponding queues respectively according to the final data features of the two modalities, and to correct the similarity obtained by the similarity calculator.
[0045] The momentum distillation phase maintains two queues, Qi and Qt, to store the image representation i and text representation t recently extracted by the teacher model. The teacher model, with the same learning structure, is initialized by the student model and continuously updated using an exponential moving average (EMA) strategy. It serves as an additional supervisory signal in the loss calculation during similarity calculations. For example, the momentum queue length is set to 10,000, the teacher model's distillation coefficient is set to 0.4, and the momentum update rate is 0.99.
[0046] Taking image text retrieval as an example, in this embodiment, the text-image pre-training model (Constrastive Language-Image Pre-training, CLIP) ViT-B / 16 is used as the backbone network to obtain visual representation. Each image input is divided into 16 non-overlapping image blocks, and then converted into a one-dimensional image label using linear projection. N v is the number of image tags. For each image tag block, a position feature vector is added to establish the position correlation between the disordered blocks, and the interaction between each block of the input image is modeled using the Transformer architecture, and its encoder is converted into a visual feature representation. Among them, D v is the dimension of the image tag. CLIP's BERT-LIKETransformer is used as the text encoder, and the text input After being converted into word vectors, the joint position encoding is input to the encoder, and the text encoder projects the token embedding into a common subspace to generate text feature representation where N t and D t is the number and dimension of text tags. When calculating similarity, the [class] tag in the text feature representation and image feature representation is used as the entire text w∈R D and image v∈R D Because the attention mechanism in Transformer is to interact with all elements pairwise, the [class] tag can contain the information of all other tags.
[0047] Operation S4, with the goal of minimizing the KL divergence between the reconstructed joint feature label and the original joint feature label, constructs the modal interaction loss, and constructs a total loss function including the modal interaction loss, the fusion denoising loss and the cross-modal contrastive learning loss.
[0048] Taking image-text retrieval as an example, in mid-term self-supervised denoising, language-based methods can usually only recognize the high-level semantics of visual content, making it difficult to accurately reconstruct image features. The lightweight cross-modal decoder Transformer does not directly regress mask feature values, but instead predicts the semantic class distribution of the corresponding text and image areas. Therefore, it uses the joint feature label before denoising as the true label, with the goal of minimizing the KL divergence between the two distributions.
[0049] According to an embodiment of the present invention, the modal interaction loss is:
[0050]
[0051] Among them, L msd is the modal interaction loss, KL[] is the KL divergence calculation function, cat() is the connection vector function, trans() is the transformer function, w i is the first modal data feature vector of the i-th layer, v i is the feature vector of the second modal data of the i-th layer, is the feature vector of the first modal data of the i-th layer after noise addition, is the feature vector of the second modal data of the i-th layer after noise addition.
[0052] The cross-modal decoders used in the mid-term self-supervised denoising training task and the late fusion denoising training task are initialized using some layers of the CLIP text encoder. Preferably, the Adam optimizer is used, the learning rate is set to 1e-6, and it is decayed by half every 10 epochs. The training epoch is set to 40, the batch size of each training is 128, and single precision is used in the unimodal encoder.
[0053] During training, a cross-modal contrastive learning loss L is used ITC The calculation of cosine distance is done by dot product of vectors. Taking image text retrieval as an example, given a batch of M image language pairs, the best matching image language pair is predicted from M×M possible image language pairs. Where M is the given batch size, so there are M in a training batch. 2 ×M negative visual-language pairs. Special image tagging using the dual encoder output [CLS V ] indicates visual I, special text mark [CLS T ] represents the text T, calculates the normalized softmax image-to-text similarity and text-to-image similarity, and maintains two queues to store the momentum single encoder's most recent K image-text representations v c ′ ls and w c ′ls , calculate the momentum label and Finally, the cross entropy loss function is used for updating.
[0054] The total loss function L is:
[0055] L=L itc +L msd +L mlm
[0056]
[0057] Among them, L itc is the cross-modal contrastive learning loss, L mlm is the fusion denoising loss, E (I,T)~D [] represents all matching and non-matching sample pairs, H() represents the cross entropy loss, and y msk () is the one-hot encoding vocabulary distribution of the masked modal data, p msk () is the predicted probability of the masked modal data, I is the masked modal data, The remaining modal data.
[0058] Operation S5 trains the cross-modal retrieval model with the goal of convergence of the total loss function.
[0059] Preferably, in this embodiment, when training samples are initially generated, the training sample set can be expanded by adding noise. When the first modality is text and the second modality is an image, the training sample set is expanded by: using one or more of synonym replacement, random insertion, random swapping, and random deletion to expand the text sample set; using random masking to filter out portions of the image sample to form new image samples to expand the image sample set. When text noise is added, a dynamic processing method is adopted, randomly modifying the text input each time a minimum batch is constructed.
[0060] Synonym replacement involves randomly selecting a non-stop word in a sentence and then randomly selecting a synonym from that word's set to replace the original word. Random insertion involves randomly selecting a non-stop word in a sentence and then randomly selecting a synonym from its set of synonyms and inserting it into a random position in the original sentence. Random replacement involves swapping the positions of two randomly selected words in a sentence. Random deletion involves randomly deleting each word in a sentence with a certain probability. In implementation details, for each sentence input in each batch, there is a 50% probability of adding noise and a 50% probability of leaving the original text unchanged. For sentences requiring noise, a random noise addition method is selected with equal probability to modify the text. The visual noise addition module randomly masks image blocks at a large percentage, then uses a visual encoder to encode the unmasked image blocks. This reduces the encoder's computational complexity and memory requirements, thereby enabling larger batch sizes.
[0061] Accordingly, in operation S5, with the goal of convergence of the total loss function L, the cross-modal retrieval model is trained using the expanded text sample set and image sample set.
[0062] Taking image-text retrieval as an example, the cross-modal retrieval model construction method based on denoising and momentum distillation in this embodiment includes a multi-level fusion denoising training phase and a momentum distillation training phase. The multi-level fusion denoising training phase includes: early noise addition, which can utilize limited annotated data to form more training data and train a model with stronger generalization capabilities; image noise addition based on prior knowledge of image information redundancy, through random masking to filter out some image blocks for input into the visual encoder, reducing the computational cost of training the visual encoder; mid-term self-supervised denoising, which fuses the intermediate feature outputs of the image and text encoders and aligns the local fine-grained features of the visual and text through self-supervised denoising training tasks; and late fusion denoising, which uses the image feature representation and the remaining visible text feature representation to predict the masked words after the visual and text encoders complete the feature representation of the image and sentence.
[0063] The embodiment of the present invention also provides a cross-modal retrieval method based on denoising and momentum distillation, which adopts Figure 1-Figure 5 In the illustrated embodiment, the trained cross-modal retrieval model obtained by the cross-modal retrieval model construction method based on denoising and momentum distillation retrieves second modal data that matches the first modal data.
[0064] In this embodiment, the principle of cross-modal retrieval using a cross-modal retrieval model is similar to Figure 1-Figure 5 The principles of the cross-modal retrieval model construction method based on denoising and momentum distillation in the illustrated embodiment are the same and will not be repeated here.
[0065] The embodiment of the present invention further provides a computer readable storage medium on which a computer program is stored. When the program is executed by a processor, Figure 1-Figure 5 The illustrated embodiment provides a method for constructing a cross-modal retrieval model based on denoising and momentum distillation, or implements the above-mentioned cross-modal retrieval method based on denoising and momentum distillation.
[0066] It will be easily understood by those skilled in the art that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for constructing a cross-modal retrieval model based on denoising and momentum distillation, characterized in that: include: Constructing an encoding unit, wherein the encoding unit includes N cascaded first modality data encoders and N cascaded second modality data encoders, where N>1; An i-th self-supervised denoising unit is provided between the output terminals of the i-th first modal data encoder and the i-th second modal data encoder, for sequentially performing denoising and decoding reconstruction on the original joint feature labels of the i-th layer to obtain a reconstructed joint feature label of the i-th layer, i∈(1,N-1); Constructing a cross-modal retrieval model, the cross-modal retrieval model comprising: a fusion denoising unit, a momentum distillation unit, and a similarity calculator arranged at an output end of an encoding unit, as well as the encoding unit and the self-supervised denoising unit; With the goal of minimizing the KL divergence between the reconstructed joint feature labeling and the original joint feature labeling, a modality interaction loss is constructed, and a total loss function including the modality interaction loss, the fusion denoising loss, and the cross-modal contrastive learning loss is constructed; Training the cross-modal retrieval model with the goal of convergence of the total loss function; The first modality is text, the second modality is image; The i-th self-supervised denoising unit is specifically used for: Concatenate the output of the i-th first modality data encoder and the output of the i-th second modality data encoder to obtain the original joint feature label of the i-th layer; Noise is added to the original joint feature signature of the i-th layer in a masking manner, and the joint feature signature containing the noise is decoded and reconstructed through a lightweight cross-modal decoder to obtain the reconstructed joint feature signature of the i-th layer.
2. The method for constructing a cross-modal retrieval model based on denoising and momentum distillation according to claim 1, wherein: Also includes: Expand the text sample set by using one or more of the following methods: synonym replacement, random insertion, random exchange, and random deletion; Random masking is used to filter out some images from the image samples to form new image samples to expand the image sample set; The training of the cross-modal retrieval model includes: using the expanded text sample set and image sample set to train the cross-modal retrieval model.
3. The method for constructing a cross-modal retrieval model based on denoising and momentum distillation according to claim 1, wherein: The modal interaction loss is: Among them, L msd is the modal interaction loss, KL[ ] is the KL divergence calculation function, cat( ) is the connection vector function, trans( ) is the transformer function, w i is the first modal data feature vector of the i-th layer, v i is the feature vector of the second modal data of the i-th layer, is the feature vector of the first modal data of the i-th layer after noise addition, is the feature vector of the second modal data of the i-th layer after noise addition.
4. The method for constructing a cross-modal retrieval model based on denoising and momentum distillation according to claim 1, wherein: The fusion denoising unit is used for: The data feature labels of the two modalities output by the last layer of the encoding unit are sequentially connected and decoded and reconstructed to obtain the data features of the two modalities; The visible features in the data features of the two modalities are used to predict the mask features in the data features of the two modalities, and the final data features of the two modalities are obtained.
5. The method for constructing a cross-modal retrieval model based on denoising and momentum distillation according to claim 4, wherein: The similarity calculator is used to calculate the similarity between the final data features of the two modalities obtained by the fusion denoising unit, so as to output the second modal data having the highest similarity with the first modal data.
6. The method for constructing a cross-modal retrieval model based on denoising and momentum distillation according to claim 4, wherein: The momentum distillation unit is used to accumulate corresponding queues according to the final data features of the two modalities, and correct the similarity obtained by the similarity calculator.
7. A cross-modal retrieval method based on denoising and momentum distillation, characterized in that: Using the trained cross-modal retrieval model obtained by the cross-modal retrieval model construction method based on denoising and momentum distillation as described in any one of claims 1-6, second modal data matching the first modal data is retrieved.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, it implements the cross-modal retrieval model construction method based on denoising and momentum distillation as described in any one of claims 1-6, or implements the cross-modal retrieval method based on denoising and momentum distillation as described in claim 7.
Citation Information
Patent Citations
Dynamic weighted cross-modal fusion network retrieval method and system, and electronic equipment
CN115827954A
Visual-semantic representation learning via multi-modal contrastive training
US20220284321A1