Deep cross-modal retrieval method and system based on multi-level structure self-sensing

By employing a multi-level structured self-aware deep cross-modal retrieval method and using a single-modal model as a teacher for cross-modal distillation training, the problem of single-modal performance degradation in cross-modal learning is solved, achieving consistent representation of cross-modal retrieval and improving single-modal retrieval performance.

CN119415725BActive Publication Date: 2025-11-11NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411355845.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-11-11
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

Existing visual language retrieval models affect unimodal structures during cross-modal learning, leading to a decline in unimodal retrieval performance and making it difficult to maintain consistent representations in cross-modal retrieval.

Method used

We employ a deep cross-modal retrieval method with a multi-level self-aware structure. We use an excellent single-modal model as a teacher model to perform multi-level distillation on the cross-modal retrieval model. We train the model through cross-modal contrastive loss and representational distillation loss, thus preserving the performance of single-modal retrieval and improving cross-modal retrieval capabilities.

Benefits of technology

It effectively solves the problem of single-modal performance degradation in cross-modal learning, improves the consistency representation capability of cross-modal retrieval, and maintains the performance of single-modal retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119415725B_ABST
    Figure CN119415725B_ABST
Patent Text Reader

Abstract

This invention discloses a deep cross-modal retrieval method and system based on multi-level structure self-awareness. The method includes: processing the original samples into original sample sequences, which include paired images and text entities; inputting the original sample sequences into a cross-modal network and calculating cross-modal contrast loss and cross-modal matching loss; inputting the original sample sequences into a unimodal teacher network and calculating representation distillation loss to enable the cross-modal network to learn the optimal cross-modal representation at the unimodal representation level; calculating the sample similarity matrices of the unimodal network and the cross-modal network respectively, fusing the similarity matrices of the two unimodal networks and calculating the structure self-awareness relation distillation loss with the cross-modal network similarity matrix; and updating the parameters of the overall model to train the model. This invention learns excellent cross-modal consistency representations by using cross-modal matching loss and multi-level structure self-awareness distillation loss, preserving the structural integrity of the unimodal network and improving the cross-modal retrieval performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and in particular to a deep cross-modal retrieval method and system based on multi-level structure self-awareness. Background Technology

[0002] Visual language retrieval aims to search for similar instances within a modality based on queries from another modality. However, in practice, it's more common to search for both unimodal and cross-modal instances simultaneously. Large information retrieval websites typically present users with both cross-modal and unimodal instances; that is, for a given text / image query, they return similar text and visual instances. In experiments with various benchmark models, we observed a widespread performance degradation in unimodal retrieval compared to independent unimodal training. The main reason for this performance degradation is that cross-modal joint learning affects the unimodal structure. Therefore, we propose a distillation model to learn cross-modal representations while maintaining the integrity of the unimodal structure. Summary of the Invention

[0003] The purpose of this invention is to provide a deep cross-modal retrieval method and system based on multi-level structure self-awareness. It utilizes excellent single-modal models as teacher models to perform multi-level distillation on cross-modal retrieval models, thereby preserving single-modal retrieval performance and promoting cross-modal retrieval capabilities.

[0004] The technical solution for achieving the objective of this invention is as follows: Firstly, this invention provides a deep cross-modal retrieval method based on multi-level structure self-awareness, comprising the following steps:

[0005] Step 1: Process the original samples into original sample sequences. The original samples include paired images and text entities.

[0006] Step 2: Input the original sample sequences into the cross-modal network and calculate the cross-modal contrast loss and cross-modal matching loss;

[0007] Step 3: Input the original sample sequence into the unimodal teacher network and calculate the representation distillation loss so that the cross-modal network learns the optimal cross-modal representation at the unimodal representation level;

[0008] Step 4: Calculate the sample similarity matrix for each of the unimodal network and the cross-modal network. After fusing the similarity matrices of the two unimodal networks, calculate the structural self-awareness relationship distillation loss with the cross-modal network similarity matrix.

[0009] Step 5: Update the parameters of the overall model according to the above loss function to train the model.

[0010] In a second aspect, the present invention provides a deep cross-modal retrieval system based on multi-level structure self-awareness, for implementing the method described in the first aspect, the system comprising:

[0011] The first module is used to process the original samples into a sequence of original samples, which includes paired images and text entities.

[0012] The second module is used to input the original sample sequences into the cross-modal network and calculate the cross-modal contrast loss and cross-modal matching loss.

[0013] The third module is used to input the original sample sequence into the unimodal teacher network and calculate the representation distillation loss so that the cross-modal network learns the optimal cross-modal representation at the unimodal representation level.

[0014] The fourth module is used to calculate the sample similarity matrix of the unimodal network and the cross-modal network, and to calculate the structural self-awareness relationship distillation loss by fusing the similarity matrices of the two unimodal networks and the cross-modal network similarity matrix.

[0015] The fifth module updates the parameters of the overall model based on the above loss function to train the model.

[0016] Thirdly, the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the method described in the first aspect.

[0017] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0018] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0019] Compared with existing technologies, the significant advantages of this invention are as follows: This invention proposes a comprehensive visual-language retrieval method that preserves the unimodal structure while learning cross-modal consistent representations. By obtaining the optimal unimodal model through contrastive loss, and using the unimodal model to perform multi-level distillation on the cross-modal retrieval model, a novel cross-modal retrieval model is proposed. This effectively solves the problem that the representations learned by traditional cross-modal learning methods suffer from performance degradation in unimodal retrieval and difficulty in learning good cross-modal consistent representations. Attached Figure Description

[0020] Figure 1 This is an overall flowchart of the deep cross-modal retrieval method based on multi-level structure self-awareness of the present invention.

[0021] Figure 2 This is a network framework diagram for a deep cross-modal retrieval method based on multi-level self-aware structure.

[0022] Figure 3 This is a flow chart for relational distillation.

[0023] Figure 4 This is a flowchart of the model training steps. Detailed Implementation

[0024] Combination Figures 1-4 This invention provides a deep cross-modal retrieval method based on multi-level structure self-awareness, comprising the following steps:

[0025] (1) The original samples are processed into original sample sequences, which include paired images and text entities;

[0026] ① Initialize the text and images, specifically in the following form:

[0027]

[0028] Where L T Let d be the number of texts, d be the dimension of the text vector, and L be the length of the text. I =HW / P 2 (P×P) represents the size of each block of the image, T is the original text, I is the original image, and C, H, and W are the number of channels, width, and height of the image, respectively.

[0029] (2) Input the original sample sequences into the cross-modal network and calculate the cross-modal contrast loss and cross-modal matching loss.

[0030] ① Encode the image using an attention-based image encoder to obtain the first feature representation of the image modality, specifically in the following form:

[0031]

[0032] Where X I Image modal features;

[0033] ② The image is encoded using an attention-based text encoder to obtain the first feature representation of the text modality, specifically in the following form:

[0034]

[0035] Where X T Text modal features;

[0036] ③ Use a cross-modal attention mechanism and a feedforward network to obtain the first feature representation of the cross-modal module, specifically in the following form:

[0037]

[0038] MultiAtt(X)=[Att(X)1,...,Att(X)M W M

[0039] FFN=(MultiAtt(X))=max(0,MultiAtt(X)W1+b1)W2+b2

[0040] Where x I With x T The dimensions are the same; M is the number of multi-head attention heads, d M = d / M, where d is the dimension of the first feature shared subspace; X represents X I or X T W M W1 and W2 are two learning parameters of the multi-head attention and feedforward networks, respectively; Att is the attention flag; Q T Q I For image and text queries, V is the key for images and text. I and V T The values ​​are: MultiAtt represents multi-head attention, FFN represents a feedforward network, max represents taking the maximum value, b1 and b2 are the two biases of the feedforward network, and softmax is the activation function.

[0041] ④ Calculate the cross-modal contrast loss, specifically in the form of:

[0042]

[0043] Where d(I,T)=cos(g) i (I CLS ),g t (T CLS )) and d(T,I)=cos(g t (T CLS ),g i (I CLs )) is the similarity function, g represents the encoder, p is the model prediction, CE is the cross-entropy loss, τ is the temperature scale coefficient, and I CLS and T CLS For the [CLS] token of the visual and language encoders, exp stands for expectation. For the image-text loss expectation, y i2t (I) retrieves the text truth value for image I, y t2i (T) represents the ground truth value of the image retrieved from the text T, p i2t (I) and p t2i (T) represents the image retrieval text prediction and the text retrieval image prediction, respectively, and J represents the number of images or text in a batch. and These represent the retrieval predictions for the k-th image or text in a batch.

[0044] ⑤ Calculate the cross-modal matching loss, specifically in the form of:

[0045] l itm =[CE(y match ,p matc )]

[0046] Where y match For a one-dimensional one-hot encoding representing the truth value, p match =φ(IT CLS ) represents the matching prediction probability output by the two-dimensional classifier, and φ represents the binary classifier.

[0047] (3) Input the original sample sequence and train the optimal single-modal model using contrast loss and distance maximum regularization loss.

[0048]

[0049] R v =||min((D v -△ + )⊙(D v -△ - ),0)||0

[0050] R w =||min((D w -△ + )⊙(D w -△ - ),0)||0

[0051] Where s(·) is used to measure the similarity between two instances; D v (I i ,I j )=s(I i ,I j ) and D w (I i ,I j )=s(I i ,I j () is a similarity measure; negative example I - T - Samples were taken from the same batch, positive example I + T + R is obtained by perturbing each instance. v With R w For distance regularization; K is the batch number, Δ + =γ + ·1 N×N With △ - =γ - ·1 N×NLet γ be the threshold parameter, where γ is the marginal value that distinguishes the similarity of data pairs, and l ince For the nce loss of the image, l tnce For the nce loss of the text.

[0052] (4) Input the original sample sequence into the single-modal teacher network and calculate the representation distillation loss so that the cross-modal network learns the optimal cross-modal representation at the single-modal representation level.

[0053] ① Calculate the visual modality to characterize distillation loss:

[0054]

[0055] in Indicates visual similarity among individuals within the same batch. As expected, y i2i (I) is the truth value of the image retrieved, p i2i (I) is the predicted value of the image retrieved from the image.

[0056] ② Calculate the distillation loss of text modal characterization:

[0057]

[0058] in This indicates the text-to-text similarity within the same batch.

[0059] (5) Calculate the sample similarity matrix of the unimodal network and the cross-modal network respectively. After fusing the similarity matrices of the two unimodal networks, calculate the structural self-awareness relationship distillation loss with the cross-modal network similarity matrix, as follows:

[0060] Calculate the similarity matrix S between the two unimodal models and the cross-modal model within the same batch. I S T With S IT ;

[0061] Based on the hyperparameter γ, the mixture similarity matrix S of the two single modes is obtained. o :

[0062] S o =γS I +(1-λ)S T

[0063] λ is a hyperparameter;

[0064] Calculate the distillation loss between the mixture similarity matrix and the cross-modal model similarity matrix:

[0065]

[0066] (6) Calculate the cross-modal loss and distillation loss, and update the model.

[0067] The specific loss function for the overall model is as follows:

[0068] ① Calculate the cross-modal loss:

[0069] l cr =l itc +l itm

[0070] ② Calculate the distillation loss for multiple particle sizes:

[0071] l md =l iic +l ttc +l rd

[0072] ③ Overall loss function:

[0073] l = l cr +l md

[0074] Based on the same inventive concept, this invention also proposes a deep cross-modal retrieval system based on multi-level structure self-awareness, used to implement the above-mentioned deep cross-modal retrieval method. The system includes:

[0075] The first module is used to process the original samples into a sequence of original samples, which includes paired images and text entities.

[0076] The second module is used to input the original sample sequences into the cross-modal network and calculate the cross-modal contrast loss and cross-modal matching loss.

[0077] The third module is used to input the original sample sequence into the unimodal teacher network and calculate the representation distillation loss so that the cross-modal network learns the optimal cross-modal representation at the unimodal representation level.

[0078] The fourth module is used to calculate the sample similarity matrix of the unimodal network and the cross-modal network, and to calculate the structural self-awareness relationship distillation loss by fusing the similarity matrices of the two unimodal networks and the cross-modal network similarity matrix.

[0079] The fifth module updates the parameters of the overall model based on the above loss function to train the model.

[0080] The specific implementation methods of modules one through five are the same as those described above, and will not be repeated here.

[0081] This invention uses an attention-based network as the backbone network and trains a unimodal model to extract image text features using contrastive loss and distance maximum regularization loss. By using the unimodal model as the teacher and the cross-modal retrieval model as the student for relation distillation, it achieves excellent performance in both unimodal and cross-modal retrieval tasks.

[0082] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A deep cross-modal retrieval method based on multi-level structure self-awareness, characterized in that, Includes the following steps: Step 1: Process the original samples into original sample sequences. The original samples include paired images and text entities. Obtaining the original sample sequence involves the following steps: Initialize text and images, specifically in the following form: Where L T Let d be the number of texts, d be the dimension of the text vector, and L be the length of the text. I =HW / P 2 (P×P) represents the size of each block of the image, T is the original text, I is the original image, and C, H, and W are the number of channels, width, and height of the image, respectively. Step 2 involves inputting the original sample sequences into the cross-modal network and calculating the cross-modal contrast loss and cross-modal matching loss, specifically including: ① Encode the image using an attention-based image encoder to obtain the first feature representation of the image modality, specifically in the following form: Where X I Image modal features; ② The image is encoded using an attention-based text encoder to obtain the first feature representation of the text modality, specifically in the following form: Where X T Text modal features; ③ Use a cross-modal attention mechanism and a feedforward network to obtain the first feature representation of the cross-modal module, specifically in the following form: MultiAtt(X)=[Att(X)1,...,Att(X) M ]W M FFN=(MultiAtt(X))=max(0,MultiAtt(X)W1+b1)W2+b2 Where x I With x T The dimensions are the same; M is the number of multi-head attention heads, d M = d / M, where d is the dimension of the first feature shared subspace; X represents X I or X T W M W1 and W2 are two learning parameters of the multi-head attention and feedforward networks, respectively; Att is the attention flag; Q T Q I For image and text queries, V is the key for images and text. I and V T The values ​​are: MultiAtt is multi-head attention, FFN is a feedforward network, max means take the maximum value, b1 and b2 are the two biases of the feedforward network, and softmax is the activation function. ④ Calculate the cross-modal contrast loss, specifically in the form of: Where d(I,T)=cos(g) i (I CLS ),g t (T CLS )) and d(T,I)=cos(g t (T CLS ),g i (I CLS )) is the similarity function, g represents the encoder, p is the model prediction, CE is the cross-entropy loss, τ is the temperature scale coefficient, and I CLS and T CLS For the [CLS] token of the visual and language encoders, exp stands for expectation. For the image-text loss expectation, y i2t (I) retrieves the text truth value for image I, y t2i (T) represents the ground truth value of the image retrieved from the text T, p i2t (I) and p t2i (T) represents the image retrieval text prediction and the text retrieval image prediction, respectively, and J represents the number of images or text in a batch. and These represent the retrieval predictions for the k-th image or text in a batch; ⑤ Calculate the cross-modal matching loss, specifically in the form of: it itm =[CE(y match ,p match )] Where y match For a one-dimensional one-hot encoding representing the truth value, p match =φ(IT CLS ) represents the matching prediction probability output by the two-dimensional classifier, and φ represents the binary classifier; Step 3 involves inputting the original sample sequence into the unimodal teacher network and calculating the representation distillation loss to enable the cross-modal network to learn the optimal cross-modal representation at the unimodal representation level. This includes the following steps: ① Inputting the original sample sequence into a single-modal network to obtain the optimal single-modal model includes the following steps: R v =||min((D v -△+)⊙(D v -△ - ),0)||0 R w =||min((D w -△+)⊙(D w -△ - ),0)||0 Where s(·) is used to measure the similarity between two instances; D v (I i ,I j )=s(I i ,I j ) and D w (I i ,I j )=s(I i ,I j () is a similarity measure; negative example I - T - Samples were taken from the same batch, positive example I + T + R is obtained by perturbing each instance. v With R w For distance regularization; K is the batch number, Δ + =γ + ·1 N×N With △ - =γ - ·1 N×N Let γ be the threshold parameter, where γ is the marginal value that distinguishes the similarity of data pairs, and l ince For the nce loss of the image, l tnce For the nce loss of the text; ② Calculate the visual modality to characterize the distillation loss: in Indicates visual similarity among individuals within the same batch. As expected, y i2i (I) is the truth value of the image retrieved, p i2i (I) is the predicted value of the image retrieved from the image; ③ Calculate the text modal characterization distillation loss: in Indicates the text-to-text similarity within the same batch; Step 4: Calculate the sample similarity matrices for both the unimodal network and the cross-modal network. After fusing the similarity matrices of the two unimodal networks, calculate the structural self-awareness relationship distillation loss by combining it with the cross-modal network similarity matrix. This includes the following steps: ① Calculate the similarity matrix S between the two unimodal models and the cross-modal model within the same batch. I S T With S IT ; ②Based on the hyperparameter λ, the mixture similarity matrix S of the two single modes is obtained. O : S O =λS I +(1-λ)S T λ is a hyperparameter; ③ Calculate the distillation loss between the mixture similarity matrix and the cross-modal model similarity matrix: Step 5: Update the parameters of the overall model according to the above loss function to train the model: ① Calculate the cross-modal loss: l cr =l itc +l itm ② Calculate the distillation loss for multiple particle sizes: l md =l iic +l ttc +l rd ③ Overall loss function: l=l cr +l md 。 2. A deep cross-modal retrieval system based on multi-level self-aware structure, characterized in that, The system for implementing the method of claim 1 includes: The first module is used to process the original samples into a sequence of original samples, which includes paired images and text entities. The second module is used to input the original sample sequences into the cross-modal network and calculate the cross-modal contrast loss and cross-modal matching loss. The third module is used to input the original sample sequence into the unimodal teacher network and calculate the representation distillation loss so that the cross-modal network learns the optimal cross-modal representation at the unimodal representation level. The fourth module is used to calculate the sample similarity matrix of the unimodal network and the cross-modal network, and to calculate the structural self-awareness relationship distillation loss by fusing the similarity matrices of the two unimodal networks and the cross-modal network similarity matrix. The fifth module updates the parameters of the overall model based on the above loss function to train the model.

3. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method of claim 1.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method of claim 1.

5. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 1.

Citation Information

Patent Citations

  • Depth cross-modal hash image retrieval method based on joint semantic matrix

    CN113177132A

  • Cross-modal retrieval method, system and equipment based on supervised comparison

    CN113239214A