Cross-modal image-text retrieval method and system based on adaptive comparative learning
By introducing adaptive contrast learning and Gaussian discriminant analysis in cross-modal graphic and text retrieval, the boundary parameters are dynamically adjusted to distinguish cloned negative samples, which solves the problem of indistinguishable graphic and text pairs in the prior art, and improves the accuracy and robustness of matching results.
Patent Information
- Application Number
- CN202510211202.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The existing cross-modal graphic search methods are difficult to distinguish semantic-related graphic pairs during training, which makes cloned negative samples difficult to process, affecting the accuracy of the matching results.
Using an adaptive contrast learning method, by introducing two boundary parameters m1 and m2, semantic consistency is dynamically adjusted in the probability distribution of each pair of positive samples and distinguishing the distribution characteristics of cloned negative samples. Gaussian discriminant analysis is used to predict potential cloned negative samples and dynamic comparison learning is achieved by iteratively updating boundary parameters.
Effectively distinguish between cloned negative samples, enhance the tightness of positive samples, improve the accuracy and robustness of picture and text matching, and optimize the supervision effect during training.
Smart Images

Figure CN120045736A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cross-modal, and specifically, relates to a cross-modal image-text retrieval method and system based on adaptive contrastive learning. Background Art
[0002] Cross-modal image-text retrieval has become one of the most basic multimodal tasks in recent years. Its tasks include searching images using text queries and retrieving corresponding sentences using images. Existing image-text matching methods usually learn a robust cross-modal fusion representation through various fusion paradigms (such as target-level or global-level alignment). At the same time, contrastive learning is widely used as the objective function of cross-modal tasks, which shortens the distance between similar positive samples in the latent space while expanding the distance between negative samples. Among these methods, triplet ranking loss and contrastive loss have achieved remarkable success in image-text contrastive learning. Specifically, the triple ranking loss maintains the relative distance between samples within a batch through a manually set fixed boundary (see Faghri, Fartash, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler. "Vse++: Improving visual-semantic embeddings with hard negatives." arXiv preprint arXiv: 1707.05612 (2017).). The contrast loss is defined as the cross entropy of the softmax-normalized similarity between each pair of image and text samples (see Radford, Alec, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry et al. "Learning transferable visual models from natural language supervision." In International conference on machine learning, pp. 8748-8763. PMLR, 2021.).In self-supervised learning and multi-modal tasks, InfoNCE and its variants are the main forms of contrastive loss (see Oord, Aaron van den, Yazhe Li, and Oriol Vinyals. "Representation learning with contrastive predictive coding." arXiv preprint arXiv:1807.03748 (2018). Cheng, Pengyu, Weituo Hao, Shuyang Dai, Jiachang Liu, Zhe Gan, and Lawrence Carin. "Club: A contrastive log-ratio upper bound of mutual information." In International conference on machine learning, pp. 1779-1788. PMLR, 2020.).
[0003] Despite some progress, current contrastive learning methods for cross-modal image-text retrieval face a severe challenge: during training, image-text pairs often contain semantically related text annotations and similar visual cues, making them difficult to distinguish in practice (reflected in their very close similarity scores). Undoubtedly, retrieving such samples is a sub-optimal result because the ground truth annotations manually labeled show more precise and fine-grained semantics and closer corresponding relationships. Different from false negative samples that should be regarded as "positive samples", samples with "consistent semantics but not optimal" are called clone negative samples, which are extremely challenging to match in such scenarios. It is extremely difficult to distinguish clone negative samples through cross-modal fusion because clone negative samples are essentially also "positive samples". Current image-text matching still mainly relies on basic triplet ranking loss or contrastive loss as the learning objective, which sets fixed margins during training, making it very difficult to dynamically handle the problem of clone negative samples during training because they have predefined labels for all image-text pairs (positive or negative). Summary of the Invention
[0004] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide a cross-modal image-text retrieval method and system based on adaptive contrastive learning.
[0005] According to the first aspect of the present invention, there is provided a cross-modal image-text retrieval method based on adaptive contrastive learning, including:
[0006] Using an image encoder and a text encoder Extract the visual modality feature v and the text modality feature w of the image I and the text T respectively, and calculate the similarity to obtain the cosine similarity s of the two in the same subspace;
[0007] Introduce two boundary parameters m 1 and m 2 into the probability distribution of each pair of positive samples, and combine the cosine similarity s to calculate the probability of the image I corresponding to the text T Adjust the semantic consistency of each pair of positive samples and distinguish the distribution characteristics of cloned negative samples;
[0008] Define and calculate the saliency scores of each batch of samples, and select the sequences of the negative samples with the largest and smallest saliency scores in each batch of samples as the observed cloned negative samples S cln and the significant negative samples S sln ;
[0009] Based on the observed cloned negative samples S cln and the significant negative samples S sln , according to the Gaussian discriminant analysis method, predict all potential cloned negative samples of each batch of samples and generate anchor samples;
[0010] According to the two boundary conditions of the similarity of the anchor samples, iteratively update the values of the two boundary parameters m 1 and m 2 ;
[0011] After updating the values of the two boundary parameters m 1 and m 2 , adaptively contrastive learning loss, optimize batch by batch to achieve dynamic contrastive learning between positive and negative samples, and train the image encoder and the text encoder
[0012] Use the trained image encoder and the text encoder to extract the visual modality features and text modality features of the image to be detected and the text to be detected, and calculate the similarity to obtain the cosine similarity of the two in the same subspace, as the basis for cross-modal image-text retrieval.
[0013] Preferably, the use of the image encoder and the text encoder to extract the visual modality feature v and the text modality feature w of the image I and the text T respectively, and calculate the similarity to obtain the cosine similarity s of the two in the same subspace, includes:
[0014] The image encoder converts the RGB pixels of the image I into a feature representation v;
[0015] The text encoder converts the text T into tokens, and then converts the tokens into a text embedding w;
[0016] Performs cosine similarity calculation on the feature representation v and the text embedding w to obtain s(I, T), which is used to measure the similarity degree of the image-text pair.
[0017] Preferably, two boundary parameters m 1 and m 2 are introduced into the probability distribution of each pair of positive samples, and combined with the cosine similarity s, the probability p i (I) of the image I corresponding to the text T is calculated to adjust the semantic consistency of each pair of the positive samples and distinguish the distribution characteristics of the cloned negative samples, including:
[0018]
[0019] In the formula, i represents the subscript index of the text T corresponding to the image I, that is, (I, T i ) is a pair of image-text pairs, M represents the number of negative samples in a batch, m 1 and m 2 are two introduced boundary parameters, and the denominator consists of a positive sample m 1 (s(I, T i ) - m 2 ) and M negative samples .
[0020] Preferably, an adaptive contrast learning objective function is established through the probability p i (I) of the image I corresponding to T, specifically:
[0021]
[0022] Among them, E I~D represents the expected value, indicating the expectation of summing after sampling all images I from the data distribution D; H represents the cross-entropy loss function; N represents the number of samples in the batch; the subscript i represents the target label of the i-th sample, and y i (I) is a one-hot encoded vector representing the true class or text of the image I.
[0023] Preferably, the significance score of each batch of samples is defined and calculated, and the negative sample sequences with the largest and smallest significance scores in each batch of samples are selected as the observed cloned negative sample S cln and the significant negative sample S sln , including:
[0024] Define the significance score as the difference between the average of the cosine similarities of positive samples and the cosine similarities of negative samples:
[0025]
[0026] s ii and s ij are both cosine similarities. s ii represents the positive sample similarity, that is, the value along the main diagonal in the similarity matrix. s ij represents the negative sample similarity, that is, the value of the coordinate (i, j) in the similarity matrix, that is, the similarity between the i-th image or text sample and the j-th text or image sample;
[0027] Calculate the significance score of each batch of samples according to the above definition;
[0028] Select the sample pair with the highest significance score as the significant negative sample S sln , which contains 1 positive sample and M negative samples;
[0029] Select the sample pair with the lowest significance score as the observed cloned negative sample S cln , which contains 1 positive sample and M potential cloned negative samples.
[0030] Preferably, the significance score is the update criterion of the boundary parameters m 1 and m 2 , and is calculated and iterated during the training of each batch.
[0031] Preferably, based on the observed cloned negative sample S cln and the significant negative sample S sln , according to the Gaussian discriminant analysis method, predict all potential cloned negative samples of each batch of samples and generate anchor samples, including:
[0032] Calculate the class probability that each pair of image and text pairs is a cloned negative sample, and its distribution expression is:
[0033]
[0034] In the formula: Using the binary classification method, respectively represent whether the pairwise similarity is / is not a cloned negative sample; s represents the cosine similarity score; represents the original score of the similarity without probability output, represents the original score of the similarity without probability output;
[0035] Based on the properties of Gaussian discrimination: Gaussian distributions with the same covariance for features, and the similarity score being a univariate variable, simplify the distribution expression of the class probability to a univariate distribution:
[0036] And wherein, represents a Gaussian distribution with a mean of u0 and a variance of ; represents a Gaussian distribution with a mean of μ 1 and a variance of ;
[0037] Arrange the distribution expression as:
[0038]
[0039] wherein, π 1 and π 0 are respectively and prior probabilities; Given that S sln and S cln are respectively composed of the most representative significant negative samples and potential clone negative samples in a training batch, use their respective negative samples to represent the empirical means and variances μ 1 , μ 0 , σ 1 and σ 0 ;
[0040] Based on the following criteria, select the potential clone negative samples in each batch of samples:
[0041]
[0042] S * represents the predicted clone negative samples in each batch of samples;
[0043] Define the median sample in the set of the said S * as the anchor sample.
[0044] Preferably, iteratively update the values of the two boundary parameters m 1 and m 2 according to the two boundary conditions of the similarity of the anchor sample, including:
[0045] Calculate the probability that the anchor sample is correctly retrieved:
[0046]
[0047] wherein, anchor represents the anchor; I represents the image, T represents the text, the subscript u represents the index of the anchor sample, that is, the index of the text or image corresponding to the anchor image or text, and the subscript k represents all indices except u in the batch, that is, the negative sample index of the anchor;
[0048] Propose the two boundary conditions, which are respectively:
[0049]
[0050] ∈ is a constant; represents the sum of all similarity log odds except for the index u.
[0051] Combining the two boundary conditions, the updated m in each iteration process is derived 1 and m 2 .
[0052] Preferably, after updating the values of the two boundary parameters m 1 and m 2 , the adaptive contrast learning loss is optimized batch by batch to achieve dynamic contrast learning between positive and negative samples, and the image encoder and the text encoder include:
[0053] Obtain the updated boundary parameters m 1 and m 2 ;
[0054] Based on the updated boundary parameters m 1 and m 2 , use the adaptive contrast learning objective function to calculate the loss, perform adaptive contrast learning, and obtain the trained image encoder and the text encoder
[0055] According to the second aspect of the present invention, there is provided a cross-modal image-text retrieval system based on adaptive contrast learning, including:
[0056] Similarity module: Use the image encoder and the text encoder to extract the visual modality features v and text modality features w of the image I and the text T respectively, and calculate the cosine similarity s between them in the same subspace;
[0057] Boundary parameter module: Introduce two boundary parameters m 1 and m 2 into the probability distribution of each pair of positive samples, and combine the cosine similarity s to calculate the probability of the image I corresponding to the text T Adjust the semantic consistency of each pair of positive samples and distinguish the distribution characteristics of cloned negative samples;
[0058] Significance Score Module: Define and calculate the significance scores of each batch of samples, and select the negative sample sequences with the largest and smallest significance scores in each batch of samples as the observed cloned negative samples S cln and the significant negative sample S cln ;
[0059] Anchor Module: Based on the observed cloned negative sample S cln and the significant negative sample S cln , according to the Gaussian discriminant analysis method, predict all potential cloned negative samples of each batch of samples and generate anchor samples;
[0060] Parameter Update Module: According to two boundary conditions of the similarity of the anchor samples, iteratively update the values of the two boundary parameters m 1 and m 2 ;
[0061] Contrastive Learning Module: After updating the values of the two boundary parameters m 1 and m 2 , adaptively contrastive learning loss, and optimize batch by batch to achieve dynamic contrastive learning between positive and negative samples, and train the image encoder and the text encoder
[0062] Inference Module: Use the trained image encoder and the text encoder to extract the visual modality features and text modality features of the image to be detected and the text to be detected, and calculate the similarity to obtain the cosine similarity of the two in the same subspace, as the basis for cross-modal image-text retrieval.
[0063] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects:
[0064] The cross-modal image-text retrieval method and system based on adaptive contrastive learning in the embodiments of the present invention, aiming at the cloned negative samples that are likely to appear in the image-text dataset, use Gaussian discriminant analysis to predict the potential cloned negative samples in each training batch without additional training, laying a foundation for further learning the potential high-dimensional key semantics of the cloned negative samples.
[0065] The cross-modal image-text retrieval method and system based on adaptive contrastive learning in the embodiments of the present invention jointly improve the traditional contrastive learning method, and invent an adaptive contrastive learning method. Through two dynamically fine-tuned boundary parameters, the semantics of the cloned negative samples are propagated during the training process, enhancing the tightness of the positive samples, and effectively solving the influence brought by the cloned negative samples. Brief Description of the Drawings
[0066] Figure 1Flowchart of a cross-modal image-text retrieval method based on adaptive contrastive learning in an embodiment of the present invention;
[0067] Figure 2 Comparison of cloned negative samples, observed positive samples, and significant negative samples in an embodiment of the present invention;
[0068] Figure 3 Schematic diagram of the degree of sample discrimination in the joint subspace in an embodiment of the present invention;
[0069] Figure 4 Schematic diagram of the framework structure of a cross-modal image-text retrieval method based on adaptive contrastive learning in an embodiment of the present invention;
[0070] Figure 5 Schematic diagram of the final result of image-text retrieval of a pair of cloned negative samples based on adaptive contrastive learning in a specific embodiment of the present invention;
[0071] Figure 6 Schematic diagram of the visualization result of the training process of the adaptive boundary parameter introduced by the network in a specific embodiment of the present invention;
[0072] Figure 7 Schematic diagram of the visualization result of the visual-text joint feature clustering in a specific embodiment of the present invention;
[0073] Figure 8 Comparison result of the final retrieval accuracy of different retrieval methods in a specific embodiment of the present invention. Detailed implementation manners
[0074] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings: These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.
[0075] In an embodiment of the present invention, a cross-modal image-text retrieval method based on adaptive contrastive learning is provided. As Figure 1 shown, it includes the following steps:
[0076] S100: Use an image encoder and a text encoder to extract the visual modality feature v of the image I and the text modality feature w of the text T respectively, and calculate the cosine similarity s between them in the same subspace;
[0077] S200: Introduce two boundary parameters m 1 and m 2 into the probability distribution of each pair of positive samples, and combine the cosine similarity s to calculate the probability Adjust the semantic consistency of each pair of positive samples and distinguish the distribution characteristics of cloned negative samples;
[0078] S300: Define and calculate the significance scores of each batch of samples, and select the negative sample sequences with the maximum and minimum significance scores in each batch of samples as the observed cloned negative sample S cln and the significant negative sample S sln ;
[0079] S400: Based on the observed cloned negative sample S cln and the significant negative sample S sln , according to the Gaussian discriminant analysis method, predict all potential cloned negative samples of each batch of samples and generate anchor samples;
[0080] S500: According to two boundary conditions of the similarity of the anchor samples, iteratively update the values of the two boundary parameters m 1 and m 2 ;
[0081] S600, after updating the values of the two boundary parameters m 1 and m 2 , adaptively contrastive learning loss, and optimize batch by batch to achieve dynamic contrastive learning between positive and negative samples, and train the image encoder and the text encoder
[0082] S700, use the image encoder and the text encoder trained by S100 - S600 to extract the visual modality features and text modality features of the image to be detected and the text to be detected, and calculate the similarity to obtain the cosine similarity of the two in the same subspace as the basis for cross-modal image-text retrieval.
[0083] The above embodiments are used to solve the problems of uncertain matching results and insufficient optimization caused by cloned negative samples in existing methods. By introducing two adaptive boundary parameters (scaling parameter m 1 and offset parameter m 2 ), dynamically adjust the tightness of positive samples and the distinguishability of cloned negative samples, thereby effectively improving the accuracy and robustness of image-text matching; by using the Gaussian discriminant analysis method that does not require additional training to judge anchor samples, the influence of cloned negative samples on the training process can be dynamically adjusted to optimize the training process.
[0084] In addition, it can further be applied to the weakly supervised image-text retrieval scenario, replacing manual annotation with pseudo-label descriptions generated by images, thereby improving the retrieval performance of the system on low-quality image-text datasets, laying a foundation for large-scale image-text contrastive learning, and having general applicability.
[0085] In a preferred embodiment of the present invention, in step S100, based on the image-text pair, image and text encoders are used to extract image and text features and calculate the similarity degree. Specifically:
[0086] S101, the image encoder converts the RGB pixels of the input image into a high-dimensional feature representation v.
[0087] S102, the text encoder converts the text into tokens and then into a high-dimensional text embedding w.
[0088] S103, cosine similarity calculation is performed on the feature representations v and w of the image I and the text T to obtain s(I,T), which is used to measure the similarity degree of the image-text pair.
[0089] Of course, the order of S101 and S102 is not uniquely limited. In other embodiments, the implementation steps can be S102 / S101 or implemented simultaneously.
[0090] During the training process, it is necessary to select the clone negative samples and the significant negative sample queue from each training batch and obtain the anchor samples. In a preferred embodiment of the present invention, the target anchor samples are obtained by sequentially implementing S200\S300\S400. Specifically:
[0091] S200, two boundary parameters m 1 and m 2 are introduced into the probability distribution of each pair of positive samples, and combined with the cosine similarity s, the probability of the image I corresponding to T is calculated to adjust the semantic consistency of each pair of positive samples and distinguish the distribution characteristics of the clone negative samples. The specific implementation process is as follows:
[0092] S201, the similarity p(I) after passing through the softmax transfer function can be expressed as:
[0093]
[0094] where m 1 and m 2 are two introduced boundary parameters for adaptively adjusting the semantic consistency of the positive sample image-text pair and distinguishing the distribution characteristics of the clone negative sample pair.
[0095] The denominator consists of M + 1 samples, specifically: one positive sample and M negative samples. If the momentum memory is not used, M is numerically equal to the training batch size N. τ is a temperature parameter that controls the overall supervision information. p i (I) represents the probability of the image I corresponding to T.
[0096] Furthermore, the adaptive contrast learning objective function can be expressed as:
[0097]
[0098] Among them, y(I)y i (I) is a one-hot encoded vector representing the true class (or text) of image I.
[0099] Pull the distances between positive sample image-text pairs closer, while pushing the distances between predefined negative sample pairs farther apart. E I~D represents the expected value, which is the expectation of the sum after sampling all images I from the data distribution D; H represents the cross-entropy loss function, which measures the difference between the probability distribution p(I) predicted by the model and the target label y(I); N represents the number of samples in the batch; the subscript i represents the target label of the i-th sample.
[0100] In some specific embodiments, if the momentum memory is not used, M is numerically equal to the training batch size N. If the momentum memory is used, the value of M can be set by hyperparameters. For example, M can be set to 4096 to cover as many negative samples as possible.
[0101] In the above embodiments, two boundary parameters are specifically introduced, as scaling and offset factors respectively, for synchronously enhancing the tightness of positive samples and including the supervision of cloned negative samples.
[0102] S300, Define and calculate the significance scores of each batch of samples, and select the negative sample sequences with the maximum and minimum significance scores in each batch of samples as the observed cloned negative sample S cln and the significant negative sample S sln . The specific process is as follows:
[0103] S301, To predict the potential cloned negative samples in each batch of image-text pairs, first define the significance score, which can be expressed as:
[0104]
[0105] s ii and s ij are both cosine similarities. s ii represents the positive sample similarity, that is, the value along the main diagonal in the similarity matrix. s ij represents the negative sample similarity, that is, the value at the coordinate (i, j) in the similarity matrix, that is, the similarity between the i-th image or text sample and the j-th text or image sample;
[0106] The significance score better reflects the degree of significance by considering the similarity of negative sample pairs.
[0107] S302, Take the sample pair with the highest significance score as the significant negative sample Ssln , which contains 1 positive sample and M significantly negative samples. Define the sample pair with the lowest significance score as the observed clone negative sample S cln , which contains 1 positive sample and M negative samples;
[0108] Figure 2 Shows the difference between the clone negative samples and the relatively common significantly negative samples. Specifically, image-text pairs often contain semantically related text annotations and similar visual cues, making them difficult to distinguish in practice (reflected in their very close similarity scores). Without a doubt, retrieving such samples is a suboptimal result because the true annotations manually labeled exhibit more precise and fine-grained semantics and a closer corresponding relationship.
[0109] S400, based on the observed clone negative sample S cln and the significantly negative sample S sln , according to the Gaussian discriminant analysis method, predict all potential clone negative samples of each batch of samples and generate anchor samples. The specific process is as follows:
[0110] S401, First, starting from the class probability, it can be expressed as:
[0111]
[0112] In the formula: Using the binary classification method, respectively represent whether the pairwise similarity is / is not a clone negative sample; s represents the cosine similarity score; represents the original score of the similarity without probability output, represents the original score of the similarity without probability output. A clone negative sample predictor that does not require additional training can be obtained by analyzing the sample distribution of clone negative samples and Gaussian discriminant analysis.
[0113] S402, In Gaussian discriminant analysis, features are usually assumed to follow a Gaussian distribution with the same covariance. Since the similarity score is a univariate variable, the above distribution expression can be simplified to a univariate distribution. Among them and Combining this assumption with we can get:
[0114]
[0115] where, π 1 and π 0 are respectively and prior probabilities; Given S sln and S clnIt consists of the most representative significant negative samples and potential clone negative samples within a training batch respectively, and uses their respective negative samples to represent the empirical mean and variance μ 1 、μ 0 、σ 1 and σ 0 ;
[0116] S403: The selection of potential in-batch clone negative samples is based on the following criteria:
[0117]
[0118] S * represents the predicted clone negative samples in each batch of samples.
[0119] S404, the anchor is defined as the median sample in the set of S * to obtain the anchor sample.
[0120] In the above embodiment, in order to gradually adjust the boundary parameters, based on Gaussian discriminant analysis, an anchor sample is selected from the similarity scores of each batch without introducing additional training. The anchor sample can effectively reflect the intensity of the in-batch clone negative samples and impose penalties through the boundary parameters, thereby adaptively expanding the distance between the positive samples and the clone negative samples.
[0121] In a preferred embodiment of the present invention, in step S500, according to two boundary conditions of the similarity of the anchor sample, the values of two boundary parameters m 1 and m 2 are iteratively updated, and the specific process is as follows:
[0122] S501, calculate the anchor probability.
[0123] In order to calculate the boundary parameters m 1 and m 2 introduced in step S2 in each batch learning, a specific analysis is carried out on the similarity p(I) passing through the softmax transfer function. First, the corresponding probability can be calculated from the anchor obtained in the fourth step of the equation (referring to the category label sample, that is, the probability of the correct matching of the anchor image-text pair):
[0124]
[0125] where, anchor represents the anchor; I represents the image, T represents the text, the subscript u represents the index of the anchor sample, that is, the index of the text or image corresponding to the anchor image or text, and the subscript k represents all indices except u in the batch, that is, the negative sample index of the anchor;
[0126] S502, propose boundary conditions.
[0127] Since the anchor represents the set-average probability of potential clone negative samples within each batch, it correspondingly reflects to a great extent the approximate convergence degree of the model. Assuming that the probability of can be controlled, the overall supervision within each batch can be well regulated. Therefore, based on gradually adjust m 1 and m 2 scheme. Analyzed the boundary conditions: Since the cross-entropy loss is expressed as When the similarity score s(I_u, T_u) of the anchor continuously approaches 1 (indicating high similarity), a relatively small m 1 will impose an unnecessary penalty loss on the correctly matched image-text pairs. Therefore, when s(I_u, T_u) is close to 1, should also be as close to 1 as possible to enhance the ability to distinguish clone negative samples. This condition can be expressed as:
[0128]
[0129] Combined with the expression, another boundary condition:
[0130]
[0131] ∈ is a constant used to avoid invalid calculations when m 1 reaches 0; represents the sum of all similarity log-odds except for the index u.
[0132] S503, derive the new m 1 and m 2 .
[0133] Combining the two boundary conditions, the updated m 1 and m 2 in each iteration process can be derived, and adaptively contrastive learning loss during training. In this way, can effectively propagate the semantics of clone negative samples through m 1 and m 2 .
[0134] The above embodiments consider and handle the clone negative samples appearing in the dataset, and effectively promote the model to distinguish them during training through an adaptive image-text contrastive learning method.
[0135] Based on the same inventive concept, in other embodiments of the present invention, there is provided a cross-modal image-text retrieval system based on adaptive contrastive learning, including:
[0136] Similarity module: Using an image encoder and a text encoder to extract the visual modality feature v of the image I and the text modality feature w of the text T respectively, and calculate the similarity to obtain the cosine similarity s of the two in the same subspace;
[0137] Boundary parameter module: Introduce two boundary parameters m 1 and m 2 into the probability distribution of each pair of positive samples, and combine with the cosine similarity s to calculate the probability of the image I corresponding to the text T to adjust the semantic consistency of each pair of positive samples and distinguish the distribution characteristics of cloned negative samples;
[0138] Significance score module: Define and calculate the significance score of each batch of samples, and select the negative sample sequences with the largest and smallest significance scores in each batch of samples as the observed cloned negative sample S cln and the significant negative sample S cln ;
[0139] Anchor module: Based on the observed cloned negative sample S cln and the significant negative sample S sln predict all potential cloned negative samples of each batch of samples and generate anchor samples according to the Gaussian discriminant analysis method;
[0140] Parameter update module: Iteratively update the values of the two boundary parameters m 1 and m 2 according to two boundary conditions of the similarity of the anchor samples;
[0141] Contrastive learning module: After updating the values of the two boundary parameters m 1 and m 2 adaptively contrastive learning loss, optimize batch by batch to achieve dynamic contrastive learning between positive and negative samples, and train the image encoder and the text encoder
[0142] Inference module: Use the trained image encoder and the text encoder to extract the visual modality feature and the text modality feature of the image to be detected and the text to be detected, and calculate the similarity to obtain the cosine similarity of the two in the same subspace, as the basis for cross-modal image-text retrieval.
[0143] In the above examples of the present invention, the specific implementation of each module / unit can refer to the implementation technology of the corresponding steps of the cross-modal image-text retrieval method based on adaptive contrastive learning in the above embodiments, which will not be elaborated here.
[0144] To verify the feasibility and effectiveness of the above embodiments, in a specific embodiment of the present invention, the image frames used are from the image-text pairs in the database Flickr30K for cross-modal image-text retrieval performance evaluation. Among them, Flickr30K contains 31,783 images, and each image corresponds to 5 different sentences. It is divided into a training set / test set / validation set in the ratio of 29,783 / 1,000 / 1,000. MS-COCO contains 123,287 images, and each image corresponds to 5 sentences. It is divided into a training set / test set / validation set according to 113,287 / 5,000 / 5,000 images. The test set is further divided into MS-COCO 1K (the average result of 5 test sets) and MS-COCO 5K (the result of 5,000 test images).
[0145] Figure 3 It shows that under the adaptive contrast learning method in the above embodiments, the discrimination distance between negative samples is more significant than that of the widely used triplet ranking loss and contrast loss, verifying its effectiveness.
[0146] Figure 4 It shows the composition of the cross-modal image-text retrieval network structure. Specifically, significant negative samples and observed clone negative samples are selected through significance scores. Subsequently, based on Gaussian discriminant analysis, two introduced boundary parameters are dynamically adjusted in the loss function to achieve a more refined contrast learning goal. The adaptive contrast learning method effectively adjusts the distance between positive samples and clone negative samples by gradually adjusting the boundary parameters. Different from traditional contrast learning methods, this network structure can adaptively adjust parameters according to the intensity of clone negative samples, thus enhancing robustness in challenging scenarios.
[0147] In summary, the method of this embodiment effectively improves the accuracy and robustness of image-text matching by introducing two adaptive boundary parameters (scaling parameter and offset parameter) to dynamically adjust the tightness of positive samples and the discrimination of clone negative samples; the Gaussian discriminant analysis method that does not require additional training is used to judge anchor samples, enabling the influence of clone negative samples on the training process to be dynamically adjusted to optimize the training process.
[0148] As Figure 5As shown, it is a schematic diagram of the final result of the image text retrieval based on this embodiment, with the retrieval result ranking (Rank@K) as the evaluation index (see Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He. Stacked cross attention for image-text matching. In Proceedings of the European conference on computer vision (ECCV), pp. 201–216, 2018.). Specifically, given a query image / text, the text / image in the database is sorted according to the similarity with the query image / text. If at least one correct text / image appears in the top K positions of the ranking, the retrieval is considered correct, otherwise it is wrong. Figure 5 The effect of the adaptive contrastive learning method in the embodiment of the present invention on processing cloned negative samples is analyzed, the matching results are sorted according to the similarity score, and compared with two other widely used objective functions, revealing the following conclusions: Compared with contrast loss and triple ranking loss, adaptive contrastive learning loss performs better on R@1 and can successfully exclude other potential negative samples, proving that the effect of the embodiment of the present invention on processing cloned negative samples is significant; at the same time, under adaptive contrastive learning loss, the cross-modal fusion representation learned by the model performs better in identifying cloned negative samples. For example, for the cloned negative sample of "a group of people running or participating in a marathon in the city", its similarity score is only 0.30. In contrast, although the model retrieved the true label under contrast loss and triple ranking loss, their similarity scores for cloned negative samples still showed suboptimal solutions. The retrieval method based on adaptive contrastive learning has a better distinction effect for more challenging samples.
[0149] Figure 6 This is the convergence process of the boundary parameters introduced by the present invention during model training. 2 gradually increases and stabilizes around 0.36, while m 1 After experiencing a certain amount of oscillation, it stabilized around the 10th iteration cycle, with a stable value of about 38.1, showing an excellent convergence speed.
[0150] Figure 7 The distribution of the visual-text joint embedding in the t-SNE space at the 10th iteration cycle is shown respectively. Compared with the contrast loss and triple ranking loss, the adaptive contrast loss shows better clustering performance, further proving its effectiveness in accelerating the training convergence process of the cross-modal image-text retrieval model.
[0151] Figure 8 It is the comparison result of the final retrieval accuracy based on the performance obtained in this embodiment. It includes two datasets, Flickr30K and MS-COCO. A variety of competitive cross-modal image-text retrieval methods are selected as the baseline models, and the plug-and-play effect of the embodiment of the present invention, that is, the image-text retrieval method based on adaptive contrast learning (AdaCL), is verified. It can be observed that the embodiment of the present invention outperforms the original baseline method in each type of backbone network. On the Flickr30K and MS-COCO datasets, the embodiment of the present invention has achieved very significant absolute improvements in (R@1, R@5, R@10) in the image-text retrieval task, indicating that the embodiment of the present invention can be applied to more general network architectures without the need to additionally design a cross-modal fusion mechanism.
[0152] In summary, the cross-modal image-text retrieval method based on adaptive contrast learning provided in this embodiment considers and solves the problems of uncertainty of matching results and insufficient optimization caused by cloned negative samples, and reduces the impact and challenges on the cross-modal image-text retrieval model. Adaptive contrast learning is invented. By introducing two tunable boundary parameters and an anchor point, the tightness between image-text pairs is dynamically adjusted, and the semantic information of potential cloned negative samples is learned; the anchor point is selected according to the distribution of negative samples without explicit training, so as to achieve progressive adjustment and enhance the supervision effect in a small batch; the adaptive contrast learning method shows excellent robustness, and the superior performance demonstrates the great potential of adaptive contrast learning in reducing the dependence on manual annotation.
[0153] Although the content of the present invention has been introduced in detail through the above preferred embodiments, it should be recognized that the above description should not be considered as a limitation of the present invention. After those skilled in the art have read the above content, various modifications and alternatives to the present invention will be obvious. Therefore, the protection scope of the present invention should be defined by the appended claims.
Claims
1. A cross-modal image-text retrieval method based on adaptive contrastive learning, characterized in that: include: Using Image Encoder and text encoder Extract the visual modality features v and text modality features w of image I and text T respectively, and perform similarity calculation to obtain the cosine similarity s of the two in the same subspace; Introduce two boundary parameters m1 and m2 into the probability distribution of each pair of positive samples, and combine the cosine similarity s to calculate the probability that the image I corresponds to the text T Adjusting the semantic consistency of each pair of positive samples and distinguishing the distribution characteristics of cloned negative samples; Define and calculate the significance score of each batch of samples, and select the negative sample sequences with the largest and smallest significance scores in each batch of samples as observed clone negative samples S cln and significant negative samples S sln ; Based on the observed clone negative sample S cln and the significant negative sample S sln , according to the Gaussian discriminant analysis method, predict all potential clone negative samples of each batch of samples and generate anchor samples; Iteratively updating the values of the two boundary parameters m1 and m2 according to the two boundary conditions of the similarity of the anchor point samples; After updating the values of the two boundary parameters m1 and m2, adaptive contrast learning loss is used to optimize batch by batch to achieve dynamic contrast learning between positive and negative samples and train the image encoder. and text encoder Use the trained image encoder and the text encoder The visual modal features and text modal features of the image to be detected and the text to be detected are extracted, and the similarity is calculated to obtain the cosine similarity of the two in the same subspace as the basis for cross-modal image and text retrieval.
2. The cross-modal image-text retrieval method based on adaptive contrastive learning according to claim 1, characterized in that: The use of an image encoder and text encoder The visual modality features v and text modality features w of image I and text T are extracted respectively, and the similarity is calculated to obtain the cosine similarity s of the two in the same subspace, including: The image encoder Convert the RGB pixels of the image I into feature expressions v; The text encoder Convert the text T into word-grams, and then convert the word-grams into text embeddings w; The cosine similarity calculation of the feature expression v and the text embedding w is performed to obtain s(I, T), which is used to measure the similarity between the image and text pairs.
3. The cross-modal image-text retrieval method based on adaptive contrastive learning according to claim 1, characterized in that: The two boundary parameters m1 and m2 are introduced into the probability distribution of each pair of positive samples, and the probability p of the image I corresponding to the text T is calculated by combining the cosine similarity s. i (I) adjusting the semantic consistency of each pair of positive samples and distinguishing the distribution characteristics of cloned negative samples, including: Where i represents the subscript index of the text T corresponding to the image I, that is, (I, T i ) is a pair of image-text pairs, M represents the number of negative samples in a batch, m1 and m2 are two introduced boundary parameters, and the denominator is a positive sample m1(s(I,T i )-m2) and M negative samples composition.
4. The cross-modal image-text retrieval method based on adaptive contrastive learning according to claim 3, characterized in that: The probability p of the image I corresponding to T i (I) Establish an adaptive contrastive learning objective function, specifically: Among them, E I~D represents the expected value, which represents the expectation of the sum of all images I sampled from the data distribution D; H represents the cross entropy loss function; N represents the number of samples in the batch; the subscript i represents the target label of the i-th sample, y i (I) is a one-hot encoded vector representing the true category or text of image I.
5. The cross-modal image-text retrieval method based on adaptive contrastive learning according to claim 1, characterized in that: The definition and calculation of the significance score of each batch of samples, and the negative sample sequences with the largest and smallest significance scores in each batch of samples are selected as the observed clone negative samples S cln and significant negative samples S sln ,include: The significance score is defined as the difference between the average cosine similarity of the positive sample and the average cosine similarity of the negative sample: s ii and ij Both are cosine similarity, s ii Represents the similarity of positive samples, that is, the value along the main diagonal in the similarity matrix, s ij Represents the negative sample similarity, that is, the value of the coordinate (i, j) in the similarity matrix, that is, the similarity between the i-th image or text sample and the j-th text or image sample; Calculate the significance score of each batch of samples according to the above definition; Select the sample pair with the highest significance score as the significant negative sample S sln , which contains 1 positive sample and M negative samples; Select the sample pair with the lowest significance score as the observed clone negative sample S cln , which contains 1 positive sample and M potential clone negative samples.
6. The cross-modal image-text retrieval method based on adaptive contrastive learning according to claim 5, characterized in that: The significance score is the update criterion for the boundary parameters m1 and m2, which is calculated and iterated during each batch training.
7. The cross-modal image-text retrieval method based on adaptive contrastive learning according to claim 1, characterized in that: The negative sample S is cloned based on the observation cln and the significant negative sample S sln , according to the Gaussian discriminant analysis method, predict all potential clone negative samples of each batch of samples and generate anchor samples, including: Calculate the category probability that each pair of image and text is a cloned negative sample, and its distribution expression is: In the formula: using the binary classification method, They respectively indicate whether the pairwise similarity is a clone negative sample or not; s indicates the cosine similarity score; Represents the raw score of similarity without probability output, Represents the raw score of similarity without probability output; Based on the properties of Gaussian discriminant: the features have Gaussian distribution with the same covariance, and the similarity score is a univariate variable, the distribution expression of the category probability is simplified to a univariate distribution: and in, The representative mean is u0 and the variance is Gaussian distribution, The representative mean is μ1 and the variance is Gaussian distribution of The distribution expression is arranged as follows: Among them, π1 and π0 are and The prior probability of S sln and S cln They are respectively composed of the most representative significant negative samples and potential clone negative samples in a training batch, and their respective negative samples are used to represent the empirical mean and variance μ1, μ0, σ1 and σ0; Potential clonal negative samples in each batch of samples are selected based on the following criteria: S * represents the predicted clone negative samples in each batch of samples; Define the S * The median sample in the set is the anchor sample.
8. The cross-modal image-text retrieval method based on adaptive contrastive learning according to claim 1, characterized in that: The iterative updating of the values of the two boundary parameters m1 and m2 according to the two boundary conditions of the similarity of the anchor point samples comprises: Calculate the probability that the anchor point sample is correctly retrieved: in, anchor represents the anchor point; I represents the image, T represents the text, the subscript u represents the index of the anchor point sample, that is, the index of the text or image corresponding to the anchor point image or text, and the subscript k represents all indexes in the batch except u, that is, the negative sample index of the anchor point; Put forward the The two boundary conditions are: ∈ is a constant; represents the sum of the log-odds of all similarities except index u; Combining the two boundary conditions, the updated m1 and m2 in each iteration process are derived.
9. The cross-modal image-text retrieval method based on adaptive contrastive learning according to claim 4, characterized in that: After updating the values of the two boundary parameters m1 and m2, adaptive contrast learning loss is used to optimize batch by batch to achieve dynamic contrast learning between positive and negative samples and train the image encoder. and text encoder include: Get updated boundary parameters m1 and m2; Based on the updated boundary parameters m1 and m2, the adaptive contrast learning objective function is used to calculate the loss, and adaptive contrast learning is performed to obtain a trained image encoder. and text encoder 10. A cross-modal image-text retrieval system based on adaptive contrastive learning, characterized in that: include: Similarity module: using image encoder and text encoder Extract the visual modality features v and text modality features w of image I and text T respectively, and perform similarity calculation to obtain the cosine similarity s of the two in the same subspace; Boundary parameter module: introduces two boundary parameters m1 and m2 into the probability distribution of each pair of positive samples, and calculates the probability that the image I corresponds to the text T in combination with the cosine similarity s Adjusting the semantic consistency of each pair of positive samples and distinguishing the distribution characteristics of cloned negative samples; Significance score module: defines and calculates the significance score of each batch of samples, and selects the negative sample sequences with the largest and smallest significance scores in each batch of samples as observed clone negative samples S cln and significant negative samples S sln ; Anchor module: clone negative samples S based on the observations cln and the significant negative sample S sln , according to the Gaussian discriminant analysis method, predict all potential clone negative samples of each batch of samples and generate anchor samples; Parameter updating module: iteratively updating the values of the two boundary parameters m1 and m2 according to the two boundary conditions of the similarity of the anchor point samples; Contrastive learning module: After updating the values of the two boundary parameters m1 and m2, adaptive contrast learning loss is used to optimize batch by batch to achieve dynamic contrast learning between positive and negative samples and train the image encoder and text encoder Inference module: Use the trained image encoder and the text encoder The visual modal features and text modal features of the image to be detected and the text to be detected are extracted, and the similarity is calculated to obtain the cosine similarity of the two in the same subspace as the basis for cross-modal image and text retrieval.
Citation Information
Patent Citations
OSCAR-based image-text retrieval model training method and image-text retrieval realization method
CN117390213A
Comparative learning unsupervised cross-modal hash retrieval algorithm based on graph attention mechanism
CN119377462A
Image-text retrieval method based on comparative learning and modal fusion
CN119441512A
Visual question answering method and apparatus, electronic device and storage medium
WO2024164616A1
Cited By
Cross-modal image-text retrieval method and device based on decoupling learning and medium
CN121524371A
Non-supervision comparison cross-modal hash retrieval method based on lifelong learning
CN121560932A
Unsupervised contrastive cross-modal hashing retrieval method based on lifelong learning
CN121560932B