Graphical and Textual Cross-Modal Retrieval Method and System Based on Soft Distance Constraint of False Negative Samples

Through the cross-modal search method of graphic and text based on soft distance constraints of false negative samples, the problem of false negative samples in cross-modal search of marine remote sensing graphic and text is solved, and the accuracy and robustness of the search is improved.

CN120104825BActive Publication Date: 2025-07-18OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510569986.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-18
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

There are a large number of false negative samples in the existing cross-modal search methods for marine remote sensing graphics and text, which makes it difficult to fit or fit errors in the model, reducing the search accuracy.

Method used

The cross-modal search method of graphic and text based on soft distance constraints of false negative samples is adopted. Through prompt perception Transformer decoder, mixed distribution false negative sample mining module and loss calculation module, the cross-modal semantic gap is narrowed, single-modal data noise is removed, and retrieval accuracy is improved.

Benefits of technology

Effectively mine false negative samples in the dataset, improve the accuracy of cross-modal retrieval, reduce the cost of calculation, enhance the robustness of feature representation, and improve the credibility and data utilization of sample judgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104825B_ABST
    Figure CN120104825B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of cross-modal image-text retrieval, and discloses an image-text cross-modal retrieval method and system based on soft distance constraint of false negative samples. The method includes the steps of feature extraction; the step of mining false negative samples with a mixed distribution: designing a sample pool for model fitting, where the sample pool contains affinity matrices of cross-modal features with the same semantics and different semantics in the same proportion, and using the expectation-maximization algorithm to fit the Gaussian mixture distribution of the affinity matrix; calculating the affinity matrix of the current input image features and text features, and using the maximum a posteriori method to determine the sub-population distribution to which the negative samples belong; classifying the negative samples into "false negative samples", "true negative samples" and "fuzzy samples" by comparing the cumulative probability with the significance level; calculating the soft distance constraint triplet loss. By means of the present invention, the cross-modal semantic gap is narrowed, the noise in the unimodal data is removed, the "false negative" samples in the data set are efficiently mined, and the accuracy of cross-modal retrieval is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cross-modal image and text retrieval, and in particular to an image and text cross-modal retrieval method and system based on soft distance constraints of false negative samples. Background Art

[0002] As one of the important means of cross-modal retrieval of marine remote sensing data, image-text retrieval has attracted more and more attention from researchers in recent years. At present, cutting-edge image-text retrieval methods are committed to using saliency mining methods such as attention mechanisms and graph convolutional networks to filter out redundant background information in images and texts, and improve the matching accuracy of image-text retrieval.

[0003] However, the above methods still have the following problems when applied to the ocean:

[0004] Existing remote sensing image and text cross-modal retrieval methods use triplet loss to train the network. Triplet loss involves three elements, namely, anchor points, positive samples corresponding to anchor points, and negative samples corresponding to anchor points. Triplet loss aims to reduce the distance between anchor points and positive samples, and increase the distance between anchor points and negative samples. However, there are a large number of "false negative" samples in the marine remote sensing cross-modal retrieval dataset, that is, there are many negative samples that match anchor points. There are three main reasons for this phenomenon. First, the high similarity between marine remote sensing data. Compared with ordinary images, the similarity between marine remote sensing data is higher, resulting in similar or even identical descriptions of many data. For example, image 1 and image 2 can both be described as "two cargo ships parked on the sea." Therefore, when the anchor point is image 1, text 2 corresponding to image 2 is a "false negative" sample. Second, the modal gap between images and texts. Compared with texts, images contain richer semantic information; compared with images, texts give more concise descriptions. This modal gap will produce many "false negative" samples. Third, errors caused by human labeling. The texts in the dataset are all manually labeled based on the images, and there are too many human interference factors.

[0005] Existing methods still use triplet loss for modeling, and the distance of "false negative" samples is widened, making the model difficult to fit or fitting incorrectly, thereby reducing the robustness of feature representation and ultimately reducing the accuracy of cross-modal image and text retrieval. Summary of the invention

[0006] In view of the shortcomings of the prior art, the present invention provides a method and system for cross-modal retrieval of images and texts based on soft distance constraints of false negative samples, and designs a prompt-aware Transformer decoder, a mixed distribution false negative sample mining module, and a loss calculation module. The present invention narrows the cross-modal semantic gap and removes the noise in the unimodal data; it mines the "false negative" samples in the data set in a more efficient and reliable way, and improves the accuracy of cross-modal retrieval.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0008] A cross-modal retrieval method for images and texts based on soft distance constraints of false negative samples, comprising the following steps:

[0009] Feature extraction step: For image data I and text data T, extract image feature X and text feature F through an encoder-decoder.

[0010] Step of mining false negative samples with mixed distribution: Design a sample pool for model fitting. The sample pool contains affinity matrices of cross-modal features with the same semantics and different semantics in the same proportion. Use the expectation-maximization algorithm to fit the Gaussian mixture distribution of the affinity matrix; based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image feature and text feature, and use the maximum a posteriori method to determine the sub-population distribution to which the negative samples belong; then calculate the cumulative probability, and by comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples" and "fuzzy samples", and store them in the corresponding sets.

[0011] Loss calculation step: On the basis of the triplet loss, calculate the soft distance constraint triplet loss, keep the calculation of the true negative samples, delete the calculation of the false negative samples, and modify the calculation of the fuzzy samples.

[0012] Further, the feature extraction step is specifically as follows: For image data I and text data T, first use an image encoder and a text encoder to extract image feature V and text feature U respectively; then generate robust image feature X and text feature F through a prompt-aware Transformer decoder.

[0013] Further, the prompt-aware Transformer decoder includes two decoders, which are respectively applied to image and text features; the two decoders share parameters. When processing image and text features, use image feature V and text feature U as the key and value respectively, and use a prompt P as the query to implement the training of the network, and generate robust image feature X and text feature F respectively.

[0014] Further, in the step of mining false negative samples with mixed distribution, design a sample pool for model fitting. The sample pool contains affinity matrices of cross-modal features with the same semantics and different semantics in the same proportion , and use the expectation-maximization algorithm to fit the Gaussian mixture distribution of the affinity matrix, which is realized in the following way:

[0015] For the elements in make a frequency histogram, and use the expectation-maximization algorithm for The frequency histogram is fitted to a mixture of Gaussian distributions, and the density function of this distribution is

[0016] ;

[0017] where, is the mixing coefficient, is the probability density function of the Gaussian distribution, and are the parameters of the Gaussian distribution, represents the parameter space; g represents the sub - population index, g = 1, 2. When g = 1, it indicates the distribution of the sub - population with a small mean, and when g = 2, it indicates the distribution of the sub - population with a large mean; the distribution with a large mean is defined as the "same - semantics distribution", and the one with a small mean is defined as the "different - semantics distribution".

[0018] Furthermore, based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image features and text features, and use the maximum a posteriori method to determine the sub - population distribution to which the negative samples belong; on this basis, calculate the cumulative probability, and by comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples" and "ambiguous samples", and store them in the corresponding sets; it is achieved in the following way:

[0019] The judgments of "false negative samples", "true negative samples" and "ambiguous samples" are relative to the anchor points. When performing image - to - text retrieval, for image anchor point i, among all negative text samples j, j = 1,..., B, ; based on the image feature X and the negative text feature F, calculate their affinity matrix C; based on the fitted Gaussian mixture distribution, calculate the maximum a posteriori for the elements in C to determine the sub - population distribution to which the negative text sample j belongs, that is, determine the value of g. g = 2 represents the "same - semantics distribution", and g = 1 represents the "different - semantics distribution";

[0020] On this basis, calculate the cumulative probability, and by comparing the cumulative probability with the significance level, classify the negative text samples into "false negative samples", "true negative samples" and "ambiguous samples", and store them in the corresponding sets; the calculation formula for the cumulative probability is

[0021] ;

[0022] x is the independent variable, and its value ranges from 0 to , , are the parameters of the updated Gaussian distribution, is the cumulative probability; then, define the significance level , when , and g = 1, the negative text sample is a "true negative sample" and is stored in the set , when , and g = 2, the corresponding negative sample text is "false negative sample", which is stored in the set , and other negative sample texts are "ambiguous samples", which are stored in the set ;

[0023] Similarly to image-to-text retrieval, text-to-image retrieval obtains the image "true negative sample" set , "false negative sample" set , and "ambiguous sample" set in the same way.

[0024] Further, the triplet loss is:

[0025] ;

[0026] The processing of the "true negative sample" by the soft distance-constrained triplet loss is

[0027] ;

[0028] The processing of the "ambiguous sample" by the soft distance-constrained triplet loss is

[0029] ;

[0030] The processing of the "false negative sample" by the soft distance-constrained triplet loss is to directly delete it. Therefore, the soft distance-constrained triplet loss is

[0031] ;

[0032] Among them, represents the soft distance-constrained triplet loss, represents the similarity between the image and the positive sample pair of the text, represents the similarity between the image and the corresponding "true negative" text, represents the similarity between the image and the corresponding "ambiguous" text, represents the similarity between the image and the corresponding "negative" text, represents the similarity between the text and the corresponding "true negative" image, represents the similarity between the text and the corresponding "ambiguous" image, represents the similarity between the text and the corresponding "negative" image; and respectively represent the "true negative" image sample and the "true negative" text sample; and respectively represent the "ambiguous" image sample and the "ambiguous" text sample; and respectively represent the "negative" image sample and the "negative" text sample; , represents the interval parameter, which is used to widen the gap between the image and the positive text pair and the image and the negative text pair; is the weight of the "blurred" sample corresponding to the image, is the weight of the "blurred" sample corresponding to the text;

[0033] ;

[0034] represents the adjustment factor, using the same calculation method.

[0035] The present invention also provides a cross-modal retrieval method for images and texts based on soft distance constraints of false negative samples to implement the method described above. The system includes a feature extraction module, a mixed distribution false negative sample mining module, and a loss calculation module.

[0036] The feature extraction module includes an image encoder and a text encoder, which respectively extract the image feature V and the text feature U; it also includes two prompt-aware Transformer decoders, which are respectively applied to the image feature V and the text feature U; the two decoders share parameters. When processing the image and text features, the image feature V and the text feature U are respectively used as the key and the value, and the prompt P is used as the query to implement the training of the network, and the robust image feature X and the text feature F are respectively generated.

[0037] The mixed distribution false negative sample mining module is designed for a sample pool for model fitting. The sample pool contains affinity matrices of cross-modal features with the same semantics and different semantics in the same proportion, and uses the expectation-maximization algorithm to fit the Gaussian mixture distribution of the affinity matrix; based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image feature and text feature, and use the maximum a posteriori method to determine the sub-population distribution to which the negative sample belongs; on this basis, calculate the cumulative probability, and by comparing the cumulative probability with the significance level, classify the negative sample into "false negative sample", "true negative sample" and "blurred sample", and store them in the corresponding sets.

[0038] The loss calculation module is used to calculate the soft distance constraint triplet loss.

[0039] Compared with the prior art, the advantages of the present invention are:

[0040] First, the third-party medium of prompts is used to narrow the cross-modal gap while reducing the computational cost. In the prompt-aware Transformer decoder module, the dimensions of the generated features F and Z are consistent with the prompt, and the dimensions of the prompt are smaller than the original image and text features V and U, which reduces the computational cost for the subsequent modeling process; in addition, the prompt is similar to a medium, that is, the original features U and V are compared with the medium P and converted into a third-party representation, which not only narrows the semantic gap, but also removes the noise in the unimodal data.

[0041] Second, the "false negative", "true negative" and "fuzzy" samples in the data set are mined in a more efficient and reliable way. In the mixed distribution false negative sample mining module, the sample pool is used to fit the Gaussian mixture distribution. The calculation method of this distribution is more stable because the sample pool contains affinity matrices of the same semantic and different semantic cross-modal features in the same proportion. In this case, one distribution will not be covered by another distribution, such as the false negative distribution will be covered by the true negative distribution. As the network is trained, the features are continuously optimized, and the distribution is also continuously optimized, so as to fit a more reliable distribution. Based on this distribution, the affinity matrix of the current input image feature and text feature is calculated, and the maximum a posteriori method is used to determine the sub-population distribution g=1 or g=2 to which the negative sample belongs. In addition, by comparing the cumulative probability with the significance level, the negative samples are classified into "false negative samples", "true negative samples" and "fuzzy samples", and stored in the corresponding sets. The interference of noise in the samples is eliminated by using statistical methods. On the other hand, since the set is obtained by comparing the cumulative probability with the significance level, it is regarded as a kind of credibility and can be adaptively adjusted according to the distribution shape, which improves the credibility of sample judgment.

[0042] Third, the valuable information in negative samples is fully utilized in the form of soft distance constraints. In the triple loss module of soft distance constraints, the calculation of true negative samples is maintained, the calculation of false negative samples is deleted, and the calculation of fuzzy samples is modified. Different calculation strategies are adopted for different samples, which not only improves the utilization of data, but also facilitates the identification of positive and negative samples, so that the network can learn the potential association between image and text pairs more deeply. In addition, in the calculation process of fuzzy samples, the cumulative probability is combined with the similarity. By adjusting the weight of the fuzzy samples, a larger weight is given to the distribution boundary of the same semantics and different semantics similarity (affinity matrix), which increases the boundary of the same semantics and different semantics. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0044] Figure 1 It is a schematic diagram of the method flow of the present invention;

[0045] Figure 2 It is a visual verification experiment of the sample pool in the embodiments of the present invention. Among them, (a) is the Gaussian mixture distribution fitted by the similarity calculated from the cross-modal data in the pool in the 0th round, (b) is the Gaussian mixture distribution fitted by the similarity calculated from the cross-modal data in the pool in the 1st round, (c) is the Gaussian mixture distribution fitted by the similarity calculated from the cross-modal data in the pool in the 2nd round, and (d) is the Gaussian mixture distribution fitted by the similarity calculated from the cross-modal data in the pool in the 3rd round. Specific embodiments

[0046] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments.

[0047] Combined with Figure 1 As shown, this embodiment provides a cross-modal retrieval method of image and text based on soft distance constraint of false negative samples, following the main framework of TGKT. In the feature extraction part, for image data I and text data T, the parameters pre-trained in VIT and BERT are still used to extract features V and U. Different from TGKT, the method proposed by the present invention no longer adopts the text-guided image feature extraction module, but innovatively designs respectively in the feature extraction stage and the similarity matching stage: a prompt-aware Transformer decoder, mining of false negative samples with a mixed distribution, and a soft distance constraint triplet loss.

[0048] Combined with Figure 1 As shown, the following details the detailed steps of the present invention.

[0049] 1. Steps of feature extraction: For image data I and text data T, image features X and text features F are extracted through an encoder-decoder.

[0050] The steps of feature extraction are specifically as follows: For image data I and text data T, first, an image encoder and a text encoder are used to extract image features V and text features U respectively; then, a prompt-aware Transformer decoder is used to generate robust image features X and text features F.

[0051] The aforementioned prompt-aware Transformer decoder: It contains two decoders, which are respectively applied to the image feature V and the text feature U; the two decoders share parameters. When processing the image and text features, the image feature V and the text feature U are respectively used as the key and value, and through a prompt P as the query to realize the training of the network, and robust image feature X and text feature F are respectively generated.

[0052] Specifically, this module designs a self-training prompt P, which is transmitted as a query to the two decoders. At the same time, for the Transformer decoder of the image, the image feature V is used as the key and value to realize the training of the Transformer network, and a robust image feature F is generated.

[0053] ;

[0054] Among them, represents the Transformer decoder, , , are the parameters to be trained in the network, converting P into a query and converting V into a key and value.

[0055] Similarly, for the Transformer decoder of the text, the text feature U is used as the key and value to generate a robust text feature X.

[0056] ;

[0057] Among them, , , are the parameters to be trained in the network, converting P into a query and converting U into a key and value.

[0058] The two decoders finally generate the robust image feature F and the text feature X. This prompt is updated by backpropagation. Its advantages are that, on the one hand, the dimensions of the generated features F and X are the same as those of the prompt, and the dimension of the prompt is smaller than the original image and text features V and U, reducing the computational cost for the subsequent modeling process; on the other hand, the prompt is similar to a medium, that is, the original features U and V are both compared with the medium P and converted into a representation of a third party, not only narrowing the semantic gap but also removing the noise in the unimodal data.

[0059] 2. Steps for mining false negative samples with mixed distributions: Design a sample pool for model fitting. The sample pool contains affinity matrices of cross-modal features with the same semantics and different semantics in the same proportion. Use the expectation-maximization algorithm to fit the Gaussian mixture distribution of the affinity matrix. Based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image features and text features, and use the maximum a posteriori method to determine the sub-population distribution to which the negative samples belong. On this basis, calculate the cumulative probability. By comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples", and "ambiguous samples", and store them in the corresponding sets.

[0060] Specifically, the steps for mining false negative samples with mixed distributions are as follows:

[0061] (1) Design a sample pool for model fitting. The sample pool contains affinity matrices of cross-modal features with the same semantics and different semantics in the same proportion , and use the expectation-maximization algorithm to fit the Gaussian mixture distribution of the affinity matrix, which is achieved in the following way:

[0062] For the elements make a frequency histogram, and use the expectation-maximization algorithm to fit the frequency histogram of with a mixture Gaussian distribution. The density function of this distribution is

[0063] ;

[0064] where i, j represent sample indices, , B is the total number of texts and images. is the mixing coefficient, is the probability density function of the Gaussian distribution, and are the parameters of the Gaussian distribution, represents the parameter space; g represents the sub-population index, g = 1, 2. When g is 1, it indicates the sub-population distribution with a small mean. When g is 2, it indicates the sub-population distribution with a large mean. Define the distribution with a large mean as the "same-semantics distribution" and the distribution with a small mean as the "different-semantics distribution".

[0065] (2) Based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image features and text features, and use the maximum a posteriori method to determine the sub-population distribution to which the negative samples belong; on this basis, calculate the cumulative probability. By comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples", and "ambiguous samples", and store them in the corresponding sets.

[0066] Specifically, it is achieved in the following way:

[0067] Note that the judgments of "false negative samples", "true negative samples" and "fuzzy samples" are relative to the anchor points. When performing image-to-text retrieval, for image anchor point i, among all negative text samples j, where j = 1,..., B, . Based on the image feature X and the negative text feature F, calculate their affinity matrix C; based on the fitted Gaussian mixture distribution, for the elements in C calculate the maximum a posteriori to determine the sub-population distribution to which the negative text sample j belongs, that is, determine the value of g, where g = 2 represents "same semantic distribution" and g = 1 represents "different semantic distribution".

[0068] Based on this, calculate the cumulative probability. By comparing the cumulative probability with the significance level, classify the negative text samples into "false negative samples", "true negative samples" and "fuzzy samples", and store them in the corresponding sets; the calculation formula for the cumulative probability is

[0069] ;

[0070] x is the independent variable, and its value ranges from 0 to , , are the parameters of the updated Gaussian distribution, is the cumulative probability; then, define the significance level . When , and g = 1, the negative text sample is a "true negative sample" and is stored in the set . When , and g = 2, the corresponding negative text sample is a "false negative sample" and is stored in the set . Other negative text samples are "fuzzy samples" and are stored in the set .

[0071] Similarly to image-to-text retrieval, for text-to-image retrieval, obtain the sets of image "true negative samples" , "false negative samples" , and "fuzzy samples" in the same way.

[0072] In this step, the advantage of establishing the distribution based on the affinity matrix is that the Gaussian mixture distribution is fitted with the sample pool. As the network is trained, the features are continuously optimized, and the distribution is also continuously optimized, thus fitting a more reliable distribution. Based on this distribution, the affinity matrix of the current input image features and text features is calculated, and the maximum a posteriori method is used to determine the sub-population distribution g = 1 or g = 2 to which the negative samples belong. In addition, by comparing the cumulative probability with the significance level, the negative samples are classified into "false negative samples", "true negative samples", and "ambiguous samples", and stored in the corresponding sets. Statistical methods are used to eliminate the interference of noise in the samples. On the other hand, since the sets are obtained by comparing the cumulative probability with the significance level, they are regarded as a kind of credibility, which can be adaptively adjusted according to the distribution shape, improving the credibility of sample judgment.

[0073] 3. Steps for loss calculation: Based on the triplet loss, calculate the soft distance-constrained triplet loss, keep the calculation of true negative samples, delete the calculation of false negative samples, and modify the calculation of ambiguous samples.

[0074] Specifically, the triplet loss is:

[0075] ;

[0076] The processing of "true negative samples" in the soft distance-constrained triplet loss is similar to the original triplet loss, which is

[0077] ;

[0078] The processing of "ambiguous samples" in the soft distance-constrained triplet loss is

[0079] ;

[0080] The processing of "false negative samples" in the soft distance-constrained triplet loss is to directly delete them. Therefore, the soft distance-constrained triplet loss is

[0081] ;

[0082] Among them, represents the soft distance-constrained triplet loss, represents the similarity between the image and the positive text sample pair, represents the similarity between the image and the corresponding "true negative" text, represents the similarity between the image and the corresponding "ambiguous" text, represents the similarity between the image and the corresponding "negative" text, represents the similarity between the text and the corresponding "true negative" image, represents the similarity between the text and the corresponding "ambiguous" image, Represents the similarity between the representative text and the corresponding "negative" image; and respectively represent the "true negative" image samples and the "true negative" text samples; and respectively represent the "blurred" image samples and the "blurred" text samples; and respectively represent the "negative" image samples and the "negative" text samples; , denotes the interval parameter, which serves to widen the gap between the image and positive text pairs and the image and negative text pairs; is the weight of the "blurred" sample corresponding to the image, is the weight of the "blurred" sample corresponding to the text;

[0083] ;

[0084] represents the adjustment factor, and is also calculated in the above manner.

[0085] Training:

[0086] For each dataset, 80% of the samples are used as the training set, 10% of the samples are used as the validation set, and the remaining 10% are used as the test set. Two evaluation metrics, R@K (K = 1, 5, and 10) and mR, are used to evaluate the model. R@K represents the proportion of successful matches among the top 10 highest similarity results. The experiments are conducted on a single NVIDIA Titan RTX GPU. The network is trained for 150 epochs using the Adam optimizer, and the minimum batch size is set to 128. During training, the learning rate is adjusted to 1e-4, and it decreases by 0.7 every 20 epochs.

[0087] As another embodiment of the present invention, this embodiment provides a cross-modal image-text retrieval system based on soft distance constraints of false negative samples for implementing the retrieval method described above. The system includes a feature extraction module, a false negative sample mining module with a mixed distribution, and a loss calculation module.

[0088] The feature extraction module includes an image encoder and a text encoder, which extract the image feature V and the text feature U respectively; it also includes two prompt-aware Transformer decoders, which are respectively applied to the image feature V and the text feature U; the two decoders share parameters. When processing the image and text features, the image feature V and the text feature U are respectively used as the key and value, and the prompt P is used as the query to implement the training of the network, and robust image features X and text features F are respectively generated.

[0089] The mixed distribution false negative sample mining module is designed for a sample pool for model fitting. The sample pool contains affinity matrices of cross-modal features with the same semantics and different semantics in the same proportion, and the Gaussian mixture distribution of the affinity matrix is fitted using the expectation-maximization algorithm. Based on the fitted Gaussian mixture distribution, the affinity matrix of the current input image features and text features is calculated, and the negative sample sub-population distribution is determined using the maximum a posteriori method. On this basis, the cumulative probability is calculated, and by comparing the cumulative probability with the significance level, the negative samples are classified into "false negative samples", "true negative samples", and "fuzzy samples", and stored in the corresponding sets.

[0090] The loss calculation module is used to calculate the soft distance constraint triplet loss.

[0091] The functions and implementation methods of each module can be seen in the previous description and will not be elaborated here.

[0092] Through the foregoing method for model training, the trained cross-modal image-text retrieval system based on soft distance constraint of false negative samples can be used for cross-modal image-text retrieval, which can retrieve the corresponding text for the input image or retrieve the corresponding image for the input text. For example, it can automatically retrieve the text data that accurately describes the satellite remote sensing image based on the satellite remote sensing image, or automatically retrieve the matching remote sensing image in the database based on the given text data.

[0093] Experimental verification

[0094] In this embodiment, the technical effect of the present invention is verified through ablation experiments. The results of the ablation experiments are shown in Table 1.

[0095] Table 1 Ablation experiments

[0096]

[0097] Table 1 shows the results of each module in the present invention. Among them, Experiments 1, 2, and 3 represent the results of the baseline method, the baseline method + feature extraction module, and the baseline method + mixed distribution false negative sample mining module + soft distance constraint triplet loss, respectively. It can be seen from Table 1 that the feature extraction module and the mixed distribution false negative sample mining module + soft distance constraint triplet loss are improved by 1.31% and 3.23% respectively compared with the baseline method, which shows the effectiveness of the proposed method.

[0098] In addition, to verify the effectiveness of the sample pool, a sample pool visualization verification experiment is also carried out in this embodiment.

[0099] Figure 2 The sample pool visualization verification experiment during the training process is shown, and this process is simulated by the similarity calculated from the cross-modal data in the pool.Figure 2 Epoch in it represents the training round, Figure 2 The four figures in it respectively correspond to 4 Epochs, showing the sample similarity distribution under different Epochs, Figure 2 (a) in it is the Gaussian mixture distribution fitted to the similarity calculated from the cross-modal data in the pool in the 0th round (Epoch: 0), Figure 2 (b) in it is the Gaussian mixture distribution fitted to the similarity calculated from the cross-modal data in the pool in the 1st round (Epoch: 1), Figure 2 (c) in it is the Gaussian mixture distribution fitted to the similarity calculated from the cross-modal data in the pool in the 2nd round (Epoch: 2), Figure 2 (d) in it is the Gaussian mixture distribution fitted to the similarity calculated from the cross-modal data in the pool in the 3rd round (Epoch: 3). The abscissa in the figure is the cross-modal similarity value (i.e., the similarity value of the affinity matrix C), and the ordinate is the number of elements corresponding to this similarity value.

[0100] Figure 2 The black solid line in it represents the negative sample similarity, the blue dashed line represents the similarity distribution of different semantic negative sample pairs, the red dashed line represents the distribution of the same semantic negative sample pairs, and the yellow line represents the average similarity of the positive sample pairs. Means represents the mean. In the first training cycle, the model mainly expands the distance between different semantic and the same semantic negative samples, increasing the mean difference from 0.01 to 0.28; in the second cycle, it focuses on constructing a significant interval between negative samples and positive samples, increasing the distance between the mean and the average similarity of positive samples from 0.03 to 0.08; subsequent training cycles continuously optimize this separation effect, indicating that the feature representation in the embedding space has been effectively improved.

[0101] In summary, the present invention (1) designs a prompt-aware Transformer decoder, which includes two Transformer decoders, respectively applied to the image feature V and the text feature U, and the two decoders share parameters. Specifically, this module designs a self-trained prompt P, which is transmitted as a query to the two decoders. At the same time, for the image decoder, the image feature V is used as the key and value to implement the training of the Transformer network. Similarly, for the text decoder, the text feature U is used as the key and value, and the two decoders finally generate robust image features F and text features X. (2) Designs a hybrid distribution false negative sample mining module, which designs a sample pool for model fitting. The sample pool contains affinity matrices of cross-modal features with the same semantics and different semantics in the same proportion, and uses the expectation-maximization algorithm to fit the Gaussian mixture distribution of the affinity matrix; based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image feature and text feature, and use the maximum a posteriori method to determine the sub-population distribution to which the negative sample belongs; on this basis, calculate the cumulative probability, and by comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples" and "ambiguous samples", and store them in the corresponding sets. (3) Designs a triplet loss with soft distance constraints, keeps the calculation of true negative samples, deletes the calculation of false negative samples, and modifies the calculation of ambiguous samples.

[0102] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Those of ordinary skill in the art in the technical field of the present invention, within the scope of the essence of the present invention, any changes, modifications, additions or substitutions should fall within the protection scope of the present invention.

Claims

1. A cross-modal retrieval method for images and texts based on soft distance constraints of false negative samples, characterized in that Including the following steps: The step of feature extraction: For the image data I and the text data T, the image feature X and the text feature F are extracted through an encoder-decoder; The step of mining false negative samples with a mixture distribution: Design a sample pool for model fitting. The sample pool contains an affinity matrix of cross-modal features with the same semantics and different semantics in the same proportion. The expectation-maximization algorithm is used to fit the Gaussian mixture distribution of the affinity matrix; Based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image feature and text feature, and use the maximum a posteriori method to determine the sub-population distribution of the negative samples; Then calculate the cumulative probability. By comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples" and "fuzzy samples", and store them in the corresponding sets; The step of loss calculation: Based on the triplet loss, calculate the soft-distance-constrained triplet loss, keep the calculation of the true negative samples, delete the calculation of the false negative samples, and modify the calculation of the fuzzy samples; Among them, the processing of the soft-distance-constrained triplet loss for the "fuzzy samples" is: Among them, S(X, F) represents the similarity between the image and the positive sample pair of the text, S(X, F″) represents the similarity between the image and the corresponding "blurred" text, S(X″, F) represents the similarity between the text and the corresponding "blurred" image, and X″ and F″ respectively represent the "blurred" image sample and the "blurred" text sample; [x] + ≡ max(x, 0), where α represents the margin parameter, which is used to widen the gap between the image and the positive text pair and the image and the negative text pair; w F is the weight of the "blurred" sample corresponding to the image, and w X is the weight of the "blurred" sample corresponding to the text; is the set of blurred samples in the image modality, is the set of blurred samples in the text modality.

2. The cross-modal retrieval method of text and image based on soft distance constraint of false negative samples according to claim 1, wherein The step of feature extraction is specifically as follows: For the image data I and the text data T, first use an image encoder and a text encoder to extract the image feature V and the text feature U respectively; Then generate robust image features X and text features F through a prompt-aware Transformer decoder.

3. The method for cross-modal retrieval of images and texts based on soft distance constraints of false negative samples according to claim 2, wherein The prompt-aware Transformer decoder includes two decoders, which are respectively applied to the image and text features; The two decoders share parameters. When processing the image and text features, the image feature V and the text feature U are used as the key and value respectively, and a prompt P is used as the query to realize the training of the network, and generate robust image features X and text features F respectively.

4. The method for cross-modal retrieval of graphics and texts based on soft distance constraints of false negative samples according to claim 1, wherein In the step of mining false negative samples with a mixture distribution, design a sample pool for model fitting. The sample pool contains an affinity matrix C′ of cross-modal features with the same semantics and different semantics in the same proportion. The expectation-maximization algorithm is used to fit the Gaussian mixture distribution of the affinity matrix, which is realized in the following way: For an element c′ in C′ ij make a frequency histogram, and use the Expectation-Maximization algorithm to fit a mixture Gaussian distribution to the frequency histogram of c′ ij The density function of this distribution is where, π g is the mixing coefficient, Φ(·|μ g , σ g ) is the probability density function of the Gaussian distribution, μ g and σ g are the parameters of the Gaussian distribution, Θ represents the parameter space; g represents the sub-population index, g = 1, 2. When g is 1, it indicates the sub-population distribution with a small mean. When g is 2, it indicates the sub-population distribution with a large mean. The distribution with a large mean is defined as the "same semantic distribution", and the distribution with a small mean is defined as the "different semantic distribution".

5. The method for cross-modal retrieval of graphics and text based on soft distance constraint of false negative samples according to claim 4, characterized in that, Based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image feature and text feature, and use the maximum a posteriori method to determine the sub-population distribution of the negative samples; On this basis, calculate the cumulative probability. By comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples" and "fuzzy samples", and store them in the corresponding sets; It is realized in the following way: The judgments of "false negative samples", "true negative samples" and "fuzzy samples" are relative to the anchor points. When performing image-to-text retrieval, for the image anchor point i, among all negative sample texts j, where j = 1,..., B and j ≠ i; based on the image feature X and the negative text feature F, calculate their affinity matrix C; based on the fitted Gaussian mixture distribution, for the element c in C ij calculate the maximum a posteriori to determine the sub-population distribution to which the negative sample text j belongs, that is, determine the value of g, where g = 2 is the "same semantic distribution" and g = 1 is the "different semantic distribution"; on this basis, calculate the cumulative probability, and by comparing the cumulative probability with the significance level, classify the negative text samples as "false negative samples", "true negative samples" and "fuzzy samples", and store them in the corresponding sets; The calculation formula for the cumulative probability is x is the independent variable, and its value ranges from 0 to c ij and are the parameters of the updated Gaussian distribution, and τ is the cumulative probability; then, a significance level β is defined. When τ < β and g = 1, the negative sample text is a "true negative sample" and is stored in the set When 1 - τ < β and g = 2, the corresponding negative sample text is a "false negative sample" and is stored in the set Other negative sample texts are "fuzzy samples" and are stored in the set ; Similarly to image-to-text retrieval, text-to-image retrieval obtains the set of image "true negative samples" in the same way. The set of "false negative samples" and the set of "ambiguous samples".

6. The method for cross-modal retrieval of images and texts based on soft distance constraints of false negative samples according to claim 5, wherein The triplet loss is: The processing of the soft-distance-constrained triplet loss for the "true negative samples" is The processing of the soft-distance-constrained triplet loss for the "false negative samples" is to directly delete them. Therefore, the soft-distance-constrained triplet loss is L′(I, T) = L(I, T)0 + L(I, T)1; Among them, L′(I, T) represents the soft distance constraint triplet loss, and S(X, F′) represents the similarity between the image and the corresponding "true negative" text. represents the similarity between the image and the corresponding "negative" text, and S(X′, F) represents the similarity between the text and the corresponding "true negative" image. represents the similarity between the text and the corresponding "negative" image; X′ and F′ respectively represent the "true negative" image sample and the "true negative" text sample. and respectively represent the "negative" image sample and the "negative" text sample. ∈ represents a regulatory factor, w X The same calculation method is adopted.

7. A cross-modal retrieval system for images and texts based on soft distance constraints of false negative samples, characterized in that For implementing the method according to any one of claims 1-6, it includes a feature extraction module, a mixed-distribution false negative sample mining module, and a loss calculation module. The feature extraction module includes an image encoder and a text encoder, which respectively extract an image feature V and a text feature U; it also includes two prompt-aware Transformer decoders, which are respectively applied to the image feature V and the text feature U; The two decoders share parameters. When processing the image and text features, the image feature V and the text feature U are respectively used as the key and the value, and the prompt P is used as the query to implement the training of the network, and robust image features X and text features F are respectively generated; The mixed-distribution false negative sample mining module is designed for a sample pool for model fitting. The sample pool contains affinity matrices of cross-modal features with the same semantics and different semantics in the same proportion, and uses the expectation-maximization algorithm to fit the Gaussian mixture distribution of the affinity matrix; based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image feature and text feature, and use the maximum a posteriori method to determine the sub-population distribution to which the negative sample belongs; on this basis, calculate the cumulative probability, and by comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples", and "fuzzy samples", and store them in the corresponding sets; The loss calculation module is used to calculate the soft distance constraint triplet loss.

Citation Information

Patent Citations

  • Rapid random model correction method for active learning type ocean platform structure

    CN119494245A

  • Method and apparatus for training cross-modal retrieval model, device and storage medium

    EP4053751A1