Image-text cross-modal retrieval method and system based on false negative sample soft distance constraint

By adopting a method based on soft distance constraint of false negative samples in cross-modal search of marine remote sensing graphics and text, using prompt perception Transformer decoder and mixed distribution false negative sample mining module, the "false negative" sample problem is solved, and the accuracy and robustness of the search is improved.

CN120104825AActive Publication Date: 2025-06-06OCEAN UNIV OF CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510569986.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-06
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

The existing cross-modal search method for marine remote sensing graphics and texts has a large number of "false negative" samples when processing marine remote sensing data, which makes it difficult to fit the model and reduces the robustness of feature representation and retrieval accuracy.

Method used

Using the cross-modal search method of graphic and text based on soft distance constraints of false negative samples, a prompt-perceived Transformer decoder, a mixed distribution false negative sample mining module and a loss calculation module are designed. Through these modules, the cross-modal semantic gap is narrowed, the noise in single-modal data is removed, and the retrieval accuracy is improved.

Benefits of technology

By narrowing the cross-modal divide and removing noise, the accuracy and reliability of cross-modal retrieval of graphics and text are improved, and the model's ability to learn potential associations of images and text is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104825A_ABST
    Figure CN120104825A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of cross-modal image-text retrieval, and discloses an image-text cross-modal retrieval method and system based on false negative sample soft distance constraint. A mixed distribution false negative sample mining step: designing a sample pool for model fitting, the sample pool containing affinity matrixes of cross-modal features of the same semantics and different semantics with the same proportion, and fitting Gaussian mixed distribution of the affinity matrixes by adopting an expectation maximization algorithm; calculating a current input image feature and text feature affinity matrix, and determining subpopulation distribution to which a negative sample belongs by using a maximum posterior method; the negative samples are classified into false negative samples, true negative samples and fuzzy samples by comparing the cumulative probability with the significance level; and soft distance constraint triple loss is calculated. According to the method, a cross-modal semantic gap is reduced, noise in single-modal data is removed, 'false negative 'samples in a data set are efficiently mined, and the accuracy of cross-modal retrieval is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cross-modal image and text retrieval, and in particular to an image and text cross-modal retrieval method and system based on soft distance constraints of false negative samples. Background Art

[0002] As one of the important means of cross-modal retrieval of marine remote sensing data, image-text retrieval has attracted more and more attention from researchers in recent years. At present, cutting-edge image-text retrieval methods are committed to using saliency mining methods such as attention mechanisms and graph convolutional networks to filter out redundant background information in images and texts, and improve the matching accuracy of image-text retrieval.

[0003] However, the above methods still have the following problems when applied to the ocean: Existing remote sensing image and text cross-modal retrieval methods use triplet loss to train the network. Triplet loss involves three elements, namely, anchor points, positive samples corresponding to anchor points, and negative samples corresponding to anchor points. Triplet loss aims to reduce the distance between anchor points and positive samples, and increase the distance between anchor points and negative samples. However, there are a large number of "false negative" samples in the marine remote sensing cross-modal retrieval dataset, that is, there are many negative samples that match anchor points. There are three main reasons for this phenomenon. First, the high similarity between marine remote sensing data. Compared with ordinary images, the similarity between marine remote sensing data is higher, resulting in similar or even identical descriptions of many data. For example, image 1 and image 2 can both be described as "two cargo ships parked on the sea." Therefore, when the anchor point is image 1, text 2 corresponding to image 2 is a "false negative" sample. Second, the modal gap between images and texts. Compared with texts, images contain richer semantic information; compared with images, texts give more concise descriptions. This modal gap will produce many "false negative" samples. Third, errors caused by human labeling. The texts in the dataset are all manually labeled based on the images, and there are too many human interference factors.

[0004] Existing methods still use triplet loss for modeling, and the distance of "false negative" samples is widened, making the model difficult to fit or fitting incorrectly, thereby reducing the robustness of feature representation and ultimately reducing the accuracy of cross-modal image and text retrieval. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a method and system for cross-modal retrieval of images and texts based on soft distance constraints of false negative samples, and designs a prompt-aware Transformer decoder, a mixed distribution false negative sample mining module, and a loss calculation module. The present invention narrows the cross-modal semantic gap and removes the noise in the unimodal data; it mines the "false negative" samples in the data set in a more efficient and reliable way, and improves the accuracy of cross-modal retrieval.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: The image-text cross-modal retrieval method based on false negative sample soft distance constraint includes the following steps: Feature extraction steps: for image data I and text data T, extract image features X and text features F through the encoder-decoder; The steps of mixed distribution false negative sample mining are as follows: design a sample pool for model fitting, which contains affinity matrices of homosemantic and heterosemantic cross-modal features in the same proportion, and use the maximum expectation algorithm to fit the Gaussian mixture distribution of the affinity matrix; based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image features and text features, and use the maximum a posteriori method to determine the sub-population distribution to which the negative samples belong; then calculate the cumulative probability, and by comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples" and "fuzzy samples", and store them in the corresponding collections; Loss calculation steps: Based on the triplet loss, calculate the soft distance constrained triplet loss, keep the calculation of true negative samples, delete the calculation of false negative samples, and modify the calculation of blurred samples.

[0007] Furthermore, the feature extraction steps are as follows: for image data I and text data T, the image encoder and text encoder are first used to extract image features V and text features U respectively; then the cue-aware Transformer decoder is used to generate robust image features X and text features F.

[0008] Furthermore, the prompt-aware Transformer decoder includes two decoders, which are applied to image and text features respectively; the two decoders share parameters, and when processing image and text features, they use image features V and text features U as keys and values ​​respectively, and use a prompt P as a query to implement network training, and generate robust image features X and text features F respectively.

[0009] Furthermore, in the mixed distribution false negative sample mining step, a sample pool for model fitting is designed, which contains affinity matrices with the same semantics and different semantics cross-modal features in the same proportion. , the Gaussian mixture distribution of the affinity matrix is ​​fitted using the maximum expectation algorithm, which is achieved in the following way: right Elements in Make a frequency histogram and use the maximum expectation algorithm to The frequency histogram of is fitted with a mixed Gaussian distribution, the density function of which is ; in, is the mixing coefficient, is the probability density function of the Gaussian distribution, and are the parameters of the Gaussian distribution, Represents the parameter space; g represents the subgroup index, g=1,2, when g is 1, it indicates a subgroup distribution with a small mean, and when g is 2, it indicates a subgroup distribution with a large mean; the distribution with a large mean is defined as "same semantic distribution", and the distribution with a small mean is defined as "different semantic distribution".

[0010] Furthermore, based on the fitted Gaussian mixture distribution, the affinity matrix of the current input image features and text features is calculated, and the sub-population distribution to which the negative samples belong is determined using the maximum a posteriori method; on this basis, the cumulative probability is calculated, and by comparing the cumulative probability with the significance level, the negative samples are classified into "false negative samples", "true negative samples" and "fuzzy samples", and stored in the corresponding sets; this is achieved in the following way: The judgment of "false negative samples", "true negative samples" and "fuzzy samples" is relative to the anchor point. In image-to-text retrieval, for image anchor point i, among all negative sample texts j, j=1,...,B, ; Based on the image feature X and the negative text feature F, calculate its affinity matrix C; based on the fitted Gaussian mixture distribution, the elements in C Calculate the maximum posterior to determine the subgroup distribution to which the negative sample text j belongs, that is, determine the value of g, where g=2 means "same semantic distribution" and g=1 means "different semantic distribution"; On this basis, the cumulative probability is calculated. By comparing the cumulative probability with the significance level, the negative text samples are classified into "false negative samples", "true negative samples" and "fuzzy samples" and stored in the corresponding sets; the calculation formula of the cumulative probability is: ; x is the independent variable, ranging from 0 to between, , is the parameter of the updated Gaussian distribution, is the cumulative probability; then, define the significance level ,when , and when g=1, the negative sample text is a "true negative sample" and is stored in the collection ,when , and g=2, the corresponding negative sample text is "false negative sample", stored in the collection , other negative sample texts are "fuzzy samples" and stored in the collection middle; Similar to image-to-text retrieval, text-to-image retrieval obtains a set of “true negative samples” of images in the same way , "false negative sample" set , and the Fuzzy Samples collection .

[0011] Furthermore, the triplet loss is: ; The soft distance constraint triplet loss handles “true negative samples” as follows: ; The soft distance constraint triplet loss handles the “blurred samples” as ; The soft distance constraint triplet loss directly deletes the “false negative samples”, so the soft distance constraint triplet loss is ; in, represents the soft distance constrained triplet loss, Represents the similarity between the image and text positive sample pairs, represents the similarity between the image and the corresponding “true negative” text, Represents the similarity between the image and the corresponding "fuzzy" text, Represents the similarity between the image and the corresponding “negative” text, represents the similarity between the text and the corresponding “true negative” image, represents the similarity between the text and the corresponding "blurred" image, Represents the similarity between the text and the corresponding “negative” image; and Represent the “true negative” image samples and “true negative” text samples respectively; and Respectively represent the “blurred” image sample and the “blurred” text sample; and Respectively represent "negative" image samples and "negative" text samples; , represents the spacing parameter, which is used to increase the gap between the image and positive text pairs and the image and negative text pairs; is the weight of the “blurred” sample corresponding to the image, is the weight of the “fuzzy” sample corresponding to the text; ; represents the regulating factor, The same calculation method is used.

[0012] The present invention also provides a method for cross-modal image and text retrieval based on soft distance constraints of false negative samples, which is used to implement the above method. The system includes a feature extraction module, a mixed distribution false negative sample mining module, and a loss calculation module. The feature extraction module includes an image encoder and a text encoder, which extract image features V and text features U respectively; and also includes two prompt-aware Transformer decoders, which are applied to the image features V and the text features U respectively; the two decoders share parameters, and when processing image and text features, the image features V and the text features U are used as keys and values ​​respectively, and the prompt P is used as a query query to implement network training, and respectively generate robust image features X and text features F; The mixed distribution false negative sample mining module is designed for a sample pool for model fitting, the sample pool contains affinity matrices of homosemantic and heterosemantic cross-modal features in the same proportion, and the maximum expectation algorithm is used to fit the Gaussian mixture distribution of the affinity matrix; based on the fitted Gaussian mixture distribution, the affinity matrix of the current input image feature and text feature is calculated, and the sub-population distribution to which the negative sample belongs is determined using the maximum a posteriori method; on this basis, the cumulative probability is calculated, and by comparing the cumulative probability with the significance level, the negative samples are classified into "false negative samples", "true negative samples" and "fuzzy samples", and stored in the corresponding sets; The loss calculation module is used to calculate the soft distance constraint triplet loss.

[0013] Compared with the prior art, the present invention has the advantages of: First, the third-party medium of prompts is used to narrow the cross-modal gap while reducing the computational cost. In the prompt-aware Transformer decoder module, the dimensions of the generated features F and Z are consistent with the prompt, and the dimensions of the prompt are smaller than the original image and text features V and U, which reduces the computational cost for the subsequent modeling process; in addition, the prompt is similar to a medium, that is, the original features U and V are compared with the medium P and converted into a third-party representation, which not only narrows the semantic gap, but also removes the noise in the unimodal data.

[0014] Second, the "false negative", "true negative" and "fuzzy" samples in the data set are mined in a more efficient and reliable way. In the mixed distribution false negative sample mining module, the sample pool is used to fit the Gaussian mixture distribution. The calculation method of this distribution is more stable because the sample pool contains affinity matrices of the same semantic and different semantic cross-modal features in the same proportion. In this case, one distribution will not be covered by another distribution, such as the false negative distribution will be covered by the true negative distribution. As the network is trained, the features are continuously optimized, and the distribution is also continuously optimized, so as to fit a more reliable distribution. Based on this distribution, the affinity matrix of the current input image feature and text feature is calculated, and the maximum a posteriori method is used to determine the sub-population distribution g=1 or g=2 to which the negative sample belongs. In addition, by comparing the cumulative probability with the significance level, the negative samples are classified into "false negative samples", "true negative samples" and "fuzzy samples", and stored in the corresponding sets. The interference of noise in the samples is eliminated by using statistical methods. On the other hand, since the set is obtained by comparing the cumulative probability with the significance level, it is regarded as a kind of credibility and can be adaptively adjusted according to the distribution shape, which improves the credibility of sample judgment.

[0015] Third, the valuable information in negative samples is fully utilized in the form of soft distance constraints. In the triple loss module of soft distance constraints, the calculation of true negative samples is maintained, the calculation of false negative samples is deleted, and the calculation of fuzzy samples is modified. Different calculation strategies are adopted for different samples, which not only improves the utilization of data, but also facilitates the identification of positive and negative samples, so that the network can learn the potential association between image and text pairs more deeply. In addition, in the calculation process of fuzzy samples, the cumulative probability is combined with the similarity. By adjusting the weight of the fuzzy samples, a larger weight is given to the distribution boundary of the same semantics and different semantics similarity (affinity matrix), which increases the boundary of the same semantics and different semantics. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0017] Figure 1 It is a schematic diagram of the method flow of the present invention; Figure 2This is a sample pool visualization verification experiment in an embodiment of the present invention, wherein (a) is the Gaussian mixture distribution of the similarity fit calculated for cross-modal data in the pool in round 0, (b) is the Gaussian mixture distribution of the similarity fit calculated for cross-modal data in the pool in round 1, (c) is the Gaussian mixture distribution of the similarity fit calculated for cross-modal data in the pool in round 2, and (d) is the Gaussian mixture distribution of the similarity fit calculated for cross-modal data in the pool in round 3. DETAILED DESCRIPTION

[0018] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0019] Combination Figure 1 As shown, this embodiment provides a cross-modal image-text retrieval method based on soft distance constraints of false negative samples, following the main architecture of TGKT. In the feature extraction part, for image data I and text data T, the parameters pre-trained in VIT and BERT are still used to extract features V and U. Unlike TGKT, the method proposed in the present invention no longer uses a text-guided image feature extraction module, but innovatively designs the following in the feature extraction stage and similarity matching stage: prompt-aware Transformer decoder, mixed distribution false negative sample mining, and soft distance constraint triplet loss.

[0020] Combination Figure 1 As shown, the detailed steps of the present invention are described in detail below.

[0021] 1. Feature extraction steps: For image data I and text data T, extract image features X and text features F through the encoder-decoder.

[0022] The specific steps of feature extraction are as follows: for image data I and text data T, first use the image encoder and text encoder to extract image features V and text features U respectively; then generate robust image features X and text features F through the cue-aware Transformer decoder.

[0023] The prompt-aware Transformer decoder comprises two decoders, which are respectively applied to image features V and text features U; the two decoders share parameters, and when processing image and text features, the image features V and text features U are respectively used as keys and values, and a prompt P is used as a query query to implement network training, and generate robust image features X and text features F respectively.

[0024] Specifically, this module designs a self-training prompt P, which is transmitted as a query to the two decoders. At the same time, for the Transformer decoder of the image, the image feature V is used as the key and value to implement the training of the Transformer network and generate a robust image feature F.

[0025] ;

[0026] in, represents the Transformer decoder, , , For the parameters that need to be trained in the network, convert P into query and convert V into key and value.

[0027] Similarly, for the Transformer decoder of text, the text feature U is used as the key and value to generate a robust text feature X.

[0028] ; in, , , For the parameters that need to be trained in the network, P is converted into query and U is converted into key and value.

[0029] The two decoders finally generate robust image features F and text features X. The prompt relies on back propagation for updating, which has the advantages that, on the one hand, the dimensions of the generated features F and X are consistent with the prompt, and the dimensions of the prompt are smaller than the original image and text features V and U, which reduces the computational cost for the subsequent modeling process; on the other hand, the prompt is similar to a medium, that is, the original features U and V are compared with the medium P and converted into a representation of a third party, which not only narrows the semantic gap, but also removes the noise in the unimodal data.

[0030] 2. Steps for mining false negative samples with mixed distribution: Design a sample pool for model fitting, which contains affinity matrices of homosemantic and cross-modal features with different semantics in the same proportion, and use the maximum expectation algorithm to fit the Gaussian mixture distribution of the affinity matrix. Based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image features and text features, and use the maximum a posteriori method to determine the sub-population distribution to which the negative samples belong. On this basis, calculate the cumulative probability, and by comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples" and "fuzzy samples", and store them in the corresponding collections.

[0031] Specifically, the steps of mixed distribution false negative sample mining are as follows: (1) Design a sample pool for model fitting, which contains affinity matrices with the same semantics and different semantics cross-modal features in the same proportion. , the Gaussian mixture distribution of the affinity matrix is ​​fitted using the maximum expectation algorithm, which is achieved in the following way: right Elements in Make a frequency histogram and use the maximum expectation algorithm to The frequency histogram of is fitted with a mixed Gaussian distribution, the density function of which is ; Among them, i, j represent sample indexes, , B is the total number of texts and images. is the mixing coefficient, is the probability density function of the Gaussian distribution, and are the parameters of the Gaussian distribution, Represents the parameter space; g represents the subgroup index, g=1,2, when g is 1, it indicates a subgroup distribution with a small mean, and when g is 2, it indicates a subgroup distribution with a large mean; the distribution with a large mean is defined as "same semantic distribution", and the distribution with a small mean is defined as "different semantic distribution".

[0032] (2) Based on the fitted Gaussian mixture distribution, the affinity matrix of the current input image features and text features is calculated, and the maximum a posteriori method is used to determine the subpopulation distribution to which the negative samples belong. On this basis, the cumulative probability is calculated, and by comparing the cumulative probability with the significance level, the negative samples are classified into "false negative samples", "true negative samples" and "fuzzy samples" and stored in the corresponding sets.

[0033] This is achieved through the following methods: It should be noted that the judgment of "false negative samples", "true negative samples" and "fuzzy samples" is relative to the anchor point. In image-to-text retrieval, for image anchor point i, among all negative sample texts j, j=1,...,B, Based on the image feature X and the negative text feature F, the affinity matrix C is calculated; based on the fitted Gaussian mixture distribution, the elements in C are Calculate the maximum posterior to determine the subgroup distribution to which the negative sample text j belongs, that is, determine the value of g, g=2 means "same semantic distribution", g=1 means "different semantic distribution".

[0034] On this basis, the cumulative probability is calculated. By comparing the cumulative probability with the significance level, the negative text samples are classified into "false negative samples", "true negative samples" and "fuzzy samples" and stored in the corresponding sets; the calculation formula of the cumulative probability is: ; x is the independent variable, ranging from 0 to between, , is the parameter of the updated Gaussian distribution, is the cumulative probability; then, define the significance level ,when , and when g=1, the negative sample text is a "true negative sample" and is stored in the collection ,when , and g=2, the corresponding negative sample text is "false negative sample", stored in the collection , other negative sample texts are "fuzzy samples" and stored in the collection middle.

[0035] Similar to image-to-text retrieval, text-to-image retrieval obtains a set of “true negative samples” of images in the same way , "false negative sample" set , and the Fuzzy Samples collection .

[0036] In this step, the advantage of establishing a distribution based on an affinity matrix is ​​that a Gaussian mixture distribution is fitted with a sample pool. As the network is trained, features are continuously optimized, and distribution is also continuously optimized, thereby fitting a more reliable distribution. Based on this distribution, the affinity matrix of the current input image features and text features is calculated, and the subpopulation distribution g=1 or g=2 to which the negative sample belongs is determined using the maximum a posteriori method. In addition, by comparing the cumulative probability with the significance level, negative samples are classified into "false negative samples", "true negative samples" and "fuzzy samples", and stored in corresponding sets, and statistical methods are used to eliminate the interference of noise in the samples. On the other hand, since the set is obtained by comparing the cumulative probability with the significance level, it is regarded as a kind of credibility, which can be adaptively adjusted according to the distribution shape, thereby improving the credibility of sample judgment.

[0037] 3. Loss calculation steps: Based on the triplet loss, calculate the soft distance constraint triplet loss, keep the calculation of true negative samples, delete the calculation of false negative samples, and modify the calculation of fuzzy samples.

[0038] In detail, the triplet loss is: ; The treatment of “true negative samples” by the soft distance constraint triplet loss is similar to the original triplet loss, which is ; The soft distance constraint triplet loss handles the “blurred samples” as ; The soft distance constraint triplet loss directly deletes the “false negative samples”, so the soft distance constraint triplet loss is ; in, represents the soft distance constrained triplet loss, Represents the similarity between the image and text positive sample pairs, represents the similarity between the image and the corresponding “true negative” text, Represents the similarity between the image and the corresponding "fuzzy" text, Represents the similarity between the image and the corresponding “negative” text, represents the similarity between the text and the corresponding “true negative” image, represents the similarity between the text and the corresponding "blurred" image, Represents the similarity between the text and the corresponding “negative” image; and Represent the “true negative” image samples and “true negative” text samples respectively; and Respectively represent the “blurred” image sample and the “blurred” text sample; and Respectively represent "negative" image samples and "negative" text samples; , represents the spacing parameter, which is used to increase the gap between the image and positive text pairs and the image and negative text pairs; is the weight of the “blurred” sample corresponding to the image, is the weight of the “fuzzy” sample corresponding to the text; ; represents the regulating factor, The above method is also used for calculation.

[0039] train: For each dataset, 80% of the samples are used as training sets, 10% of the samples are used as validation sets, and the remaining 10% are used as test sets. Two evaluation indicators, R@K (K = 1, 5, and 10) and mR, are used to evaluate the model. R@K represents the proportion of successful matches that appear in the top 10 highest similarity results. The experiments were conducted on a single NVIDIA Titan RTX GPU, and the network was trained for 150 epochs using the Adam optimizer, with a minimum batch size of 128. During training, the learning rate was adjusted to 1e-4, and the learning rate was decreased by 0.7 every 20 epochs.

[0040] As another embodiment of the present invention, this embodiment provides a cross-modal image and text retrieval system based on soft distance constraints of false negative samples, which is used to implement the retrieval method as described above. The system includes a feature extraction module, a mixed distribution false negative sample mining module, and a loss calculation module.

[0041] The feature extraction module includes an image encoder and a text encoder, which extract image features V and text features U respectively; it also includes two prompt-aware Transformer decoders, which are applied to the image features V and the text features U respectively; the two decoders share parameters, and when processing image and text features, the image features V and the text features U are used as keys and values ​​respectively, and the prompt P is used as a query query to implement network training, and respectively generate robust image features X and text features F.

[0042] The mixed distribution false negative sample mining module is designed for a sample pool for model fitting, which contains affinity matrices of same-semantic and different-semantic cross-modal features in the same proportion, and uses the maximum expectation algorithm to fit the Gaussian mixture distribution of the affinity matrix; based on the fitted Gaussian mixture distribution, the affinity matrix of the current input image feature and text feature is calculated, and the sub-population distribution to which the negative sample belongs is determined using the maximum a posteriori method; on this basis, the cumulative probability is calculated, and by comparing the cumulative probability with the significance level, the negative samples are classified into "false negative samples", "true negative samples" and "fuzzy samples", and stored in the corresponding sets.

[0043] The loss calculation module is used to calculate the soft distance constraint triplet loss.

[0044] The functions and implementation methods of each module can be found in the previous records and will not be repeated here.

[0045] The model is trained using the aforementioned method, and the trained image-text cross-modal retrieval system based on soft distance constraints of false negative samples can be used for cross-modal image-text retrieval. The corresponding text can be retrieved for the input image, or the corresponding image can be retrieved for the input text. For example, text data that accurately describes the image can be automatically retrieved based on a satellite remote sensing image, or remote sensing images that match the image can be automatically retrieved based on given text data.

[0046] Experimental verification In this embodiment, an ablation experiment is performed to verify the technical effect of the present invention. The ablation experiment results are shown in Table 1.

[0047] Table 1 Ablation experiment Table 1 shows the results of each module in the present invention, where experiments 1, 2 and 3 represent the results of the baseline method, the baseline method + feature extraction module and the baseline method + mixed distribution false negative sample mining module + soft distance constraint triple loss, respectively. It can be seen from Table 1 that the feature extraction module and the mixed distribution false negative sample mining module + soft distance constraint triple loss are improved by 1.31% and 3.23% respectively compared with the baseline method, which shows the effectiveness of the proposed method.

[0048] In addition, in order to verify the effectiveness of the sample pool, this embodiment also conducted a sample pool visualization verification experiment.

[0049] Figure 2 A visualization validation experiment of sample pooling during training is shown, which is simulated by calculating the similarity of cross-modal data in the pool. Figure 2 Epoch in represents the training round. Figure 2 The four figures in the figure correspond to four Epochs, showing the sample similarity distribution under different Epochs. Figure 2 (a) in the figure is the Gaussian mixture distribution of the similarity fit calculated for cross-modal data in the pool in Epoch 0. Figure 2 (b) is the Gaussian mixture distribution of the similarity fit calculated for cross-modal data in the pool in round 1 (Epoch: 1). Figure 2 (c) in the figure is the Gaussian mixture distribution of the similarity fit calculated for the cross-modal data in the pool in the second round (Epoch: 2). Figure 2 (d) in the figure is the Gaussian mixture distribution of the similarity fit calculated for the cross-modal data in the pool in Epoch 3. The horizontal axis in the figure is the cross-modal similarity value (i.e. the similarity value of the affinity matrix C), and the vertical axis is the number of elements corresponding to the similarity value.

[0050] Figure 2 The black solid line in the middle represents the similarity of negative samples, the blue dotted line represents the similarity distribution of negative sample pairs with different semantics, the red dotted line represents the distribution of negative sample pairs with the same semantics, and the yellow line represents the average similarity of positive sample pairs, where means represents the mean. In the first training cycle, the model mainly expands the distance between negative samples with different semantics and negative samples with the same semantics, increasing the mean difference from 0.01 to 0.28; the second cycle focuses on building a significant interval between negative samples and positive samples, increasing the distance between the mean and the average similarity of positive samples from 0.03 to 0.08; subsequent training cycles continue to optimize this separation effect, indicating that the feature representation in the embedding space has been effectively improved.

[0051] In summary, the present invention (1) designs a prompt-aware Transformer decoder, which includes two Transformer decoders, which are applied to image features V and text features U respectively, and the two decoders share parameters. Specifically, the module designs a self-training prompt P, which is transmitted to the two decoders as a query. At the same time, for the image decoder, the image feature V is used as the key and value to implement the training of the Transformer network. Similarly, for the text decoder, the text feature U is used as the key and value. The two decoders finally generate robust image features F and text features X. (2) A mixed distribution false negative sample mining module was designed. A sample pool for model fitting was designed. The sample pool contained affinity matrices of homosemantic and different semantic cross-modal features of the same proportion. The maximum expectation algorithm was used to fit the Gaussian mixture distribution of the affinity matrix. Based on the fitted Gaussian mixture distribution, the affinity matrix of the current input image features and text features was calculated, and the maximum a posteriori method was used to determine the sub-population distribution to which the negative samples belonged. On this basis, the cumulative probability was calculated, and by comparing the cumulative probability with the significance level, the negative samples were classified into "false negative samples", "true negative samples" and "fuzzy samples", and stored in the corresponding sets. (3) A soft distance constrained triple loss was designed to keep the calculation of true negative samples, delete the calculation of false negative samples, and modify the calculation of fuzzy samples.

[0052] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Any changes, modifications, additions or substitutions made by ordinary technicians in this technical field within the essential scope of the present invention should fall within the protection scope of the present invention.

Claims

1. A cross-modal image-text retrieval method based on soft distance constraints of false negative samples, characterized in that: The following steps are involved: Feature extraction steps: for image data I and text data T, extract image features X and text features F through the encoder-decoder; The steps of mixed distribution false negative sample mining are as follows: design a sample pool for model fitting, which contains affinity matrices of homosemantic and heterosemantic cross-modal features in the same proportion, and use the maximum expectation algorithm to fit the Gaussian mixture distribution of the affinity matrix; based on the fitted Gaussian mixture distribution, calculate the affinity matrix of the current input image features and text features, and use the maximum a posteriori method to determine the sub-population distribution to which the negative samples belong; then calculate the cumulative probability, and by comparing the cumulative probability with the significance level, classify the negative samples into "false negative samples", "true negative samples" and "fuzzy samples", and store them in the corresponding collections; Loss calculation steps: Based on the triplet loss, calculate the soft distance constrained triplet loss, keep the calculation of true negative samples, delete the calculation of false negative samples, and modify the calculation of blurred samples.

2. The image-text cross-modal retrieval method based on false negative sample soft distance constraint according to claim 1 is characterized in that: The specific steps of feature extraction are as follows: for image data I and text data T, first use the image encoder and text encoder to extract image features V and text features U respectively; then generate robust image features X and text features F through the cue-aware Transformer decoder.

3. The image-text cross-modal retrieval method based on false negative sample soft distance constraint according to claim 2 is characterized in that: The prompt-aware Transformer decoder includes two decoders, which are applied to image and text features respectively; the two decoders share parameters, and when processing image and text features, they use image features V and text features U as keys and values ​​respectively, and use a prompt P as a query to implement network training, and generate robust image features X and text features F respectively.

4. The image-text cross-modal retrieval method based on false negative sample soft distance constraint according to claim 1 is characterized in that: In the mixed distribution false negative sample mining step, a sample pool is designed for model fitting, which contains affinity matrices with the same semantics and different semantics cross-modal features in the same proportion. , the Gaussian mixture distribution of the affinity matrix is ​​fitted using the maximum expectation algorithm, which is achieved in the following way: right Elements in Make a frequency histogram and use the maximum expectation algorithm to The frequency histogram of is fitted with a mixed Gaussian distribution, the density function of which is ; in, is the mixing coefficient, is the probability density function of the Gaussian distribution, and are the parameters of the Gaussian distribution, Represents the parameter space; g represents the subgroup index, g=1,2, when g is 1, it indicates the subgroup distribution with a small mean, and when g is 2, it indicates the subgroup distribution with a large mean; the distribution with a large mean is defined as "same semantic distribution", and the distribution with a small mean is defined as "different semantic distribution".

5. The image-text cross-modal retrieval method based on false negative sample soft distance constraint according to claim 4 is characterized in that: Based on the fitted Gaussian mixture distribution, the affinity matrix of the current input image features and text features is calculated, and the maximum a posteriori method is used to determine the subpopulation distribution to which the negative samples belong. On this basis, the cumulative probability is calculated, and by comparing the cumulative probability with the significance level, the negative samples are classified into "false negative samples", "true negative samples" and "fuzzy samples", and stored in the corresponding sets. This is achieved in the following way: The judgment of "false negative samples", "true negative samples" and "fuzzy samples" is relative to the anchor point. During image-to-text retrieval, for image anchor point i, among all negative sample texts j, j=1,...,B, ; Based on the image feature X and the negative text feature F, calculate its affinity matrix C; based on the fitted Gaussian mixture distribution, the elements in C Calculate the maximum posterior to determine the subgroup distribution to which the negative sample text j belongs, that is, determine the value of g, where g=2 means "same semantic distribution" and g=1 means "different semantic distribution"; On this basis, the cumulative probability is calculated, and by comparing the cumulative probability with the significance level, the negative text samples are classified into "false negative samples", "true negative samples" and "fuzzy samples" and stored in the corresponding sets; The cumulative probability is calculated as ; x is the independent variable, ranging from 0 to between, , is the parameter of the updated Gaussian distribution, is the cumulative probability; then, define the significance level ,when , and when g=1, the negative sample text is a "true negative sample" and is stored in the collection ,when , and g=2, the corresponding negative sample text is "false negative sample", stored in the collection , other negative sample texts are "fuzzy samples" and stored in the collection middle; Similar to image-to-text retrieval, text-to-image retrieval obtains a set of "true negative samples" of images in the same way , "false negative sample" set , and the "fuzzy sample" collection .

6. The image-text cross-modal retrieval method based on false negative sample soft distance constraint according to claim 1, characterized in that: The triplet loss is: ; The soft distance constraint triplet loss handles "true negative samples" as follows: ; The soft distance constraint triplet loss handles "fuzzy samples" as follows: ; The soft distance constraint triplet loss directly deletes the "false negative samples", so the soft distance constraint triplet loss is ; in, represents the soft distance constrained triplet loss, Represents the similarity between the image and text positive sample pairs, Represents the similarity between the image and the corresponding "true negative" text, Represents the similarity between the image and the corresponding "fuzzy" text, Represents the similarity between the image and the corresponding "negative" text, represents the similarity between the text and the corresponding "true negative" image, Represents the similarity between the text and the corresponding "blurred" image, Represents the similarity between the text and the corresponding "negative" image; and Respectively represent "true negative" image samples and "true negative" text samples; and Respectively represent "blurred" image samples and "blurred" text samples; and Respectively represent "negative" image samples and "negative" text samples; , represents the spacing parameter, which is used to increase the gap between the image and positive text pairs and the image and negative text pairs; is the weight of the "blurred" sample corresponding to the image, is the weight of the "fuzzy" sample corresponding to the text; ; represents the regulating factor, The same calculation method is used.

7. A cross-modal image-text retrieval system based on soft distance constraints for false negative samples, characterized by: The method for realizing the cross-modal image-text retrieval method based on false negative sample soft distance constraint as claimed in any one of claims 1 to 6 comprises a feature extraction module, a mixed distribution false negative sample mining module, and a loss calculation module. The feature extraction module includes an image encoder and a text encoder, which extract image features V and text features U respectively; and also includes two cue-aware Transformer decoders, which are applied to the image features V and the text features U respectively; The two decoders share parameters. When processing image and text features, they use image features V and text features U as keys and values, and prompts P as queries to implement network training and generate robust image features X and text features F respectively. The mixed distribution false negative sample mining module is designed for a sample pool for model fitting, the sample pool contains affinity matrices of homosemantic and heterosemantic cross-modal features in the same proportion, and the maximum expectation algorithm is used to fit the Gaussian mixture distribution of the affinity matrix; based on the fitted Gaussian mixture distribution, the affinity matrix of the current input image feature and text feature is calculated, and the sub-population distribution to which the negative sample belongs is determined using the maximum a posteriori method; on this basis, the cumulative probability is calculated, and by comparing the cumulative probability with the significance level, the negative samples are classified into "false negative samples", "true negative samples" and "fuzzy samples", and stored in the corresponding sets; The loss calculation module is used to calculate the soft distance constraint triplet loss.

Citation Information

Patent Citations

  • Universal interactive constraint graph layout system and layout method

    CN115017367A

  • Medical image segmentation method combining curve structure prompts and deep neural network

    CN118314121A

  • Rapid random model correction method for active learning type ocean platform structure

    CN119494245A

  • Method and apparatus for training cross-modal retrieval model, device and storage medium

    EP4053751A1

  • Methods and systems for generating polycubes and all-hexahedral meshes of an object

    US20140028673A1