Text-to-image retrieval model training method and retrieval method

By employing multimodal feature extraction and classification, combined with a training method incorporating bidirectional Kullback-Leibler divergence loss function and total loss function, the problem of coupling effect between noisy samples and true hard samples is solved, thereby improving the robustness and accuracy of the text-to-image retrieval model.

CN120929707APending Publication Date: 2025-11-11UESTC (SHENZHEN) ADVANCED RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510965706.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing text-to-image retrieval models suffer from coupling effects when dealing with noisy samples and real hard samples, leading to reduced retrieval accuracy.

Method used

Refined local and global features of sample pairs are obtained through multimodal feature extraction, and the samples are classified into noisy, fuzzy, and clean sample sets. The bidirectional Kullback-Leibler divergence loss function is used in the pre-training stage, and the total loss function is used for iterative training in subsequent stages to select difficult samples for learning.

Benefits of technology

It improves the robustness and accuracy of the text-to-image retrieval model, enhances the model's generalization ability, and improves the accuracy and precision of the retrieval method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929707A_ABST
    Figure CN120929707A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text image processing, in particular to a text-to-image retrieval model training method and retrieval method.The method comprises the steps that N text-image sample pairs are obtained; performing feature extraction processing on the sample pair to obtain refined local features and global features of the sample pair; and based on the refined local features and global features of the sample pairs, carrying out classification processing on the N sample pairs to obtain different sample sets and prediction labels of the samples, and obtaining difficult negative samples and difficult positive samples from the N sample pairs. In the preheating training stage, loss processing is carried out on prediction labels of all samples through a total bidirectional KL divergence loss function. Then, in a conventional training stage, a text-to-image retrieval model is obtained through the total loss function, the difficult positive samples and the difficult negative samples; according to the method, images containing noise and real difficult images can be efficiently distinguished, and the robustness, the accuracy and the precision of a text-to-image retrieval model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text image processing technology, and in particular to a text-to-image retrieval model training method and retrieval method. Background Technology

[0002] Text-to-image retrieval technology is a challenging task that links visual and linguistic patterns, aiming to retrieve target images from large-scale image collections based on natural language descriptions. This technology has significant practical applications in video surveillance, public safety, and retrieval systems. Currently, existing text-to-image retrieval models treat all high-loss samples as hard samples during the construction process—the most challenging samples or those difficult to classify or predict correctly. Training the retrieval model using these hard samples leads to low accuracy. This ignores the fundamental difference between pseudo-hard samples (noisy samples) and true hard samples (correctly labeled but poorly learned samples, such as low-quality images affected by occlusion, blurring, or low resolution, but correctly associated with the text), thus hindering effective optimization of retrieval models and methods. Noisy samples are those where the image and text do not match, or where the text incorrectly describes the image information. Furthermore, the dataset inevitably contains noisy labels caused by noisy samples and low-quality images, further complicating the process of identifying true hard samples.

[0003] Therefore, existing retrieval models or methods suffer from reduced retrieval accuracy due to the coupling effect between noisy samples and true hard samples. Summary of the Invention

[0004] This application provides a text-to-image retrieval model training method and retrieval method, which solves the technical problem of reduced retrieval accuracy caused by the coupling effect between noisy samples and real hard samples in the prior art. It achieves the technical effect of efficiently distinguishing noisy images from real hard images by decoupling noisy samples and real hard samples, thereby improving the robustness, accuracy and precision of the text-to-image retrieval model, and thus improving the accuracy and precision of the text-to-image retrieval method.

[0005] In a first aspect, embodiments of the present invention provide a method for training a text-to-image retrieval model, comprising:

[0006] Obtain N text-image sample pairs, where N is an integer not less than 1;

[0007] For each sample pair, feature extraction is performed to obtain refined local and global features for each sample pair;

[0008] Based on the refined local and global features of each sample pair, the N sample pairs are classified to obtain a noisy sample set, a fuzzy sample set, and a clean sample set, as well as the predicted labels of the noisy samples, the fuzzy samples, and the clean samples. Difficult negative samples are selected from the N sample pairs, and difficult positive samples are selected from the N sample pairs based on the noisy sample set.

[0009] During the warm-up training phase, the refined local and global features of each sample pair are processed by the total bidirectional Kullback-Leibler divergence loss function to achieve iterative training until the first preset number of iterations is reached, after which the regular training phase begins.

[0010] During the regular training phase, loss processing is applied to the predicted labels of the noisy samples, the ambiguous samples, and the clean samples using the total loss function, as well as the difficult positive samples and the difficult negative samples, to achieve iterative training until a second preset number of iterations is reached, resulting in a text-to-image retrieval model.

[0011] Based on the same inventive concept, in a second aspect, the present invention also provides a text-to-image retrieval method, comprising:

[0012] The target text and image set are input into a text-to-image retrieval model, wherein the text-to-image retrieval model is a retrieval model obtained by the text-to-image retrieval model training method as described in the first aspect;

[0013] The text-to-image retrieval model obtains refined local and global features for each matching pair based on the matching pairs formed between the target text and each image in the image set.

[0014] In each matching pair, a score matrix of local features is obtained based on the refined local features, and a score matrix of global features is obtained based on the global features. Then, a target score matrix is ​​obtained based on the score matrix of local features and the score matrix of global features, thereby obtaining the target score matrix of each matching pair.

[0015] Based on the target score matrix of each matching pair, the predicted probability of each matching pair is obtained, and based on the predicted probability of each matching pair, the sorting number of each image is obtained.

[0016] One or more technical solutions in the embodiments of the present invention have at least the following technical effects or advantages:

[0017] In this embodiment of the invention, after acquiring N sample pairs, feature extraction is performed on the text and image of each sample pair, i.e., multimodal feature extraction is performed. This allows for visual alignment of the multimodal features, resulting in local and global features of the text and image for each sample pair, and further refined local and global features for each sample pair. Refined local features capture fine-grained information about the sample pair, helping to identify subtle differences within the samples and effectively filtering out background noise or irrelevant information, highlighting key parts relevant to the target task. Global features capture the overall semantic information of the sample, crucial for understanding the overall picture and context of the sample. Extraction of global features ensures that the model finds the main correlation between text and image at the overall level, avoiding the one-sidedness of local matching. Thus, fine-grained cross-modal association is achieved, improving the retrieval accuracy and precision of the retrieval model.

[0018] Based on the refined local and global features of each sample pair, the samples are divided into different categories: noisy sample set, fuzzy sample set, and clean sample set. This facilitates subsequent filtering of noisy samples, refinement of fuzzy samples, and accurate learning to identify clean samples. The refined local and global features of each sample pair accurately separate the noisy, fuzzy, and clean sample sets, and yield predicted labels for each sample pair: the predicted labels for noisy samples, fuzzy samples, and clean samples. This prevents the model from fitting noisy samples, which could lead to incorrect optimization. The separated fuzzy and clean samples are then used for subsequent mining of positive and negative hard instances. Furthermore, hard negative samples can be selected from N sample pairs, and hard positive samples can be selected from N sample pairs based on the noisy sample set. In this way, by decoupling noisy samples from real hard samples, images containing noise and real hard images can be efficiently distinguished. Furthermore, the robustness, accuracy, and precision of the text-to-image retrieval model can be improved by accurately learning and recognizing hard positive samples and hard negative samples, thereby enhancing the accuracy and precision of the text-to-image retrieval method.

[0019] During the pre-training phase, to avoid sharp, overconfident predictions from the model, most normalized losses are made close to zero, making samples difficult to distinguish. This embodiment of the invention adds a negative entropy term (i.e., a bidirectional Kullback-Leibler divergence loss function) to penalize overconfident predictions and produce a more uniform loss distribution, which helps in effective modeling in subsequent stages. During the regular training phase, the total loss function incorporates the predicted labels of difficult positive samples, difficult negative samples, as well as noisy samples, blurred samples, and clean samples into a unified optimization framework, achieving collaborative learning of diverse samples. The model can effectively handle complex and low-quality samples while processing high-quality samples, improving the accuracy of fine-grained matching, reducing the negative impact of noisy samples, and enhancing the model's generalization ability. Therefore, by learning to process difficult samples through the total loss function, the robustness, accuracy, and precision of the text-to-image retrieval model are improved, thereby enhancing the accuracy and precision of the text-to-image retrieval method. Attached Figure Description

[0020] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0021] Figure 1 A flowchart illustrating the steps of the text-to-image retrieval model training method in an embodiment of the present invention is shown.

[0022] Figure 2 A flowchart illustrating the steps of a text-to-image retrieval method according to an embodiment of the present invention is shown. Detailed Implementation

[0023] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0024] Example 1

[0025] The first embodiment of the present invention provides a method for training a text-to-image retrieval model, such as... Figure 1 As shown, it includes:

[0026] S101, Obtain N text-image sample pairs, where N is an integer not less than 1;

[0027] S102, perform feature extraction processing on each sample pair to obtain the refined local and global features of each sample pair;

[0028] S103, based on the refined local and global features of each sample pair, classify the N sample pairs to obtain a noisy sample set, a fuzzy sample set and a clean sample set, as well as the predicted labels of the noisy samples, the predicted labels of the fuzzy samples and the predicted labels of the clean samples, and select hard negative samples from the N sample pairs, and select hard positive samples from the N sample pairs based on the noisy sample set.

[0029] S104, in the warm-up training phase, the refined local and global features of each sample pair are processed by the total bidirectional Kullback-Leibler divergence loss function to achieve iterative training until the first preset number of iterations is reached, after which the regular training phase begins.

[0030] S105. During the regular training phase, loss processing is applied to the predicted labels of noisy samples, ambiguous samples, and clean samples using the total loss function, as well as the predicted labels of difficult positive samples and difficult negative samples, to achieve iterative training. This process continues until the second preset number of iterations is reached, resulting in a text-to-image retrieval model.

[0031] It should be noted that a sample pair u is a pair of text and its corresponding image (this sample pair can be simply referred to as a sample), where the text is a textual description of the image. In this embodiment, the true labels of all sample pairs are set to 1, indicating that the text of the sample pair corresponds to the image. Furthermore, the text-to-image retrieval model training method can be specifically implemented using an electronic device. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements any of the steps of the text-to-image retrieval model training method in Embodiment 1.

[0032] In this embodiment, after acquiring N sample pairs, feature extraction is performed on each sample pair to obtain refined local and global features. Here, feature extraction is performed separately on the text and image of each sample pair, i.e., multimodal feature extraction is performed to visually align the multimodal features, resulting in local and global features for the text and image of each sample pair, and thus refined local and global features for each sample pair. Refined local features capture fine-grained information about the sample pair, helping to identify subtle differences in the samples and effectively filtering out background noise or irrelevant information, highlighting key parts relevant to the target task. Global features capture the overall semantic information of the sample, crucial for understanding the overall picture and context of the sample. Extraction of global features ensures that the model finds the main correlation between text and image at the overall level, avoiding the one-sidedness of local matching. Thus, fine-grained cross-modal association is achieved, improving the retrieval accuracy and precision of the retrieval model.

[0033] Based on the refined local and global features of each sample pair, the N sample pairs are then classified to filter out noisy, fuzzy, and clean sample sets. This division of samples into different categories facilitates subsequent filtering of noisy samples, refinement of fuzzy samples, and accurate learning to identify clean samples. During the classification of these sets, predicted labels for noise, fuzzy, and clean samples are obtained. This precise separation of the three sets, based on the refined local and global features of each sample pair, and the acquisition of predicted labels for each pair, prevents the model from fitting noisy samples and leading to incorrect optimization. The separated fuzzy and clean samples are then used for subsequent mining of positive and negative hard instances. Furthermore, hard negative samples can be filtered from the N sample pairs, and hard positive samples can be filtered from the N sample pairs based on the noise sample set. In this way, by decoupling noisy samples from real hard samples, images containing noise and real hard images can be efficiently distinguished. Furthermore, the robustness, accuracy, and precision of the text-to-image retrieval model can be improved by accurately learning and recognizing hard positive samples and hard negative samples, thereby enhancing the accuracy and precision of the text-to-image retrieval method.

[0034] During the warm-up training phase, the refined local and global features of each sample pair are processed using the total bidirectional Kullback-Leibler divergence loss function to achieve iterative training until the first preset number of iterations is reached, after which the regular training phase begins. Since the retrieval model in this embodiment has high requirements for fine-grained matching, the bidirectional Kullback-Leibler divergence loss function is used to process the separated sample set during the model's warm-up training phase. This allows the model's performance to stabilize initially and avoids overfitting to noise during the warm-up period, which could lead to overconfident (i.e., low-entropy) predictions. Compared to using cross-entropy loss, in the warm-up training phase, to avoid sharp, overconfident predictions that make most normalized losses close to zero and thus difficult to distinguish samples, this embodiment adds a negative entropy term (i.e., the bidirectional Kullback-Leibler divergence loss function) to penalize confident predictions and produce a more uniform loss distribution, which helps in effective modeling in subsequent stages.

[0035] After the initial training phase, the model transitions to the regular training phase, where it begins processing difficult samples. In this phase, the predicted labels for noisy, ambiguous, and clean samples are processed using the total loss function, along with the predicted labels for difficult positive and negative samples, to achieve iterative training. This process continues until the second predetermined number of iterations is reached, resulting in the text-to-image retrieval model. The total loss function incorporates the predicted labels for difficult positive, difficult negative, noisy, ambiguous, and clean samples into a unified optimization framework, enabling collaborative learning across diverse samples. The model can effectively handle both high-quality and complex / low-quality samples, improving the accuracy of fine-grained matching, reducing the negative impact of noisy samples, and enhancing the model's generalization ability. Therefore, by using the total loss function to learn and process difficult samples, the robustness, accuracy, and precision of the text-to-image retrieval model are improved, thereby enhancing the accuracy and precision of the text-to-image retrieval method.

[0036] Below, in conjunction with Figure 1 This embodiment details the specific implementation steps of the text-to-image retrieval model training method provided:

[0037] First, step S101 is executed to obtain N text-image sample pairs, where N is an integer not less than 1. Specifically, text-image sample pairs can be obtained through a camera device or an existing database, etc.

[0038] Next, step S102 is executed to perform feature extraction processing on each sample pair, obtaining the refined local features and global features of each sample pair.

[0039] Specifically, firstly, the image encoder of the CLIP (Contrastive Language-Image Pre-training) encoder encodes the images of each sample pair, resulting in an image sequence for each sample pair. Each image sequence carries a spatial token to identify the encoded spatial information of each sample pair. Then, the text encoder of the CLIP encoder encodes the text of each sample pair, resulting in a text sequence for each sample pair. Each image sequence carries a start-of-text token and an end-of-text token.

[0040] Specifically, the process of encoding the image of each sample pair using the CLIP encoder's image encoder is as follows: for each sample pair, the CLIP encoder's image encoder obtains the image I∈R of the sample pair. H×W×C H represents the image height, W represents the image width, and C represents the number of channels. The image is divided into... Non-overlapping small image patches. p represents the size of each small image patch. Each of the partitioned small image patches is linearly projected onto a feature vector, generating a vector sequence. i represents the index of the small image patch, v i This represents the feature vector corresponding to the i-th small block, i.e., the block image feature. Positional information is embedded in this vector sequence to encode the spatial information of the image, and a learnable [CLS] token, i.e., a spatial token, is added before the vector sequence to form the image sequence. That is, where d is the dimension of the shared latent space. [CLS] token v cls (i.e., spatial tokens) serve as global features of the image, identifying the spatial information of each sample pair's image sequence. As a local feature of the image.

[0041] The above image encoding process is performed on the image of each sample pair to obtain the local and global features of each sample pair.

[0042] The process of encoding the text for each sample pair using the CLIP encoder's text encoder is as follows: For each sample pair image, the CLIP encoder obtains the text T of the sample pair, encodes the text using bytes of a preset vocabulary size, adds an [SOS] token (i.e., a text start token) to indicate the start of the text sequence, and appends an [EOS] token (i.e., a text end token) to mark the end of the text sequence, thus forming the text sequence. The preset vocabulary size can be set according to actual needs; for example, the preset vocabulary size is 49152 lowercase bytes. t It is the number of word tags. It is the Nth t The feature vectors of each word, i.e., word features. After being processed by the CLIP encoder's text encoder, the [EOS] token t is used. eos Word-level features serve as global features of the text. As a local feature of the text. t eos Also serves as the end-of-text token, t sos For the text start token.

[0043] The above text encoding process is performed on the text of each sample pair to obtain the local and global features of the text of each sample pair.

[0044] In this embodiment, the image encoder and text encoder of the CLIP encoder encode the image and text of each sample pair respectively, obtaining the global and local features of the image and the global and local features of the text of the sample pair. Thus, in each sample pair, the local features of the image capture fine-grained and key information, while the global features capture the overall picture and semantic information of the image. Similarly, the local features of the text capture fine-grained and key information, while the global features capture the overall picture and semantic information of the text. In this way, multimodal feature extraction is performed on each sample pair to visually align the multimodal features and provide a foundation for constructing high-precision refined local and global features for subsequent sample pairs.

[0045] After obtaining the global and local features of the image and the text of each sample pair, the target image sequence of each sample pair is selected from the image sequence of each sample pair and the target text sequence of each sample pair is selected from the text sequence of each sample pair through the self-attention weight selection mechanism.

[0046] Specifically, in each sample pair, the image sequence In Each block of image features in the image is associated with v. cls Perform weight value calculation, i.e., v1 and v cls Calculate the weight values, v2 and v cls Weight values ​​are calculated, and so on. Therefore, the image features of each block are obtained along with v. cls The weights are assigned to a set of values, sorted in descending order, and a first preset number of image features corresponding to each weight value are selected, starting with the first weight value. This first preset number can be set according to actual needs. Regularization is then applied to the image features corresponding to this first preset number of weight values ​​to obtain the target image sequence.

[0047] For example, Each block of image features in the image is associated with v. cls Weight values ​​are calculated, and the sorted weight values ​​are 0.99, 0.98, 0.87, 0.77, 0.76, ..., 0.11. The block image features corresponding to the first preset number of weight values ​​(4) are selected, resulting in the block image features corresponding to 0.99, 0.98, 0.87, and 0.77. Regularization is then performed on these block image features to obtain the regularized image feature vector, which is the target image sequence.

[0048] Similarly, in each sample pair, the text sequence In Each word feature in the text is related to t. eos Perform weight value calculation, i.e., t1 and t eos Calculate the weight values, t2 and t eos Weights are calculated, and so on. Therefore, the features of each word and t are obtained. eos The weights are assigned to each word, and these weights are sorted in descending order. A second preset number of word features corresponding to the first weight value are selected. This second preset number can be set according to actual needs. Regularization is then applied to the word features corresponding to the second preset number of weight values ​​to obtain the regularized text feature vector, i.e., the target text sequence.

[0049] In this embodiment, optimizing the feature similarity of global features alone may fail to capture the fine-grained interactions between two modalities, thus increasing the model's performance limitations. To address this issue, local features are utilized for learning to achieve more discriminative embedding representations, thereby uncovering fine-grained correspondences. In CLIP, the global features represented by [CLS] tokens and [EOS] tokens are obtained through a weighted aggregation of features from all local tokens. These weights reflect the correlation between the global features ([CLS] tokens and [EOS] tokens) and each local token. This allows for the aggregation of important local features based on these correlation weights to embed more representative local features, i.e., refining local features. Therefore, this embodiment utilizes a self-attention weight selection mechanism to retain block-specific features and word features with the highest attention weights, improving the model's recognition and retrieval accuracy and precision.

[0050] After obtaining the target image sequence and target text sequence for each sample pair, max pooling is performed on the target image sequence and target text sequence for each sample pair to obtain the refined local features of each sample pair. Furthermore, the global features of each sample pair are obtained based on the spatial token and text end token of each sample pair.

[0051] Specifically, in each sample pair, the max pooling process is as shown in equations (1) and (2):

[0052]

[0053] in, For the target image sequence, For the target text sequence, MLP(·) represents a multilayer perceptron layer, FC(·) represents a linear transformation, MaxPool(·) represents max pooling, and Φ v Φ represents the improved local features of the image. t This represents the improved local features of the text. Φ v and Φ t Forming refined local features

[0054] In each sample pair, the space token v of the sample pair cls and text end token t eos Forming global features

[0055] In this embodiment, refined local features for each sample pair are obtained through max pooling, and global features for each sample pair are obtained based on the spatial token and text end token. By learning from the local features, a more discriminative representation is embedded, revealing the correspondence between fine-grained and global features, thus forming refined local features. And based on the [CLS] token v... cls and [EOS] token t eos This yields global features, providing a solid foundation for subsequent loss calculations and sample classification. It also improves the model's precision, accuracy, and generalization ability, as well as the robustness and precision of the training method.

[0056] In the process of obtaining the refined local and global features of each sample pair, the matching probability of each sample pair is also obtained: based on N sample pairs, N text features and N image features are obtained. For each text feature, a matching operation is performed between the text feature and each of the N image features to obtain the matching probability between the text feature and each image feature. The above matching operation is performed on each text feature, and the matching probability between each text feature and each image feature is calculated.

[0057] Specifically, there are N sample pairs u, each consisting of a text and an image. Based on these N sample pairs, N text features and N image features are obtained. These N text features and N image features are then recombined to obtain the recombined sample u. 重组 (V i T j V i T represents the image features of the i-th image. j This represents the text features of the j-th text. Reconstructed sample u 重组 Authentic Labels Defined as: if V i and T j If they belong to the same target object, it means V i and T j Belonging to the same identity, V i and T j If they are a matching pair, then This indicates that the sample represents the true label p of u. u This also indicates that sample pair u is a positive sample; otherwise V represents i and T j Not belonging to the same target object, indicating V i and T j Belonging to different identities, V i and T j This is a non-matching pair. The target object is an object with an identifier. For example, the target object could be a pedestrian, and the target object could be a car.

[0058] The entire retrieval model takes N sample pairs u as input. These N sample pairs u form an N×N input matrix. Each element in this input matrix is ​​a recombined sample u. 重组 Recombined samples whose text and image features belong to the same target object are identified as positive samples. Recombined samples whose text and image features do not belong to the same target object are identified as negative samples. The true label for each positive sample is 1, and the true label for each negative sample is 0. The elements on the diagonal starting from the element in the first row and first column of the input matrix represent these N sample pairs u.

[0059] For each text feature, a matching operation is performed between the text feature and each of the N image features to obtain the matching probability between the text feature and each image feature. The matching operation process is as follows: reconstruct sample u 重组 Matching probability The calculation formula for the softmax function in formula (3) is shown below:

[0060]

[0061] Among them, sim(Vi T j ) is V i and T j The similarity between them, sim(V) i T n ) is V i and T n The similarity between them, where τ is the temperature coefficient.

[0062] The above matching operation is performed on each text feature, and the matching probability of each text feature with each image feature is, i.e., the probability of matching each recombined sample u. 重组 The matching probability.

[0063] In this embodiment, the matching operation is used to obtain the matching probability, which is then used to calculate the loss based on matching local and global features. The matching probability between each text feature and all image features is calculated using the Softmax function, accurately measuring the similarity between text and image features. This approach provides a clear matching probability for each recombined sample, thereby quantifying the degree of association between text and images. For each text feature, a matching operation is performed with all image features to form a complete probability distribution. This many-to-many matching method can comprehensively characterize the relationship between text and images.

[0064] Furthermore, based on the principle of the softmax function calculation formula (3), the improved local feature Φ is calculated. v and Φ t The cosine similarity between the samples yields the refined local features of the sample pair u. Matching probability (Also known as the local feature matching probability of sample u), and the global feature v cls and t eos The cosine similarity between the samples yields the global features of the sample pair u. Matching probability (Also known as the global feature matching probability of sample u).

[0065] In step S102 of this embodiment, feature extraction is performed on the text and image of each sample pair, i.e., multimodal feature extraction is performed to visually align the multimodal features, obtaining local and global features of the text and the image of each sample pair, and further obtaining refined local and global features of each sample pair. Refined local features capture fine-grained information of the sample pair, helping to identify subtle differences in the samples and effectively filtering out background noise or irrelevant information, highlighting key parts relevant to the target task. Global features capture the overall semantic information of the sample, which is crucial for understanding the overall picture and context of the sample. Extraction of global features ensures that the model finds the main correlation between text and image at the overall level, avoiding the one-sidedness of local matching. Thus, fine-grained cross-modal association is achieved, improving the retrieval accuracy and precision of the retrieval model.

[0066] Next, step S103 is executed, which classifies the N sample pairs based on the refined local and global features of each sample pair to obtain a noisy sample set, a fuzzy sample set, and a clean sample set, as well as the predicted labels of the noisy samples, the fuzzy samples, and the clean samples. Difficult negative samples are selected from the N sample pairs, and difficult positive samples are selected from the N sample pairs based on the noisy sample set.

[0067] Specifically, the Fuzzy-c means (FCM) clustering module performs clustering on the refined local and global features of each sample pair, obtaining the local feature membership probability and the global feature membership probability of each sample pair. Both the local and global feature membership probabilities represent the probability of belonging to the low-loss cluster.

[0068] Specifically, the refined local features of each sample are processed through the Fuzzy-c means (FCM) clustering module. Clustering is performed to group samples and obtain the local feature membership probabilities for each sample. Furthermore, through the fuzzy C-means clustering module, the global features of each sample are analyzed. Perform clustering and grouping to obtain the global feature membership probability of each sample. Therefore, the fuzzy C-means clustering module assigns a membership probability to each sample pair. That is, the local feature membership probability of a sample pair. and global feature membership probability Here, the membership probability represents the likelihood of belonging to a low-loss cluster, such as the local feature membership probability of a sample pair. This represents the probability that a sample pair's refined local features belong to a low-loss cluster in the refined local feature clustering classification, and the global feature membership probability. This indicates the probability that the global features of a sample pair belong to a low-loss cluster in the clustering classification of global features.

[0069] It should be noted that the following text All operations or formulas represent Operations or formulas that are consistent with the principles of summation.

[0070] To prevent error accumulation in the early stages, the number of clusters is determined based on the dynamic variance of the loss values ​​calculated from all pairs of samples participating in the clustering process. This ensures adaptive data separation. The number of clusters is shown in formula (4):

[0071]

[0072] Where, γ var It is the variance threshold.

[0073] After obtaining the membership probability of each sample pair Then, based on the local feature membership probability of each sample pair, the predicted local feature value of each sample pair is obtained. And based on the global feature membership probability of each sample pair, the predicted global feature value of each sample pair is obtained.

[0074] Specifically, the local feature prediction value and global feature prediction value of each sample pair are obtained through formula (5).

[0075]

[0076] in, This represents the state of sample u in the current round t, i.e. This represents the local feature prediction value of the sample for u in the current round t. This represents the global feature prediction value of the sample pair u in the current round t. It is the probability that sample pair u belongs to the low-loss cluster during k rounds of training, i.e. It is the sample pair u during the k rounds of training. The probability of belonging to the low-loss cluster. Also known as the local feature membership probability of sample pair u in k rounds. It is the probability that sample pair u belongs to the low-loss cluster during k rounds of training. Also known as the global feature membership probability of sample pair u in k rounds. t is the current round of training. min(t,z) ensures that an appropriate window is selected in the initial stage of training, that is, when t < z, a sliding window of size t is selected, and after t ≥ z, a sliding window of fixed size z is selected. It sets a probability threshold, which can be set according to actual needs.

[0077] Formula (5) means that if the probability that a sample pair belongs to the low-loss cluster exceeds a set probability threshold, then the sample pair is determined to be clean.

[0078] This embodiment considers that the instantaneous loss of a single sample pair (i.e., the loss in the current round) is an unstable signal. Due to the randomness of the training process of the model in this embodiment, it can fluctuate rapidly, and a sliding window of size z is used to smooth the signal.

[0079] It should be noted that one round of training means performing steps S101-S105 once for N sample pairs. Since the data for N sample pairs is enormous, in practice, the N sample pairs are divided into several batches, and steps S101-S105 are performed sequentially on each batch to complete one round of training. This speeds up the model training process and achieves high precision and accuracy. For example, if the N sample pairs are divided into 3 batches, in practice, steps S101-S105 are performed once for the first batch, then once for the second batch, and so on for the third batch, thus completing one round of training.

[0080] After obtaining the local and global feature prediction values ​​for each sample pair, the target sample category corresponding to each sample pair is obtained based on the local and global feature prediction values. The sample pairs are then assigned to the target sample categories to form a noise sample set, a fuzzy sample set, and a clean sample set. The target sample categories include: noise sample category, model sample category, and clean sample category.

[0081] Specifically, the target sample category corresponding to each sample pair is obtained through formula (6):

[0082]

[0083] in, Indicates the target sample category. This represents the noise sample set, i.e., the noise sample categories. This represents a fuzzy sample set, i.e., fuzzy sample categories. This represents the clean sample set, i.e., the clean sample category. v represents the logical OR operator. This assigns each sample pair to its corresponding sample category, forming the noisy sample set, the fuzzy sample set, and the clean sample set.

[0084] For example, to illustrate formula (6), if and or and but Corresponding sample pairs If a sample belongs to the fuzzy sample category, it is considered a fuzzy sample. This sample pair is then classified into the fuzzy sample category, forming a fuzzy sample set with other fuzzy samples.

[0085] In this embodiment, the FCM clustering algorithm is used to group the refined local feature loss and global feature loss of each sample pair, obtaining the local feature membership probability and the global feature membership probability of each sample pair. This assigns sample pairs to different categories in the form of soft membership. Subsequently, fuzzy samples and boundary samples (i.e., difficult positive samples) can be labeled based on this membership probability, effectively addressing the complexity of fuzzy and boundary samples. Next, based on the local feature membership probability of each sample pair, the predicted local feature value of each sample pair is obtained, and based on the global feature membership probability of each sample pair, the predicted global feature value of each sample pair is obtained. Then, based on the predicted local and global feature values ​​of each sample pair, the target sample category corresponding to each sample pair is obtained, and the sample pairs are assigned to the target sample category, forming a noisy sample set, a fuzzy sample set, and a clean sample set. By combining the predicted local and global feature values, the sample pairs are accurately assigned to the noisy sample, fuzzy sample, and clean sample categories. The resulting noisy, fuzzy, and clean sample sets provide high-quality input data for subsequent model training, ensuring continuous optimization of model performance.

[0086] After obtaining the noisy sample set, the fuzzy sample set, and the clean sample set, the labels of elements in each sample set are updated. Specifically, for each clean sample in the clean sample set, the true label of the clean sample is used as the predicted label for the clean sample. For clean samples, since the global features and refined local features of the clean sample are consistent, the label is... The labels of clean samples remain unchanged.

[0087] In each noise sample within the noise sample set, the true label of the noise sample is modified to a preset label value, and this preset label value is used as the predicted label of the noise sample. The preset label value can be set according to actual needs, for example, it can be set to 0. For noise samples, since noise samples indicate the presence of noise and are therefore unreliable, the predicted label of the noise sample is corrected to the preset label value.

[0088] In each fuzzy sample in the fuzzy sample set, the predicted label of the fuzzy sample is obtained based on the refined local and global features of the fuzzy sample.

[0089] Specifically, based on the refined local features of the fuzzy samples, the local feature matching probability of the fuzzy samples is obtained. Based on the global features of the fuzzy samples, the global feature matching probability of the fuzzy samples is obtained.

[0090] The average matching probability of the fuzzy sample is obtained by considering the local feature matching probability and the global feature matching probability of the fuzzy sample. As shown in formula (7):

[0091]

[0092] Where u represents a sample pair belonging to the fuzzy sample category, i.e., u is a fuzzy sample in this case. u can also be represented as u 模糊 replace.

[0093] Based on the local feature membership probabilities and global feature membership probabilities of the fuzzy sample, the average prediction probability F of the fuzzy sample is obtained. u,t As shown in formula (8):

[0094]

[0095] Based on the average matching probability of fuzzy samples and average prediction probability F u,t The membership soft labels of the fuzzy samples are obtained, as shown in formula (9). The membership labels of the fuzzy samples are then determined as the predicted labels of the fuzzy samples.

[0096]

[0097] in, This is a soft label representing membership, also known as a predicted label for fuzzy samples. u The true label for the fuzzy sample.

[0098] In this embodiment, for each fuzzy sample, the membership soft label of the fuzzy sample is used as the predicted label of the fuzzy sample. The membership soft label balances the true label and the model predicted label of the fuzzy sample to reduce noise.

[0099] Therefore, the predicted labels for each sample set (i.e., the predicted labels for each sample pair) are modified as shown in Equation (10):

[0100]

[0101] in, The predicted label for u for each sample.

[0102] By forcibly correcting the labels of noisy samples (i.e., setting them to 0), the interference of low-quality samples on model performance can be significantly reduced, ensuring that the model focuses more on high-quality samples and important ambiguous samples. Through the label update mechanism, the true label of each sample pair is adjusted to a predicted label that better matches its category, providing a more accurate and efficient supervision signal for model training.

[0103] Furthermore, the specific processes for selecting difficult negative samples from N sample pairs, and for selecting difficult positive samples from N sample pairs based on a noisy sample set, are as follows:

[0104] Combine the N sample pairs to obtain N 2 The recombined samples were divided into three groups. Recombined samples whose text features and image features belonged to the same target object were identified as positive samples, while recombined samples whose text features and image features did not belong to the same target object were identified as negative samples.

[0105] For example, suppose the model inputs three sample pairs: N1(V1,T1), N2(V2,T2), and N3(V3,T3). The textual features T and graphical features V of these three sample pairs are recombined to obtain nine recombined sample pairs: N1'(V1,T1), N2'(V1,T2), N3'(V1,T3), N4'(V2,T1), N5'(V2,T2), N6'(V2,T3), N7'(V3,T1), N8'(V3,T2), and N9'(V3,T3). Of these nine reconstructed sample pairs, if the text features and image features of the reconstructed sample pair belong to the same target object, pedestrian A, then reconstructed sample pairs N1'(V1,T1), N5'(V2,T2), N8'(V3,T2), and N9'(V3,T3) are identified as positive samples. The text features and image features of the other reconstructed sample pairs do not belong to the same target object, pedestrian A, then the other reconstructed sample pairs are identified as negative samples, i.e., reconstructed sample pairs N2', N3', N4', N6', and N7' are negative samples.

[0106] In positive samples, if a positive sample does not belong to the noisy sample set, and the current local feature membership probability and the current global feature membership probability of the positive sample meet the set conditions, then the positive sample is identified as a difficult positive sample.

[0107] Specifically, as shown in formula (11), the difficult positive samples are obtained:

[0108]

[0109] in, For the set of difficult positive samples, Let be the local feature membership probability of sample pair u in the current round t, i.e., the current local feature membership probability. Let be the global feature membership probability of sample pair u in round t, i.e., the current global feature membership probability, and ^ be the logical AND operator. Then, the condition Φ(·) is defined as:

[0110]

[0111] Where |·| represents the absolute value.

[0112] The truly difficult positive samples are identified based on a dual standard: 1. There is a significant difference θ between the probabilities of belonging to the low-loss cluster calculated from refined local features and global features. diff The difference parameter θ diff It can be set according to actual needs, such as θ diff =0.3. 2. The probability falls within the first probability threshold θ. low =0.4 and the second probability threshold θ high Within a moderate uncertainty range with a threshold of 0.6, this ensures that the selected samples are truly difficult samples, avoiding noise or overly confident simple samples. Wherein, the first probability threshold θ... low Second probability threshold θ high It can be set according to actual needs.

[0113] Among the negative samples, the negative sample with the highest matching probability is identified as the difficult negative sample.

[0114] Specifically, as shown in formula (13), the difficult negative sample Ψ is obtained:

[0115]

[0116] Where Ψ represents a difficult negative sample, For the first Line 1 The matching probability of sample pairs in a column.

[0117] Then, step S104 is executed. In the warm-up training phase, the refined local and global features of each sample pair are processed by the total bidirectional Kullback-Leibler divergence loss function to achieve iterative training until the first preset number of iterations is reached, after which the regular training phase begins.

[0118] Specifically, because the retrieval model in this embodiment has high requirements for fine-grained matching, an adaptive bidirectional progressive hard sample mining process is designed. The purpose of this design is to enhance the retrieval model's ability to learn from truly hard samples. To better understand the loss function processing in the retrieval model and training method of this embodiment, the Cross Entropy Loss (CE) function will be explained first.

[0119] In N text-image sample pairs, the CE loss function from image to text is shown in Equation (14):

[0120]

[0121] in, p represents the positive CE loss value from image to text. i,j For image-text sample pairs (V) i ,T j ) real label, q i,j The image-text sample pairs (V) predicted by the retrieval model in this embodiment. i ,T j The matching probability of q i,j Further indicating the V in it i and T j The refined local features of the image and the text are calculated once, that is, the refined local features of the corresponding sample pair are calculated once, and then the global features of the image and the text are calculated once, that is, the global features of the corresponding sample pair are calculated once. ∈ is a constant to prevent numerical instability. V in formula (3) i and T j The exchange represents the probability q of a text-to-image (text-image) match. j,i To further calculate The same applies to subsequent cases, so I will not repeat them here.

[0122] For the labels of noisy samples (i.e., sample pairs where images and text are incorrectly associated, and sample pairs where text incorrectly describes image information), p i,j It may not always accurately represent the true distribution of the data. Instead, q i,j This can provide a better approximation. Therefore, considering the inverse relation of CE is equally important. Inverse CE term This is represented as shown in formula (15):

[0123]

[0124] in, This represents the inverse CE loss value for image-to-text conversion.

[0125] Based on the principles of formulas (14) and (15), the positive CE loss value of text to image (i.e., t2i) is obtained. and reverse CE loss value The formula will not be elaborated here.

[0126] The total bidirectional cross-entropy (BCE) loss function is shown in Equation (16):

[0127]

[0128] in,

[0129] This represents the total bidirectional cross-entropy (BCE) loss value. The total loss value is the positive cross-entropy CE. This represents the total loss value of the reverse CE.

[0130] During the initial convergence of the model and training method, due to the unstable performance of the model, it is necessary to perform a first preset number of iterations of pre-training on all data using the BCE loss function to warm up the model and prevent error accumulation. The first preset number of iterations can be set according to actual needs, such as performing 5 pre-training iterations on all data using the BCE loss function. However, during the pre-training period, the model and training method quickly overfit to noise and produce overconfident (i.e., low entropy) predictions, resulting in most samples having near-zero normalized loss. Therefore, to penalize the overconfident predictions from the model and training method during the pre-training phase, the negative entropy term is... The CE loss function is incorporated. Considering the symmetry of bidirectional CE, the BCE loss function is reformulated as the Kullback-Leibler divergence (KL) loss function, as shown in equations (19) and (20):

[0131]

[0132] in, The forward Kullback-Leibler divergence loss value from image to text. The inverse Kullback-Leibler divergence loss value from image to text is given. Based on the principles of formulas (19) and (20), the forward Kullback-Leibler divergence loss value from text to image (i.e., t2i) is obtained. and the reverse Kullback-Leibler divergence loss value The formula will not be elaborated here.

[0133] Similar to the CE loss function, the total bidirectional KL divergence (BKL) loss function is shown in equation (21):

[0134]

[0135] in,

[0136] This represents the total loss value of the two-way Kullback-Leibler divergence. This represents the total loss value for the positive Kullback-Leibler divergence. This represents the total loss value of the reverse Kullback-Leibler divergence.

[0137] It should be noted that during the warm-up training phase, the execution order is step S101-step S102-step S104. Step S103 is not executed during this phase. During the regular training phase, the execution order is step S101-step S102-step S103-step S105.

[0138] Therefore, during the warm-up training phase, the refined local and global features of each sample pair are processed using the total bidirectional Kullback-Leibler divergence loss function to achieve iterative training until the first preset number of iterations is reached, after which the regular training phase begins. This ensures the accuracy, robustness, and precision of the model and training method during the initial convergence process, avoiding overconfident predictions.

[0139] Finally, step S105 is executed. During the regular training phase, loss processing is performed on the predicted labels of noisy samples, ambiguous samples, and clean samples using the total loss function, as well as the predicted labels of difficult positive samples and difficult negative samples, to achieve iterative training until the second preset number of iterations is reached, and then the text-to-image retrieval model is obtained.

[0140] Specifically, after the warm-up training phase, the model and training method transition to the phase of handling difficult samples, i.e., the regular training phase. In the regular training phase, it is necessary to obtain the difficult positive mining matrix W based on the difficult positive samples. pos ∈R N×N The difficult negative mining matrix W is obtained based on the difficult negative samples. neg ∈R N×N The difficulty lies in mining the matrix W. pos ∈R N×N and W neg ∈R N×N The specific acquisition process is shown in formulas (24) and (25):

[0141]

[0142] in, Index represents the first Line text features Index represents the first The column image features are represented by w, which represents the weight parameters that are dynamically adjusted during training.

[0143]

[0144] in, In the first Line 1 The matching probability of sample pairs in a column is shown in formula (26):

[0145] w = 1 + η·t (26);

[0146] Where η is the weighted growth rate.

[0147] To avoid incorrect optimizations during the early, unstable training phase, this embodiment introduces a target separability metric Δ. avg Target separability measure Δ avg It can measure the model's ability to distinguish between difficult positive samples and difficult negative samples. The target separability measure is Δ. avg The forms of expression are shown in formulas (27) and (28):

[0148]

[0149] Where, Δ i2t Δ is a measure of the separability of images to text. t2i This is a measure of the separability of text to images.

[0150] Meanwhile, in the total loss function during the regular training phase, this embodiment incorporates an adaptive weighting mechanism to dynamically scale the positive and negative components of the loss, reducing the influence of simple samples and improving the model's recognition accuracy and precision, making the model focus more on challenging samples. The adaptive weighting mechanism is shown in formulas (29) and (30):

[0151] α pos =1-Q (29);

[0152] β neg =Q (30);

[0153] Where, Q∈R N×N This is the matching probability matrix for all samples.

[0154] In Δ avgWhen the value is greater than δ, δ represents the threshold for focusing on difficult samples. This means that when the target separability metric is greater than the threshold for focusing on difficult samples, the model only focuses on difficult samples after mastering the easier samples. The predicted labels of noisy samples, ambiguous samples, and clean samples are processed using the first total loss function, along with the predicted labels of difficult positive and difficult negative samples. In other words, the predicted labels of all samples are processed using the first total loss function, as shown in formula (31).

[0155]

[0156] in, The loss value of the first total loss function. The label matrix formed by the predicted labels of noisy samples, blurred samples, and clean samples, α pos To lose the positive component, β neg To lose the negative component, W pos W is the hard positive mining matrix obtained from hard positive samples. neg This is the hard negative mining matrix obtained from hard negative samples. These represent the positive cross-entropy loss values ​​for text-to-image and image-to-text. represents the inverse cross-entropy loss value from text to image and the inverse cross-entropy loss value from image to text.

[0157] In Δ avg When Δ ≤ δ, i.e., when the target separability metric is greater than the hard sample attention threshold, and when the target separability metric is not greater than the hard sample attention threshold, then the predicted labels of noisy samples, ambiguous samples, and clean samples are processed using the second total loss function. The total loss function includes the first total loss function and the second total loss function. In other words, when Δ avg When ≤δ, the predicted labels of all samples are processed by the second total loss function, as shown in formula (32):

[0158]

[0159] in, This represents the loss value of the second total loss function.

[0160] In the actual process of the regular training phase, the first total loss function is usually used for loss processing in the first few rounds of training, and the second total loss function is used thereafter. After the regular training phase reaches the second preset number of iterations, the text-to-image retrieval model is obtained. The second preset number of iterations can be set according to actual needs.

[0161] In this embodiment, the combined effect of the total bidirectional Kullback-Leibler divergence loss function during the preheating training phase and the total loss function during the regular training phase gradually uncovers difficult samples while filtering out noise. This ensures the accuracy, precision, and robustness of the retrieval model and training method in this embodiment.

[0162] The training method and the retrieval model trained using this embodiment can efficiently and accurately identify noisy samples (i.e., samples where the text and image do not match or where the text incorrectly describes the image information), ambiguous samples (i.e., samples that the model has difficulty distinguishing, images with low model recognition), and difficult samples (i.e., samples where the text and image match, but are also low-quality images, such as occluded, blurred, or low-resolution images). Therefore, the retrieval model and training method of this embodiment can achieve high-precision and high-accuracy retrieval and recognition functions. Compared to other similar retrieval models, it has higher precision and accuracy, and better robustness. Furthermore, the images identified by the retrieval model trained using this embodiment can help other object detection models or other models accurately identify image targets, achieving the goal of exploring image diversity.

[0163] For example, in the field of intelligent security, the retrieval model trained using the training method of this embodiment can retrieve images matching the "suspicious person description" from massive amounts of surveillance images. These images are then input into a target detection model, which can more accurately identify key information such as the facial features and items carried by suspicious persons, assisting the police in quickly locating targets and improving security monitoring efficiency.

[0164] In the image search function of e-commerce platforms, when a user enters "red dress" as the target text, the retrieval model in this embodiment can retrieve images matching the target text from the product image database. These images are then input into an image classification model for further analysis. Based on detailed features such as the style and design of the images, the image classification model can recommend red dresses that better suit the user's needs, thus improving the shopping experience.

[0165] In the field of cultural heritage preservation, researchers can input the target text "dougong structure of ancient architecture." The retrieval model in this example can retrieve images of buildings with dougong structures from a vast database of historical building images and input them into an image recognition model specifically designed for architectural style analysis. This helps researchers to study the structural characteristics and artistic styles of ancient buildings more deeply, providing more accurate data for the restoration and preservation of cultural heritage.

[0166] Example 2

[0167] Based on the same inventive concept, the second embodiment of the present invention also provides a text-to-image retrieval method, such as... Figure 2 As shown, it includes:

[0168] S201, Input the target text and image set into the text-to-image retrieval model, wherein the text-to-image retrieval model is the retrieval model obtained by the text-to-image retrieval model training method of Example 1;

[0169] S202, the text-to-image retrieval model is based on the matching pairs formed between the target text and each image in the image set, and obtains the refined local features and global features of each matching pair;

[0170] S203, in each matching pair, the score matrix of local features is obtained based on the refined local features, and the score matrix of global features is obtained based on the global features. Then, the target score matrix is ​​obtained based on the score matrix of local features and the score matrix of global features, and thus the target score matrix of each matching pair is obtained.

[0171] S204. Based on the target score matrix of each matching pair, obtain the predicted probability of each matching pair, and based on the predicted probability of each matching pair, obtain the sorting number of each image.

[0172] It should be noted that the text-to-image retrieval method can be implemented by an electronic device, which includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, it implements the steps of any of the methods in the text-to-image retrieval method of Embodiment 2.

[0173] Below, in conjunction with Figure 2 The following details the specific implementation steps of the text-to-image retrieval method provided in this embodiment:

[0174] First, step S201 is executed, inputting the target text and image set into the text-to-image retrieval model. The text-to-image retrieval model is the retrieval model obtained using the text-to-image retrieval model training method of Example 1. The image set can be an existing image database or a set of manually collected images. Each image in the image set carries identity information.

[0175] Next, step S202 is executed, where the text-to-image retrieval model obtains the refined local and global features of each matching pair based on the matching pairs formed between the target text and each image in the image set.

[0176] Specifically, the target text is matched with the first image in the image set to form the first matching pair. The target text is matched with the second image in the image set to form the second matching pair. The target text is matched with the third image in the image set to form the third matching pair, and so on. After obtaining each matching pair, the refined local features and global features of each matching pair are obtained based on the target text and image of the matching pair. The process of obtaining the refined local features and global features of each matching pair is consistent with the principle of "obtaining the refined local features and global features of each sample pair" in the training process of the retrieval model. That is, the process of "obtaining the refined local features and global features of each matching pair" in step S202 of this embodiment is consistent with the principle of step S102 in embodiment one, and will not be repeated here. The refined local features of each matching pair are represented as follows: The global feature representation of each matching pair is as follows: (v cls , t eos ).

[0177] Next, step S203 is executed. In each matching pair, the score matrix of the local features is obtained based on the refined local features, and the score matrix of the global features is obtained based on the global features. Then, the target score matrix is ​​obtained based on the score matrix of the local features and the score matrix of the global features, and thus the target score matrix of each matching pair is obtained.

[0178] Specifically, based on the refined local features, the score matrix of the local features is obtained, that is, Φ t Multiply by Φ v The transpose of the matrix is ​​used to obtain the score matrix of the local features. Then, based on the global features, the score matrix of the global features is obtained, i.e., t. eos Multiply by v cls The transpose of the matrix is ​​used to obtain the global feature score matrix. Then, the average of the local feature score matrix and the global feature score matrix is ​​calculated to obtain the target score matrix of the matching pair.

[0179] Then, step S204 is executed to obtain the predicted probability of each matching pair based on the target score matrix of each matching pair, and to obtain the sorting number of each image based on the predicted probability of each matching pair.

[0180] Specifically, for each matching pair, the predicted probability of each matching pair is obtained based on the target score matrix of the matching pair (which is a one-dimensional matrix) and formula (3). The predicted probabilities of the matching pairs are sorted and numbered in descending order, and the corresponding images retrieved based on the target text are displayed. The smaller the sort number of the image, the higher the matching degree between the image and the target text, that is, the higher the consistency between the identity information of the image and the target text.

[0181] In this embodiment, the retrieval model obtained through the text-to-image retrieval model training method of Embodiment 1 retrieves the corresponding image based on the target text. Thus, fine-grained cross-modal association modeling in complex scenarios is achieved, decoupling noisy samples from truly difficult samples, efficiently distinguishing noisy images from truly difficult images, improving the robustness, accuracy, and precision of the text-to-image retrieval model, and consequently enhancing the accuracy and precision of the text-to-image retrieval method.

[0182] Those skilled in the art will understand that although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.

[0183] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for training a text-to-image retrieval model, characterized in that, include: Obtain N text-image sample pairs, where N is an integer not less than 1; For each sample pair, feature extraction is performed to obtain refined local and global features for each sample pair; Based on the refined local and global features of each sample pair, the N sample pairs are classified to obtain a noisy sample set, a fuzzy sample set, and a clean sample set, as well as the predicted labels of the noisy samples, the fuzzy samples, and the clean samples. Difficult negative samples are selected from the N sample pairs, and difficult positive samples are selected from the N sample pairs based on the noisy sample set. During the warm-up training phase, the refined local and global features of each sample pair are processed by the total bidirectional Kullback-Leibler divergence loss function to achieve iterative training until the first preset number of iterations is reached, after which the regular training phase begins. During the regular training phase, loss processing is applied to the predicted labels of the noisy samples, the ambiguous samples, and the clean samples using the total loss function, as well as the difficult positive samples and the difficult negative samples, to achieve iterative training until a second preset number of iterations is reached, resulting in a text-to-image retrieval model.

2. The text-to-image retrieval model training method as described in claim 1, characterized in that, The step of performing feature extraction processing on each sample pair to obtain refined local and global features for each sample pair includes: The image encoder of each sample pair is encoded by the CLIP encoder to obtain the image sequence of each sample pair, wherein the image sequence of each sample pair carries a spatial token to identify the spatial information of the image sequence of each sample pair; The text of each sample pair is encoded using the text encoder of the CLIP encoder to obtain a text sequence for each sample pair, wherein the image sequence of each sample pair carries a text start token and a text end token; By using a self-attention weight selection mechanism, the target image sequence of each sample pair is selected from the image sequence of each sample pair, and the target text sequence of each sample pair is selected from the text sequence of each sample pair. Max pooling is performed on the target image sequence and target text sequence of each sample pair to obtain the refined local features of each sample pair, and the global features of each sample pair are obtained based on the spatial token and text end token of each sample pair.

3. The text-to-image retrieval model training method as described in claim 2, characterized in that, The process of obtaining the refined local and global features of each sample pair also includes: Based on the N sample pairs, N text features and N image features are obtained; For each of the text features, a matching operation is performed between the text feature and each of the N image features to obtain the matching probability between the text feature and each image feature; Each of the text features undergoes the matching operation described above, and the matching probability of each text feature with each of the image features is specified.

4. The text-to-image retrieval model training method as described in claim 3, characterized in that, The process of classifying N sample pairs based on the refined local and global features of each sample pair yields a noisy sample set, a fuzzy sample set, and a clean sample set, including: The fuzzy C-means clustering module performs clustering and grouping processing on the refined local features and global features of each sample pair to obtain the local feature membership probability and the global feature membership probability of each sample pair. The local feature membership probability and the global feature membership probability are both probabilities of belonging to the low-loss cluster. Based on the local feature membership probability of each sample pair, the local feature prediction value of each sample pair is obtained, and based on the global feature membership probability of each sample pair, the global feature prediction value of each sample pair is obtained. Based on the local feature prediction value and global feature prediction value of each sample pair, the target sample category corresponding to each sample pair is obtained, and the sample pair is divided into the target sample category to form a noise sample set, a fuzzy sample set and a clean sample set. The target sample category includes: noise sample category, model sample category and clean sample category.

5. The text-to-image retrieval model training method as described in claim 4, characterized in that, The predicted labels for noisy samples, blurred samples, and clean samples are obtained, including: In each clean sample in the clean sample set, the true label of the clean sample is used as the predicted label of the clean sample; In each noise sample in the noise sample set, the true label of the noise sample is modified to a preset label value, and the preset label value is used as the predicted label of the noise sample; In each fuzzy sample in the fuzzy sample set, the predicted label of the fuzzy sample is obtained based on the refined local features and global features of the fuzzy sample.

6. The text-to-image retrieval model training method as described in claim 4, characterized in that, The step of selecting difficult negative samples from the N sample pairs and selecting difficult positive samples from the N sample pairs based on the noisy sample set includes: The N sample pairs are combined to obtain N 2 The recombined samples were divided into three groups, and the recombined samples whose text features and image features belonged to the same target object were identified as positive samples, while the recombined samples whose text features and image features did not belong to the same target object were identified as negative samples. Among the negative samples, the negative sample with the highest matching probability is identified as the difficult negative sample; In the positive samples, if the positive sample does not belong to the noise sample set, and the current local feature membership probability and the current global feature membership probability of the positive sample meet the set conditions, then the positive sample is determined as the difficult positive sample.

7. The text-to-image retrieval model training method as described in claim 1, characterized in that, The total bidirectional Kullback-Leibler divergence loss function is: in, This represents the total loss value of the two-way Kullback-Leibler divergence. This represents the total loss value for the positive Kullback-Leibler divergence. This represents the total loss value of the reverse Kullback-Leibler divergence. The forward Kullback-Leibler divergence loss value from image to text. The forward Kullback-Leibler divergence loss value from text to image. The inverse Kullback-Leibler divergence loss value from image to text. This represents the inverse Kullback-Leibler divergence loss value from text to image.

8. The text-to-image retrieval model training method as described in claim 1, characterized in that, In the regular training phase, loss processing is performed on the predicted labels of the noisy samples, the ambiguous samples, and the clean samples using the total loss function, as well as the difficult positive samples and the difficult negative samples. This includes: If the target separability metric is greater than the hard sample attention threshold, then loss processing is applied to the predicted labels of the noisy samples, the predicted labels of the ambiguous samples, and the predicted labels of the clean samples using the first total loss function, as well as the hard positive samples and the hard negative samples. If the target separability metric is not greater than the difficult sample attention threshold, then the predicted labels of the noisy samples, the predicted labels of the ambiguous samples, and the predicted labels of the clean samples are processed by the second total loss function, wherein the total loss function includes the first total loss function and the second total loss function.

9. The text-to-image retrieval model training method as described in claim 8, characterized in that, The first total loss function is: in, The loss value of the first total loss function. The label matrix formed by the predicted labels of the noisy samples, the predicted labels of the blurred samples, and the predicted labels of the clean samples, α pos To lose the positive component, β neg To lose the negative component, W pos W is the difficult positive mining matrix obtained from the aforementioned difficult positive samples. neg This is the hard negative mining matrix obtained from the hard negative samples. These represent the positive cross-entropy loss values ​​for text-to-image and image-to-text. These are the inverse cross-entropy loss values ​​for text to image and image to text. The second total loss function is: in, This represents the loss value of the second total loss function.

10. A text-to-image retrieval method, characterized in that, include: The target text and image set are input into a text-to-image retrieval model, wherein the text-to-image retrieval model is a retrieval model obtained by the text-to-image retrieval model training method as described in any one of claims 1 to 9; The text-to-image retrieval model obtains refined local and global features for each matching pair based on the matching pairs formed between the target text and each image in the image set. In each matching pair, a score matrix of local features is obtained based on the refined local features, and a score matrix of global features is obtained based on the global features. Then, a target score matrix is ​​obtained based on the score matrix of local features and the score matrix of global features, thereby obtaining the target score matrix of each matching pair. Based on the target score matrix of each matching pair, the predicted probability of each matching pair is obtained, and based on the predicted probability of each matching pair, the sorting number of each image is obtained.