Image-text retrieval model training method and device and image-text retrieval method and device

By extracting image and text features and calculating the loss of all negative samples, the training method for image and text retrieval models, combined with dynamic re-ranking and course learning, solves the problem of performance degradation in image and text retrieval under low-resource environments and achieves higher robustness and accuracy.

CN120974189APending Publication Date: 2025-11-18XIAN JIAOTONG LIVERPOOL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511101787.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing image retrieval models face problems of false negatives and reduced consistency of positive samples in low-resource environments, leading to performance degradation. Furthermore, existing methods are inefficient in low-resource scenarios and struggle to cope with challenges such as ambiguous descriptions and low-quality images.

Method used

Image and text features are extracted by the visual encoder and text encoder of the image-text retrieval model. The predicted similarity is calculated using the token selection embedding module and the basic global embedding module, and the loss for focusing on all negative samples is calculated. Combined with dynamic re-ranking and course learning strategies, the model training process is optimized and the robustness of the model to low-resource environments is enhanced.

Benefits of technology

It significantly improves the robustness and training performance of the model in low-resource environments, enhances the accuracy and stability of image and text retrieval, and performs particularly well in text-to-image person re-identification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974189A_ABST
    Figure CN120974189A_ABST
Patent Text Reader

Abstract

The invention discloses an image-text retrieval model training method and device and an image-text retrieval method and device. The method comprises the following steps: encoding an image in an image library through a visual encoder of an image-text retrieval model to obtain at least two visual features of the image; encoding the text in the query set through a text encoder of the image-text retrieval model to obtain at least two text features of the text; based on the visual features and the text features, determining prediction similarity of the image text pairs through a token selection embedding module and a basic global embedding module of an image-text retrieval model, and determining sample types of the image text pairs according to the prediction similarity; the sample types comprise positive samples and negative samples; and according to the sample type and the expected similarity, calculating the loss of paying attention to all negative samples, and according to the loss, carrying out back propagation on the token selection embedding module and the basic global embedding module. According to the embodiment of the invention, the training effect of the image-text retrieval model under the condition of low resources can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of cross-modal technology, and particularly relates to a training method and device of a text-image retrieval model and a text-image retrieval method and device. BACKGROUND

[0002] Text-image retrieval is a basic task in cross-modal retrieval and has important practical significance in intelligent monitoring, cross-media search and other applications, such as retrieving the identity of a pedestrian through cross-modal semantic alignment.

[0003] Although some progress has been made recently, text-image retrieval still faces challenges due to low-resource problems such as ambiguous text descriptions and image background interference, which leads to false positive samples being mistaken for true positive samples. In addition, in real-world data, low-resource problems are almost inevitable, which reduces the consistency of true positive samples, misleads the optimization trajectory, and thus leads to a decrease in the performance of the text-image retrieval model. SUMMARY

[0004] The present application provides a training method and device of a text-image retrieval model and a text-image retrieval method and device to improve the training effect of the text-image retrieval model under low-resource conditions.

[0005] According to an aspect of the present application, a training method of a text-image retrieval model is provided, comprising:

[0006] encoding an image in an image library through a visual encoder of the text-image retrieval model to obtain at least two visual features of the image;

[0007] encoding a text in a query set through a text encoder of the text-image retrieval model to obtain at least two text features of the text;

[0008] determining a predicted similarity of an image-text pair based on the visual features and the text features through a token selection embedding module and a basic global embedding module of the text-image retrieval model, and determining a sample type of the image-text pair according to the predicted similarity; the sample type includes a true positive sample and a false positive sample;

[0009] calculating a loss of paying attention to all false positive samples according to the sample type and an expected similarity, and performing back propagation on the token selection embedding module and the basic global embedding module according to the loss.

[0010] According to another aspect of the present application, a text-image retrieval method is provided, comprising:

[0011] determining a predicted similarity of a text to be retrieved and each candidate image through a text-image retrieval model;

[0012] determining a matching image of the text to be retrieved from the candidate images according to the predicted similarity.

[0013] The image-text retrieval model is trained by the training method of the image-text retrieval model according to any of the embodiments of the present application.

[0014] According to another aspect of the present application, there is provided an image-text retrieval model training device, comprising:

[0015] an image encoding module configured to encode images in an image library by a visual encoder of the image-text retrieval model to obtain at least two visual features of the images;

[0016] a text encoding module configured to encode texts in a query set by a text encoder of the image-text retrieval model to obtain at least two text features of the texts;

[0017] a text pair analysis module configured to determine a predicted similarity of an image-text pair based on the visual features and the text features by a token selection embedding module and a basic global embedding module of the image-text retrieval model, and determine a sample type of the image-text pair according to the predicted similarity; the sample type comprises positive samples and negative samples;

[0018] a back propagation module configured to calculate a loss of focusing on all negative samples according to the sample type and an expected similarity, and perform back propagation on the token selection embedding module and the basic global embedding module according to the loss.

[0019] According to another aspect of the present application, there is provided an image-text retrieval device, comprising:

[0020] a determination module configured to determine a predicted similarity of a to-be-retrieved text and each candidate image by an image-text retrieval model;

[0021] a retrieval module configured to determine a matching image of the to-be-retrieved text from the candidate images according to the predicted similarity;

[0022] The image-text retrieval model is trained by the training device of the image-text retrieval model according to any of the embodiments of the present application.

[0023] According to another aspect of the present application, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the training method of the image-text retrieval model or the image-text retrieval method according to any of the embodiments of the present application.

[0024] According to another aspect of the present application, there is provided an electronic device comprising at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the training method of the image-text retrieval model or the image-text retrieval method according to any one of the embodiments of the present application.

[0025] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for enabling a processor to implement the training method of the image-text retrieval model or the image-text retrieval method according to any one of the embodiments of the present application when executed by the processor.

[0026] The embodiments of the present application, when calculating the loss, do not focus on the most difficult negative sample as the traditional triplet alignment loss does, but focus on all negative samples, and calculate the loss accordingly, enhance the resistance of the model to low resources, improve the robustness and training effect of the model in the low resource environment.

[0027] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0029] Figure 1 is a flowchart of a training method of an image-text retrieval model according to an embodiment of the present application;

[0030] Figure 2A is a flowchart of a training method of an image-text retrieval model according to another embodiment of the present application;

[0031] Figure 2B is a schematic diagram of a training logic of an image-text retrieval model according to another embodiment of the present application;

[0032] Figure 3 is a flowchart of an image-text retrieval method according to another embodiment of the present application;

[0033] Figure 4is a structural schematic diagram of a text-image retrieval model training device according to another embodiment of the present application;

[0034] Figure 5 is a structural schematic diagram of a text-image retrieval device according to another embodiment of the present application;

[0035] Figure 6 is a structural schematic diagram of an electronic device implementing an embodiment of the present application. DETAILED DESCRIPTION

[0036] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0037] It should be noted that the terms "first", "second", and the like in the present application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0038] Before introducing the embodiments of the present application, the related concepts in the existing text-image retrieval are briefly described:

[0039] Global alignment: This method can effectively capture the global semantic correspondence between images and text, and achieve cross-modal retrieval by learning overall features. For example, it can associate the text description with the overall content of the image, thereby establishing a semantic connection. However, this method may ignore fine-grained local details, such as specific regions in the image or specific descriptive words in the text, resulting in a decrease in retrieval accuracy in complex scenarios.

[0040] Local alignment (e.g., context-aware feature retrieval): This approach focuses on retrieving local regions of images and texts, such as accurately matching certain keywords in the text description with specific parts of the image (e.g., objects, scene elements). By being context-aware, the model can better handle relationships between local features, improving the accuracy of retrieval. However, this method has high computational complexity because multiple local regions of images and texts need to be analyzed and retrieved, which may face computational resource limitations in low-resource environments.

[0041] Pre-trained model transfer: Utilize the powerful feature extraction capabilities of pre-trained models (e.g., CLIP) to adapt to the image-text retrieval task through transfer learning. These models are pre-trained on large-scale datasets and can extract robust visual and textual features. However, in low-resource scenarios, such as when text descriptions are ambiguous or image quality is low, pre-trained models may struggle to effectively cope, leading to performance degradation.

[0042] To alleviate the low-resource problem, existing methods mainly adopt the following two strategies:

[0043] Sample selection strategy: Utilize the memory effect of deep neural networks (DNNs) to dynamically filter high-quality samples, prioritizing the learning of clean data with high confidence. For example, the model can distinguish reliable image-text pairs through confidence scores, reducing the interference of ambiguous descriptions or low-quality images on training. This strategy can effectively improve the training efficiency of the model in low-resource environments, but may overlook valuable data due to improper sample selection.

[0044] Robust loss function: Modify the loss function to reduce the impact of low-resource environments, such as introducing a weighting mechanism or regularization term to reduce the interference of ambiguous samples or low-quality data on model optimization. This method aims to make the model more robust to low resources, but its effectiveness depends on the design of the loss function and may increase the complexity of training.

[0045] However, the above methods are not specifically designed for image-text retrieval tasks and may be less efficient in low-resource scenarios. For example, the sample selection strategy may fail due to uneven data distribution, while the robust loss function may not fully adapt to the complex cross-modal semantic alignment requirements in image-text retrieval. Therefore, further exploration of optimization methods for low-resource image-text retrieval is needed to improve the performance of the model in challenging scenarios such as ambiguous descriptions and low-quality images.

[0046] Course learning can achieve progressive training from simple samples to complex samples, improve training efficiency and generalization ability. In the field of cross-modal retrieval, the implementation of course learning is divided into two strategies: data level and model level. Data level method realizes progressive learning by dynamically adjusting the difficulty distribution of training data, such as classic framework and prediction confidence based method. The model level method uses the model ability by gradually expanding, such as dropout scheduling and feature smoothing technology. In addition, some experts also use difficult sample mining and other anti-course methods to improve the robustness of the model. The above methods are static in setting the difficulty of samples, and it is difficult to measure the difficulty of a sample in many cases.

[0047] Re-ranking technology significantly improves retrieval accuracy by optimizing initial retrieval results, and has made extensive progress in single modal retrieval tasks. Taking pedestrian re-identification as an example, experts have proposed a special re-ranking framework to optimize cross-scene pedestrian retrieval through joint retrieval and verification. In addition, they have also designed another method based on multi-attention mechanism to improve the accuracy of pedestrian image ranking through deep learning. In recent years, re-ranking technology has received more and more attention in multi-modal retrieval tasks. Some experts proposed a cross-modal retrieval enhancement generation framework based on visual rich elements, which can optimize the image-text retrieval result through multi-document question answering. They also proposed an asymmetric sensitive re-ranking method based on contrastive learning to improve retrieval performance. However, these methods ignore the bidirectional T2I / T2T collaboration.

[0048] In summary, existing image-text retrieval methods face multiple limitations in low-resource scenarios: lack of specialized design for ambiguous descriptions and low-quality images, making it difficult to cope with real-world challenges; Difficulty in sample difficulty assessment, resulting in limited effectiveness of course learning and re-ranking technology; Dependence on high computational resources limits the application of the method on low-resource devices; In addition, the method focuses on one-way optimization, lacks consideration of bidirectional semantic consistency from text to image and text to text, resulting in uneven retrieval performance.

[0049] Figure 1 A flowchart of a training method of an image-text retrieval model is provided for an embodiment of the present application. The embodiment can be applicable to improving training effectiveness by optimizing the training method of the image-text retrieval model without adjusting the low-resource dataset too much. The method can be executed by an image-text retrieval model training device, which can be realized in the form of hardware and / or software, and can be configured in an electronic device with corresponding data processing capability. As shown in the figure, the method comprises: Figure 1

[0050] S110, encode the images in the image library through the visual encoder of the image-text retrieval model to obtain at least two visual features of the images.

[0051] ​S120, encode the text in the query set by a text encoder of the image-text retrieval model to obtain at least two text features of the text.

[0052] S130, determine a predicted similarity of the image-text pair by a token selection embedding module and a basic global embedding module of the image-text retrieval model based on the visual features and the text features, and determine a sample type of the image-text pair according to the predicted similarity.

[0053] S140, calculate a loss of paying attention to all negative samples according to the sample type and the expected similarity, and perform back propagation on the token selection embedding module and the basic global embedding module according to the loss.

[0054] wherein the data set used for training includes an image library wherein (I i ) is the (i)th (pedestrian) image, is a person identity label, is an image identity label, (N v ) is the total number of images; a query set wherein (T i ) is the (i)th query text, is a shared image identity label, (N t ) is the total number of texts; and an image-text pair set wherein (N) is the number of image-text pairs, and each pair shares the same image identity label and a category label The present application introduces a binary correspondence label (l ij ∈{0,1}) to represent the retrieval degree of the image-text pair (I i ,T j ), when (l ij =1) is a retrieval pair (positive sample), otherwise is a non-retrieval pair (negative sample). The text features include an end text feature based on a sequence end token and at least two local text features based on word phrases. The visual features include a global visual feature based on a classification task token and at least two local visual features capturing image details. The sample type includes positive samples and negative samples.

[0055] Specifically, a pre-trained CLIP visual encoder (f v ) is used to process the images (I i ) in the image library to generate (N0+1) discrete tokens, corresponding to N0 local visual features and 1 global feature respectively:

[0056]

[0057] where (d) is the shared embedding dimension, is the global visual feature based on [classification task] tokens, capturing the overall appearance (e.g. figure outline), is the local visual feature based on image divided into (N0) fixed-size, non-overlapping patches, capturing details (e.g. clothing color or body parts).

[0058] The query set text (T t ) is processed by the visual encoder CLIP text encoder (f i ), generating a token sequence containing (N0+2) discrete tokens:

[0059]

[0060] where and are the [start of sequence] and [end of sequence] tokens, respectively, is the local text feature based on words or phrases, capturing specific description content.

[0061] The visual encoder and text encoder are pre-trained network structures and are not trained in this training. Their network structure and parameters will not change in this training process.

[0062] An image-text pair is selected, and the visual features of the image in the image-text pair and the text features of the text in the image-text pair are input into the token selection embedding module to obtain the local similarity calculated and output by the token selection embedding module; the visual features of the image in the image-text pair and the text features of the text in the image-text pair are input into the basic global embedding module to obtain the global similarity calculated and output by the token selection embedding module; the local similarity and the global similarity are weighted and fused to obtain the predicted similarity of the image-text pair.

[0063] In order to deal with low-resource samples (i.e. image and text retrieval inaccurate sample pairs) in the data set, the predicted similarity is input into the Gaussian model, the loss distribution of the image-text pair is analyzed through the Gaussian mixture model, and the sample pair is divided into one of the three categories through the expectation maximization algorithm: clean samples noise samples and uncertain samples The pair loss is defined as:

[0064]

[0065] where, measures the difference in the shared embedding space, is the cross-modal model. In addition, the types corresponding to the three types of sample pairs are recalibrated:

[0066]

[0067] where Rand(·) randomly assigns 0 or 1 label to uncertain samples. 1 is a positive sample, and 0 is a negative sample. After re-calibration, the new types of each image-text pair can be obtained.

[0068] The traditional triplet alignment loss only focuses on the hardest negative sample, which is defined as:

[0069] L triplet =max(0,m+S n -S p )

[0070] where (S p ) is the similarity of the positive sample pair, (S n ) is the similarity of the hardest negative sample pair, and (m) is the positive margin. However, this method may ignore the information of other negative samples, limiting the robustness of the model in low-resource environments. Therefore, the invention not only focuses on the hardest negative sample, but also focuses on all negative samples to enhance the model's resistance to low resources. For this purpose, the sample type and expected similarity are used to calculate the loss of focusing on all negative samples. According to the calculated loss, the token selection embedding module and the basic global embedding module are backpropagated to optimize their network structure and parameters.

[0071] After implementation, the graph-text retrieval model trained by the training method of the invention has achieved significant performance improvement on three mainstream datasets (CUHK-PEDES, ICFG-PEDES, and RSTPReid) of text-to-image pedestrian re-identification.

[0072] The CUHK-PEDES dataset contains 40,206 images and corresponding natural language descriptions, covering 13,003 pedestrian identities. The dataset is divided into a training set (about 80%), a validation set (about 10%), and a test set (about 10%) to evaluate the performance of the model under different low-resource ratios. The ICFG-PEDES dataset contains 54,522 images and corresponding text descriptions, involving 4,102 pedestrian identities, with no validation set. The RSTPReid dataset is a smaller cross-modal dataset, containing 20,505 images and corresponding text descriptions, involving 4,101 pedestrian identities.

[0073] The experimental results show that the method achieves the highest average precision (mAP) and average inverse negative penalty (mINP) on all data sets, regardless of the low resource ratio, and performs well. Without low resource conditions, the method improves the mAP of the current optimal model on CUHK-PEDES, ICFG-PEDES and RSTPReid by 6.77%, 34.24% and 12.61% respectively, proving the effectiveness of the adaptive training mechanism. Under the condition of 20% low resource ratio, the mAP is improved by 9.15%, 25.42% and 12.60% respectively. Under the condition of 50% low resource ratio, through dynamic sample weight distribution, the method realizes stable learning and is superior to the current optimal model. Even under the condition of 80% low resource ratio, the method still maintains superior performance and effectively reduces the overfitting problem caused by low resources. In addition, the ablation experiment verifies the key role of the hierarchical cross-modal embedding, curriculum-guided optimization and dynamic reordering module in improving the low resource robustness and retrieval accuracy. In addition, an ablation experiment is performed to verify the effectiveness of each module (hierarchical cross-modal embedding, curriculum-guided optimization and dynamic reordering) in the method, which proves the key role of these modules in improving the robustness and retrieval accuracy in the low resource environment.

[0074] The embodiment of the application pays attention to all negative samples when calculating the loss, unlike the traditional triple alignment loss which only focuses on the most difficult negative sample, and calculates the loss accordingly, thereby enhancing the model's resistance to low resources and improving the model's robustness and training effect in a low resource environment.

[0075] On the basis of the above-mentioned embodiment, the method further comprises:

[0076] determining, by the image-text retrieval model, relevant images in the image library that are in front of the expected similarity ranking of the target text;

[0077] ranking, according to the similarity of the target text and other texts in the query set, the other texts to obtain relevant texts in front of the ranking;

[0078] For each relevant image, calculating the similarity between the relevant image and the relevant text, and calculating the reordering score of the relevant image according to the similarity;

[0079] reordering the relevant images according to the reordering scores of the relevant images.

[0080] Specifically, the predicted feature similarity obtained by the foregoing and the trained image-text retrieval model are used for test inference. In this stage, the initial retrieval result is optimized by dynamic reordering to reduce the influence of the semantic slight difference caused by low resources. The optimization is realized by the following steps:

[0081] Initial retrieval list generation: For query text (T j ), compute the similarity S(T j , I i ) of corresponding image-text pairs using the image-text retrieval model, and determine the initial ranking list including multiple relevant images with top ranking similarity to the target text:

[0082]

[0083] where, is the top (K) relevant image set of text (T j ), and (K) is the number of relevant images reserved for initial screening.

[0084] Dynamic threshold: To adapt to the characteristics of different texts, compute the standard deviation of the similarity between the target text (T j ) and all other texts in the query set, which reflects the difference between the texts in the query set:

[0085]

[0086] The dynamic threshold (τ j ) is calculated by the standard deviation (σ j ) and the threshold scaling factor (α):

[0087] τ j = α· σ j

[0088] Relevant text filtering: Because an image may be described by multiple texts in the query set (such as different sentences describing the same person), the retrieval results of the query text (T j ) are inherently related to the retrieval results of other semantically relevant texts. Compute the text-text similarity matrix (S(T j , T t )) using the pre-trained RoBERTa model, and sort all other texts by their similarity to the target text to generate a relevant text set containing multiple relevant texts:

[0089]

[0090] where (K2) is the number of selected relevant texts.

[0091] Candidate re-ranking score calculation: For each relevant image compute its re-ranking score:

[0092] For the first relevant image (I0), obtain its top (N) relevant text set through reverse retrieval:

[0093]

[0094] Then compare the relevant texts with the text set of (Q0), find their intersection (i.e. common texts):

[0095]

[0096] where, denotes the position index of text (T r ) in set (Q0). If there is an intersection The re-ranking score of the relevant image is calculated using an exponential decay score function based on a dynamic threshold (τ j ):

[0097]

[0098] This function applies an exponential penalty for larger position differences , ensuring that texts that align closely with the query text contribute more to the score. If the intersection is empty , a compensation score is calculated by standard deviation weighting as the re-ranking score:

[0099] s0 = N · (1 + σ j ).

[0100] For the remaining relevant images , their top (N) relevant text sets are obtained through inverse indexing:

[0101]

[0102] Calculate the intersection positions:

[0103]

[0104] If the intersection is non-empty , the re-ranking score takes the minimum position value Otherwise, assign the maximum penalty value (N) to reduce the ranking of this image.

[0105] Final retrieval list generation: use a stable sorting algorithm to re-rank the relevant images based on the re-ranking scores, and preserve the initial order of relevant images with equal re-ranking scores during the sorting process to generate the final ranking:

[0106]

[0107] where (s) is the re-ranking score vector containing all relevant images.

[0108] Through dynamic weight adjustment and curriculum learning strategy, optimize the training process, preferentially learn simple samples, gradually adapt to difficult samples, and improve model stability.

[0109] Figure 2A A flowchart of a picture-text retrieval model method is provided for another embodiment of the application, which is optimized and improved on the basis of the above-mentioned embodiment. As shown in the figure, the method comprises: Figure 2A

[0110] S210, encoding images in the image library through a visual encoder of the picture-text retrieval model to obtain at least two visual features of the images; and encoding texts in the query set through a text encoder of the picture-text retrieval model to obtain at least two text features of the texts.

[0111] S220, inputting the end text feature and the global visual feature into a basic global embedding module to obtain a global similarity output by the basic global embedding module.

[0112] S230, inputting the local text feature and the local visual feature into a token selection embedding module to obtain a local similarity output by the token selection embedding module.

[0113] The token selection embedding module is configured to select the local visual feature most relevant to the global visual feature to aggregate to obtain an enhanced local visual feature, select the local text feature most relevant to the end text feature to aggregate to obtain an enhanced local text feature, and calculate the local similarity of the enhanced local visual feature and the enhanced local text feature.

[0114] Specifically, the basic global embedding module calculates the cosine similarity between the global visual feature (based on the [classification task] token) and the end text feature (based on the [end of sequence] token) as the global similarity to capture the coarse-grained cross-modal alignment of images and texts. That is, the global feature is used to achieve fast semantic alignment, which is suitable for capturing the retrieval of overall appearance and text description. The specific calculation formula is:

[0115]

[0116] In order to capture more fine-grained semantic relationships, the token selection embedding module selects the local features with the most abundant information based on the relevance of the local features and the global tokens. The specific steps include: from the patch-level local features of images and texts, selecting the local visual features highly relevant to the global visual feature of the [classification task] token, and selecting the local text features highly relevant to the end text feature of the [classification task] token. The two selected features are respectively aggregated through maxpooling, multi-layer perceptron (MLP) and full connection layer in the module to generate enhanced local visual features ​and enhance local text features The similarity of both is finally calculated as a local similarity The specific calculation formula is:

[0117]

[0118] The token selection embedding module enhances the model's ability to capture local details (such as clothing texture or body part description) through fine-grained feature enhancement, improving robustness in low-resource environments.

[0119] S240, determining a predicted similarity of the image-text pair according to the local similarity and the global similarity; determining a sample type of the image-text pair according to the predicted similarity.

[0120] On the basis of the above-mentioned embodiments, optionally, the loss calculation formula is as follows:

[0121]

[0122] wherein, represents the size of the loss, T i represents the i-th text in the query set, I i represents the i-th image in the image set, T j represents the j-th text in the query set, I j represents the j-th image in the image set, m is a positive margin parameter, τ is a temperature parameter, S(T j ,I i ) represents the predicted similarity of the image-text pair composed of the j-th text and the i-th image, q ij = 1-l ij , l ij represents the sample type of the corresponding image-text pair, the positive sample takes the value of 1, and the negative sample takes the value of 0; S(T i ,I j ) represents the predicted similarity of the image-text pair composed of the i-th text and the j-th image, q ji = 1-l ji , l ji represents the sample type of the corresponding image-text pair, the positive sample takes the value of 1, and the negative sample takes the value of 0; is the weighted average predicted similarity of the i-th text and all positive images, is the weighted average predicted similarity of the i-th image and all positive texts.

[0123] S250, determining a difficulty weight of the current course learning dynamics according to the size of the loss; performing back propagation on the token selection embedding module and the basic global embedding module according to the difficulty weight and the loss.

[0124] Specifically, to solve the problem that it is difficult to distinguish between simple and difficult samples after confidence consensus division (i.e. positive and negative sample classification), the application proposes a dynamic adaptive alignment technology based on samples. Unlike traditional static curriculum learning, this method dynamically adjusts the difficulty weight of the sample pair according to the loss size of the sample, and during the training process, simple samples (low loss) are given priority to learn, and complex samples (high loss) are gradually introduced. The allocation formula of the difficulty weight is:

[0125]

[0126] wherein (λ) is a temperature parameter for controlling the steepness of the weight distribution, is the loss of the image-text sample pair calculated in the foregoing. Gradient back propagation is represented as:

[0127]

[0128] Through dynamic weight adjustment, the model first masters simple samples and gradually processes difficult samples, significantly improving the learning efficiency and robustness in low-resource environments. Using the confidence consensus division and the curriculum-guided triplet alignment loss, this module can train a robust model that can handle low-resource data in the image-text retrieval task.

[0129] For example, Figure 2BAs shown, the hierarchical cross-modal embedding module enhances the discriminability of feature representation through global and local feature fusion; the confidence consensus division and the course-guided triplet alignment loss improve the training stability through dynamic weight adjustment; the dynamic reordering module optimizes the reasoning result through dynamic threshold and semantic relevant text filtering, reducing the interference of low resources. Experimental results show that the present application is superior to existing methods on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets. For the problem that after the confidence consensus division module is processed, the clean data, noise data and uncertain data are jointly input into the model for training, through the course-guided triplet alignment loss module combined with the dynamic weight adjustment strategy, the samples with small loss are preferentially used for training, and the interference of noise samples is inhibited. This method significantly improves the training stability of the model in a low resource environment, ensures effective differentiation between simple and difficult samples, and improves learning efficiency and generalization ability. 3. For the problem that traditional methods only consider the text-image pair matching relationship, ignoring the high-order neighbor relationship between text-text and text-image, resulting in poor retrieval performance, the dynamic context-aware reordering module of the present application optimizes the cross-modal context information by using text-text similarity (based on RoBERTa model) and dynamic threshold, combining relevant text filtering and exponential decay score function, fully excavates the high-order neighbor relationship, significantly reduces the influence of abnormal values introduced by non-uniform text similarity, narrows the training-test gap, and improves retrieval accuracy. For low resources, the present application significantly improves the adaptability of the model to ambiguous descriptions and complex backgrounds through hierarchical embedding, course learning and dynamic reordering, and enhances the generalization performance.

[0130] The embodiment of the present application optimizes the training process through dynamic weight adjustment and course learning strategy, preferentially learns simple samples, gradually adapts to difficult samples, and improves the stability of the model.

[0131] Figure 3 A flowchart of a picture-text retrieval method provided by an embodiment of the present application, the embodiment can be applicable to picture-text retrieval by a picture-text retrieval model, the method can be executed by a picture-text retrieval device, the device can be realized in the form of hardware and / or software, and the device can be configured in an electronic device with corresponding data processing capability. As shown, the method comprises: Figure 3

[0132] S310, determining the predicted similarity of the to-be-retrieved text and each candidate image through the picture-text retrieval model.

[0133] S320, determining the matching image of the to-be-retrieved text from the candidate images according to the predicted similarity.

[0134] The picture-text retrieval model is trained by the training method of the picture-text retrieval model of any embodiment of the present application. ​

[0135] Specifically, in the retrieval stage, for each candidate image in the database, the candidate image and the user-provided text to be retrieved are input into the image-text retrieval model, and the predicted similarity of the candidate image and the text to be retrieved is determined through the text encoder, the visual encoder, the token selection embedding module and the basic global embedding module of the image-text retrieval model. The process of determining the image-text similarity through the image-text retrieval model has been described in detail in the foregoing, and will not be described here. After obtaining the predicted similarity of each candidate image and the text to be retrieved, the candidate images are sorted according to the order from high to low of the predicted similarity, and one or more candidate images ranked in the front are determined as the matching images of the text to be retrieved.

[0136] The embodiment of the present application uses the image-text retrieval model trained in a targeted manner to perform image-text retrieval, and can improve the accuracy of image-text retrieval.

[0137] On the basis of the above-mentioned embodiment, optionally, the method further comprises:

[0138] The predicted similarity of the image to be retrieved and each candidate text is determined through the image-text retrieval model.

[0139] The matching text of the image to be retrieved is determined from the candidate texts according to the predicted similarity.

[0140] Specifically, the image retrieval can be used not only to determine the matching images of the text to be retrieved, but also to determine the matching text of the image to be retrieved. The specific matching process is similar to the matching process described in the foregoing, and will not be described here.

[0141] Figure 4 A structural schematic diagram of a training device of an image-text retrieval model according to another embodiment of the present application is shown in FIG. 4. As shown in the figure, the device comprises: Figure 4

[0142] The image encoding module 410 is configured to encode the images in the image library through the visual encoder of the image-text retrieval model to obtain at least two visual features of the images.

[0143] The text encoding module 420 is configured to encode the texts in the query set through the text encoder of the image-text retrieval model to obtain at least two text features of the texts.

[0144] The text pair analysis module 430 is configured to determine the predicted similarity of the image-text pair through the token selection embedding module and the basic global embedding module of the image-text retrieval model based on the visual features and the text features, and determine the sample type of the image-text pair according to the predicted similarity; the sample type comprises positive samples and negative samples.

[0145] ​The back propagation module 440 is configured to calculate a loss of paying attention to all negative samples according to the sample type and the expected similarity, and perform back propagation on the token selection embedding module and the basic global embedding module according to the loss.

[0146] The training device of the image-text retrieval model provided in the embodiments of the present application can perform the training method of the image-text retrieval model provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0147] Optionally, the loss is calculated according to the following formula:

[0148]

[0149] wherein, denotes the size of the loss, T i denotes the i th text in the query set, I i denotes the i th image in the image set, T j denotes the j th text in the query set, I j denotes the j th image in the image set, m is a positive margin parameter, and τ is a temperature parameter, S(T j ,I i ) denotes the predicted similarity of the image-text pair formed by the j th text and the i th image, q ij = 1-l ij , l ij denotes the sample type of the corresponding image-text pair, and the positive sample takes the value of 1 and the negative sample takes the value of 0; S(T i ,I j ) denotes the predicted similarity of the image-text pair formed by the i th text and the j th image, q ji = 1-l ji , l ji denotes the sample type of the corresponding image-text pair, and the positive sample takes the value of 1 and the negative sample takes the value of 0. is the weighted average predicted similarity of the i th text and all positive images, is the weighted average predicted similarity of the i th image and all positive texts.

[0150] Optionally, the back propagation module 440 comprises:

[0151] The weight calculation unit is configured to determine the difficulty weight of the current course learning according to the size of the loss;

[0152] The back propagation unit is configured to perform back propagation on the token selection embedding module and the basic global embedding module according to the difficulty weight and the loss.

[0153] Optionally, the text features include an end text feature based on a sequence end token and at least two local text features based on word phrases; the visual features include a global visual feature based on a classification task token and at least two local visual features capturing image details; and the text pair analysis module 430 includes:

[0154] a global similarity determination unit configured to input the end text feature and the global visual feature into a basic global embedding module to obtain a global similarity output by the basic global embedding module;

[0155] a local similarity determination unit configured to input the local text features and the local visual features into a token selection embedding module to obtain a local similarity output by the token selection embedding module; the token selection embedding module is configured to select a local visual feature most relevant to the global visual feature to aggregate to obtain an enhanced local visual feature, select a local text feature most relevant to the end text feature to aggregate to obtain an enhanced local text feature, and calculate a local similarity between the enhanced local visual feature and the enhanced local text feature;

[0156] a predicted similarity determination unit configured to determine a predicted similarity of the image-text pair according to the local similarity and the global similarity.

[0157] Optionally, the apparatus further includes:

[0158] an image ranking module configured to determine, by the image-text retrieval model, relevant images in an image library that are ranked in advance according to an expected similarity to a target text;

[0159] a text ranking module configured to rank, according to a similarity between the target text and other texts in a query set, the other texts to obtain relevant texts ranked in advance;

[0160] a score calculation module configured to, for each relevant image, calculate a similarity between the relevant image and the relevant texts, and calculate a re-ranking score of the relevant image according to the similarity;

[0161] an image re-ranking module configured to re-rank the relevant images according to the re-ranking scores of the relevant images.

[0162] The training apparatus of the image-text retrieval model further described can also perform the training method of the image-text retrieval model provided by any embodiment of the present application, and has the function modules and beneficial effects corresponding to the execution method.

[0163] Figure 5 A structural schematic diagram of an image-text retrieval apparatus provided by another embodiment of the present application is shown in FIG. 4B. Figure 5 As shown in the figure, the apparatus includes:

[0164] The determining module 510 is configured to determine, by using the image-text retrieval model, a predicted similarity between the to-be-retrieved text and each candidate image.

[0165] The retrieving module 520 is configured to determine, according to the predicted similarity, a matching image of the to-be-retrieved text from the candidate images.

[0166] The image-text retrieval model is trained by using the training device for the image-text retrieval model according to any one of the embodiments of the present application.

[0167] The image-text retrieval device provided in the embodiments of the present application can perform the image-text retrieval method provided in any one of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the performing method.

[0168] Figure 6 A structural schematic diagram of an electronic device 60 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0169] As shown in Figure 6 The electronic device 60 includes at least one processor 61 and a memory, such as a read-only memory (ROM) 62, a random access memory (RAM) 63, etc., which are in communication with the at least one processor 61. The memory stores computer programs that can be executed by the at least one processor 61, and the processor 61 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 62 or loaded from the storage unit 68 into the random access memory (RAM) 63. In the RAM 63, various programs and data required for the operation of the electronic device 60 can also be stored. The processor 61, the ROM 62, and the RAM 63 are connected to each other through a bus 64. An input / output (I / O) interface 65 is also connected to the bus 64.

[0170] A plurality of components in the electronic device 60 are connected to the I / O interface 65, including: an input unit 66, such as a keyboard, a mouse, etc.; an output unit 67, such as various types of displays, a speaker, etc.; a storage unit 68, such as a magnetic disk, an optical disk, etc.; and a communication unit 69, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 69 allows the electronic device 60 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0171] The processor 61 can be various general and / or special purpose processing components having processing and computing capabilities. Some examples of the processor 61 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 61 performs various methods and processes described above, such as the training method of the image-text retrieval model or the image-text retrieval method.

[0172] In some embodiments, the training method of the image-text retrieval model or the image-text retrieval method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 68. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 60 via the ROM 62 and / or the communication unit 69. When the computer program is loaded onto the RAM 63 and executed by the processor 61, one or more steps of the training method of the image-text retrieval model or the image-text retrieval method described above can be performed. Alternatively, in other embodiments, the processor 61 can be configured to perform the training method of the image-text retrieval model or the image-text retrieval method by any other appropriate means, such as by means of firmware.

[0173] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0174] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, can cause instructions defined in the flow charts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package and partially on a remote machine or entirely on a remote machine or server.

[0175] In the context of the present application, a computer readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer readable storage medium can be a machine readable signal medium. More specific examples of a machine readable storage medium will include one or more lines of a program of instructions in a transitory signal form, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0176] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0177] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0178] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0179] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the technical solutions of the present disclosure are achieved, and the present disclosure is not limited herein.

[0180] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A training method for an image-text retrieval model, characterized in that, The method comprises: encoding images in the image library through a visual encoder of the image-text retrieval model to obtain at least two visual features of the images; encoding texts in the query set through a text encoder of the image-text retrieval model to obtain at least two text features of the texts; determining a predicted similarity of the image-text pair through a token selection embedding module and a basic global embedding module of the image-text retrieval model based on the visual features and the text features, and determining a sample type of the image-text pair according to the predicted similarity; the sample type comprises positive samples and negative samples; calculating a loss of attention to all negative samples according to the sample type and an expected similarity, and performing back propagation on the token selection embedding module and the basic global embedding module according to the loss.

2. The method of claim 1, wherein, The calculation formula of the loss is as follows: wherein, denotes the size of the loss, T i denotes the i-th text in the query set, I i denotes the i-th image in the image set, T j denotes the j-th text in the query set, I j denotes the j-th image in the image set, m is a margin parameter, and τ is a temperature parameter, S(T j ,I i ) denotes the predicted similarity of the image-text pair formed by the j-th text and the i-th image, q ij = 1 - l ij , l ij denotes the sample type corresponding to the image-text pair, and takes the value 1 for a positive sample and 0 for a negative sample; S(T i ,I j ) denotes the predicted similarity of the image-text pair formed by the i-th text and the j-th image, q ji = 1 - l ji , l ji denotes the sample type corresponding to the image-text pair, and takes the value 1 for a positive sample and 0 for a negative sample; is the weighted average predicted similarity of the i-th text with all positive images, is the weighted average predicted similarity of the i-th image with all positive texts.

3. The method of claim 2, wherein, The back propagation on the token selection embedding module and the basic global embedding module according to the loss comprises: determining a difficulty weight of the current course learning dynamics according to the size of the loss; performing back propagation on the token selection embedding module and the basic global embedding module according to the difficulty weight and the loss.

4. The method of claim 1, characterized in that, The text features comprise an end text feature based on a sequence end token and at least two local text features based on word phrases; the visual features comprise a global visual feature based on a classification task token and at least two local visual features capturing image details; the determination of the predicted similarity of the image-text pair through the token selection embedding module and the basic global embedding module of the image-text retrieval model comprises: inputting the end text feature and the global visual feature into the basic global embedding module to obtain a global similarity output by the basic global embedding module; inputting the local text features and the local visual features into the token selection embedding module to obtain a local similarity output by the token selection embedding module; the token selection embedding module is used for selecting the local visual features most relevant to the global visual feature to aggregate to obtain enhanced local visual features, and selecting the local text features most relevant to the end text feature to aggregate to obtain enhanced local text features; and calculating the local similarity of the enhanced local visual features and the enhanced local text features; determining the predicted similarity of the image-text pair according to the local similarity and the global similarity.

5. The method of claim 1, wherein, The method further comprises: determining, through the image-text retrieval model, relevant images in the image library with an expected similarity ranking in front of a target text; ranking other texts in the query set according to the similarity of the target text and the other texts to obtain relevant texts with a ranking in front; for each relevant image, calculating the similarity of the relevant image and the relevant text, and calculating a reordering score of the relevant image according to the similarity; reordering the relevant images according to the reordering scores of the relevant images.

6. A method of retrieving images, characterized by, The method comprises: determining, through the image-text retrieval model, a predicted similarity of a to-be-retrieved text and each candidate image; determining a matching image of the to-be-retrieved text from the candidate images according to the predicted similarity; wherein the image-text retrieval model is trained through the training method of the image-text retrieval model in any one of claims 1-5.

7. A training device for an image-text retrieval model, characterized in that, The device comprises: An image encoding module is configured to encode images in an image library by a visual encoder of the image-text retrieval model to obtain at least two visual features of the images. A text encoding module is configured to encode texts in a query set by a text encoder of the image-text retrieval model to obtain at least two text features of the texts. A text pair analysis module is configured to determine a predicted similarity of an image-text pair based on the visual features and the text features by a token selection embedding module and a basic global embedding module of the image-text retrieval model, and determine a sample type of the image-text pair according to the predicted similarity; the sample type includes positive samples and negative samples. A back propagation module is configured to calculate a loss of focusing on all negative samples according to the sample type and an expected similarity, and perform back propagation on the token selection embedding module and the basic global embedding module according to the loss.

8. A document retrieval apparatus characterized by comprising: The device comprises: A determination module is configured to determine a predicted similarity of a to-be-retrieved text and each candidate image by an image-text retrieval model. A retrieval module is configured to determine a matching image of the to-be-retrieved text from the candidate images according to the predicted similarity. The image-text retrieval model is trained by the training device of the image-text retrieval model of claim 7.

9. An electronic device, comprising: The electronic device comprises: At least one processor; and A memory connected in communication with the at least one processor; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the training method of the image-text retrieval model of any one of claims 1-5 and the image-text retrieval method of claim 6.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the training method of the image-text retrieval model of any one of claims 1-5 and the image-text retrieval method of claim 6 when executed.

Citation Information

Cited By

  • Cross-modal image-text retrieval method and device based on decoupling learning and medium

    CN121524371A