Image-text retrieval method and system based on fine-grained alignment and reordering
By employing a fine-grained alignment and reordering method for image-text retrieval, utilizing a pre-trained model and a multi-layer Transformer module for feature fusion, and combining various loss functions for optimization, this approach addresses the semantic gap and insufficient information utilization issues in image-text retrieval, achieving more efficient image-text matching and improved retrieval performance.
Patent Information
- Application Number
- CN202510961077.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-12
- Publication Date
- 2025-11-14
AI Technical Summary
Existing image-text retrieval methods suffer from a semantic gap when processing multimodal data, making it difficult to achieve accurate matching between images and text. Furthermore, the re-ranking mechanism fails to fully utilize the structural information in the similarity matrix, resulting in mismatches and insufficient expressive power.
We employ a fine-grained alignment and reordering approach, using the pre-trained CLIP model for feature extraction. We combine a cross-modal interaction module and a reordering mechanism, and utilize a multi-head cross-attention layer and a multi-layer Transformer module for image-text feature fusion. We also introduce three loss functions to optimize model training, including InfoNCE loss, KL divergence, and supervision signals from a single-modal teacher model, thereby improving semantic alignment capabilities.
It achieves fine-grained and precise alignment between images and text, enhancing the accuracy and robustness of retrieval, improving the discriminative ability of image-text matching, and the generalization performance of the model.
Smart Images

Figure CN120950744A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information retrieval technology, and in particular to a text and image retrieval method and system based on fine-grained alignment and reordering. Background Technology
[0002] In today's internet age, multimedia data such as text, images, and audio are appearing in all aspects of life at a remarkable rate. Initially, information was conveyed through a single form of data; now, it is expressed in multiple forms—what we call multimodal data. This increase in data formats has led to ever-higher performance requirements for retrieval systems, and how to quickly and accurately locate and retrieve relevant information is a pressing problem. Images and text, as primary information retrieval methods, have become hot research topics when used as retrieval objects. Existing image and text retrieval methods face key challenges: firstly, the "semantic gap" exists between different modalities, preventing models from associating one modality with another; secondly, when processing a single modality, they focus only on prominent information while ignoring other crucial details.
[0003] Early image-text retrieval methods typically mapped both image and text features into a shared semantic space through linear projection. This shallow approach relied on handcrafted features and had limited representational capabilities. With the development of deep learning, researchers began using deep neural networks to automatically learn high-level semantic features from raw images and text in an end-to-end manner, significantly improving inter-modal matching capabilities. During this stage, dual-encoder structures were widely used in large-scale retrieval tasks due to their high computational efficiency. However, because images and text were encoded independently, the lack of inter-modal interaction mechanisms made it difficult for the models to capture complex semantic correspondences. In recent years, pre-trained models have become the mainstream framework for image-text retrieval. By performing contrastive learning or multi-task learning on large-scale image-text pairs, general multi-modal representation capabilities are obtained; subsequently, fine-tuning is performed to improve retrieval performance in specific application scenarios. To further improve accuracy, researchers introduced re-ranking mechanisms as a supplementary module. Some methods use bidirectional ranking strategies or introduce the "supervised learning re-ranking" paradigm into visual retrieval. However, these methods often fail to fully exploit the structural information in the similarity matrix, resulting in insufficient information utilization.
[0004] Furthermore, coarse-grained image-text matching methods typically assess image content from a global perspective, making it difficult to accurately align specific objects in an image with fine-grained descriptive words in the text. This can easily lead to mismatches and insufficient expressive power. Additionally, training with contrastive loss alone focuses only on the relative distance between sample pairs, neglecting the distribution of semantic structure within a single modality. Summary of the Invention
[0005] (a) Technical problems to be solved
[0006] To address the shortcomings of existing technologies, this invention combines three loss models to effectively align semantic relationships between text and images while ensuring clear and consistent internal structures within modalities.
[0007] (II) Technical Solution
[0008] To achieve the above objectives, the present invention provides the following technical solution: a text and image retrieval method based on fine-grained alignment and reordering, comprising the following steps;
[0009] S1: Input the image and text to be retrieved, and encode the input image and text separately using the powerful feature extraction capabilities of the pre-trained model CLIP; S2: Adaptively align the text representation to relevant image regions using the cross-modal interaction module, and integrate the aligned features into a globally semantically consistent representation; S3: Calculate the similarity score between the image and text to obtain preliminary matching results; S4: Perform reverse retrieval on the initial similarity matrix through a re-ranking mechanism to enhance the matching relevance between the image and text in bidirectional retrieval; S5: Train the retrieval model by combining three loss functions, and introduce the knowledge extracted offline by the single-modal pre-trained teacher model as a soft label supervision signal to optimize the similarity learning process; S6: Use KL divergence to measure the difference between the probability distribution output by the model and the soft labels provided by the teacher model to improve the semantic alignment ability between the image and text.
[0010] As a preferred embodiment, feature extraction in step S1 includes image representation:
[0011] For each input image The image embedding is obtained by using a variant of ResNet-50 (clip) as a visual encoder; basic features are extracted stepwise using three shallow convolutional layers, and semantic features are extracted stepwise using multiple deep residual modules; finally, a global aggregation is performed using an attention mechanism to obtain a high-semantic image vector representing the entire image.
[0012]
[0013]
[0014]
[0015] Then perform a two-dimensional average pooling:
[0016]
[0017] The obtained feature map Four residual modules are input sequentially, and then a multi-head attention pooling module is used to perform weighted aggregation of spatial features: This yields a global representation of the entire image.
[0018] As a preferred embodiment, the feature extraction in step S1 further includes text representation:
[0019] For the input text, the text encoder in Clip extracts the text representation. First, it uses lowercase byte pair encoding (BPE) with a vocabulary size of 49152 bytes to segment the input text description. Then, [SOS] and [EOS] markers are added before and after the text description to identify the beginning and end of the sequence. Finally, the segmented text... The input is a Transformer, and the correlation between each unit is mined through a masked self-attention mechanism; the output of the last layer of the Transformer at the [EOS] unit is linearly mapped to the image-text joint embedding space to obtain a global text representation.
[0020] Text input is performed in batch processing. Since the sentence lengths in each batch are different, a maximum text length is set for each batch. For less than Sentences within a unit are truncated to ensure consistent dimensionality; finally, both image and text embeddings are projected onto a shared embedding space through a linear layer, and the linear mapping is calculated as follows:
[0021]
[0022] As a preferred embodiment, the cross-modal interaction in step S2 includes:
[0023] S21: First, textual and visual features are fused through a cross-attention layer. Then, the fused features are input into a multi-layer Transformer module to obtain a fused multimodal table containing the context and semantic correspondence between images and text.
[0024] S22: Subsequently, a text-image alignment matrix is calculated based on the similarity between modalities. This matrix captures the fine-grained semantic correspondence between modalities.
[0025] As a preferred embodiment, the cross-modal interaction in step S2 specifically includes:
[0026] Given input text description and input image The final hidden states of the text and vision are respectively fed into their respective encoders. The input is shared into the cross-modal interaction module; first, the text representation... As query Q, the image representation As the key K and value V, the following process is bidirectional; the same steps are performed when the image representation is Q:
[0027]
[0028]
[0029]
[0030] in, It is multi-head cross-attention, defined as follows:
[0031]
[0032] in, It is the embedded dimension, with a scaling factor of 1. After obtaining the text representation that fuses the image semantics, it is input into the Transformer to capture more complex and long-distance correspondences, enhancing the global semantic understanding capability. The complete interaction of image and text representations can be represented as follows:
[0033]
[0034] in, After normalization of the representation layer, the similarity between the text vector containing global image information and the image feature vector containing global text information is calculated. This similarity is regarded as the image-text alignment score, as shown in the following formula:
[0035]
[0036] The same principle, This represents an image representation that includes global text information.
[0037] As a preferred option, the re-ranking mechanism in step S4 optimizes the retrieval results by combining the ranking information and components of image-to-text (i2t) and text-to-image (t2i) retrievals in the original similarity matrix.
[0038] As a preferred embodiment, the objective function in step S5 of training the retrieval model is as follows:
[0039] S51: Employ the InfoNCE loss function from contrastive learning to maximize the similarity between positive samples and minimize the similarity between negative samples;
[0040] S52: Meanwhile, to further optimize cross-modal semantic alignment, KL divergence is introduced to measure the difference between the probability distributions of different modalities;
[0041] S53: By minimizing the difference between the predicted distribution and the true distribution, KL divergence is used to improve the consistency of text and image features in the semantic space.
[0042] As a preferred embodiment, step S5 introduces a three-loss-function collaborative optimization model training, which includes:
[0043] First, there is the bidirectional contrast loss, specifically image-to-text (i2t) and text-to-image (t2i):
[0044]
[0045]
[0046] Where N is the batch size. The temperature parameter is a learnable parameter; the original loss is used to enhance the discriminative ability of image-text matching, and is expressed as:
[0047]
[0048] Then, an additional unimodal similarity distribution is introduced using the teacher model to guide the alignment of cross-modal probability distributions. The following defines the similarity probability distribution between images:
[0049]
[0050] in, This represents the cosine similarity between images calculated by extracting features using a pre-trained model. Probability distribution between two images; probability distribution between texts. Similarly, we can obtain;
[0051] The cross-modal probability distribution alignment loss is expressed as:
[0052]
[0053] in, By As the target distribution, the model is guided to predict a highly matching image-text similarity distribution, thereby achieving semantic distribution alignment from image to text.
[0054] As a preferred option, to achieve structural stability constraints in a single mode, the loss function is defined as:
[0055]
[0056] This constraint helps the model form clearer boundaries in the semantic spaces of image-image and text-text, enhancing the model's ability to generalize to unknown data;
[0057] The final total loss function is:
[0058]
[0059] A text and image retrieval system based on fine-grained alignment and reordering includes a feature extraction module, a cross-modal interaction module, a reordering module, and an objective function module.
[0060] (III) Beneficial Effects
[0061] Compared with existing technologies, this invention provides a text and image retrieval method and system based on fine-grained alignment and reordering, which has the following beneficial effects:
[0062] This invention employs a fine-grained alignment and re-ranking-based image and text retrieval approach. It includes a cross-modal interaction module based on a multi-head cross-attention layer and a multi-layer Transformer module, which utilizes contextual information to understand the overall scene and achieves joint modeling from fine-grained local details to global semantics. It also includes a re-ranking module that re-ranks the initial search results by adjusting multiple variable factors, thereby improving retrieval performance. The objective function module utilizes the collaborative optimization of these three modules, enabling the model to achieve accurate semantic alignment while also demonstrating stronger robustness and generalization performance. Attached Figure Description
[0063] Figure 1 This is a schematic diagram of the retrieval method of the present invention;
[0064] Figure 2 This is a schematic diagram of the overall architecture of the retrieval model of the present invention;
[0065] Figure 3 This is a schematic diagram illustrating the specific process of the reordering algorithm of this invention. Detailed Implementation
[0066] To better understand the purpose, structure, and function of this invention, the following will further describe the image and text retrieval method and system based on fine-grained alignment and reordering, in conjunction with the accompanying drawings and specific embodiments.
[0067] Example 1
[0068] refer to Figure 1-3 This invention discloses an image and text retrieval method based on fine-grained alignment and reordering, comprising the following steps:
[0069] S1: Input the image and text to be retrieved, and encode the input image and text separately using the powerful feature extraction capabilities of the pre-trained model CLIP; S2: Adaptively align the text representation to relevant image regions using the cross-modal interaction module, and integrate the aligned features into a globally semantically consistent representation; S3: Calculate the similarity score between the image and text to obtain preliminary matching results; S4: Perform reverse retrieval on the initial similarity matrix through a re-ranking mechanism to enhance the matching relevance between the image and text in bidirectional retrieval; S5: Train the retrieval model by combining three loss functions, and introduce the knowledge extracted offline by the single-modal pre-trained teacher model as a soft label supervision signal to optimize the similarity learning process; S6: Use KL divergence to measure the difference between the probability distribution output by the model and the soft labels provided by the teacher model to improve the semantic alignment ability between the image and text.
[0070] Specifically, feature extraction in step S1 of this invention includes image representation and text representation, wherein:
[0071] 1) Image representation:
[0072] For each input image The image embedding is obtained by using a variant of ResNet-50 (clip) as a visual encoder; basic features are extracted stepwise using three shallow convolutional layers, and semantic features are extracted stepwise using multiple deep residual modules; finally, a global aggregation is performed using an attention mechanism to obtain a high-semantic image vector representing the entire image.
[0073]
[0074]
[0075]
[0076] Then perform a two-dimensional average pooling:
[0077]
[0078] The obtained feature map Four residual modules are input sequentially, and then a multi-head attention pooling module is used to perform weighted aggregation of spatial features: This yields a global representation of the entire image.
[0079] 2) Text representation:
[0080] For the input text, the text encoder in Clip extracts the text representation. First, it uses lowercase byte pair encoding (BPE) with a vocabulary size of 49152 bytes to segment the input text description. Then, [SOS] and [EOS] markers are added before and after the text description to identify the beginning and end of the sequence. Finally, the segmented text... The input is a Transformer, and the correlation between each unit is mined through a masked self-attention mechanism; the output of the last layer of the Transformer at the [EOS] unit is linearly mapped to the image-text joint embedding space to obtain a global text representation.
[0081] Text input is performed in batch processing. Since the sentence lengths in each batch are different, a maximum text length is set for each batch. For less than Sentences within a unit are truncated to ensure consistent dimensionality; finally, both image and text embeddings are projected onto a shared embedding space through a linear layer, and the linear mapping is calculated as follows:
[0082]
[0083] Traditional attention mechanisms typically calculate similarity weights between all image regions and text units through dot product operations. This approach incurs significant computational overhead as sequence length increases, especially when dealing with long text descriptions or high-resolution images. To address this issue, some methods introduce a lightweight multilayer perceptron module to generate attention weights, replacing explicit dot product calculations and thus improving computational efficiency. However, such methods inherently rely on static attention, meaning the attention weights remain fixed after generation and cannot dynamically adapt to contextual information. This limitation hinders their ability to model complex and fine-grained cross-modal semantic alignment. To overcome this challenge, this invention employs a cross-modal interaction module based on a multi-head cross-attention layer and a multilayer Transformer module. Compared to other popular multimodal interaction modules, this invention offers superior computational efficiency. The cross-modal interaction module dynamically and explicitly models the many-to-many semantic correspondence between image regions and text units. Furthermore, it leverages contextual information to understand the overall scene, achieving joint modeling from fine-grained local details to global semantics.
[0084] Specifically, the core idea of this section is to dynamically utilize information from one modality and selectively focus on the most relevant content from another modality, thereby enhancing cross-modal interaction. The cross-modal interaction in step S2 includes:
[0085] S21: First, textual and visual features are fused through a cross-attention layer. Then, the fused features are input into a multi-layer Transformer module to obtain a fused multimodal table containing the context and semantic correspondence between images and text.
[0086] S22: Subsequently, a text-image alignment matrix is calculated based on the similarity between modalities. This matrix captures the fine-grained semantic correspondence between modalities.
[0087] Specifically, given the input text description and input image The final hidden states of the text and vision are respectively fed into their respective encoders. The input is shared into the cross-modal interaction module; first, the text representation... As query Q, the image representation As the key K and value V, the following process is bidirectional; the same steps are performed when the image representation is Q:
[0088]
[0089]
[0090]
[0091] in, It is multi-head cross-attention, defined as follows:
[0092]
[0093] in, It is the embedded dimension, with a scaling factor of 1. After obtaining the text representation that fuses the image semantics, it is input into the Transformer to capture more complex and long-distance correspondences, enhancing the global semantic understanding capability. The complete interaction of image and text representations can be represented as follows:
[0094]
[0095] in, After normalization of the representation layer, the similarity between the text vector containing global image information and the image feature vector containing global text information is calculated. This similarity is regarded as the image-text alignment score, as shown in the following formula:
[0096]
[0097] The same principle, This represents an image representation that includes global text information.
[0098] Furthermore, for retrieval tasks, this invention proposes a post-processing method. This module reorders the initial retrieval results by adjusting multiple variable factors, thereby improving retrieval performance. Existing retrieval tasks use traditional methods to calculate the similarity between each query text and each query image, generating a similarity matrix where each element represents the similarity between the text and the image. For each text... From the similarity matrix Select the top with the highest similarity One entry, as the text Candidate images. The same principle applies when retrieving text from images. However, this method ignores the inherent correlation between bidirectional searches, since text and images are mutually matched and should be mutually searchable. To address this issue, a cross-modal reordering algorithm has been proposed, which utilizes... Each candidate element undergoes a reverse search, and the final search result is determined based on its ranking position in the reverse search results. While this cross-modal re-ranking algorithm improves search performance to some extent, it doesn't fully utilize the information in the similarity matrix. This section optimizes the search results by combining the ranking information and components of image-to-text (i2t) and text-to-image (t2i) searches in the original similarity matrix. The following example of t2i (text-to-image retrieval) illustrates the application of the re-ranking algorithm within this framework. (See attached manual.) Figure 2 The specific process of the algorithm is shown.
[0099] To better understand the technical solution of this invention, during the training process, the number of image-text pairs in each batch is fixed. A similarity matrix can be constructed, where the diagonal elements represent matching positive sample pairs (i.e., correct image-text pairs) with the highest similarity, and the off-diagonal elements represent mismatched negative sample pairs. Step S5, training the retrieval model, includes an objective function, specifically:
[0100] S51: Employ the InfoNCE loss function from contrastive learning to maximize the similarity between positive samples and minimize the similarity between negative samples;
[0101] S52: Meanwhile, to further optimize cross-modal semantic alignment, KL divergence is introduced to measure the difference between the probability distributions of different modalities;
[0102] S53: By minimizing the difference between the predicted distribution and the true distribution, KL divergence is used to improve the consistency of text and image features in the semantic space.
[0103] Based on this, the present invention introduces a three-loss-function collaborative optimization model training method, which includes:
[0104] First, there is the bidirectional contrast loss, specifically image-to-text (i2t) and text-to-image (t2i):
[0105]
[0106]
[0107] Where N is the batch size. The temperature parameter is a learnable parameter. The original loss is used to enhance the discriminative ability of image-text matching and is expressed as:
[0108]
[0109] Then, an additional unimodal similarity distribution is introduced using the teacher model to guide the alignment of cross-modal probability distributions. The following defines the similarity probability distribution between images:
[0110]
[0111] in, This represents the cosine similarity between images calculated by extracting features using a pre-trained model. This represents the probability that two images are semantically identical. The probability distribution between texts. Similarly, we can obtain;
[0112] The cross-modal probability distribution alignment loss is expressed as:
[0113]
[0114] in, By As the target distribution, it guides the model to predict a highly matching image-text similarity distribution, thereby achieving semantic distribution alignment from image to text.
[0115] Meanwhile, in order to achieve structural stability constraints under single-mode conditions, the loss function is defined as:
[0116]
[0117] This constraint helps the model form clearer boundaries in the semantic spaces of image-image and text-text, enhancing the model's ability to generalize to unknown data;
[0118] The final total loss function is:
[0119]
[0120] Example 2
[0121] This invention also proposes an image-text retrieval system based on fine-grained alignment and re-ranking, including a feature extraction module, a cross-modal interaction module, a re-ranking module, and an objective function module. By jointly using the above three loss models, it effectively aligns the semantic relationships between images and text while ensuring the clarity and consistency of the structure within each modality. Specifically, it improves the model's ability to discriminate between images and text; it further optimizes the fine-grained semantic association across modalities through probability distribution alignment; and it enhances the stability and generalization ability of the single-modal structure. Through the synergistic optimization of these three aspects, the model achieves accurate semantic alignment while also achieving stronger robustness and generalization performance.
[0122] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.
Claims
1. A text and image retrieval method based on fine-grained alignment and reordering, characterized in that, Includes the following steps; S1: Input the image and text to be retrieved, and encode the input image and text separately using the powerful feature extraction capabilities of the pre-trained model CLIP; S2: Adaptively align the text representation to the relevant image region using the cross-modal interaction module, and integrate the aligned features into a globally semantically consistent representation; S3: Calculate the similarity score between the image and the text to obtain preliminary matching results; S4: Perform reverse retrieval on the initial similarity matrix through a re-ranking mechanism to enhance the matching relevance between the image and the text in bidirectional retrieval; S5: Combine three loss functions to train the retrieval model, and introduce knowledge extracted offline by the single-modal pre-trained teacher model as a soft label supervision signal to optimize the similarity learning process; S6: Use KL divergence to measure the difference between the probability distribution output by the model and the soft labels provided by the teacher model, thereby improving the ability to semantically align text and images.
2. The image and text retrieval method based on fine-grained alignment and reordering according to claim 1, characterized in that, The feature extraction in step S1 includes image representation: For each input image The image embedding is obtained by using a variant of ResNet-50 (clip) as a visual encoder; basic features are extracted stepwise using three shallow convolutional layers, and semantic features are extracted stepwise using multiple deep residual modules; finally, a global aggregation is performed using an attention mechanism to obtain a high-semantic image vector representing the entire image. Then perform a two-dimensional average pooling: The obtained feature map Four residual modules are input sequentially, and then a multi-head attention pooling module is used to perform weighted aggregation of spatial features: This yields a global representation of the entire image.
3. The image and text retrieval method based on fine-grained alignment and reordering according to claim 2, characterized in that, The feature extraction in step S1 also includes text representation: For the input text, the text encoder in Clip extracts the text representation. First, it uses lowercase byte pair encoding (BPE) with a vocabulary size of 49152 bytes to segment the input text description. Then, [SOS] and [EOS] markers are added before and after the text description to identify the beginning and end of the sequence. Finally, the segmented text... Input a Transformer and mine the correlations between each unit through a masked self-attention mechanism; The output of the last layer of the Transformer at the [EOS] unit is linearly mapped to the image-text joint embedding space to obtain a global text representation; Text input is performed in batch processing. Since the sentence lengths in each batch are different, a maximum text length is set for each batch. For less than Sentences within a unit are truncated to ensure consistent dimensionality; finally, both image and text embeddings are projected onto a shared embedding space through a linear layer, and the linear mapping is calculated as follows:
4. The image and text retrieval method based on fine-grained alignment and reordering according to claim 1, characterized in that... The cross-modal interaction in step S2 includes: S21: First, textual and visual features are fused through a cross-attention layer. Then, the fused features are input into a multi-layer Transformer module to obtain a fused multimodal table containing the context and semantic correspondence between images and text. S22: Subsequently, a text-image alignment matrix is calculated based on the similarity between modalities. This matrix captures the fine-grained semantic correspondence between modalities.
5. The image and text retrieval method based on fine-grained alignment and reordering according to claim 4, characterized in that... The cross-modal interaction in step S2 specifically refers to: Given input text description and input image The final hidden states of the text and vision are respectively fed into their respective encoders. The input is shared into the cross-modal interaction module; first, the text representation... As query Q, the image representation As the key K and value V, the following process is bidirectional; the same steps are performed when the image representation is Q: in, It is multi-head cross-attention, defined as follows: in, It is the embedded dimension, with a scaling factor of 1. After obtaining the text representation that fuses the image semantics, it is input into the Transformer to capture more complex and long-distance correspondences, enhancing the global semantic understanding capability. The complete interaction of image and text representations can be represented as follows: in, After normalization of the representation layer, the similarity between the text vector containing global image information and the image feature vector containing global text information is calculated. This similarity is regarded as the image-text alignment score, as shown in the following formula: The same principle, This represents an image representation that includes global text information.
6. The image and text retrieval method based on fine-grained alignment and reordering according to claim 1, characterized in that... In step S4, the re-ranking mechanism optimizes the retrieval results by combining the ranking information and components of image-to-text (i2t) and text-to-image (t2i) retrievals in the original similarity matrix.
7. The image and text retrieval method based on fine-grained alignment and reordering according to claim 1, characterized in that... The objective function for training the retrieval model in step S5 is as follows: S51: Employ the InfoNCE loss function from contrastive learning to maximize the similarity between positive samples and minimize the similarity between negative samples; S52: Meanwhile, to further optimize cross-modal semantic alignment, KL divergence is introduced to measure the difference between the probability distributions of different modalities; S53: By minimizing the difference between the predicted distribution and the true distribution, KL divergence is used to improve the consistency of text and image features in the semantic space.
8. The image and text retrieval method based on fine-grained alignment and reordering according to claim 1, characterized in that... Step S5 introduces a collaborative optimization model training using three loss functions, including: First, there is the bidirectional contrast loss, specifically image-to-text (i2t) and text-to-image (t2i): Where N is the batch size. The temperature parameter is a learnable parameter; the original loss is used to enhance the discriminative ability of image-text matching, and is expressed as: Then, an additional unimodal similarity distribution is introduced using the teacher model to guide the alignment of cross-modal probability distributions. The following defines the similarity probability distribution between images: in, This represents the cosine similarity between images calculated by extracting features using a pre-trained model. Probability distribution between two images; probability distribution between texts. Similarly, we can obtain; The cross-modal probability distribution alignment loss is expressed as: in, By As the target distribution, the model is guided to predict a highly matching image-text similarity distribution, thereby achieving semantic distribution alignment from image to text.
9. The image and text retrieval method based on fine-grained alignment and reordering according to claim 8, characterized in that, To achieve structural stability constraints in a single-mode environment, the loss function is defined as: This constraint helps the model form clearer boundaries in the semantic spaces of image-image and text-text, enhancing the model's ability to generalize to unknown data; The final total loss function is:
10. A text and image retrieval system based on fine-grained alignment and reordering, characterized in that... It includes a vital sign extraction module, a cross-modal interaction module, a reordering module, and an objective function module.