Combined image retrieval method and system suitable for data missing scene
Through self-supervised pre-training and attention matrix masking image areas, the model gap and data requirements problems in ZS-CIR are solved, which improves the accuracy and generalization ability of image retrieval and simplifies system complexity.
Patent Information
- Application Number
- CN202510410808.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-11
AI Technical Summary
There are problems in the existing zero-sample combined image retrieval (ZS-CIR) methods, which have large gaps between pre-trained vision-language models and tasks, strong dependence on additional components or modules, high demand for manual labeling data, and limited generalization capabilities.
Through the self-supervised pre-training method, the occlusion image is generated by using the graphics and text data, a triple data set is constructed, and the contrast loss between the occlusion image features and the original image features is minimized. The attention matrix is used for image area occlusion, and the query features are generated in combination with linear combination weighting calculations to reduce the contrast loss.
It reduces the cost and time of data integration, simplifies system complexity, improves the model's performance in image retrieval tasks, enhances the ability to understand detailed features and generalizes across fields and across scenarios, and improves the accuracy of retrieval.
Smart Images

Figure CN120296193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image retrieval, and in particular to a combined image retrieval method and system applicable to data - missing scenarios. Background Art
[0002] In recent years, with the development of deep learning and vision - language models (VLM), the combined image retrieval (CIR) technology has made remarkable progress. CIR allows users to adjust a reference image by specifying text modifications to retrieve target images that meet specific conditions. Traditional CIR methods usually rely on a large amount of manually annotated triple data (i.e., a dataset containing reference images, text modifications, and target images), which not only increases the cost of data collection but also limits the generalization ability of the model.
[0003] In the field of zero - shot composed image retrieval (ZS - CIR), researchers have used pre - trained vision - language models (VLM) to adapt to downstream ZS - CIR tasks. These models are pre - trained with a large amount of image - text pair data to learn the correlation between the image and text modalities. However, the ZS - CIR task requires guiding image retrieval based on text modifications, which is significantly different from the original goal of VLM, i.e., aligning image features and text features. Some current ZS - CIR methods have introduced additional modules or techniques for improvement. For example: Pic2Word: uses pre - trained VLMs plus an unlabeled image dataset to train a text inversion network that converts the input image into language tokens, enabling flexible combination of image and text queries. SEARLE: trains a text inversion network using image data and GPT - driven regularization to generate pseudo - word tokens. Context - I2W: proposes a context - dependent mapping network that dynamically maps description - related images to words. LinCIR: trains a text inversion network to improve retrieval by adding noise to the text dataset and self - masked projection.
[0004] Although the above - mentioned methods have achieved certain success, they also have certain limitations and challenges:
[0005] Dependence on additional components: These methods need to introduce additional components or networks, increasing the complexity of the system and potentially requiring additional data resources for training.
[0006] Failure to directly address the task - gap problem: These methods do not directly handle the fundamental gap between VLM and CIR tasks but instead try to bypass the problem, such as through text inversion and other means.
[0007] Annotation data requirements: Some supervised methods still require a large amount of manually annotated triple data (reference images, text modifications, target images), which increases costs and limits their scalability.
[0008] Limited generalization ability: Since VLM focuses on learning the similarity between modalities rather than specific text-guided image modifications, its generalization ability may be limited when applied to the CIR task. Summary of the Invention
[0009] The object of the present invention is to solve the problems existing in the existing zero-shot composite image retrieval (ZS-CIR) methods, such as the large gap between the pre-trained vision-language model and the task, strong dependence on additional components or modules, high demand for manually annotated data, and limited generalization ability.
[0010] To achieve the above object, the first aspect of the present invention provides a composite image retrieval method applicable to data missing scenarios, including: obtaining text-image pair data and inputting the original image and its corresponding text description therein into a pre-trained vision-language model VLM;
[0011] Based on the attention mechanism of VLM, calculate the attention matrix between the text description and each region of the original image;
[0012] Preset a masking ratio, and according to the attention matrix, mask the image regions in the order of increasing attention values until the masking ratio is satisfied, generating a masked image;
[0013] Construct a triple data set with the masked image, text description, and original image;
[0014] Perform self-supervised pre-training optimization on the VLM through the triple data set to minimize the contrast loss between the masked image features and the original image features;
[0015] In the inference stage, encode the feature vectors of the reference image and text modification instructions input by the user and perform weighted calculation on the feature vectors of the two to generate a query feature vector, and retrieve the target image in the image library that is most similar to the query feature vector.
[0016] Preferably, the masking process includes dividing the original image into non-overlapping fixed-size image patches, obtaining the attention distribution of each image patch region according to the attention matrix, arranging the attentions of each image patch in ascending order, and starting to mask from the image regions with the front attention ranking until the masking ratio is satisfied.
[0017] Preferably, the preset masking ratio is 75%, and it can also be adjusted according to actual needs.
[0018] Preferably, the step of self-supervised training and tuning includes:
[0019] Select a VLM as the encoder, extract the feature vectors of the masked image, text description, and original image, merge the feature vectors of the masked image and text description, calculate the similarity between the merged feature vector and the feature vector of the original image, and minimize the contrast loss function.
[0020] Preferably, the expression of the contrast loss function is:
[0021] L CL (f q , f t ) = B1 i=1 ∑ B -log(∑ j=1B exp(κ(f qi , f tj ))exp(κ(f qi , f ti )))
[0022] where B is the mini-batch size and κ(·) is the cosine similarity function.
[0023] Preferably, the weighted calculation is completed by linearly combining the feature vector of the reference image and the feature vector of the text modification instruction, and the calculation expression is as follows:
[0024] f r = (1 - w)f r1 + wf rT
[0025] where f r is the query feature, f rI is the image feature, f rT is the text feature, and w is the masking ratio.
[0026] To achieve the object of the present invention, a second aspect provides a combined image retrieval system applicable to the data missing scenario, which is used to apply the combined image retrieval method applicable to the data missing scenario described in the above technical solution. The system includes:
[0027] A data acquisition and input module, configured to acquire the text-image pair data and input the original image and its corresponding text description therein into the pre-trained vision-language model VLM;
[0028] An attention calculation module, configured to calculate the attention matrix between the text description and each region of the original image based on the pre-trained vision-language model VLM;
[0029] An image masking module, which is used to mask image regions according to the attention matrix in the order of increasing attention values until the masking ratio is satisfied, and generate a masked image;
[0030] A data construction module, which is used to construct the masked image, text description, and original image into a triple dataset
[0031] A training module, which is used to perform self-supervised pre-training and tuning on the VLM through the triple dataset, and minimize the contrast loss between the masked image features and the original image features;
[0032] An inference module, which is used to encode the feature vectors of the reference image and text modification instructions input by the user during the inference phase, perform weighted calculation on the feature vectors of the two, generate a query feature vector, and retrieve the target image in the image library that is most similar to the query feature vector.
[0033] Preferably, the image masking module divides the original image by using a non-overlapping fixed-size image block division method, and the masking ratio of the entire image is 75%.
[0034] Preferably, in the inference module, the weighted calculation to generate the query feature is implemented by linearly combining the feature vectors of the reference image and text modification instructions input by the user, and the weights are associated with the masking ratio.
[0035] Preferably, the system further includes a cloud database, which is used to store the parameters of the pre-trained VLM and the feature index of the image library.
[0036] Compared with the prior art, the beneficial effects of the present invention are:
[0037] Through a self-supervised pre-training method, the present invention uses publicly available image-text pairs of data for training, without the need for a large amount of manually annotated triple data, greatly reducing the dataset cost and time, and expanding the training data scale. Without introducing additional modules or networks, only by adjusting the existing VLM can its performance in the image retrieval task be improved, simplifying the system complexity and reducing the implementation difficulty. By adopting a high-proportion masking tuning strategy, the model is forced to infer the information of the complete image from a small number of visible blocks and text descriptions, enhancing the ability to understand detailed features and improving the generalization performance across domains and scenarios. During the inference phase, the distribution shift impact between training and test inputs is alleviated through a linear combination weighting method, enabling more accurate retrieval of the target image. Description of the Drawings
[0038] Figure 1 It is a step flow chart of a combined image retrieval method applicable to a data missing scenario according to Embodiment 1 of the present invention;
[0039] Figure 2 Schematic diagram of the combined image retrieval process based on the method of the present invention in Embodiment 3 of the invention of the present application. Detailed implementation manners
[0040] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0041] Embodiment 1
[0042] Please refer to Figure 1 , Embodiment 1 of the present invention provides a combined image retrieval method applicable to data missing scenarios, including the following steps:
[0043] S1: Obtain text-image pair data and input the original image and its corresponding text description therein into a pre-trained vision-language model VLM;
[0044] S2: Based on the attention mechanism of the VLM, calculate the attention matrix between the text description and each region of the original image;
[0045] S3: Preset a masking ratio, and according to the attention matrix, mask the image regions in the order of ascending attention values until the masking ratio is satisfied, generating a masked image;
[0046] S4: Construct a triple dataset with the masked image, text description, and original image;
[0047] S5: Perform self-supervised pre-training and tuning on the VLM through the triple dataset, minimizing the contrast loss between the features of the masked image and the features of the original image;
[0048] S6: In the inference stage, encode the feature vectors of the reference image and text modification instruction input by the user and perform weighted calculation on the feature vectors of the two to generate a query feature vector, and retrieve the target image in the image library that is most similar to the query feature vector.
[0049] The process of masking the image in step S3 includes dividing the original image into non-overlapping fixed-size image patches, obtaining the attention distribution of each image patch region according to the attention matrix to determine which regions are most important for a given text modification, arranging the attention of each image patch in ascending order, and starting to mask from the image regions with the front attention ranking until the masking ratio is satisfied. The attention matrix calculation formula is as follows:
[0050]
[0051] Where S is the attention score, Q is the text feature, K is the image feature, and D is the dimension of the feature. According to the selection criteria for the high-attention regions in the attention matrix (for example, selecting the image regions corresponding to the 75% lowest attention values from smallest to largest for masking), a targeted masking operation is performed to form a masked image with semantic meaning. The masking ratio w can be dynamically adjusted to adapt to different tasks.
[0052] In step S5, the feature vectors of the masked image and the text description are respectively extracted by the image encoder and the text encoder in the VLM, and these two encoded feature vectors are combined, aiming to let the model learn how to restore or adjust the image features according to the text guidance so that the masked image is as close as possible to the original image. During this training and tuning process, the contrast loss function is minimized to optimize the model parameters, ensuring that the masked image features can be accurately mapped to the original image feature space, thereby enhancing the model's ability to understand text guidance.
[0053] In the inference stage, when the user provides a reference image and specific text modification instructions, the encoder in the VLM extracts feature vectors from the reference image and the text modification instructions to generate a query feature, which is obtained by linearly combining the reference image feature vector and the feature vector of the text modification instructions. To reduce the impact of the distribution shift between training and test inputs, a weighted method is used for calculation, and the calculation expression is as follows:
[0054] f r =(1 - w)f rI +wf rT
[0055] where, f r is the query feature, f rI is the image feature, f rT is the text feature, and w is the masking ratio.
[0056] Finally, the trained and tuned VLM model is used to calculate the similarity between this query feature and the features of the candidate images in the image database, find the most matching target image, and provide the user with an accurate and relevant retrieval result. This method not only improves the learning efficiency of the model but also enhances its generalization ability and user experience in practical applications.
[0057] Embodiment 2
[0058] Based on Embodiment 1, this Embodiment 2 provides a combined image retrieval system applicable to data missing scenarios for applying the combined image retrieval method applicable to data missing scenarios described in Embodiment 1. The system includes:
[0059] A data acquisition and input module, configured to acquire image-text pair data and input the original image and its corresponding text description therein into a pre-trained vision-language model VLM;
[0060] An attention calculation module, configured to calculate an attention matrix between the text description and each region of the original image based on the pre-trained vision-language model VLM;
[0061] An image masking module, configured to mask image regions in the order of ascending attention values according to the attention matrix until a masking ratio is met, generating a masked image;
[0062] A data construction module, configured to construct the masked image, text description, and original image into a triple dataset
[0063] A training module, configured to perform self-supervised pre-training and tuning on the VLM through the triple dataset, minimizing the contrast loss between the features of the masked image and the original image;
[0064] An inference module, configured to, in the inference stage, encode the feature vectors of the reference image and text modification instruction input by the user and perform weighted calculation on the feature vectors of both to generate a query feature vector, and retrieve the target image in the image library that is most similar to the query feature vector.
[0065] In this Embodiment 2, the image masking module divides the original image by a non-overlapping fixed-size image patch division method, and the masking ratio of the entire image is 75%.
[0066] In the inference module, the weighted calculation to generate the query feature is implemented by linearly combining the feature vectors of the reference image and text modification instruction input by the user, and the weights are associated with the masking ratio.
[0067] The system further includes a cloud database, configured to store the parameters of the pre-trained VLM and the feature index of the image library.
[0068] Embodiment 3
[0069] Please refer to Figure 2 , this Embodiment 3 is based on Embodiment 1 and Embodiment 2, and specifically describes the process of pre-training the VLM model in the solution of the present invention and performing combined image retrieval using the trained model.
[0070] 1. Data Preparation
[0071] Dataset Selection: Two publicly available datasets were selected for pre-training, including LLaVA-CC3M-Pretrain-595K (595,000 image-text pairs) and ImageNet1K test split (100,000 image-text pairs). For ImageNet1K, the BLIP2-FlanT5-XL VLM was used to generate corresponding captions (text descriptions) for each image.
[0072] Data Processing: The original image was divided into non-overlapping fixed-size patches, and 75% of the patches were selected for masking according to the attention distribution to form the masked image Im. The original image I and its text description T remained unchanged, constituting the new triple data (Im, T, I).
[0073] As Figure 2 shown:
[0074] Caption(T): The text description "black shirt with bird pattern".
[0075] Image(I): The original image, a black T-shirt with a bird pattern.
[0076] Masked Image(Im): The masked image, where 75% of the regions are masked.
[0077] 2. Joint Input and Attention Calculation
[0078] Calculate the attention degree of each part of the image, i.e., the attention distribution. The text description T and the original image I were jointly input into the pre-trained VLM to obtain the joint feature representation. The multimodal encoder inside the VLM was used to process the text description and the original image to generate the attention matrix A, which reflects the attention degree of the text description to each part of the original image. According to the attention matrix A, the attention distribution of the original image was obtained, and it was masked in ascending order of attention until the masking ratio was satisfied to generate the masked image Im. The masking ratio was set to 75% according to the task requirements to ensure that the model pays more attention to the parts emphasized by the text description during the learning process.
[0079] As Figure 2 shown:
[0080] Text Encoder: Encodes the text description T into the text feature vector ft.
[0081] Image Encoder: Encodes the original image I to generate the image feature vector fi, and encodes the masked image Im to generate the masked image feature vector fm.
[0082] 3. Model Training
[0083] Encoder Selection: Use a pre-trained VLM as the image and text encoder.
[0084] Feature Extraction: The masked image Im and text description T are respectively passed through the image encoder and text encoder to extract feature vectors, and then the two feature vectors are combined.
[0085] Loss Function: Calculate the similarity between the combined feature vector and the target image feature, and minimize the contrastive loss function to optimize the model parameters, ensuring the similarity between the masked image and the original image in the feature space. The expression of the loss function is as follows:
[0086] L CL (f q ,f t )=B1 i=1 ∑ B -log(∑ j=1B exp(κ(f qi ,f tj ))exp(κ(f qi ,f ti )))
[0087] where B is the mini-batch size and κ(·) is the cosine similarity function.
[0088] 4. Inference Phase ( Figure 2 right side)
[0089] User Input: The user provides a reference image Ir and a text modification instruction Tr to generate a query feature.
[0090] As Figure 2 shown:
[0091] Query Text(Tr): The text description "Two dogs sit on the bench" provided by the user.
[0092] Reference Image(Ir): The reference image provided by the user, a photo of two dogs sitting on a bench.
[0093] Feature Combination: Use the trained model to calculate the similarity between the query feature and the candidate image features, and find the most matching target image. Considering the differences between the training and test inputs, a weighted method is adopted to calculate the query feature:
[0094] fr=(1-w)f rI +wf rT
[0095] where f r is the query feature, f rIis the image feature, f rT is the text feature, and w is the masking ratio.
[0096] Retrieval process:
[0097] Use the query feature vector fr to perform retrieval in the indexed image library, find the target image It that is most similar to the query feature vector, and return it to the user to complete the retrieval of the composite image.
[0098] In summary of the above embodiments, the present invention uses a self-supervised pre-training method to train with widely available image-text pair data without the need for a large amount of manually annotated triple data (reference image, text modification, target image). This greatly reduces the cost and time of data preparation and expands the scale of training data. The present invention does not need to introduce additional modules or networks (such as text inversion networks), and can significantly improve its performance in the zero-shot composite image retrieval (ZS-CIR) task only by adjusting the existing vision-language models (VLMs). This method simplifies the system complexity and reduces the implementation difficulty. Adopting a high ratio masking strategy (such as 75%), the model is forced to infer the information of the complete image from a small number of visible blocks and text descriptions, enhancing the ability to understand detailed features and improving the generalization performance across domains and scenarios. In the inference stage, a weighted method is proposed, that is, the query feature is a linear combination of the image feature and the text feature, reducing the influence of the distribution shift between training and test inputs and ensuring more accurate retrieval of the target image.
[0099] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A combined image retrieval method applicable to data missing scenarios, characterized in that It includes the following steps: Obtain image-text pair data and input the original image and its corresponding text description therein into a pre-trained vision-language model VLM; Based on the attention mechanism of the VLM, calculate the attention matrix between the text description and each region of the original image; Preset a masking ratio, and according to the attention matrix, mask the image regions in ascending order of attention values until the masking ratio is satisfied to generate a masked image; Construct a triple dataset from the masked image, text description, and original image; Perform self-supervised pre-training and tuning on the VLM through the triple dataset to minimize the contrast loss between the features of the masked image and the original image; In the inference stage, encode the feature vectors of the reference image and text modification instruction input by the user, perform weighted calculation on the feature vectors of the two, generate a query feature vector, and retrieve the target image in the image library that is most similar to the query feature vector.
2. The method according to claim 1, wherein The masking process includes dividing the original image into non-overlapping fixed-size image patches, obtaining the attention distribution of each image patch region according to the attention matrix, arranging the attentions of each image patch in ascending order, and starting to mask from the image regions with the earliest attention ranking until the masking ratio is satisfied.
3. The method according to claim 2, wherein The preset masking ratio is 75%, and it can also be adjusted according to actual needs.
4. The method according to claim 1, wherein The steps of the self-supervised training and tuning include: Select the VLM as the encoder, extract the feature vectors of the masked image, text description, and original image, merge the feature vectors of the masked image and text description, calculate the similarity between the merged feature vector and the feature vector of the original image, and minimize the contrast loss function.
5. The method according to claim 4, wherein The expression of the contrast loss function is: L CL (f q ,f t ) = B1 i=1 ∑ B -log(∑ j=1B exp(κ(f qi ,f tj ))exp(κ(f qi ,f ti ))) where B is the mini-batch size and κ(·) is the cosine similarity function.
6. The method according to claim 1, characterized in that, The weighted calculation is completed by linearly combining the feature vectors of the reference image and the text modification instruction, and the calculation expression is as follows: f r = (1 - w)f rI + wf rT Among them, f r is the query feature, f rI is the image feature, f rT is the text feature, and w is the masking ratio.
7. A combined image retrieval system applicable to data missing scenarios, for applying a combined image retrieval method applicable to data missing scenarios as described in claims 1-6, characterized in that, The system includes: A data acquisition and input module for obtaining image-text pair data and inputting the original image and its corresponding text description therein into a pre-trained vision-language model VLM; An attention calculation module for calculating the attention matrix between the text description and each region of the original image based on the pre-trained vision-language model VLM; An image masking module for masking the image regions in ascending order of attention values according to the attention matrix until the masking ratio is satisfied to generate a masked image; A data construction module for constructing a triple dataset from the masked image, text description, and original image A training module for performing self-supervised pre-training and tuning on the VLM through the triple dataset to minimize the contrast loss between the features of the masked image and the original image; An inference module for encoding the feature vectors of the reference image and text modification instruction input by the user in the inference stage, performing weighted calculation on the feature vectors of the two, generating a query feature vector, and retrieving the target image in the image library that is most similar to the query feature vector.
8. The system according to claim 7, wherein The image masking module divides the original image in a non-overlapping fixed-size image block division manner, and the masking ratio of the entire image is 75%.
9. The system according to claim 7, wherein In the inference module, the weighted calculation to generate the query feature is achieved by linearly combining the feature vectors of the reference image and the text modification instruction input by the user, and the weight is associated with the masking ratio.
10. The system according to claim 7, characterized in that, The system further includes a cloud database for storing the parameters of the pre-trained VLM and the feature index of the image library.
Citation Information
Cited By
Model training method and device for identifying seal characters, equipment and medium
CN121170810A