A Combinatorial Image Retrieval Method with Dynamic Robust Fusion

The dynamic robust fusion method using BLIP encoders and soft label similarity losses addresses cross-modal challenges in combination image retrieval, improving accuracy and robustness by capturing latent associations and reducing overfitting.

CN120144817BActive Publication Date: 2025-07-15LUDONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510628846.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-15
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

Existing combined image retrieval methods have challenges in cross-modal gaps and semantic fusion, making it difficult to effectively capture implicit associations between images and text, and are affected by false negative images and false labels, resulting in model learning errors and overfitting risks.

Method used

A combined image retrieval method with dynamic robust fusion is adopted, and the BLIP text feature extractor is fine-tuned, and key features are captured in combination with a lightweight global attention mechanism, and a combination of soft label similarity comparison loss and batch classification loss is introduced to improve the robustness of the model.

Benefits of technology

Effectively capture the implicit association between images and text, improves the accuracy of combined image retrieval, reduces the risk of overfitting, and improves the robustness and retrieval accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144817B_ABST
    Figure CN120144817B_ABST
Patent Text Reader

Abstract

The present invention discloses a combined image retrieval method for dynamic robust fusion, belonging to the cross-modal retrieval technology field of multimedia retrieval. The present invention fine-tunes the BLIP text feature extractor; extracts reference image features and target image features through the BLIP image encoder, and extracts modified text features using the fine-tuned BLIP text encoder; designs a lightweight attention mechanism to capture key features between the reference image and the modified text, and fuses it with the reference image features and the fine-tuned modified text features for dynamic fusion; introduces a soft label similarity contrast loss and combines it with the batch classification loss; the combined image retrieval method of the present invention is more effective, not only effectively captures the implicit association between the reference image and the modified text, but also improves the robustness, thereby further improving the accuracy of combined image retrieval, and has good application prospects and considerable market value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for combined image retrieval with dynamic robust fusion, belonging to the technical field of cross-modal retrieval in multimedia retrieval. Background Art

[0002] In recent years, with the booming development of the Internet and the popularization of mobile terminals, people have generated a vast amount of information data such as images, videos, texts, and audios on various social platforms, search tools, and video websites. At the same time, with the development of modern urban informatization, information acquisition devices such as surveillance devices, industrial Internet of Things, and intelligent interaction systems continuously sense and record the data flow of the surrounding environment. Among the vast amount of information data, image data has rich content and convenient acquisition methods, and is an important information presentation method, and the obtained images account for a large part of the total data volume. Therefore, how to perform accurate and effective information retrieval tasks on a huge amount of image data has become increasingly challenging. Different from single-modal image retrieval methods, combined image retrieval can combine the images and text information provided by users to retrieve target images, which can meet the more accurate retrieval needs of users. Therefore, the research on combined image retrieval methods has attracted extensive attention in the academic and industrial fields. This task requires in-depth understanding of image content and text content to effectively identify the connection between them and fuse image content and text content. The main challenges lie in cross-modal gaps and semantic fusion. Therefore, the fusion algorithm directly affects the accuracy of the similarity between images and texts. Currently, most methods for combined image retrieval adopt a multi-layer attention mechanism strategy; use a hierarchical refinement attention mechanism to accurately locate key regions in visual and text inputs; by aligning the most relevant cross-modal features during the retrieval process, these methods have significantly improved the accuracy of query-target image matching. However, for the combined image retrieval task, simply capturing key information is likely to ignore implicit associations that are crucial to the task. In addition, there are a large number of false-negative images in existing datasets, which greatly affects the training effect and the inference ability of the model. Traditional methods generally use discrete labels, which are prone to causing the model to learn wrong labels and increasing the risk of overfitting. Summary of the Invention

[0003] The object of the present invention is to overcome the deficiencies of the above-mentioned existing technologies and provide a method for combined image retrieval with dynamic robust fusion.

[0004] The technical solution provided by the present invention is as follows: A method for combined image retrieval with dynamic robust fusion, characterized in that it includes the following steps:

[0005] Step S1: Using the FashionIQ and CIRR datasets, establish datasets in two modalities of images and texts. The image modality includes reference images and target images, and the text modality includes modified texts. Then divide the datasets in both modalities into training sets, validation sets, and test sets;

[0006] Step S2: Use the training set to fine-tune the BLIP text feature extractor to obtain the fine-tuned BLIP text feature extractor;

[0007] It includes the following steps:

[0008] Step S21: Use the BLIP image encoder to extract reference image features and target image features, and use the BLIP text encoder to extract modified text features;

[0009] Step S22: Perform element-wise addition on the extracted reference image features and modified text features to obtain combined features;

[0010] Step S23: Use the combined features obtained after element-wise addition and the target image features for batch-based classification loss training to obtain the fine-tuned BLIP text feature extractor;

[0011] Step S3: Construct the objective function on the training set;

[0012] It includes the following steps:

[0013] Step S31: Use the BLIP image encoder to extract reference image features and target image features, and use the fine-tuned BLIP text encoder to extract fine-tuned modified text features;

[0014] Step S32: Adopt the method of lightweight global attention to capture the key features between the reference image and the fine-tuned modified text;

[0015] Step S33: Fuse the key features with the reference image features and the fine-tuned modified text features to obtain fused features;

[0016] Step S34: Calculate the similarity between the fused features and the target image features;

[0017] Step S35: Adopt the soft label similarity contrast loss and combine it with the batch classification loss to obtain the final combined loss;

[0018] Step S4: Perform image-text matching and use the recall rate index to determine the matching accuracy.

[0019] Furthermore, in the said Step S2:

[0020] Step S21: Input the reference image and the target image of the training set into the BLIP image feature extractor to extract the reference image feature and the target image feature. Input the modified text of the training set and the validation set into the text feature extractor of BLIP to extract the modified text feature. Use to represent the reference image, to represent the target image, to represent the modified text, to represent the query composed of the reference image and the modified text;

[0021] Step S22: Perform element-wise addition on the reference image feature and the modified text feature of the training set to obtain a combined feature, aiming to make where, represents the image encoder of BLIP, represents the text encoder of BLIP;

[0022] Step S23: Perform batch-based classification loss training on the combined feature obtained after element-wise addition and the target image feature. The formula for the batch-based classification (BBC) loss is as follows:

[0023] ;

[0024] where, represents the i-th combined feature in the training batch, represents the j-th combined feature in the training batch, represents the i-th target image feature in the training batch, represents the j-th target image feature in the training batch, represents the training batch size, represents the temperature parameter, set to 100, represents the similarity kernel of the cosine similarity; After this classification loss training, the fine-tuned BLIP text feature extractor is obtained.

[0025] Furthermore, in step S3:

[0026] Step S31: Input the modified text of the training set and the validation set into the fine-tuned BLIP text feature extractor to obtain the fine-tuned modified text feature, represented by ; Input the reference image and the target image of the training set into the BLIP image encoder to extract the reference image feature and the target image feature respectively. The reference image feature is represented by and the target image feature is represented by ;

[0027] Step S32: Adopt the method of lightweight global attention to edit the BLIP feature to obtain:

[0028] ;

[0029] ;

[0030] Among them, represents the attention weight of the reference image feature, represents the attention weight for fine-tuning and modifying the text feature, , It is stated that the dimension of the feature vector is D-dimensional; represents feature connection, represents the multi-layer fully connected module of the reference image feature, represents the multi-layer fully connected module for fine-tuning and modifying the text feature, represents the non-linear activation function;

[0031] The key feature is obtained through the following formula :

[0032] ;

[0033] In step S33, the key feature captured by the attention mechanism is fused with the reference image feature and the fine-tuned and modified text feature to retain the implicit information in the reference image and the fine-tuned and modified text. The final fused feature is expressed as:

[0034] ;

[0035] Among them, represents the linear layer dimension transformation, represents the dynamically learned weight parameter;

[0036] In step S34, the cosine similarity is used to calculate the similarity between the fused feature and the target image feature , represents the similarity between two features, represents the i-th fused feature in the training batch, represents the j-th fused feature in the training batch;

[0037] In step S35, the soft label similarity is used to mitigate the overfitting noise problem;

[0038] In step S351, for the i-th query and the j-th target image in the training set batch, the soft similarity label is generated by the following formula:

[0039] ;

[0040] Among them, represents the temperature factor, represents the similarity between the i-th target image and the j-th target image, represents the training batch size;

[0041] obtain all soft similarity labels ;

[0042] Step S352, combine the soft similarity labels with the similarity contrast loss to obtain the soft-label similarity contrast loss;

[0043] The specific steps are as follows:

[0044] First, transform the similarity between the fused features and the target image features into a probability distribution through softmax:

[0045] ;

[0046] Use to represent the predicted probability, and use to represent the noisy hard similarity label;

[0047] The similarity contrast loss formula is as follows:

[0048] ;

[0049] Combine the soft similarity labels , the predicted probability , the noisy hard similarity label , and the similarity contrast loss to obtain the soft-label similarity contrast loss:

[0050]

[0051] Among them, represents the label weight parameter, represents the i-th fused feature in the training batch, represents the j-th fused feature in the training batch, represents the i-th target image feature in the training batch, represents the j-th target image feature in the training batch;

[0052] Step S353, combine the soft-label similarity contrast loss with the batch classification loss;

[0053] The formula for the batch classification loss is as follows:

[0054] ;

[0055] Among them, represents the training batch size, represents the temperature parameter, represents the similarity kernel of cosine similarity;

[0056] Combine the soft label similarity contrast loss with the batch classification loss to obtain the final combined loss:

[0057] ;

[0058] where represents a hyperparameter for balancing the two loss terms.

[0059] The beneficial effects of the present invention are as follows: First, the present invention fine-tunes the text feature extractor; then, the fine-tuned BLIP text feature extractor extracts the fine-tuned modified text features, and the BLIP image feature extractor extracts the reference image features and the modified text features; a lightweight attention mechanism is designed to capture the key features between the reference image and the modified text, and dynamically fuse them with the reference image features and the fine-tuned modified text features; finally, the soft label similarity contrast loss is introduced and combined with the batch classification loss to improve the robustness of the method.

[0060] The present invention not only effectively captures the implicit association between the reference image and the modified text, but also improves the robustness of the method, thereby further improving the accuracy of combined image retrieval, having good application prospects and considerable market value. This makes the present invention more effective than most existing combined image retrieval methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is the loss curve diagram of the present invention trained on the FashionIQ dataset;

[0062] Figure 2 is the accuracy curve diagram of image retrieval of the present invention on the FashionIQ dataset;

[0063] Figure 3 is the comparison curve diagram of the present invention with BLIP4CIR on the FashionIQ dataset;

[0064] Figure 4 is the loss curve diagram of the present invention trained on the CIRR dataset;

[0065] Figure 5 is the accuracy curve diagram of image retrieval of the present invention on the CIRR dataset;

[0066] Figure 6 is the comparison curve diagram of the present invention with BLIP4CIR on the CIRR dataset. DETAILED DESCRIPTION OF THE INVENTION

[0067] The following is a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings:

[0068] A method for combined image retrieval with dynamic robust fusion, comprising the following steps:

[0069] Step S1, using the FashionIQ and CIRR datasets, establish datasets in two modalities of images and texts. The image modality includes reference images and target images, and the text modality includes modified texts. Divide the datasets in the two modalities into training sets, validation sets, and test sets.

[0070] Step S2, fine-tune the BLIP text feature extractor using the training set;

[0071] It includes the following steps:

[0072] Step S21, input the reference images and target images of the training set into the BLIP image feature extractor to extract reference image features and target image features, input the modified texts of the training set and validation set into the text feature extractor of BLIP to extract modified text features, and use to represent the reference image, to represent the target image, to represent the modified text, to represent the query composed of the reference image and the modified text.

[0073] Step S22, perform element-wise addition on the reference image features and modified text features of the training set to obtain combined features, with the aim of making , where represents the image encoder of BLIP, represents the text encoder of BLIP.

[0074] Step S23, perform batch-based classification loss training on the combined features obtained after element-wise addition and the target image features. The formula for the batch-based classification (BBC) loss is as follows:

[0075] ;

[0076] where represents the i-th combined feature in the training batch, represents the j-th combined feature in the training batch, represents the i-th target image feature in the training batch, represents the j-th target image feature in the training batch, represents the training batch size, represents the temperature parameter, set to 100, It represents the similarity kernel of cosine similarity. After being trained by this classification loss, a fine-tuned BLIP text feature extractor is obtained.

[0077] Step S3: Construct the objective function on the training set;

[0078] It includes the following steps:

[0079] Step S31: Input the modified texts of the training set and the validation set into the fine-tuned BLIP text feature extractor to obtain fine-tuned modified text features, denoted by . Input the reference images and target images of the training set into the BLIP image encoder to extract reference image features and target image features respectively. The reference image features are denoted by , and the target image features are denoted by .

[0080] Step S32: Adopt the method of lightweight global attention to edit the BLIP features, and obtain:

[0081] ;

[0082] ;

[0083] Among them, represents the attention weight of the reference image features, represents the attention weight of the fine-tuned modified text features, , indicates that the dimension of the feature vector is D-dimensional. represents feature concatenation, represents the multi-layer fully connected module of the reference image features, represents the multi-layer fully connected module of the fine-tuned modified text features, represents the non-linear activation function;

[0084] The key features are obtained through the following formula :

[0085] .

[0086] Step S33: Fuse the key features captured by the attention mechanism with the reference image features and the fine-tuned modified text features to retain the implicit information in the reference image and the fine-tuned modified text. The final fused features are represented as:

[0087] ;

[0088] Among them, represents the linear layer dimension transformation, represents the dynamically learned weight parameter.

[0089] Step S34, use cosine similarity to calculate the similarity between the fused feature and the target image feature , representing the similarity between two features, denoting the i-th fused feature in the training batch, denoting the j-th fused feature in the training batch.

[0090] Step S35, use soft label similarity to mitigate the overfitting noise problem.

[0091] Step S351, for the i-th query and the j-th target image in the training set batch, generate soft similarity labels by the following formula:

[0092] ;

[0093] where, denoting the temperature factor, representing the similarity between the i-th target image and the j-th target image, denoting the training batch size.

[0094] Obtain all the soft similarity labels .

[0095] Step S352, combine the soft similarity labels with the similarity contrast loss to obtain the soft label similarity contrast loss.

[0096] The specific steps are as follows:

[0097] First, convert the similarity between the fused feature and the target image feature into a probability distribution through softmax:

[0098] ;

[0099] Use representing the predicted probability, and use representing the noisy hard similarity label.

[0100] The similarity contrast loss formula is as follows:

[0101] ;

[0102] Combine the soft similarity labels , the predicted probability , the noisy hard similarity label , and the similarity contrast loss to obtain the soft label similarity contrast loss:

[0103]

[0104] Among them, represents the label weight parameter, represents the i-th fused feature in the training batch, represents the j-th fused feature in the training batch, represents the i-th target image feature in the training batch, represents the j-th target image feature in the training batch.

[0105] Step S353: Combine the soft label similarity comparison loss and the batch classification loss.

[0106] The formula for the batch classification loss is as follows:

[0107] ;

[0108] Among them, represents the training batch size, represents the temperature parameter, set to 100, represents the similarity kernel of the cosine similarity.

[0109] Combine the soft label similarity comparison loss with the batch classification loss to obtain the final combined loss:

[0110] ;

[0111] Among them, represents a hyperparameter for balancing the two loss terms. Through hyperparameter experiments, the hyperparameter for the Fashion-IQ dataset is set to 0.5, and for the CIRR dataset, the hyperparameter will be set to 0.1;

[0112] Step S4: Perform image-text matching and determine the matching accuracy using the recall rate metric.

[0113] This embodiment is trained, validated, and tested on the FashionIQ and CIRR datasets. The FashionIQ dataset contains three categories (Dress, Shirt, and Toptee), with approximately 46,000 training instances and 15,000 test samples. To obtain triples containing reference images, modified texts, and target images, the dataset is manually annotated with modified texts, and finally 18,000 triples for training and 6,000 triples for testing are obtained. The CIRR dataset comes from 21,000 real open-domain images of NLVR2, covering more than 36,000 real-world triples, where 80% is used for training, 10% for validation, and 10% for testing.

[0114] The recall rate (Recall, R@K) is used as the performance evaluation criterion, where K represents the top K results returned in the similarity ranking. For the three categories of the FashionIQ dataset, K is specified as 10 and 50. For the CIRR dataset, K is specified as 1, 5, 10, and 50, and Recall Subset@K is used for evaluation. The CIRR dataset provides 503 fully annotated subsets, each containing 6 visually similar images. Recall Subset@k measures the proportion of true images ranked among the top k results in their respective subsets. This case is compared with BLIP4CIR. The Recall@K results of the combined image retrieval task on the FashionIQ dataset are shown in Table 1, and the Recall@K and Recall Subset@k results of the combined image retrieval task on the CIRR dataset are shown in Table 2.

[0115] Table 1 Comparison of Recall@K for Combined Image Retrieval on the FashionIQ Dataset

[0116]

[0117] As Figure 1 shown, the loss is continuously converging on the FashionIQ dataset, while the accuracy of image retrieval in Figure 2 is continuously increasing, thus proving the effectiveness of the proposed loss.

[0118] Figure 3 It shows that the Recall@K results of the present invention on the FashionIQ dataset are improved compared with BLIP4CIR.

[0119] Table 2 Comparison of Recall@K and Recall Subset@k for Combined Image Retrieval on the CIRR Dataset

[0120]

[0121] As Figure 4 shown, the loss is continuously converging on the CIRR dataset, while the accuracy of image retrieval in Figure 5 is continuously increasing, thus proving the effectiveness of the proposed loss.

[0122] Figure 6 It shows that the Recall@K and Recall Subset@k results of the present invention on the CIRR dataset are improved compared with BLIP4CIR.

[0123] It should be understood that the parts not elaborated in this specification all belong to the prior art. The above embodiments only describe the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for combined image retrieval with dynamic robust fusion, characterized in that It includes the following steps: Step S1, using the FashionIQ and CIRR datasets, establish datasets in two modalities of images and texts. The image modality includes reference images and target images, and the text modality includes modified texts. Then divide the datasets in both modalities into a training set, a validation set, and a test set; Step S2, use the training set to fine-tune the BLIP text feature extractor to obtain the fine-tuned BLIP text feature extractor; It includes the following steps: Step S21, use the BLIP image encoder to extract reference image features and target image features, and use the BLIP text encoder to extract modified text features; Step S22, perform element-wise addition on the extracted reference image features and modified text features to obtain combined features; Step S23, perform batch-based classification loss training on the combined features obtained after element-wise addition and the target image features to obtain the fine-tuned BLIP text feature extractor; Step S3, construct the objective function on the training set; It includes the following steps: Step S31, use the BLIP image encoder to extract the features of the reference image and the target image, and use the fine-tuned BLIP text encoder to extract the fine-tuned modified text features; Step S32, adopt the method of lightweight global attention to capture the key features between the reference image and the fine-tuned modified text; Step S33, fuse the key features captured by the attention mechanism with the reference image features and the fine-tuned modified text features to obtain fused features; Step S34, calculate the similarity between the fused features and the target image features; Step S35, adopt the soft-label similarity contrast loss and combine it with the batch classification loss to obtain the final combined loss; Step S4, perform image-text matching and determine the matching accuracy using the recall rate metric.

2. The combined image retrieval method with dynamic robust fusion according to claim 1, wherein In the said Step S2: Step S21, input the reference image and the target image of the training set into the BLIP image feature extractor to extract the reference image features and the target image features, input the modified text of the training set and the validation set into the text feature extractor of BLIP to extract the modified text features, and use to represent the reference image, to represent the target image, to represent the modified text, to represent the query composed of the reference image and the modified text; Step S22, adding the reference image features and the modified text features of the training set element-wise to obtain combined features, with the purpose of making , where represents the image encoder of BLIP, represents the text encoder of BLIP; Step S23, perform batch-based classification loss training on the combined features obtained after element-wise addition and the target image features. The formula for the batch-based classification (BBC) loss is as follows: ; Among them, represents the i-th combined feature in the training batch, represents the j-th combined feature in the training batch, represents the i-th target image feature in the training batch, represents the j-th target image feature in the training batch, represents the training batch size, represents the temperature parameter, set to 100, represents the similarity kernel of cosine similarity; the fine-tuned BLIP text feature extractor is obtained through training with this classification loss.

3. A method for combined image retrieval with dynamic robust fusion according to claim 1, characterized in that, In the said Step S3: Step S31, input the modified texts of the training set and the validation set into the fine-tuned BLIP text feature extractor to obtain fine-tuned modified text features, denoted by ; input the reference images and target images of the training set into the BLIP image encoder to extract reference image features and target image features respectively. The reference image features are denoted by , and the target image features are denoted by . Step S32, adopt the method of lightweight global attention to edit the BLIP features to obtain: ; ; Among them, represents the attention weight of the reference image features, represents the attention weight for fine-tuning and modifying the text features, , indicates that the dimension of the feature vector is D-dimensional; represents feature connection, represents the multi-layer fully-connected module for the reference image features, represents the multi-layer fully-connected module for fine-tuning and modifying the text features, represents the non-linear activation function; The key features are obtained through the following formula : ; Step S33, fuse the key features captured by the attention mechanism with the reference image features and the fine-tuned modified text features to retain the implicit information in the reference image and the fine-tuned modified text. The final fused feature is expressed as: ; Among them, represents the linear layer dimension transformation, represents the dynamically learned weight parameters; Step S34, use cosine similarity to calculate the similarity between the fused feature and the target image feature , represents the similarity between two features, denotes the i-th fused feature in the training batch, denotes the j-th fused feature in the training batch; Step S35, use soft-label similarity to mitigate the overfitting noise problem; Step S351, for the i-th query and the j-th target image in the training set batch, generate soft similarity labels by the following formula: ; Among them, represents the temperature factor, represents the similarity between the i-th target image and the j-th target image, represents the training batch size; Obtain all soft similarity tags ; Step S352, combine the soft similarity labels with the similarity contrast loss to obtain the soft-label similarity contrast loss; The specific steps are as follows: First, convert the similarity between the fused features and the target image features into a probability distribution through softmax: ; Use represents the predicted probability, and use to represent the noisy hard similarity label; The formula for the similarity contrast loss is as follows: ; Combine the soft similarity label , prediction probability , noisy hard similarity label , and the similarity contrast loss to obtain the soft label similarity contrast loss: ; Among them, represents the label weight parameter, represents the i-th fused feature in the training batch, represents the j-th fused feature in the training batch, represents the i-th target image feature in the training batch, represents the j-th target image feature in the training batch; Step S353, combine the soft-label similarity contrast loss with the batch classification loss; The formula for the batch classification loss is as follows: ; Among them, represents the training batch size, represents the temperature parameter, represents the similarity kernel of cosine similarity; Combine the soft label similarity contrast loss with the batch classification loss to obtain the final combined loss: ; Among them, represents a hyperparameter for balancing two loss terms.

Citation Information

Patent Citations

  • Unsupervised cross-modal hash retrieval method based on CLIP and attention fusion mechanism

    CN118861327A

  • Data compatibility for text-enhanced visual retrieval

    US20230073843A1