Visual Re-ranking Method, System, Device and Storage Medium for Image Retrieval

By using the Transformer encoder structure to aggregate and update image features, the problem of inefficient existing image retrieval technology is solved, efficient image reordering is achieved, and real-time requirements are met.

CN114090816BActive Publication Date: 2025-05-30UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111394403.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-23
Publication Date
2025-05-30
Estimated Expiration
2041-11-23

AI Technical Summary

Technical Problem

Existing image retrieval technology is inefficient in the face of massive data, especially the reordering method based on K nearest neighbors and spatial verification will affect the retrieval efficiency.

Method used

The Transformer encoder structure is used to aggregate the correlation features of the top-k image, and the correlation features of each top-k image are updated through the full connection layer, and the similarity with the query image is calculated and the reordered.

Benefits of technology

It improves the reordering performance of image retrieval, and meets the real-time requirements of the retrieval system, so that image retrieval can be efficiently performed in large-scale data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114090816B_ABST
    Figure CN114090816B_ABST
Patent Text Reader

Abstract

The present invention discloses a visual re-ranking method, system, device and storage medium for image retrieval. The related method includes: obtaining a list of initial retrieval results, selecting the top k images with higher rankings for re-ranking, which are called top-k images, and selecting the top L images as anchor images; calculating the similarity between each top-k image and the anchor images respectively to obtain corresponding associated features; using the encoder structure of Transformer to aggregate the associated features of all top-k images, and updating a new associated feature for each top-k image; using the new associated features of each top-k image to calculate the similarity with the query image, and obtaining a re-ranked list according to the similarity magnitude. The above solution can ensure the real-time performance required by the retrieval system while guaranteeing the re-ranking performance, and this method can be applied to general retrieval systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image retrieval, and in particular, to a visual re-ranking method, system, device, and storage medium for image retrieval. Background Art

[0002] With the rapid development of the Internet and the popularization of mobile intelligent terminals, multimedia data, especially visual data, has shown an explosive growth. Billions of people share and browse photos and videos on the Internet. In order to enable users to quickly find the content they are interested in from these multimedia data, multimedia retrieval technology has received extensive attention and has developed rapidly. In the face of massive data, how to design efficient retrieval algorithms has always been a research hotspot in the academic and industrial circles at home and abroad.

[0003] Images, as an important part of multimedia data, have become the focus of attention in the field of information retrieval. Different from the early text-based image retrieval technology, content-based image retrieval uses the visual content of images as the basis for searching, which more directly expresses the user's search intention and can also be used as an important supplement to text search to further improve the search performance.

[0004] The image retrieval task is to retrieve images related to the query image from a large-scale image database according to the given query image. However, the effect of the original retrieval results may be less than satisfactory, and image retrieval re-ranking is a post-processing method that can effectively improve the retrieval performance.

[0005] Retrieval re-ranking mainly includes three types of methods: query expansion, K-nearest neighbor-based re-ranking method, and space verification-based re-ranking method.

[0006] 1) The query expansion method adds the features of the top results in the retrieval result list to the query image features to obtain a new query feature, and then performs a secondary retrieval based on this. The defect of this method is that it needs to perform a secondary retrieval or even multiple retrievals. When the database scale is very large, this method will seriously affect the efficiency of image retrieval.

[0007] 2) The K-nearest neighbor-based re-ranking method obtains a new relevance ranking by mining the K-nearest neighbor structural relationship between images in the database. The defect of this method is that the re-ranking method based on K-nearest neighbor reconstruction needs to construct the K-nearest neighbor relationship and perform similarity propagation during the query, which will also affect the retrieval efficiency.

[0008] 3) The space verification method usually uses the spatial context information of the local invariant features of the image to delete incorrect matches, but this scheme will seriously affect the retrieval efficiency. Summary of the Invention

[0009] The object of the present invention is to provide a visual re-ranking method, system, device and storage medium for image retrieval, which can ensure the real-time performance required by the retrieval system while ensuring the re-ranking performance, and this method can be applied to general retrieval systems.

[0010] The object of the present invention is achieved by the following technical solutions:

[0011] A visual re-ranking method for image retrieval, comprising:

[0012] Obtain the initial retrieval result list, select the top k images with higher rankings for re-ranking, called top-k images, and select the top L images as anchor images;

[0013] Calculate the similarity between each top-k image and the anchor images respectively to obtain the corresponding associated features;

[0014] Use the encoder structure of Transformer to aggregate the associated features of all top-k images and update a new associated feature for each top-k image;

[0015] Utilize the new associated features of each top-k image to calculate the similarity with the query image, and obtain the re-ranked list according to the similarity.

[0016] A visual re-ranking system for image retrieval, used to implement the foregoing method, the system includes:

[0017] A list acquisition and image screening unit, used to obtain the initial retrieval result list, select the top k images with higher rankings for re-ranking, called top-k images, and select the top L images as anchor images;

[0018] An associated feature calculation unit, used to calculate the similarity between each top-k image and the anchor images respectively to obtain the corresponding associated features;

[0019] An associated feature update unit, used to use the encoder structure of Transformer to aggregate the associated features of all top-k images and update a new associated feature for each top-k image;

[0020] A re-ranking unit, used to utilize the new associated features of each top-k image to calculate the similarity with the query image, and obtain the re-ranked list according to the similarity.

[0021] A processing device, comprising: one or more processors; a memory for storing one or more programs;

[0022] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing method.

[0023] A readable storage medium stores a computer program, characterized in that when the computer program is executed by a processor, the foregoing method is implemented.

[0024] As can be seen from the technical solutions provided by the present invention above, the encoder structure of the Transformer is used to aggregate the correlation features of the top-k images and then re-rank them. It is a re-ranking method based on context similarity aggregation, which can not only have high re-ranking performance, but also meet the real-time requirements of the retrieval system. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0026] Figure 1 Schematic diagram of a visual re-ranking method for image retrieval provided by an embodiment of the present invention;

[0027] Figure 2 Schematic diagram of a visual re-ranking system for image retrieval provided by an embodiment of the present invention;

[0028] Figure 3 Schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0030] First, the following explanations will be made for the terms that may be used in this article:

[0031] Descriptions using terms such as "including", "comprising", "containing", "having", or other similar semantics shall be construed as non-exclusive inclusion. For example, including a technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) shall be construed as not only including the explicitly listed technical feature element, but also including other technical feature elements well-known in the art that are not explicitly listed.

[0032] The following provides a detailed description of a visual re-ranking method for image retrieval provided by the present invention. Content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art. For conditions not specified in the embodiments of the present invention, they are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. Reagents or instruments not specified in the embodiments of the present invention for the manufacturer are all conventional products that can be obtained through commercial purchase.

[0033] As Figure 1 shown, a visual re-ranking method for image retrieval mainly includes the following steps:

[0034] Step 1: Obtain the initial retrieval result list, select the top k images with higher rankings for re-ranking, which are called the top-k images of the original retrieval results (hereinafter referred to as top-k images), and select the top L images as anchor images.

[0035] Step 2: Calculate the similarity between each top-k image and the anchor images respectively to obtain the corresponding associated features.

[0036] Step 3: Use the encoder structure of Transformer to aggregate the associated features of all re-ranked images and update a new associated feature for each top-k image;

[0037] Step 4: Calculate the similarity between each top-k image and the query image using the new associated feature of each top-k image, and obtain the re-ranked list according to the similarity.

[0038] To more clearly show the technical solution provided by the present invention and the resulting technical effects, the following uses specific embodiments to provide a detailed description of a visual re-ranking method for image retrieval provided by the embodiments of the present invention.

[0039] 1. Obtain the initial retrieval result list through initial retrieval.

[0040] In the embodiments of the present invention, the image retrieval database is defined as D = {I 1 , I 2 , …, I N}, each image I can be mapped to a high-dimensional feature space through the feature extraction function φ(·), f = φ(I) ∈ R d , obtaining the feature f of each image = {f 1 , f 2 , …, f N}, where d is the feature dimension, f is the feature of the image, the subscript is the serial number of the image, and N is the number of images in the database.

[0041] Given a query image, the initial retrieval result list R = [r 1 , r 2 , …, r k is obtained using the features of the extracted images, where r is the image in the initial retrieval result list, the subscript is the serial number of the image, and K is the number of images in the initial retrieval result list.

[0042] Exemplarily, the initial retrieval result list can be obtained by calculating the cosine similarity between features and sorting them in descending order of cosine similarity.

[0043] II. Image Screening.

[0044] In the embodiments of the present invention, the first k images in the initial retrieval results are re-sorted. Therefore, the first k images in the list are selected as the top-k images; in addition, the first L images are also selected as the anchor images.

[0045] In the embodiments of the present invention, it is assumed that the query image is returned first. Therefore, the query image is the first in the initial retrieval result list, that is, the query image is included in the image sets formed by the top-k images and the anchor images respectively.

[0046] III. Calculating Associated Features.

[0047] In the embodiments of the present invention, the similarity between each top-k image and the anchor image is calculated respectively, and the corresponding associated feature is obtained and expressed as:

[0048]

[0049] where represents the similarity between a single top-k image and a single anchor image, l = 1, …, L.

[0050] IV. Updating Associated Features.

[0051] For each top-k image, after passing through the fully connected layer (FC), the encoder structure of the Transformer is used to aggregate the associated features of other top-k images. The associated feature sequence A = {a 1 , …, a k} is input to the Transformer encoder structure, which can dynamically learn to aggregate the weights of each top-k image and output the updated new associated feature sequence Y = {y 1 ,…,y k}, where a and y are the associated features before and after the update, respectively, and the subscript is the image number.

[0052] 5. Image reordering.

[0053] In the embodiment of the present invention, the similarity between each top-k image and the query image is recalculated using the new associated features obtained by the above update. As mentioned above, the top-k images include the query image, so the query image also uses the updated new associated features. After the calculation is completed, the corresponding images are sorted in descending order according to the similarity to obtain a re-sorted list.

[0054] Exemplarily, a cosine similarity calculation may be performed.

[0055] 6. Model training.

[0056] In the embodiment of the present invention, the model training is mainly used to train the parameters in the fully connected layer and the encoder of the Transformer.

[0057] 1. Construct a loss function.

[0058] 1) Contrastive loss function.

[0059] Related images should have a larger cosine similarity, and unrelated images should have a smaller cosine similarity. Therefore, the contrast loss function is calculated by using the similarity between the new associated features of the query image and the new associated features of other top-k images. It is expressed as:

[0060]

[0061]

[0062] Among them, 1() represents the indicator function, if r 1 With r i If the indicator function is 1, then it is 0. 1 With r i Whether it is related is determined by whether the images belong to the same category. Whether they belong to the same category is reflected by the annotation information of the training set. τ is the temperature coefficient and sim() is the similarity function. 1 is the query image r 1 The new associated features, y i represents the top-k image r iNew associated features, where 2 ≤ i ≤ k; when the query image is similar to the relevant images and dissimilar to the irrelevant images, the contrastive loss reaches its minimum value.

[0063] 2) Mean Squared Error (MSE) loss function.

[0064] To retain the information in the original associated features (i.e., the associated features before update) and increase the stability of training, the MSE loss function is calculated using the associated features of the top-k images before and after update. It is expressed as:

[0065]

[0066] where a and y are the associated features before and after update respectively, with the subscript being the image serial number, and MLP() represents a multi-layer perceptron that maps the updated associated features back to the original associated feature space.

[0067] Combining the above two parts of the loss function to construct the final loss function, which is expressed as:

[0068]

[0069] where λ is the weight coefficient of the loss term.

[0070] 2. Data augmentation.

[0071] To enhance the robustness of the network and avoid overfitting, data augmentation methods are used. For the same training sample, by using different feature extractors {φ(·) 1 , φ(·) 2 , …, φ(·) M}, different initial retrieval result lists are obtained to expand the training set, where M is the number of feature extractors. Through data augmentation, the training set can be expanded to three times its original size.

[0072] The operations described in the above two to four are performed on the initial retrieval result lists obtained using different feature extractors, and they are brought into the above loss function for model training.

[0073] In the embodiments of the present invention, the remaining processes of the model training can be implemented with reference to conventional techniques, and the present invention will not elaborate. After completing the model training, the reordered list is directly generated using the updated associated features in the manner described in the above five and fed back to the query user.

[0074] In the above solution of the embodiment of the present invention, the information input during reordering is the associated features of the top-k images, which are the results of calculating the similarity between the top-k images and the anchor image, and have no direct association with the original visual features. Therefore, a kind of visual features can be used to obtain the initial retrieval result, and the associated features are calculated to train the network to obtain the reordering model (corresponding to the aforementioned third to fifth parts), and it is applied to other different visual features. Therefore, the present invention has good robustness and versatility. In addition, the present invention ensures the real-time performance required by the retrieval system while ensuring the reordering performance.

[0075] Another embodiment of the present invention also provides a visual reordering system for image retrieval, which is mainly used to implement the method provided by the foregoing embodiment, such as Figure 2 shown, the system mainly includes:

[0076] A list acquisition and image screening unit, configured to obtain the initial retrieval result list, select the top k images with the top rankings for reordering, which are called top-k images, and select the top L images as anchor images;

[0077] An associated feature calculation unit, configured to calculate the similarity between each top-k image and the anchor image respectively, and obtain the corresponding associated features;

[0078] An associated feature update unit, configured to aggregate the associated features of all top-k images using the encoder structure of Transformer, and update a new associated feature for each top-k image;

[0079] A reordering unit, configured to calculate the similarity between each top-k image's new associated feature and the query image, and obtain the reordered list according to the similarity magnitude.

[0080] Another embodiment of the present invention also provides a processing device, such as Figure 3 shown, which mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided by the foregoing embodiment.

[0081] Furthermore, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.

[0082] In the embodiment of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:

[0083] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.;

[0084] The output device can be a display terminal;

[0085] The memory can be a Random Access Memory (RAM), or a non-volatile memory, such as a disk memory.

[0086] Another embodiment of the present invention further provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiment when executed by a processor.

[0087] In the embodiment of the present invention, the readable storage medium, as a computer-readable storage medium, can be disposed in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a Read-Only Memory (ROM), a magnetic disk, or an optical disc.

[0088] As described above, the foregoing is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A visual re-ranking method for image retrieval, characterized in that, it includes: Obtain the initial retrieval result list, select the top k images with higher rankings for re-ranking, called top-k images, and select the top L images as anchor images; Calculate the similarity between each top-k image and the anchor images respectively to obtain the corresponding correlation features; Use the encoder structure of Transformer to aggregate the correlation features of all top-k images, and update a new correlation feature for each top-k image; Utilize the new correlation features of each top-k image to calculate the similarity with the query image, and obtain the re-ranked list according to the similarity; The encoder structure using Transformer aggregates the correlation features of all top-k images and updates a new correlation feature for each top-k image, including: forming a correlation feature sequence A = {a 1 , …, a k} from the correlation features of all top-k images, after passing through a fully connected layer, inputting it into the encoder structure of Transformer, and the Transformer encoder can dynamically learn to aggregate the weights of each top-k image, outputting a new updated correlation feature sequence Y = {y 1 , …, y k}, where a and y are the correlation features before and after update respectively, and the subscript is the image serial number; The query image is the image ranked first in the initial retrieval result list, and the query image is included in each of the image sets formed by the top-k images and the anchor images; During the training phase, the contrast loss function is calculated using the similarity between the new associated features of the query image and the new associated features of the other top-k images and the least mean square error loss function is calculated using the associated features before and after updating the top-k images Using a contrastive loss function and the least mean square error loss function construct the final loss function, expressed as: where λ is the weight coefficient of the loss term.

2. The visual re-ranking method for image retrieval according to claim 1, characterized in that, The calculation of the similarity between each top-k image and the anchor images respectively to obtain the corresponding correlation features is expressed as: Among them, represents the similarity between a single top-k image and a single anchor image, where l = 1, …, L.

3. The visual re-ranking method for image retrieval according to claim 1, characterized in that, The calculated contrastive loss function is expressed as: where 1() represents the indicator function, which is equal to 1 if r 1 is related to r i , and 0 otherwise, τ is the temperature coefficient, and sim() is the similarity function; y 1 is the new associated feature of the query image r 1 , y i represents the new associated feature of the top-k image r i , 2 ≤ i ≤ k; the contrastive loss reaches the minimum when the query image is similar to the relevant images and dissimilar to the irrelevant images.

4. The visual re-ranking method for image retrieval according to claim 1, characterized in that, Minimum Mean Square Error Loss Function It is expressed as: where a and y are the correlation features before and after update respectively, the subscript is the image serial number, and MLP() represents a multi-layer perceptron.

5. The visual re-ranking method for image retrieval according to any one of claims 1 to 4, characterized in that, Obtaining the initial retrieval result list includes: For the image retrieval database D = {I 1 , I 2 , …, I N}, mapping it to the feature space through the feature extraction function to obtain the features f = {f 1 , f 2 , …, f N} of each image, where I is the image in the image retrieval database, f is the feature of the image, the subscript is the serial number of the image, and N is the number of images in the database; Given a query image, using the features of the extracted images to obtain the initial retrieval result list R = [r 1 , r 2 , …, r K , where r is the image in the initial retrieval result list, the subscript is the serial number of the image, and K is the number of images in the initial retrieval result list; During the training phase, by using different feature extractors {φ(·) 1 , φ(·) 2 , …, φ(·) M}, the training set is augmented; that is, by using different feature extractors {φ(·) 1 , φ(·) 2 , …, φ(·) M}, different initial retrieval result lists are obtained, where M is the number of feature extractors.

6. A visual re-ranking system for image retrieval, characterized in that, used to implement the method according to any one of claims 1 to 5, and the system includes: A list acquisition and image screening unit for obtaining the initial retrieval result list, selecting the top k images with higher rankings for re-ranking, called top-k images, and selecting the top L images as anchor images; A correlation feature calculation unit for calculating the similarity between each top-k image and the anchor images respectively to obtain the corresponding correlation features; A correlation feature update unit for using the encoder structure of Transformer to aggregate the correlation features of all top-k images, and updating a new correlation feature for each top-k image; A re-ranking unit for utilizing the new correlation features of each top-k image to calculate the similarity with the query image, and obtaining the re-ranked list according to the similarity.

7. A processing device, characterized in that, it includes: One or more processors; A memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image retrieval reordering method and system based on arrangement fusion

    CN106708943A

  • Image retrieval method, electronic equipment and related products

    CN112765381A