Asymmetric Image Retrieval Method, System, Device and Storage Medium

By using the anchor vector generated by the product quantizer in asymmetric image retrieval and imposing consistency constraints in query model training, the problem of unutilized training settings and structural information in the prior art is solved, and more efficient image retrieval performance is achieved.

CN114969422BActive Publication Date: 2025-05-30UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210676297.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-05-30
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

The existing asymmetric image retrieval scheme has dependence on the training settings, and does not fully consider the structural information embedded in the gallery model, resulting in poor retrieval performance.

Method used

By using a product quantizer to generate anchor vectors embedded in the gallery model, and in query model training, consistency constraints are applied to the similarity between these anchor vectors and image feature vectors to guide the training of the query model.

Benefits of technology

This method can weaken the requirements for training settings while considering the structural information of the embedded space of the gallery model, thereby improving the performance and robustness of image retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114969422B_ABST
    Figure CN114969422B_ABST
Patent Text Reader

Abstract

The present invention discloses an asymmetric image retrieval method, system, device and storage medium. In related methods, first, the product quantization method is used to generate a large number of anchor vectors in the embedding space of the image library to describe its spatial structure; then, these anchor vectors are shared between the query model and the image library model. During training, the relationship between the feature vector of each image and the corresponding anchor vector is regarded as structural similarity and is constrained to be consistent between the query model and the image library model. This allows the query model to ignore the feature details of the image library model and pay more attention to the overall spatial structure, enabling the embedding spaces of the query model and the image library model to be aligned, which is crucial for asymmetric retrieval and can improve the retrieval performance; in addition, the present invention does not utilize the annotations of images and can be trained using a large amount of unlabeled image data. Therefore, the present invention has good robustness and versatility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image retrieval, and in particular, to an asymmetric image retrieval method, system, device and storage medium. Background Art

[0002] With the rapid development of the Internet and the popularization of mobile intelligent terminals, multimedia data, especially visual data, has shown an explosive growth. Billions of people share and browse photos and videos on the Internet. In order to enable users to quickly find the content they are interested in from these multimedia data, multimedia retrieval technology has received extensive attention and has developed rapidly. In the face of massive data, how to design efficient retrieval algorithms has always been a research hotspot in the academic and industrial circles at home and abroad. As an important part of multimedia data, images have become the focus of attention in the field of information retrieval. Different from the early text-based image retrieval technology, content-based image retrieval uses the visual content of images as the basis for searching, which more directly expresses the user's search intention and can also be used as an important supplement to text search to further improve the search performance.

[0003] In most traditional methods, the query and database images are processed by the same model, which is called symmetric retrieval. To obtain higher retrieval performance, a large model with high performance is usually simply selected, which is very expensive in terms of computing resources. In some practical applications, the query image is usually processed on a terminal device with limited resources, while the database image undergoes offline feature extraction on a server with rich resources. Due to the limitation of computing resources, it is impossible to deploy a large model on the terminal device. Because of the low latency and low resource occupancy rate of small models, they are a better choice. Thanks to feature compatibility learning, it can use lightweight and large models to process the query model and the gallery model respectively, which allows us to enjoy the excellent feature extraction ability of the large model on the server side while maintaining low resource consumption on the query side. This asymmetric setting is called asymmetric retrieval. In some existing methods, the query model and the gallery model are used to extract the embeddings of the anchor image and its positive / negative sample images respectively. Then, asymmetric metric learning is adopted to guide the learning of the query model. In addition, direct feature regression has been proven to be effective. Some other methods consider from the perspectives of model parameter compatibility and model structure compatibility, and propose a compatibility-aware neural network search framework to search for the optimal query model structure and parameters.

[0004] However, the existing asymmetric retrieval solutions mainly have the following technical problems:

[0005] Technical problem 1 existing in the prior art: Using a compatibility-aware neural network search framework, which depends on the classifier of the gallery model and assumes that the dataset for training the gallery model can be obtained to train the lightweight query model.

[0006] Technical problem 2 existing in the prior art: Using asymmetric metric learning to guide the learning of the query model, only imposing instance-level constraints on the lightweight query model without considering the rich structural information in the embedding space. Summary of the Invention

[0007] The object of the present invention is to provide an asymmetric image retrieval method, system, device and storage medium, which can consider the structural information in the embedding space of the gallery model while weakening the requirements for training settings, thereby improving the retrieval performance.

[0008] The object of the present invention is achieved by the following technical solutions:

[0009] An asymmetric image retrieval method, comprising:

[0010] Using the gallery model to extract features from each image in the image database respectively, and offline training a product quantizer, and generating anchor vectors in the embedding space of the gallery model by using the trained product quantizer;

[0011] Inputting each image into the gallery model and the query model respectively to obtain the first feature vector extracted by the image model and the second feature vector extracted by the query model; calculating the similarities between the first feature vector and the second feature vector and the corresponding anchor vectors respectively, obtaining the first similarity corresponding to the first feature vector and the anchor vector, and the second similarity corresponding to the second feature vector and the anchor vector; imposing a consistency constraint on the first similarity and the second similarity to guide the training of the query model;

[0012] Inputting the image to be retrieved into the trained query model, and extracting the corresponding feature vector by the trained query model for retrieval.

[0013] An asymmetric image retrieval system, comprising:

[0014] An anchor vector generation unit, configured to use the gallery model to extract features from each image in the image database respectively, and offline train a product quantizer, and generate anchor vectors in the embedding space of the gallery model by using the trained product quantizer;

[0015] The query model training unit inputs each image into the gallery model and the query model respectively to obtain the first feature vector extracted by the image model and the second feature vector extracted by the query model; calculates the similarities between the first feature vector and the second feature vector and the corresponding anchor vectors respectively to obtain the first similarity corresponding to the first feature vector and the anchor vector, and the second similarity corresponding to the second feature vector and the anchor vector; applies a consistency constraint to the first similarity and the second similarity to guide the training of the query model;

[0016] The image retrieval unit is used to input the image to be retrieved into the trained query model, and the trained query model extracts the corresponding feature vector and performs retrieval.

[0017] A processing device includes: one or more processors; a memory for storing one or more programs;

[0018] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing method.

[0019] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.

[0020] As can be seen from the technical solutions provided by the present invention above, a structure similarity preservation learning method is designed to achieve feature consistency between the query model and the gallery model. In the present invention, first, the product quantization method is adopted to generate a large number of anchor vectors in the embedding space of the gallery to describe its space structure; then, these anchor vectors are shared between the query model and the gallery model. During training, the relationship between the feature vector of each image and the corresponding anchor vector is regarded as the structure similarity and is constrained to be consistent between the query model and the gallery model. This allows the query model to ignore the feature details of the gallery model and pay more attention to the overall space structure, so that the embedding spaces of the query model and the gallery model are aligned, which is crucial for asymmetric retrieval and can improve the retrieval performance; in addition, the present invention does not utilize the annotations of images and can use a large amount of unlabeled image data for training, so the present invention has good robustness and versatility. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0022] Figure 1Flowchart of an asymmetric image retrieval method provided by an embodiment of the present invention;

[0023] Figure 2 Flowchart of anchor vector production and query model training provided by an embodiment of the present invention;

[0024] Figure 3 Schematic diagram of an asymmetric image retrieval system provided by an embodiment of the present invention;

[0025] Figure 4 Schematic diagram of a processing device provided by an embodiment of the present invention. Detailed implementation manners

[0026] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] First, the following explanations are made for the terms that may be used in this article:

[0028] Descriptions with semantic meanings such as "including", "comprising", "containing", "having" or other similar ones shall be interpreted as non-exclusive inclusion. For example: including a certain technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, sizes, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) shall be interpreted as not only including the clearly listed certain technical feature element, but also including other well-known technical feature elements in the art that are not clearly listed.

[0029] The following provides a detailed description of an asymmetric image retrieval method, system, device and storage medium provided by the present invention. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those of ordinary skill in the art. For those conditions not specified in the embodiments of the present invention, they are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer.

[0030] Embodiment 1

[0031] An embodiment of the present invention provides an asymmetric image retrieval method, as Figure 1 shown, which mainly includes the following steps:

[0032] Step 1: Offline train a product quantizer and generate anchor vectors.

[0033] A direct way to generate anchor vectors is to generate them in the feature embedding space of the gallery model using the flat k-means clustering method. However, a large number of anchors are required to finely characterize the structure of the gallery model's embedding space. For k-means clustering, the required training samples and computational complexity are several times the number of center points. When this number is large, the cost of clustering is unaffordable. Therefore, the present invention turns to product quantization to quickly expand the number of anchors in the space at a lower cost. The main process is as follows: First, use the gallery model to extract features from each image in the image database respectively, and offline train the product quantizer. Then, use the center vectors of the offline-trained product quantizer as the anchor vectors in the gallery model's embedding space.

[0034] In the embodiment of the present invention, the product quantizer adopts an offline training method, and the training method can be implemented with reference to conventional techniques, which is not limited in the present invention. Anchor vectors are generated by the trained product quantizer. Specifically, there is a component in the product quantizer called the center vector, and the present invention uses the center vector as the anchor vector in the gallery model's embedding space.

[0035] In the embodiment of the present invention, the gallery model is a pre-trained model, which can be implemented using existing models, and the present invention does not limit the model structure.

[0036] In the embodiment of the present invention, the image database is denoted as X = {x 1 , x 2 , …, x n}, and the features of the image database are denoted as G = {g 1 , g 2 , …, g n}, where x i represents the i-th image, and g i represents the feature vector of the i-th image x i , and i = 1, 2, …, n, where n represents the total number of images.

[0037] Since the gallery model φ g (·) is frozen during the training process of the query model, the features G of the image database are offline-extracted features. The generation method of anchor vectors generated by the offline-trained product quantizer is as follows: After extracting the features G of the image database, each feature vector g i is respectively split into M sub-feature vectors. The same-segment sub-feature vectors of all feature vectors (i.e., the sub-feature vectors of the same serial number segment of n feature vectors) are clustered to obtain a corresponding set of anchor vectors. The set of anchor vectors corresponding to the j-th sub-feature vector is denoted as C j ∈ R K×d, where \(j = 1,\ldots,M\), \(j\) is the serial number of a segment of feature vectors, which is also equivalent to the serial number of the corresponding group of anchor vectors, \(\mathbb{R}\) represents the set of real numbers, \(K\) represents the number of center points, and the \(k\)-th center point corresponds to a sub-anchor vector \(d\) represents the dimension of a single sub-anchor vector; finally, \(M\) groups of anchor vectors are obtained, each group of anchor vectors corresponds to a sub-embedding space of the gallery model, and each group of anchor vectors contains \(K\) sub-anchor vectors.

[0038] As Figure 2 shown in the upper part, the main process of training the product quantizer to generate anchor vectors is shown. The dots in each subspace in the upper right part represent the elements in a segment of sub-feature vectors, and the \(X\) symbols represent the center points; of course, the number of dots (i.e., the number of elements) shown in the figure is only for illustration and does not represent the actual number.

[0039] In the embodiments of the present invention, product quantization has two main advantages. First, it is very easy to generate a large number of anchor vectors. Second, there is no need to directly store huge anchor vectors. Through the above processing, only \(M\times K\) sub-anchor vectors need to be stored. At the same time, during the training process of the query model, a segmentation mechanism is also adopted, rather than directly calculating the similarity between the feature vectors and all sub-anchor vectors, thereby greatly reducing the training cost.

[0040] Step 2: Train the query model.

[0041] In the embodiments of the present invention, each image is respectively input into the gallery model and the query model to obtain the first feature vector extracted by the image model and the second feature vector extracted by the query model; the similarities between the first feature vector and the second feature vector and the corresponding anchor vectors are respectively calculated to obtain the first similarity corresponding to the first feature vector and the anchor vector, and the second similarity corresponding to the second feature vector and the anchor vector; a consistency constraint is imposed on the first similarity and the second similarity to guide the training of the query model.

[0042] In the embodiments of the present invention, the first feature vector extracted by the image model \(\varphi\) g (·) is denoted as \(g\), and the second feature vector extracted by the query model \(\varphi\) q (·) is denoted as \(q\); the first feature vector \(g\) and the second feature vector \(q\) are respectively split to obtain \(M\) segments of sub-feature vectors, which are expressed as:

[0043] \(g\rightarrow u\) 1 (\(g\)), \(u\) 2 (\(g\)), \(\ldots\), \(u\) M (\(g\))

[0044] \(q\rightarrow u\) 1 (\(q\)), \(u\) 2 (\(q\)), \(\ldots\), \(u\) M (\(q\))

[0045] wherein, u j (g), u j (q) respectively represent the j-th sub-feature vector in the first feature vector g and the second feature vector q, where j = 1, …, M.

[0046] As Figure 2 shown, the above splitting operation is implemented by a product quantizer after offline training.

[0047] Calculate the similarity S between the M sub-feature vectors of the first feature vector g and the corresponding anchor vectors respectively g , and calculate the similarity S between the M sub-feature vectors of the second feature vector q and the corresponding anchor vectors respectively q . Among them:

[0048] The formula for calculating the similarity between the j-th sub-feature vector of the first feature vector g and the corresponding anchor vector is expressed as:

[0049]

[0050] The formula for calculating the similarity between the j-th sub-feature vector of the second feature vector q and the corresponding anchor vector is expressed as:

[0051]

[0052] Among them, represents the similarity between the j-th sub-feature vector of the first feature vector g and the corresponding anchor vector, represents the similarity between the j-th sub-feature vector of the second feature vector q and the corresponding anchor vector; C j represents the j-th group of anchor vectors, corresponding to the j-th sub-embedding space of the gallery model, K represents the number of center points, and the k-th center point corresponds to the sub-anchor vector

[0053] Exemplarily, a cosine similarity function can be selected to calculate the similarity, and the calculated similarities all belong to the structural similarity.

[0054] For asymmetric image retrieval, an ideal query model φ q(·) Not only maintains feature compatibility but also preserves the structural similarity in the embedding space of the gallery model. To achieve this, a consistency constraint is imposed on the first similarity and the second similarity. Specifically: both the first similarity and the second similarity are transformed into probability distributions, and the KL divergence is used to measure the consistency of the two probability distributions, and the query model is constrained by the consistency difference between the two probability distributions. It should be noted that each similarity calculated previously is based on each sub-feature vector. Therefore, here, the consistency constraint is also imposed on the similarity corresponding to each sub-feature vector. By imposing a consistency constraint on the structural similarity, the feature vectors extracted by the trained query model are located in the feature embedding space of the gallery model.

[0055] Such as Figure 2 The lower part of shows the main training process of the query model, and the first similarity will be used as a pseudo-label to constrain the second similarity.

[0056] In the embodiments of the present invention, the query model can also be implemented with reference to conventional techniques, and the present invention does not limit the structure of the query model. At the same time, the training process of the query model after imposing the consistency constraint can also be referred to conventional techniques, which will not be elaborated here.

[0057] Step 3, Image retrieval.

[0058] In the embodiments of the present invention, the database side uses a large model with strong expressive ability (that is, the gallery model described above) to extract image features and establish an index structure. The user side adopts a trained lightweight small model (that is, the query model trained above). The image to be retrieved is input into the trained query model, and the trained query model extracts the corresponding feature vectors and performs retrieval. The retrieval process involved here can be referred to conventional techniques. For example, the feature vectors extracted by the query model are calculated for similarity with the feature vectors extracted by the gallery model, and after sorting in descending order according to the similarity, multiple images corresponding to the top-ranked similarities are selected to generate a retrieval result and feedback it to the user.

[0059] The above solutions in the embodiments of the present invention mainly obtain the following beneficial effects:

[0060] 1. The relationship between the feature vector of each image and the corresponding anchor vector during training is regarded as structural similarity and is constrained to be consistent between the query model and the gallery model. This allows the query model to ignore the feature details of the gallery model and pay more attention to the overall spatial structure, aligning the embedding spaces of the query model and the gallery model, which is crucial for asymmetric retrieval and can improve the retrieval performance.

[0061] 2. Without using the annotations of images, large-scale unlabeled image data can be used for training, so it has good robustness and generality.

[0062] Example 2

[0063] The present invention also provides an asymmetric image retrieval system, which is mainly implemented based on the method provided in the foregoing embodiment. As Figure 3 shown, the system mainly includes:

[0064] An anchor vector generation unit, configured to extract features from each image in the image database by using the gallery model, and offline train a product quantizer, and generate an anchor vector in the embedding space of the gallery model by using the trained product quantizer;

[0065] A query model training unit, configured to input each image into the gallery model and the query model respectively, and obtain a first feature vector extracted by the image model and a second feature vector extracted by the query model; calculate the similarities between the first feature vector and the second feature vector and the corresponding anchor vectors respectively, obtain a first similarity corresponding to the first feature vector and the anchor vector, and a second similarity corresponding to the second feature vector and the anchor vector; impose a consistency constraint on the first similarity and the second similarity to guide the training of the query model;

[0066] An image retrieval unit, configured to input the image to be retrieved into the trained query model, and extract the corresponding feature vector by the trained query model for retrieval.

[0067] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.

[0068] Example 3

[0069] The present invention also provides a processing device, as Figure 4 shown, which mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiment.

[0070] Further, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.

[0071] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:

[0072] The input device can be a touch screen, an image acquisition device, physical buttons, a mouse, etc.;

[0073] The output device can be a display terminal;

[0074] The memory can be a Random Access Memory (RAM), or a non-volatile memory, such as a disk memory.

[0075] Embodiment 4

[0076] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiments when executed by a processor.

[0077] In the embodiments of the present invention, as a computer-readable storage medium, the readable storage medium can be disposed in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a Read-Only Memory (ROM), a magnetic disk, or an optical disc.

[0078] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An asymmetric image retrieval method, characterized in that, comprising: Performing feature extraction on each image in the image database using a gallery model, and offline training a product quantizer, and generating anchor vectors in the embedding space of the gallery model using the trained product quantizer; Inputting each image into the gallery model and the query model respectively to obtain a first feature vector extracted by the image model and a second feature vector extracted by the query model; calculating the similarities between the first feature vector and the second feature vector and the corresponding anchor vectors respectively, obtaining a first similarity corresponding to the first feature vector and the anchor vector, and a second similarity corresponding to the second feature vector and the anchor vector; imposing a consistency constraint on the first similarity and the second similarity to guide the training of the query model; Inputting the image to be retrieved into the trained query model, and extracting the corresponding feature vector by the trained query model for retrieval; The generating of the anchor vectors in the embedding space of the gallery model using the trained product quantizer includes: The image database is denoted as X = {x 1 , x 2 , …, x n}, and the features of the image database are denoted as G = {g 1 , g 2 , …, g n}, where x i represents the i-th image, and g i represents the feature vector of the i-th image x i , i = 1, 2, …, n, and n represents the total number of images; Each feature vector is split into M sub-feature vectors respectively. The same sub-feature vectors of all feature vectors are clustered to obtain a corresponding set of anchor vectors. Denote the set of anchor vectors corresponding to the j-th sub-feature vector as C j ∈R K×d , j = 1, …, M, where j is the serial number of a sub-feature vector, R represents the set of real numbers, K represents the number of center points, and the k-th center point corresponds to the sub-anchor vector d represents the dimension of a single sub-anchor vector; finally, M sets of anchor vectors are obtained. Each set of anchor vectors corresponds to a sub-embedding space of the gallery model, and each set of anchor vectors contains K sub-anchor vectors; The calculating of the similarities between the first feature vector and the second feature vector and the corresponding anchor vectors respectively includes: Denoting the first feature vector extracted by the image model as g, and denoting the second feature vector extracted by the query model as q; Splitting the first feature vector g and the second feature vector q respectively to obtain M sub-feature vectors each, expressed as: g→u 1 (g),u 2 (g),…,u M (g) q→u 1 (q),u 2 (q),…,u M (q) where, u j (g), u j (q) respectively represent the j-th sub-feature vector in the first feature vector g and the second feature vector q, where j = 1, …, M; Calculating the similarities between the M sub-feature vectors of the first feature vector g and the corresponding anchor vectors respectively, and calculating the similarities between the M sub-feature vectors of the second feature vector q and the corresponding anchor vectors respectively; The formula for calculating the similarity between the j-th sub-feature vector of the first feature vector g and the corresponding anchor vector is expressed as: Among them, represents the similarity between the j-th sub-feature vector of the first feature vector g and the anchor vector of the corresponding image; C j represents the j-th group of anchor vectors, corresponding to the j-th sub-embedding space of the gallery model, K represents the number of center points, and the k-th center point corresponds to the sub-anchor vector The formula for calculating the similarity between the j-th sub-feature vector of the second feature vector q and the corresponding anchor vector is expressed as: Among them, represents the similarity between the j-th sub-feature vector of the second feature vector q and the anchor vector of the corresponding image; C j represents the j-th group of anchor vectors, corresponding to the j-th sub-embedding space of the gallery model, K represents the number of center points, and the k-th center point corresponds to the sub-anchor vector 2. The asymmetric image retrieval method according to claim 1, characterized in that, The imposing of the consistency constraint on the first similarity and the second similarity includes: Converting both the first similarity and the second similarity into probability distributions, using the KL divergence to measure the consistency of the two probability distributions, and constraining the query model with the consistency difference of the two probability distributions.

3. The asymmetric image retrieval method according to claim 1, characterized in that, Both the first similarity and the second similarity belong to the structural similarity, and by imposing a consistency constraint on the structural similarity, the feature vectors extracted by the trained query model are located in the feature embedding space of the gallery model.

4. An asymmetric image retrieval system, characterized in that, Implemented based on the method according to any one of claims 1 to 3, and the system includes: An anchor vector generation unit for performing feature extraction on each image in the image database using a gallery model, and offline training a product quantizer, and generating anchor vectors in the embedding space of the gallery model using the trained product quantizer; A query model training unit inputs each image into a gallery model and a query model respectively to obtain a first feature vector extracted by the image model and a second feature vector extracted by the query model; calculates the similarities between the first feature vector and the second feature vector and their corresponding anchor vectors respectively to obtain a first similarity corresponding to the first feature vector and the anchor vector, and a second similarity corresponding to the second feature vector and the anchor vector; and imposes a consistency constraint on the first similarity and the second similarity to guide the training of the query model. An image retrieval unit is configured to input an image to be retrieved into the trained query model, and the trained query model extracts a corresponding feature vector and performs retrieval.

5. A processing device characterized in that it includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 3.

6. A readable storage medium storing a computer program characterized in that when the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Visual reordering method, system and equipment for image retrieval and storage medium

    CN114090816A

  • Method and apparatus for multimedia content indexing and retrieval based on product quantization

    EP3115908A1