General image retrieval method based on image quantization

CN118939828BActive Publication Date: 2026-09-29UESTC (SHENZHEN) ADVANCED RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411142461.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-09-29
Estimated Expiration
2044-08-20

AI Technical Summary

Technical Problem

然而,在图像检索中尤其是在新领域和新类别的现实应用中,方法的表现并不理想

Benefits of technology

[0027]1)本发明设计了一个二阶代码本,包含已知类别的代表性信息,通过对所有二阶码字进行组合投影来提取新图像,提高了对未知类别的泛化能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118939828B_ABST
    Figure CN118939828B_ABST
Patent Text Reader

Abstract

The application discloses a general image retrieval method based on image quantization, which comprises the following steps: firstly, constructing a training sample set; then, constructing an image quantization model comprising a visual feature extraction module, a first-order quantization cross-domain module, a combined feature extraction module and a second-order quantization cross-class module; in the training process, a first-order codebook is randomly initialized, cross-domain consistent features are learned by the visual feature extraction module and the first-order codebook through a contrast learning mechanism, a second-order codebook is obtained based on the first-order codebook update, new class-aware features are obtained by the combined feature extraction module and the second-order codebook, and quantized combined features are obtained by the second-order quantization cross-class module; in the retrieval process, each image in a search image library and a to-be-detected image are input into the image quantization model to obtain the combined features of each image, and a retrieval result is obtained based on the combined features. The application can effectively solve the cross-domain and cross-class challenges in the image retrieval task, and significantly improve the accuracy and efficiency of image retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image retrieval technology, and more specifically, relates to a general image retrieval method based on image quantization. Background Technology

[0002] Image quantization is a method to reduce the precision of image data. By statistically analyzing and mapping the pixel values ​​of an image, it compresses multi-bit pixel values ​​into fewer bits, thereby reducing image file size and image quality. Currently, existing image quantization methods have been applied in various fields, achieving significant success, particularly in computer vision tasks such as image recognition, detection, and retrieval. However, in image retrieval, especially in real-world applications involving new domains and categories, the performance of these methods is less than ideal. This is because traditional image quantization methods require continuous annotation of new data and model retraining when processing images from unknown domains or categories, leading to high resource consumption and necessitating further improvement. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide a general image retrieval method based on image quantization. It uses first-order and second-order codebooks to generate quantization results that are consistent across domains and perceptive across categories, thereby solving the cross-domain and cross-category challenges in image retrieval tasks and significantly improving the accuracy and efficiency of image retrieval.

[0004] To achieve the above-mentioned objectives, the general image retrieval method based on image quantization of the present invention includes the following steps:

[0005] S1: Construct a training sample set X according to actual needs, including images of N types of items and their corresponding category texts;

[0006] S2: Construct an image quantization model, including a visual feature extraction module, a first-order quantization cross-domain module, a combined feature extraction module, and a second-order quantization cross-class module, wherein:

[0007] The visual feature extraction module is used to extract the visual embedding features of the input image x. L represents the preset feature dimension, and the visual embedding feature Z is sent to the first-order quantization cross-domain module and the combined feature extraction module;

[0008] The first-order quantization cross-domain module is used to compute the visual embedding feature Z and the first-order codebook. Each first-order code word The similarity is used to quantize the visual embedding features, with the first-order code word having the highest similarity.

[0009] The combined feature extraction module extracts the combined feature representation ζ of the input image x and sends it to the second-order quantization module. The combined feature extraction module includes an attention module, a feature fusion module, a linear layer, and a feedforward neural network, wherein:

[0010] The attention module is used to treat the visual embedding feature Z as the query in the attention mechanism, and the second-order codebook subset is used. As the key and value in the attention mechanism, K represents the number of item types contained in the second-order codebook subset. The attention mechanism is used to extract the image to obtain the attention vector v of the input image and send it to the feature fusion module.

[0011] The feature fusion module is used to superimpose the attention vector v onto the visual embedding feature Z to obtain the feature z = Z + v and send it to the linear layer;

[0012] The linear layer is used to perform a linear transformation on the feature z to obtain the feature LN(z) and send it to the feedforward neural network;

[0013] A feedforward neural network is used to process the feature LN(z) to obtain the combined feature representation ζ of the input image x;

[0014] The second-order quantization cross-class module is used to calculate the visual embedding feature Z and the second-order codebook. Each second-order code word E i The similarity is used to determine the code word E, where i = 1, 2, ..., N. i weight w i The greater the similarity, the higher the weight, and then the combined features are obtained.

[0015] S3: Randomly initialize N code words D with dimension L. i ′, for N code words D i Normalization is performed to obtain the code word D. i This constitutes the initialization first-order codebook.

[0016] S4: Uniformly sample the current batch of training samples X from the different classes of the training sample set X. B Let each training sample be x. i,j , i = 1, 2, ..., N, j = 1, 2, ..., M, where M represents the number of samples in each category;

[0017] S5: Subset X of training samples B Each training sample x i,j The input image quantization model extracts visual embedding features Z from the visual feature extraction module and the first-order quantization module. i,j And quantized visual embedding features are

[0018] S6: Quantize the visual embedding features using N×M training samples. The second-order codebook Ω is calculated based on the first-order codebook Ω. (2) Second-order codebook Ω (2) Each second-order code word E i The calculation formula is:

[0019]

[0020] S7: For the training sample subset X B Each training sample x i,j From the second-order codebook Ω (2) Filter the subset of second-order codebooks that do not contain their actual item types. Visual embedding features Z i,j and a subset of second-order codebooks The combined feature extraction module in the input image quantization model obtains the combined feature ζ. i,j Then input the second-order quantization module based on the second-order codebook Ω (2) Obtain quantized combination features

[0021] S8: Calculate the loss function and then update the parameters of the image quantization model and the first-order codebook;

[0022] S9: Determine whether the training termination condition has been met. If it has, the training ends; otherwise, return to step S4.

[0023] S10: Transform the entire second-order codebook Ω (2) As a subset of the second-order codebook Then, each image in the image library is input into the image quantization model, based on the trained first-order codebook Ω and second-order codebook Ω. (2) and a subset of second-order codebooks Obtain the combined features of each image;

[0024] S11: Transform the entire second-order codebook Ω (2) As a subset of the second-order codebook Then, the image to be retrieved is input into the image quantization model, based on the first-order codebook Ω and the second-order codebook Ω obtained during training. (2) and a subset of second-order codebooks The combined features are obtained, and then the similarity between the combined features and the combined features of each image in the image database is calculated. The R images with the highest similarity are selected as the search results and returned. The value of R is set according to actual needs.

[0025] This invention presents a general image retrieval method based on image quantization. First, a training sample set is constructed. Then, an image quantization model is built, comprising a visual feature extraction module, a first-order quantization cross-domain module, a combined feature extraction module, and a second-order quantization cross-category module. During training, a first-order codebook is randomly initialized. Then, cross-domain consistency features are learned through a contrastive learning mechanism using the visual feature extraction module and the first-order codebook. A second-order codebook is updated based on the first-order codebook. New category-aware features are obtained through the combined feature extraction module and the second-order codebook. Finally, quantized combined features are obtained by the second-order quantization cross-category module. During retrieval, each image in the search image library is input into the image quantization model to obtain the combined features of each image. The image to be retrieved is also input into the image quantization model to obtain its combined features. Then, the similarity between the combined features and the combined features of each image in the image library is calculated to obtain the retrieval results.

[0026] The present invention has the following beneficial effects:

[0027] 1) This invention designs a second-order codebook containing representative information of known categories. New images are extracted by combining and projecting all second-order codewords, which improves the generalization ability for unknown categories.

[0028] 2) This invention proposes a cross-alignment contrastive learning objective in model training to simultaneously reduce quantization error and inter-domain differences, encourage the model to generate domain-invariant, new category-aware quantized code words, thereby improving retrieval performance. Attached Figure Description

[0029] Figure 1 This is a flowchart illustrating a specific implementation of the general image retrieval method based on image quantization according to the present invention.

[0030] Figure 2 This is a structural diagram of the image retrieval model in this invention;

[0031] Figure 3 This is an example diagram of the image quantization process in this invention;

[0032] Figure 4 This is a schematic diagram of multiple codebooks in this embodiment. Detailed Implementation

[0033] The specific embodiments of the present invention will now be described with reference to the accompanying drawings to enable those skilled in the art to better understand the invention. It should be particularly noted that in the following description, detailed descriptions of known functions and designs that might obscure the main content of the invention will be omitted here.

[0034] Example

[0035] Figure 1This is a flowchart illustrating a specific implementation of the general image retrieval method based on image quantization according to the present invention.

[0036] like Figure 1 As shown, the specific steps of the general image retrieval method based on image quantization of the present invention include:

[0037] S101: Obtain the training sample set:

[0038] Construct a training sample set X according to actual needs, including images of N types of items and their corresponding category labels.

[0039] S102: Constructing an image quantization model:

[0040] To achieve image retrieval, this invention constructs an image quantization model. Figure 2 This is a structural diagram of the image retrieval model in this invention. For example... Figure 2 As shown, the image quantization model in this invention includes a visual feature extraction module, a first-order quantization cross-domain module, a combined feature extraction module, and a second-order quantization cross-category module. Each module will be described in detail below.

[0041] The visual feature extraction module is used to extract the visual embedding features of the input image x. L represents the preset feature dimension, and the visual embedding feature Z is sent to the first-order quantization cross-domain module and the combined feature extraction module. In this embodiment, the visual feature extraction module includes a visual Transformer neural network, a fully connected layer, and a normalization module, wherein:

[0042] The visual Transformer neural network is used to extract the raw features of an image and output them to a fully connected layer;

[0043] Fully connected layers are used to process the original features to obtain initial visual embedding features. And output to the normalization module;

[0044] The normalization module is used to normalize the initial visual embedding features Z′ to obtain the visual embedding features. The specific method of normalization can be set according to actual needs. In this embodiment, max-min normalization is used.

[0045] The first-order quantization cross-domain module is used to compute the visual embedding feature Z and the first-order codebook. Each first-order code word The similarity is calculated, i = 1, 2, ..., N, and the first-order code word with the highest similarity is used as the quantized visual embedding feature.

[0046] The combined feature extraction module extracts the combined feature representation ζ of the input image x and sends it to the second-order quantization module. The combined feature extraction module includes an attention module, a feature fusion module, a linear layer, and a feedforward neural network, where:

[0047] The attention module is used to treat the visual embedding feature Z as the query in the attention mechanism, and the second-order codebook subset is used. In the attention mechanism, K represents the number of item types in the second-order codebook subset, acting as both the key and value. The attention mechanism extracts the attention vector v from the input image and sends it to the feature fusion module. The attention mechanism is a commonly used image processing mechanism, and its specific process will not be elaborated here. In this invention, the attention mechanism can be represented as:

[0048]

[0049] Where σ() represents the activation function, F Q (), F K (), F V () represent the linear projection functions for query, key, and value, respectively.

[0050] The feature fusion module is used to superimpose the attention vector v onto the visual embedded feature Z to obtain the feature z = Z + v and send it to the linear layer.

[0051] The linear layer is used to perform a linear transformation on the feature z to obtain the feature LN(z) and send it to the feedforward neural network.

[0052] A feedforward neural network is used to process the feature LN(z) to obtain the combined feature representation ζ of the input image x.

[0053] It can be seen that the combined feature representation ζ can be expressed as:

[0054] ζ=FFN(LN(Z+v))

[0055] Where LN() represents a linear layer and FFN() represents a feedforward neural network.

[0056] The second-order quantization cross-class module is used to calculate the visual embedding feature Z and the second-order codebook. Each second-order code word E i The similarity is used to determine the code word E, where i = 1, 2, ..., N. i weight w i The greater the similarity, the higher the weight, and then the combined features are obtained.

[0057] As can be seen from the above process, the image quantization model in this invention takes the input image as an "unknown class" sample, and then performs attention calculation on the visual embedding features and the second-order codebook containing the known class to obtain the quantized combination features of the "unknown class" sample on the known class, thereby realizing image quantization. Figure 2 This is an example diagram of the image quantization process in this invention. Figure 2 This invention demonstrates the projection process from known class samples to a first-order codebook and then to a second-order codebook, as well as the generation of quantified combination features from unseen class samples through weighted combination of second-order codebooks. It is evident that this invention improves the ability to identify unknown class samples by weighted combination of code words from known categories.

[0058] S103: Initialization Codebook

[0059] Randomly initialize N code words D with dimension L i ′, for N code words D i Normalization is performed to obtain the code word D. i This constitutes the initialization first-order codebook.

[0060] S104: Select training samples for the current batch:

[0061] The current batch of training sample subsets X is obtained by uniformly sampling from different categories of the training sample set X. B Let each training sample be x. i,j Let i = 1, 2, ..., N, j = 1, 2, ..., M, where M represents the number of samples in each category. Therefore, the total number of images in a single batch is N × M.

[0062] S105: Extracting quantized visual embedding features:

[0063] Subset of training samples X B Each training sample x i,j The input image quantization model extracts visual embedding features Z from the visual feature extraction module and the first-order quantization module. i,j And quantized visual embedding features are

[0064] S106: Computing the second-order codebook:

[0065] Quantized visual embedding features using N×M training samples The second-order codebook Ω is calculated based on the first-order codebook Ω. (2) Second-order codebook Ω (2) Each second-order code word E i The calculation formula is:

[0066]

[0067] S107: Extracting quantized combined features:

[0068] For the training sample subset X B Each training sample x i,j From the second-order codebook Ω (2) Filter the subset of second-order codebooks that do not contain their actual item types. Visual embedding features Z i,j and a subset of second-order codebooks The combined feature extraction module in the input image quantization model obtains the combined feature ζ. i,j Then input the second-order quantization module based on the second-order codebook Ω (2) Obtain quantized combination features

[0069] S108: Update image quantization model parameters:

[0070] Calculate the loss function and then update the parameters of the image quantization model and the first-order codebook.

[0071] In this embodiment, to improve the training effect of the image quantization model, the loss function is calculated by comprehensively considering embedded features and combined features. The specific method is as follows:

[0072] The cross-domain loss L of the embedded features is calculated using the following formula. Z :

[0073]

[0074] Where o() represents the contrastive learning loss, <> represents the similarity calculation, e represents the natural constant, and τ represents the temperature coefficient.

[0075] The cross-class loss L of the combined features is calculated using the following formula. ζ :

[0076]

[0077] Then, the loss function L is calculated using the following formula:

[0078] L = L Z +λL ζ

[0079] Where λ represents the weight.

[0080] S109: Determine if the training termination condition has been met. If it has, training ends and proceed to step S110; otherwise, return to step S104. The training termination condition can be set according to actual needs, typically by setting it to the preset maximum number of training iterations or the convergence of model parameters.

[0081] After multiple rounds of training, the image quantization model acquires domain invariance and class-aware capabilities, thereby enabling image retrieval.

[0082] S110: Extract combined features from the image library:

[0083] The entire second-order codebook Ω (2) As a subset of the second-order codebook Then, each image in the image library is input into the image quantization model, based on the trained first-order codebook Ω and second-order codebook Ω. (2) and a subset of second-order codebooks The combined features of each image are obtained.

[0084] S111: Image Retrieval

[0085] The entire second-order codebook Ω (2) As a subset of the second-order codebook Then, the image to be retrieved is input into the image quantization model, based on the first-order codebook Ω and the second-order codebook Ω obtained during training. (2) and a subset of second-order codebooks The combined features are obtained, and then the similarity between the combined features and the combined features of each image in the image database is calculated. The R images with the highest similarity are selected as the search results and returned. The value of R is set according to actual needs.

[0086] In practical applications, when the image library contains too many image categories, the second-order codebook may become too large, reducing retrieval efficiency. To address this, a multiple codebook approach can be used. Figure 3 This is a schematic diagram of multiple codebooks in this embodiment. For example... Figure 3 As shown, multiple codebooks refer to the segmentation of the codebook, specifically dividing the original L-dimensional second-order codebook into Q blocks. During the image quantization model operation, the visual embedding features Z extracted by the visual feature extraction module are also segmented into Q blocks, corresponding one-to-one with the segmented codebooks. Each sub-block of the visual embedding features Z undergoes independent quantization learning within its corresponding sub-codebook. Finally, the final quantized combined features are obtained by concatenating the quantized combined features of multiple sub-blocks.

[0087] To better illustrate the technical effects of the present invention, specific examples are used to experimentally verify the present invention.

[0088] Example 1

[0089] The experimental conditions set in this embodiment are as follows: System: Ubuntu 20.04, Software: Python 3.8, Processor: CPU E5-2678v3@2.50GHz×2, Memory: 256GB.

[0090] To more accurately evaluate the effectiveness of this invention in image retrieval tasks, this embodiment conducts image retrieval experiments on the DomainNet dataset. Using SE-ResNet50 and ViT-Base as the base models, this invention is compared with several state-of-the-art methods. The comparison methods include:

[0091] For details on DPgQ, please refer to the literature "Gao L, Zhu X, Song J, et al. Beyond product quantization: Deep progressive quantization for image retrieval[J]. arXiv preprint arXiv:1906.06698,2019.";

[0092] For details on SnMpQ, please refer to the literature "Paul S, Dutta T, Biswas S. Universal cross-domain retrieval: Generalizing across classes and domains[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision.2021:12056-12064.";

[0093] For details on CLIP, please refer to the literature "Learning transferable visual models from natural language supervision".

[0094] Table 1 is a comparison table of the image retrieval performance of the present invention and the comparative method in Example 1 on multi-domain average and bit length.

[0095]

[0096]

[0097] Table 1

[0098] As shown in Table 1, the present invention achieves the best mean accuracy (mAP) across all bit lengths, reaching a maximum of 40.1%. Notably, the present invention performs particularly well in retrieval tasks involving unknown categories and unknown domains, improving mAP by up to 8.0% in retrieval tasks involving unknown categories.

[0099] Furthermore, to further verify the performance of this invention in complex scenarios containing distractors, experimental results show that even when the search set contains distractors of known categories, this invention can still significantly improve retrieval accuracy, with mAP increasing by up to 6.9%. In contrast, current state-of-the-art methods perform poorly when handling scenarios containing distractors, demonstrating the robustness and accuracy of this invention in complex retrieval tasks.

[0100] Example 2

[0101] To more accurately evaluate the effectiveness of this invention in a continuous learning environment, this embodiment performs sequence-based image classification learning on the SketchyExtended dataset. Using SE-ResNet50 and ViT-Base as the original models, this invention is compared with several existing methods, including DPgQ, SnMpQ, and CLIP, as well as the following:

[0102] For details on CSQ, please refer to the literature "Yuan L, Wang T, Zhang X, et al. Central similarity quantization for efficient image and video retrieval[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition.2020:3083-3092.";

[0103] For details on GrdHash, please refer to the literature "Su S, Zhang C, Han K, et al. Greedy hash: Towardsfast optimization for accurate hash coding in cnn[J]. Advances in neural information processing systems, 2018, 31."

[0104] For details on DSH, please refer to the literature "Liu L, Shen F, Shen Y, et al. Deep sketch hashing: Fastfree-hand sketch-based image retrieval[C] / / Proceedings of the IEEE conferenceon computer vision and pattern recognition.2017:2862-2871.";

[0105] PSKD, for details, please refer to the document "Wang K, Wang Y, Xu

[0106] For details on TCN, please refer to the literature "Wang H, Deng C, Liu T, et al. Transferable coupled network for zero-shot sketch-based image retrieval[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2021,44(12):9181-9194.";

[0107] Table 2 is a comparison table of the image classification performance of the present invention and the comparative method in Example 2.

[0108] CSQ 41.3 45.9 48.0 - GrdHash 42.3 47.2 48.9 - DSH 40.7 46.1 47.7 - DPgQ 43.2 48.4 51.0 53.7 SnMpQ 40.4 45.1 50.2 53.1 PSKD 43.3 49.1 51.6 54.4 TCN 43.3 49.0 51.6 54.2 CLIP 50.2 52.8 54.5 56.9 This invention (Hash) 46.9 48.8 52.5 - This invention 49.8 52.1 53.4 53.9 This invention (ViT) 58.2 59.0 59.3 60.1

[0109] Table 2

[0110] As shown in Table 2, this invention achieved the best mean accuracy (mAP) across all bit lengths, reaching a maximum of 59.3%. Notably, this invention performs particularly well in retrieval tasks involving unknown categories, improving mAP by up to 8.0%, demonstrating its strong generalization ability. Furthermore, this invention significantly improves retrieval accuracy even when handling complex scenarios containing distractors, proving its robustness and accuracy in complex retrieval tasks.

[0111] By comparing the performance of different methods at various bit lengths, the overall advantages of this invention in different retrieval tasks can be seen. In particular, compared with the most advanced methods currently available, this invention performs exceptionally well when handling unknown categories and domains.

[0112] Through the above data analysis and comparison, it can be seen that the present invention is superior to other quantization-based image retrieval methods. These results verify the effectiveness and superiority of the present invention.

[0113] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.

Claims

1. A general image retrieval method based on image quantization, characterized in that, Includes the following steps: S1: Construct a training sample set X according to actual needs, including images of N types of items and their corresponding category texts; S2: Construct an image quantization model, including a visual feature extraction module, a first-order quantization cross-domain module, a combined feature extraction module, and a second-order quantization cross-class module, wherein: The visual feature extraction module is used to extract the visual embedding features of the input image x. L represents the preset feature dimension, and the visual embedding feature Z is sent to the first-order quantization cross-domain module and the combined feature extraction module; The first-order quantization cross-domain module is used to compute the visual embedding feature Z and the first-order codebook. Each first-order code word The similarity is used to quantize the visual embedding features, with the first-order code word having the highest similarity. The combined feature extraction module extracts the combined feature representation ζ of the input image x and sends it to the second-order quantization module. The combined feature extraction module includes an attention module, a feature fusion module, a linear layer, and a feedforward neural network, where: The attention module is used to treat the visual embedding feature Z as the query in the attention mechanism, and the second-order codebook subset is used. As the key and value in the attention mechanism, K represents the number of item types contained in the second-order codebook subset. The attention mechanism is used to extract the attention vector v of the input image and send it to the feature fusion module. The feature fusion module is used to superimpose the attention vector v onto the visual embedding feature Z to obtain the feature z = Z + v and send it to the linear layer; The linear layer is used to perform a linear transformation on the feature z to obtain the feature LN(z) and send it to the feedforward neural network; A feedforward neural network is used to process the feature LN(z) to obtain the combined feature representation ζ of the input image x; The second-order quantization cross-class module is used to calculate the visual embedding feature Z and the second-order codebook. Each second-order code word E i The similarity is used to determine the code word E, where i = 1, 2, ..., N. i weight w i The greater the similarity, the higher the weight, and then the combined features are obtained. S3: Randomly initialize N code words D with dimension L. i ′, for N code words D i Normalization is performed to obtain the code word D. i This constitutes the initialization first-order codebook. S4: Uniformly sample the current batch of training samples X from the different classes of the training sample set X. B Let each training sample be x. i,j , i = 1, 2, ..., N, j = 1, 2, ..., M, where M represents the number of samples in each category; S5: Subset X of training samples B Each training sample x i,j The input image quantization model extracts visual embedding features Z from the visual feature extraction module and the first-order quantization module. i,j And quantized visual embedding features are S6: Quantize the visual embedding features using N×M training samples. The second-order codebook Ω is calculated based on the first-order codebook Ω. (2) Second-order codebook Ω (2) Each second-order code word E i The calculation formula is: S7: For the training sample subset X B Each training sample x i,j From the second-order codebook Ω (2) Filter the subset of second-order codebooks that do not contain their actual item types. Visual embedding features Z i,j and a subset of second-order codebooks The combined feature extraction module in the input image quantization model obtains the combined feature ζ. i,j Then input the second-order quantization module based on the second-order codebook Ω (2) Obtain quantized combination features S8: Calculate the loss function and then update the parameters of the image quantization model and the first-order codebook; S9: Determine whether the training termination condition has been met. If it has, the training ends; otherwise, return to step S4. S10: Transform the entire second-order codebook Ω (2) As a subset of the second-order codebook Then, each image in the image library is input into the image quantization model, based on the trained first-order codebook Ω and second-order codebook Ω. (2) and a subset of second-order codebooks Obtain the combined features of each image; S11: Transform the entire second-order codebook Ω (2) As a subset of the second-order codebook Then, the image to be retrieved is input into the image quantization model, based on the first-order codebook Ω and the second-order codebook Ω obtained during training. (2) and a subset of second-order codebooks The combined features are obtained, and then the similarity between the combined features and the combined features of each image in the image database is calculated. The R images with the highest similarity are selected as the search results and returned. The value of R is set according to actual needs.

2. The general image retrieval method according to claim 1, characterized in that, The visual feature extraction module in step S2 includes a visual Transformer neural network, a fully connected layer, and a normalization module, wherein: The visual Transformer neural network is used to extract the raw features of an image and output them to a fully connected layer; Fully connected layers are used to process the original features to obtain initial visual embedding features. And output to the normalization module; The normalization module is used to normalize the initial visual embedding features Z′ to obtain the visual embedding features.

3. The general image retrieval method according to claim 1, characterized in that, The method for calculating the loss function in step S8 is as follows: The cross-domain loss L of the embedded features is calculated using the following formula. Z : Where o() represents the contrastive learning loss, <> represents the similarity calculation, e represents the natural constant, and τ represents the temperature coefficient; The cross-class loss L of the combined features is calculated using the following formula. ζ : Then, the loss function L is calculated using the following formula: L=L Z +λL ζ Where λ represents the weight.

4. The general image retrieval method according to claim 1, characterized in that, The second-order codebook is divided into Q blocks. During the image quantization model operation, the visual embedding features Z extracted by the visual feature extraction module are also divided into Q blocks, which correspond one-to-one with the divided codebook. Each sub-block of the visual embedding feature Z is independently quantized and learned in its corresponding sub-codebook. Finally, the final quantized combined feature is obtained by splicing the quantized combined features of multiple sub-blocks.

Citation Information

Patent Citations

  • Data retrieval method based on Hash learning of dimension analysis quantizer

    CN112241475A

  • Method for acquiring image retrieval model, image retrieval method, device and equipment

    CN114299306A