Image hash retrieval method based on category center

Through the interaction between local features of the image and the category center, fine-grained features are extracted, which solves the problem that global features cannot capture subtle differences in the prior art, and significantly improves the accuracy of image retrieval.

CN120030181APending Publication Date: 2025-05-23UNIT 32002 OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411995486.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing supervised deep hashing technology relies on global features and is unable to effectively capture subtle differences and fine-grained features in images, resulting in a decrease in retrieval accuracy in multi-category image retrieval.

Method used

The image hash retrieval method based on category center is used to extract the specific position information that each category center focuses on through the interaction between the local features of the image and the category center, and weight it according to the position importance to generate a fine-grained hash representation.

Benefits of technology

It significantly improves the image retrieval accuracy in the fine-grained retrieval task, can deeply explore the fine-grained image content related to categories, and enhances the expression ability of hash code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030181A_ABST
    Figure CN120030181A_ABST
Patent Text Reader

Abstract

The invention discloses an image hash retrieval method based on a category center, and belongs to the technical field of image retrieval. The method comprises the following steps: acquiring an image of a to-be-generated hash code, and extracting a high-dimensional feature vector of the image; generating a first hash code corresponding to the high-dimensional feature vector; inputting the first hash code into the trained attention module; determining a plurality of categories corresponding to the image; inputting the features corresponding to all the category-concerned spatial positions corresponding to the image into a trained multi-category information fusion module, fusing the features corresponding to all the category-concerned spatial positions corresponding to the image, and inputting the fused features into a trained hash code generation module; generating a second hash code corresponding to the image; storing the second hash code in a hash code database; and querying the hash code database according to an input query instruction. According to the method, the image retrieval accuracy in the fine-grained retrieval task is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of image retrieval, and in particular relates to an image hash retrieval method based on category center. Background Art

[0002] With the rapid growth of digital image data, hash-based retrieval methods have been widely used in the field of large-scale image retrieval due to their efficient storage and computing performance. Hashing technology significantly reduces storage space and speeds up retrieval by mapping high-dimensional image features to low-dimensional hash codes. Deep hashing methods have significant advantages over traditional hashing methods. They can extract high-dimensional deep features, achieve end-to-end training, have good flexibility and adaptability, and can easily meet the needs of large-scale data sets.

[0003] Supervised deep hashing methods combine the feature extraction capabilities of deep learning with the efficiency of hash coding, and use labeled training data to learn more discriminative hash codes. These methods usually extract global features of images through deep neural networks and optimize hash codes through specific loss functions, so that the hash codes of similar images are as close as possible, while the hash codes of dissimilar images are as far apart as possible.

[0004] Existing supervised deep hashing technology mainly relies on global features extracted by deep neural networks to learn hash codes. This method improves the efficiency of image retrieval to a certain extent, but it also has obvious shortcomings. Global features only represent the overall information of the image and cannot effectively capture subtle differences and fine-grained features in the image. When an image contains multiple categories, a single global feature may cause information loss or confusion, thereby affecting retrieval accuracy. In addition, hash codes generated based on global information cannot fuse information from multiple categories, resulting in a decrease in the discriminative ability of hash codes when processing multi-category images. These factors limit the application effect of existing methods in complex scenes and fine-grained retrieval tasks. Summary of the invention

[0005] The present invention provides an image hash retrieval method based on category center, which is used to solve the above problem or at least partially solve the above problem.

[0006] In a first aspect, a category center-based image hash retrieval method is disclosed, the method comprising:

[0007] Step S1: Obtain an image for which a hash code is to be generated, and extract a high-dimensional feature vector of the image;

[0008] Step S2: inputting the high-dimensional feature vector into a trained hash code generation module to generate a first hash code corresponding to the high-dimensional feature vector;

[0009] Step S3: inputting the first hash code into the trained attention module; determining multiple categories corresponding to the image;

[0010] For each category: determining a category center vector corresponding to the category, outputting a spatial attention coefficient of the category corresponding to the image based on the category center vector and the first hash code, and generating a feature corresponding to the spatial position of the category;

[0011] Step S4: inputting the features corresponding to the spatial positions of all the categories concerned corresponding to the image into the trained multi-category information fusion module, the multi-category information fusion module fuses the features corresponding to the spatial positions of all the categories concerned corresponding to the image, and inputting the fused features into the trained hash code generation module to generate a second hash code corresponding to the image;

[0012] Step S5: storing the second hash code in a hash code database; querying the hash code database according to the input query instruction.

[0013] Preferably, the step S3: inputting the first hash code into a trained attention module; determining multiple categories corresponding to the image;

[0014] For each category: determine the category center vector corresponding to the category, output the spatial attention coefficient of the category corresponding to the image based on the category center vector and the first hash code, and generate features corresponding to the spatial position of the category attention, wherein:

[0015] The first hash code is input into the trained attention module, and the first hash code is recorded as G, G = Z(F)∈(-1,+1) H×W×K , where F is the high-dimensional feature vector of the image, H, W and D represent the height, width and dimension of the high-dimensional feature vector of the image, respectively; Z(·) is a hash function; K represents the length of the first hash code; represents the field of real numbers;

[0016] Pre-setting a plurality of categories and determining a plurality of categories corresponding to the image;

[0017] For each corresponding category:

[0018] Determine a category center vector corresponding to the category, wherein the category center vector is determined by mapping a number of preset samples to [-1, 1] based on the distribution state of the preset samples, and the category center vector serves as a representative of the corresponding category;

[0019] Get the category center vector P∈[-1,+1] of the category to which the image belongs C×K , C represents the number of all categories corresponding to the image;

[0020] Output the spatial attention coefficient A of the category to which the image belongs,

[0021] A=softmax(P⊙G)∈(0,1) C×H×W

[0022] Generate the feature J corresponding to the spatial position of interest to this category, is the inner product symbol.

[0023] Preferably, in step S4, wherein:

[0024] The fusion formula is

[0025] F fusion is the fused feature;

[0026] Generate a second hash code Q, Q = Z (F fusion ).

[0027] Preferably, extracting the high-dimensional feature vector of the image includes: using a convolutional neural network to extract local features and high-dimensional features at different positions of the image, and generating the high-dimensional feature vector of the image based on the local features and the high-dimensional features.

[0028] Preferably, the hash code generation module includes an optimized hash function, the expression of the hash function is Z(x)=tanh(x), x is the input, and Z(x) is the generated hash code.

[0029] In a second aspect, a category center-based image hash retrieval device is disclosed, the device comprising:

[0030] Feature extraction module: configured to obtain an image to be used for generating a hash code, and extract a high-dimensional feature vector of the image;

[0031] A first hash code generation module: configured to input the high-dimensional feature vector into the trained hash code generation module to generate a first hash code corresponding to the high-dimensional feature vector;

[0032] A second feature acquisition module: configured to input the first hash code into the trained attention module; determine multiple categories corresponding to the image;

[0033] For each category: determining a category center vector corresponding to the category, outputting a spatial attention coefficient of the category corresponding to the image based on the category center vector and the first hash code, and generating a feature corresponding to the spatial position of the category;

[0034] A second hash code generation module: configured to input the features corresponding to the spatial positions of all the categories concerned corresponding to the image into a trained multi-category information fusion module, wherein the multi-category information fusion module fuses the features corresponding to the spatial positions of all the categories concerned corresponding to the image, and inputs the fused features into the trained hash code generation module to generate a second hash code corresponding to the image;

[0035] Query module: configured to store the second hash code in a hash code database; and query the hash code database according to an input query instruction.

[0036] In a third aspect, an electronic device is disclosed, the electronic device comprising:

[0037] at least one processor; and

[0038] a memory communicatively connected to the at least one processor; wherein,

[0039] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0040] In a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is disclosed, wherein the computer instructions are used to cause the computer to execute the method as described above.

[0041] The present invention has the following technical effects:

[0042] The present invention enhances the fine-grained image representation capability of hash codes through the interaction between local features of the image and the category center. The method extracts the specific location information of each category center through the interaction between the local features of the image and the category center vector, and applies different weights according to the importance of each location, extracts different features for each category, and finally generates a fine-grained hash representation. This invention can deeply explore the fine-grained image content related to the category and significantly improve the image retrieval accuracy in fine-grained retrieval tasks.

[0043] This application uses a fine-grained deep image hash retrieval method based on category centers, using local features at different locations extracted by convolutional neural networks and interacting with category centers to generate fine-grained hash representations. Since this technical means can deeply capture subtle features related to categories and apply different weights according to the importance of positions, thereby enhancing the expressive power of hash codes, it solves the problem that traditional hash methods cannot effectively distinguish fine-grained features in multi-class labeled image retrieval, and significantly improves the image retrieval accuracy in fine-grained retrieval tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flowchart of the image hash retrieval method based on category center;

[0045] Figure 2 The schematic diagram of the architecture of the image hash retrieval method based on category center;

[0046] Figure 3 It is a structural diagram of an image hash retrieval device based on category center. DETAILED DESCRIPTION

[0047] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0048] like Figure 1-Figure 2 As shown, the present invention provides an image hash retrieval method based on category center, the method comprising:

[0049] Step S1: Obtain an image for which a hash code is to be generated, and extract a high-dimensional feature vector of the image;

[0050] Step S2: inputting the high-dimensional feature vector into a trained hash code generation module to generate a first hash code corresponding to the high-dimensional feature vector;

[0051] Step S3: inputting the first hash code into the trained attention module; determining multiple categories corresponding to the image;

[0052] For each category: determining a category center vector corresponding to the category, outputting a spatial attention coefficient of the category corresponding to the image based on the category center vector and the first hash code, and generating a feature corresponding to the spatial position of the category;

[0053] Step S4: inputting the features corresponding to the spatial positions of all the categories concerned corresponding to the image into the trained multi-category information fusion module, the multi-category information fusion module fuses the features corresponding to the spatial positions of all the categories concerned corresponding to the image, and inputting the fused features into the trained hash code generation module to generate a second hash code corresponding to the image;

[0054] Step S5: storing the second hash code in a hash code database; querying the hash code database according to the input query instruction.

[0055] The present invention utilizes a convolutional neural network to extract local features at different positions of an image; a first hash code generation module is used to generate hash features of lower dimensions; an attention module is used to interact with hash features of different category centers and different spatial positions, generate a spatial attention coefficient of features for each category and weight them to obtain features of each category; a multi-category information fusion module is responsible for fusing information of multiple categories to obtain fused global features.

[0056] The extracting of the high-dimensional feature vector of the image includes: using a convolutional neural network to extract local features and high-dimensional features at different positions of the image, and generating the high-dimensional feature vector of the image based on the local features and the high-dimensional features.

[0057] In the present invention, a convolutional neural network (CNN) is used to extract local features at different positions of the image. By selecting a suitable network architecture (such as ResNet or VGG) and using a pre-trained model to extract high-dimensional features at different positions. This process lays the foundation for subsequent hash code generation and the interaction between features and category centers, ensuring that the extracted features can fully reflect the fine-grained information of the image.

[0058] In step S2, the hash code generation module includes an optimized hash function, the expression of the hash function is Z(x)=tanh(x), x is the input, and Z(x) is the generated hash code.

[0059] In the present invention, the hash function reduces the dimension of the input high-dimensional feature vector to obtain a first hash code of a specified length, the first hash code is responsible for interacting with each category center, and does not calculate the similarity between images. The cross entropy loss acts on the second hash code and the category center.

[0060] Step S3: inputting the first hash code into the trained attention module; determining multiple categories corresponding to the image;

[0061] For each category: determine the category center vector corresponding to the category, output the spatial attention coefficient of the category corresponding to the image based on the category center vector and the first hash code, and generate features corresponding to the spatial position of the category attention, wherein:

[0062] The first hash code is input into the trained attention module, and the first hash code is recorded as G, G = Z(F)∈(-1,+1) H×W×K , where F is the high-dimensional feature vector of the image, H, W and D represent the height, width and dimension of the high-dimensional feature vector of the image, respectively; Z(·) is the hash function; K represents the length of the first hash code; represents the field of real numbers;

[0063] Pre-setting a plurality of categories and determining a plurality of categories corresponding to the image;

[0064] For each corresponding category:

[0065] Determine a category center vector corresponding to the category, wherein the category center vector is determined by mapping a number of preset samples to [-1, 1] based on the distribution state of the preset samples, and the category center vector serves as a representative of the corresponding category;

[0066] Get the category center vector P∈[-1,+1] of the category to which the image belongs C×K , C represents the number of all categories corresponding to the image;

[0067] Output the spatial attention coefficient A of the category to which the image belongs,

[0068] A=softmax(P⊙G)∈(0,1) C×H×W

[0069] Generate the feature J corresponding to the spatial position of interest to this category, is the inner product symbol.

[0070] In the present invention, the characterization ability of hash codes for fine-grained information is enhanced by interacting with local image features and category center vectors. First, the feature center of each category is calculated as the representative of the category. Then, the attention mechanism is used to generate a spatial attention coefficient for each category to emphasize the important areas related to the category in the image. By weighting local features, the importance of different positions can be highlighted, thereby extracting feature representations for each category. This process enhances the model's ability to capture fine-grained features by focusing on the specific position information that the category center is concerned about.

[0071] The step S4: inputting the features corresponding to the spatial positions of all the categories concerned corresponding to the image into the trained multi-category information fusion module, the multi-category information fusion module fuses the features corresponding to the spatial positions of all the categories concerned corresponding to the image, and inputting the fused features into the trained hash code generation module to generate a second hash code corresponding to the image, wherein:

[0072] The fusion formula is

[0073] F fusion is the fused feature;

[0074] Generate a second hash code Q, Q = Z (F fusion ).

[0075] The first hash code generation module, the attention module, the multi-category information fusion module, and the second hash code generation module are jointly trained through a loss function.

[0076] In the present invention, the multi-category information fusion module is responsible for integrating the feature representations of different categories to generate the final global features. The module adopts fusion strategies such as weighted averaging and splicing to integrate features from different categories. In this way, the module can generate a global feature representation containing comprehensive information of each category, and then input the obtained new global feature into the hash generation module to obtain the final hash code. The loss function between the category center and the hash code realizes the multi-category constraint of the hash code by combining label smoothing and cross entropy loss. Specifically, by assigning the same probability value to each category, the representation ability of the hash code for multi-category images is learned, thereby improving the fine-grained retrieval performance of the image.

[0077] like Figure 3 As shown, the present invention provides an image hash retrieval device based on category center, the device comprising:

[0078] Feature extraction module: configured to obtain an image to be used for generating a hash code, and extract a high-dimensional feature vector of the image;

[0079] A first hash code generation module: configured to input the high-dimensional feature vector into the trained hash code generation module to generate a first hash code corresponding to the high-dimensional feature vector;

[0080] A second feature acquisition module: configured to input the first hash code into the trained attention module; determine multiple categories corresponding to the image;

[0081] For each category: determining a category center vector corresponding to the category, outputting a spatial attention coefficient of the category corresponding to the image based on the category center vector and the first hash code, and generating a feature corresponding to the spatial position of the category;

[0082] A second hash code generation module is configured to input the features corresponding to the spatial positions of all the categories concerned corresponding to the image into a trained multi-category information fusion module, the multi-category information fusion module fuses the features corresponding to the spatial positions of all the categories concerned corresponding to the image, and inputs the fused features into the trained hash code generation module to generate a second hash code corresponding to the image;

[0083] Query module: configured to store the second hash code in a hash code database; and query the hash code database according to an input query instruction.

[0084] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the technical solutions described in the above embodiments can still be modified, or some or all of the technical features therein can be replaced by equivalents, and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A category-center-based image hash retrieval method, characterized in that: The method comprises the following steps: Step S1: Obtain an image for which a hash code is to be generated, and extract a high-dimensional feature vector of the image; Step S2: inputting the high-dimensional feature vector into a trained hash code generation module to generate a first hash code corresponding to the high-dimensional feature vector; Step S3: inputting the first hash code into the trained attention module; Determining a plurality of categories corresponding to the image; For each category: determining a category center vector corresponding to the category, outputting a spatial attention coefficient of the category corresponding to the image based on the category center vector and the first hash code, and generating a feature corresponding to the spatial position of the category; Step S4: inputting the features corresponding to the spatial positions of all the categories concerned corresponding to the image into the trained multi-category information fusion module, the multi-category information fusion module fuses the features corresponding to the spatial positions of all the categories concerned corresponding to the image, and inputting the fused features into the trained hash code generation module to generate a second hash code corresponding to the image; Step S5: storing the second hash code in a hash code database; querying the hash code database according to the input query instruction.

2. The method according to claim 1, characterized in that The step S3: inputting the first hash code into the trained attention module; determining multiple categories corresponding to the image; For each category: determine the category center vector corresponding to the category, output the spatial attention coefficient of the category corresponding to the image based on the category center vector and the first hash code, and generate features corresponding to the spatial position of the category attention, wherein: The first hash code is input into the trained attention module, and the first hash code is recorded as G, G = Z(F)∈(-1,+1) H×W×K , where F is the high-dimensional feature vector of the image, H, W and D represent the height, width and dimension of the high-dimensional feature vector of the image, respectively; Z(·) is a hash function; K represents the length of the first hash code; represents the field of real numbers; Pre-setting a plurality of categories and determining a plurality of categories corresponding to the image; For each corresponding category: Determine a category center vector corresponding to the category, wherein the category center vector is determined by mapping a number of preset samples to [-1, 1] based on the distribution state of the preset samples, and the category center vector serves as a representative of the corresponding category; Get the category center vector P∈[-1,+1] of the category to which the image belongs C×K , C represents the number of all categories corresponding to the image; Output the spatial attention coefficient A of the category to which the image belongs, A=softmax(P⊙G)∈(0,1) C×H×W Generate the feature J corresponding to the spatial position of interest to this category, is the inner product symbol.

3. The method according to claim 2, characterized in that The step S4, wherein: The fusion formula is F fusion is the fused feature; Generate a second hash code Q, Q = Z (F fusion ).

4. The method according to claim 3, characterized in that The extracting of the high-dimensional feature vector of the image includes: using a convolutional neural network to extract local features and high-dimensional features at different positions of the image, and generating the high-dimensional feature vector of the image based on the local features and the high-dimensional features.

5. The method according to claim 3, characterized in that The hash code generation module includes an optimized hash function, the expression of the hash function is Z(x)=tanh(x), x is the input, and Z(x) is the generated hash code.

6. An image hash retrieval device based on category center, characterized in that: The device comprises: Feature extraction module: configured to obtain an image to be used for generating a hash code, and extract a high-dimensional feature vector of the image; A first hash code generation module: configured to input the high-dimensional feature vector into the trained hash code generation module to generate a first hash code corresponding to the high-dimensional feature vector; A second feature acquisition module: configured to input the first hash code into the trained attention module; determine multiple categories corresponding to the image; For each category: determining a category center vector corresponding to the category, outputting a spatial attention coefficient of the category corresponding to the image based on the category center vector and the first hash code, and generating a feature corresponding to the spatial position of the category; A second hash code generation module: configured to input the features corresponding to the spatial positions of all the categories concerned corresponding to the image into a trained multi-category information fusion module, wherein the multi-category information fusion module fuses the features corresponding to the spatial positions of all the categories concerned corresponding to the image, and inputs the fused features into the trained hash code generation module to generate a second hash code corresponding to the image; Query module: configured to store the second hash code in a hash code database; and query the hash code database according to an input query instruction.

7. A computer-readable storage medium, wherein a plurality of instructions are stored in the storage medium; the plurality of instructions are used for a processor to load and execute the method as claimed in any one of claims 1 to 5.

8. An electronic device, characterized in that: The electronic device comprises: A processor, which is used to execute multiple instructions; A memory for storing a plurality of instructions; The plurality of instructions are used to be stored in the memory and loaded and executed by the processor according to any one of claims 1 to 5.