Long-tail data hash code generation method and system based on mutual information
By introducing feature enhancement and mutual information optimization in the deep hashing method, the problem of poor retrieval effect of the deep hashing method on the long-tail data set is solved, and fast and accurate image retrieval in the long-tail data set is achieved.
Patent Information
- Application Number
- CN202510078411.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The search results of existing deep hashing methods on long-tail data sets are poor, mainly due to the difference in feature extraction capabilities of head classes and tail classes.
By inputting image data into a feature-enhanced feature extraction network, feature-enhanced image data features are generated and inputted to the hash activation layer to generate a binary hash code. At the same time, based on the mutual information between the category label and the image data features, the feature extraction network and hash activation layer are optimized, and the long-tail hash model that has been trained is obtained.
It improves the correlation between features and labels, reduces the influence of long-tail distribution, and can quickly and accurately search in image data that obeys long-tail distribution, and overcomes the impact of uneven number of different categories.
Smart Images

Figure CN120011582A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and more specifically, to a method and system for generating long-tail data hash codes based on mutual information. Background Art
[0002] With the development of the Internet and media, the amount of image data has exploded. We can obtain a large amount of image data from network devices and media every day. These images are huge in number, and most of the time users only need to use a part of these numerous image data, which may be a certain type of image or a few specific types of images. Therefore, how to efficiently and accurately retrieve specific content from such a huge image database according to user needs is one of the challenging tasks facing image retrieval.
[0003] As an important branch of image retrieval methods, deep hashing methods have received increasing attention. Image retrieval methods based on deep hashing use deep neural networks to extract features from input images, convert them into high-dimensional feature vectors, and map the vectors to low-dimensional binary hash codes through learned hash functions. This hash-based method has a fast retrieval speed and ensures the accuracy of retrieval results. However, image data in reality often presents a long-tail distribution, and the number of images in each category is unevenly distributed. In the data with a long-tail distribution, the category with a large number is called the head class, and the category with a small number is called the tail class. For the head class, due to the sufficient number of samples, the trained network has a strong feature extraction ability for these class data, and the retrieval effect is good; while for the tail class, due to the insufficient number of samples, the network cannot extract good features, resulting in poor retrieval results. On the entire data set, the interaction of different feature extraction capabilities of the head class and the tail class leads to poor retrieval results of existing deep hashing methods on long-tail data sets. Summary of the invention
[0004] In order to overcome the defect that the existing deep hashing method has poor retrieval results on long-tail data sets, the present invention provides a long-tail data hash code generation method and system based on mutual information.
[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0006] A method for generating a long-tail data hash code based on mutual information comprises the following steps:
[0007] Obtain image data and its category labels in the long-tail dataset;
[0008] Inputting the image data into a feature extraction network combined with feature enhancement to obtain image data features after feature enhancement;
[0009] Inputting the image data features into a hash activation layer to generate a corresponding binary hash code;
[0010] A loss function is constructed based on maximizing the mutual information between the category label and the image data features, which is used to optimize the feature extraction network and the hash activation layer, and a trained long-tail hash model is obtained to generate a long-tail data hash code.
[0011] Furthermore, the present invention also proposes a long-tail data hash code generation system based on mutual information, which applies the long-tail data hash code generation method of the present invention. The system includes:
[0012] Data collection module, used to obtain image data and its category labels in the long-tail dataset;
[0013] A feature extraction module, used for inputting the image data into a feature extraction network combined with feature enhancement to obtain image data features after feature enhancement;
[0014] A hash activation module, used for inputting the image data features into a hash activation layer to generate a corresponding binary hash code;
[0015] The optimization module is used to construct a loss function based on maximizing the mutual information between the category label and the image data feature, to optimize the feature extraction network and the hash activation layer, to obtain a trained long-tail hash model, and to generate a long-tail data hash code.
[0016] Furthermore, the present invention also proposes a device, including a memory and a processor, wherein the memory stores computer-readable instructions, wherein when the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the long-tail data hash code generation method as described in the present invention.
[0017] Furthermore, the present invention also proposes a storage medium on which computer-readable instructions are stored, wherein the computer-readable instructions, when executed by a processor, implement all or part of the steps of the long-tail data hash code generation method as described in the present invention.
[0018] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0019] The present invention integrates the theory of mutual information into the deep hashing method. By maximizing the mutual information between image features and labels, the correlation between features and labels is improved, and the impact of long-tail distribution is reduced. It can perform fast and accurate retrieval in image data that obeys long-tail distribution. Users can overcome the impact of the imbalanced number of images of different categories and perform image retrieval quickly and accurately. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 The figure is a flow chart of a method for generating a long-tail data hash code according to an embodiment of the present invention.
[0021] Figure 2 The figure is a flow chart of a method for generating a long-tail data hash code according to an embodiment of the present invention.
[0022] Figure 3 The diagram is an architecture diagram of a system for generating long-tail data hash codes according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0024] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0025] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0026] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] Example 1
[0028] See also Figure 1 , 2 This embodiment proposes a method for generating a long-tail data hash code based on mutual information, which includes the following steps:
[0029] S1. Obtain image data and its category labels in the long-tail dataset;
[0030] S2, inputting the image data into a feature extraction network combined with feature enhancement to obtain image data features after feature enhancement;
[0031] S3, inputting the image data features into a hash activation layer to generate a corresponding binary hash code;
[0032] S4. Constructing a loss function based on maximizing the mutual information between the category label and the image data features, which is used to optimize the feature extraction network and the hash activation layer, and obtaining a trained long-tail hash model for generating a long-tail data hash code.
[0033] In this embodiment, the theory of mutual information is integrated into the deep hashing method. By maximizing the mutual information between image features and labels, the correlation between features and labels is improved, and the impact of long-tail distribution is reduced, which enables fast and accurate retrieval in image data that obeys long-tail distribution.
[0034] The trained long-tail hash model includes a trained feature extraction network and a hash activation layer. In practical applications, the processed image data is sequentially passed through the trained feature extraction network and the hash activation layer to generate a corresponding hash code.
[0035] In an optional embodiment, step S1 further includes the following steps:
[0036] The category labels corresponding to the images in the acquired long-tail dataset are preprocessed to obtain the one-hot vector of the category labels, which is used to convert the category labels into binary vectors that can be understood by the machine learning algorithm.
[0037] In an optional embodiment, in step S2, when the image data is input into a feature extraction network combined with feature enhancement, the following steps are performed:
[0038] Input image I into the feature extraction network to obtain the image feature f(I) of image I;
[0039] The image feature f(I) is input into the feature enhancement module to obtain the feature enhanced image feature f′(I).
[0040] Furthermore, in an optional embodiment, the feature extraction network includes a ResNet-50 architecture pre-trained on the ImageNet dataset.
[0041] Among them, ResNet-50 is a deep convolutional neural network architecture with a 50-layer deep network structure, including an input layer, a residual block, a global average pooling layer, and a fully connected layer. It can learn rich feature representations through multi-scale feature extraction and residual connections, and has strong generalization capabilities.
[0042] Furthermore, in an optional embodiment, the feature enhancement module includes a cross-attention feature enhancement (CAFE) module based on ACHNet or a dynamic meta-embedding (DME) module based on LTHNet.
[0043] ACHNet (Attention-guided Contrastive Hashing for Long-tailed Image Retrieval) is an attention-guided contrastive hashing method for long-tailed image retrieval, which enhances features through a cross-attention mechanism to improve the retrieval performance of the model under long-tail data distribution.
[0044] LTHNet (Long-Tail Hashing Network) is a deep hashing network used to process long-tail data distribution. It aims to enrich feature representation by combining direct features and memory features, which can effectively improve the recognition performance of the model under long-tail data distribution.
[0045] In an optional embodiment, in step S3, the image data features are input into a hash activation layer to generate a corresponding binary hash code, including the following steps:
[0046] Inputting the image data features after feature enhancement into a hash activation layer, and reducing the dimension of the image data features through a plurality of fully connected layers;
[0047] The reduced-dimensional image data features are converted into binary hash codes through an activation function.
[0048] Among them, the hash activation layer contains multiple fully connected layers that map high-dimensional features into low-dimensional binary codes with fixed code lengths, and the features after dimensionality reduction are output as binary hash codes composed of -1 and +1 through the activation function.
[0049] Further optionally, the activation function includes a Sign function, a Tanh function and a Sigmoid function.
[0050] Exemplarily, in this embodiment, the Tanh function is selected to convert the binary hash code.
[0051] In an optional embodiment, in step S4, the loss function constructed based on maximizing the mutual information between the category label and the image data feature includes InfoNCE loss; its expression is:
[0052]
[0053] Where (f′, y) represents the joint distribution between the feature-enhanced image data feature f′ and its category label y; W is a trainable weight matrix used to achieve the bilinear transformation from the image data feature f′ to the category label y; f′ j represents the jth feature-enhanced image data feature in the joint distribution (f′, y);
[0054] The mutual information between the category label and the image data features is maximized using the InfoNCE loss, and the feature extraction network and the hash activation layer are optimized using a small batch gradient descent method to obtain a trained long-tail hash model.
[0055] In this embodiment, InfoNCE loss is used to maximize mutual information, which has the advantages of convenience, ease of use and excellent effect.
[0056] As an example, InfoNCE loss is a loss function used for self-supervised learning, which learns model parameters by comparing the similarity between positive and negative samples, thereby improving the discrimination of features. It is defined as:
[0057]
[0058] Where X represents the sample set, and f(x,y) is the similarity function used to measure the similarity between the positive sample x and the negative sample y.
[0059] In this embodiment, the image data features after feature enhancement are taken as x samples, the labels are taken as y samples, and Using the properties of logarithms to simplify, we get:
[0060]
[0061] In practical applications, F(x,y) often uses a trainable weight matrix W to implement the bilinear transformation from x samples to y samples, that is:
[0062] F(x,y)=xWy
[0063] Combining the above formula, we can get the expression of the InfoNCE loss function used to maximize the mutual information between the category label and the image data features. Using this loss function and optimizing the feature extraction network and hash activation layer using the mini-batch gradient descent method, we can get a trained long-tail hash model for image retrieval in long-tail datasets. In practical applications, users can overcome the impact of the imbalanced number of images of different categories and perform image retrieval quickly and accurately.
[0064] Example 2
[0065] See also Figure 3This embodiment applies the long-tail data hash code generation method proposed in Embodiment 1 and proposes a long-tail data hash code generation system based on mutual information.
[0066] The long-tail data hash code generation system based on mutual information proposed in this embodiment includes:
[0067] Data collection module, used to obtain image data and its category labels in the long-tail dataset;
[0068] A feature extraction module, used for inputting the image data into a feature extraction network combined with feature enhancement to obtain image data features after feature enhancement;
[0069] A hash activation module, used for inputting the image data features into a hash activation layer to generate a corresponding binary hash code;
[0070] The optimization module is used to construct a loss function based on maximizing the mutual information between the category label and the image data feature, to optimize the feature extraction network and the hash activation layer, to obtain a trained long-tail hash model, and to generate a long-tail data hash code.
[0071] In an optional embodiment, the data acquisition module is further used to preprocess the category labels corresponding to the images in the acquired long-tail data set to obtain a one-hot vector of the category label.
[0072] In an optional embodiment, the feature extraction module is equipped with a ResNet-50 architecture, which is pre-trained on the ImageNet dataset.
[0073] In an optional embodiment, the feature extraction module also includes a cross-attention feature enhancement module based on ACHNet or a dynamic meta-embedding module based on LTHNet, which is used as a feature enhancement module to enhance image features.
[0074] In an optional embodiment, the feature extraction module performs the following steps when inputting the image data into the feature extraction network combined with feature enhancement:
[0075] Input image I into the feature extraction network to obtain the image feature f(I) of image I;
[0076] The image feature f(I) is input into the feature enhancement module to obtain the feature enhanced image feature f′(I).
[0077] In an optional embodiment, the hash activation module performs the following steps when inputting the image data features into the hash activation layer to generate the corresponding binary hash code:
[0078] Inputting the image data features after feature enhancement into a hash activation layer, and reducing the dimension of the image data features through a plurality of fully connected layers;
[0079] The reduced-dimensional image data features are converted into binary hash codes through an activation function.
[0080] In an optional embodiment, in the optimization module, the loss function constructed based on maximizing the mutual information between the category label and the image data feature includes InfoNCE loss; its expression is:
[0081]
[0082] Where (f′, y) represents the joint distribution between the feature-enhanced image data feature f′ and its category label y; W is a trainable weight matrix used to achieve the bilinear transformation from the image data feature f′ to the category label y; f′ j represents the jth feature-enhanced image data feature in the joint distribution (f′, y);
[0083] The mutual information between the category label and the image data features is maximized using the InfoNCE loss, and the feature extraction network and the hash activation layer are optimized using a small batch gradient descent method to obtain a trained long-tail hash model.
[0084] It can be understood that the system of this embodiment corresponds to the method of the above-mentioned embodiment 1, and the options in the above-mentioned embodiment 1 are also applicable to this embodiment, so they will not be described repeatedly here.
[0085] Example 3
[0086] This embodiment proposes a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the long-tail data hash code generation method proposed in Embodiment 1.
[0087] Example 4
[0088] This embodiment proposes a storage medium on which computer-readable instructions are stored, wherein the computer-readable instructions, when executed by a processor, implement all or part of the steps of the long-tail data hash code generation method proposed in Embodiment 1.
[0089] Exemplarily, the storage medium includes, but is not limited to, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0090] Exemplarily, the instructions, programs, code sets or instruction sets may be implemented using conventional programming languages.
[0091] Exemplarily, the processor includes but is not limited to a smart phone, a personal computer, a server, a network device, etc., and is used to execute all or part of the steps of the long-tail data hash code generation method described in Example 1.
[0092] Each embodiment of the present invention is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely exemplary, in which the modules described as separate components may or may not be physically separated, and the functions of each module can be implemented in the same or one or more software and / or hardware when implementing the scheme of the present invention. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the scheme of this embodiment.
[0093] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A method for generating long-tail data hash codes based on mutual information, characterized in that: The following steps are involved: Obtain image data and its category labels in the long-tail dataset; Inputting the image data into a feature extraction network combined with feature enhancement to obtain image data features after feature enhancement; Inputting the image data features into a hash activation layer to generate a corresponding binary hash code; A loss function is constructed based on maximizing the mutual information between the category label and the image data features, which is used to optimize the feature extraction network and the hash activation layer, and a trained long-tail hash model is obtained to generate a long-tail data hash code.
2. The long-tail data hash code generation method according to claim 1, characterized in that: The image data is input into a feature extraction network combined with feature enhancement, comprising the following steps: Input image I into the feature extraction network to obtain the image feature f(I) of image I; The image feature f(I) is input into the feature enhancement module to obtain the feature enhanced image feature f′(I).
3. The long-tail data hash code generation method according to claim 2, characterized in that: The feature extraction network includes a ResNet-50 architecture pre-trained on the ImageNet dataset.
4. The method for generating a long-tail data hash code according to claim 2, characterized in that: The feature enhancement module includes a cross-attention feature enhancement module based on ACHNet or a dynamic meta-embedding module based on LTHNet.
5. The method for generating a long-tail data hash code according to claim 1, wherein: Inputting the image data features into a hash activation layer to generate a corresponding binary hash code comprises the following steps: Inputting the image data features after feature enhancement into a hash activation layer, and reducing the dimension of the image data features through a plurality of fully connected layers; The reduced-dimensional image data features are converted into binary hash codes through an activation function.
6. The long-tail data hash code generation method according to claim 1, characterized in that: The loss function constructed based on maximizing the mutual information between the category label and the image data features includes InfoNCE loss; Its expression is: Wherein, (f′, y) represents the joint distribution between the feature-enhanced image data feature f′ and its category label y; W is a trainable weight matrix used to realize the bilinear transformation from the image data feature f′ to the category label y; f′j represents the jth feature-enhanced image data feature in the joint distribution (f′, y); The mutual information between the category label and the image data features is maximized using the InfoNCE loss, and the feature extraction network and the hash activation layer are optimized using a small batch gradient descent method to obtain a trained long-tail hash model.
7. The method for generating a long-tail data hash code according to any one of claims 1 to 6, characterized in that: The method further comprises the following steps: preprocessing the category labels corresponding to the images in the acquired long-tail data set to obtain a one-hot vector of the category labels.
8. A long-tail data hash code generation system based on mutual information, applying the long-tail data hash code generation method according to any one of claims 1 to 7, characterized in that: The system comprises: Data collection module, used to obtain image data and its category labels in the long-tail dataset; A feature extraction module, used for inputting the image data into a feature extraction network combined with feature enhancement to obtain image data features after feature enhancement; A hash activation module, used for inputting the image data features into a hash activation layer to generate a corresponding binary hash code; The optimization module is used to construct a loss function based on maximizing the mutual information between the category label and the image data feature, to optimize the feature extraction network and the hash activation layer, to obtain a trained long-tail hash model, and to generate a long-tail data hash code.
9. A device comprising a memory and a processor, wherein the memory stores computer-readable instructions, characterized in that: When the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the long-tail data hash code generation method according to any one of claims 1 to 7.
10. A storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, all or part of the steps of the long-tail data hash code generation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Hashing method and device for image features and processing equipment
CN114898104A
Cross-domain long-tail image classification method based on self-supervised learning and self-training mechanism
CN117115547A