Image Retrieval Method for Hash Model Training and Adaptive Binary Quantization in Noisy Environments

By introducing adaptive regular loss and similarity retention loss in deep hash learning and adjusting model parameters, the problems of poor generalization ability and large binary quantization error of hash model in the noisy label environment are solved, and higher hash encoding value accuracy and image retrieval performance are achieved.

CN115858841BActive Publication Date: 2025-06-20XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211242280.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2025-06-20
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

The existing deep hash learning methods have poor generalization capabilities in noise label environments and large binary quantization errors, resulting in low hash encoding value accuracy and affecting image retrieval performance.

Method used

A hash model training method is proposed. By obtaining a training set and initial model containing multiple sample images, the first similarity reservation loss and adaptive regular loss are used to adjust the model parameters, and adaptive binary quantization in a noisy environment is realized.

Benefits of technology

Effectively alleviate the negative impact of noise labels on the hash model, improve the robustness of the model and the accuracy of hash encoding values, and improve the accuracy and performance of image retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115858841B_ABST
    Figure CN115858841B_ABST
Patent Text Reader

Abstract

The present invention discloses an image retrieval method for hash model training and adaptive binary quantization in a noisy environment, including: obtaining multiple sample images, each sample image having its own class label; obtaining an initial model including an initial feature extraction network, an activation function, and an initial classification network: inputting the sample images into the initial model during the z-th training to obtain the z-th hash code value and the z-th prediction value of each sample image; determining the z-th first similarity retention loss according to the number of input sample images, the z-th prediction value, and the class label of each sample image; determining the z-th adaptive regularization loss according to the hyperparameters of the initial model, the number of input sample images, the z-th prediction value, and the (z-1)-th prediction value of each sample image; adjusting the network parameters of the model obtained from the (z-1)-th training according to the two obtained losses until obtaining a first pre-trained feature extraction network, an activation function, and a first pre-trained classification network, thereby obtaining a pre-trained hash model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an image retrieval method for hash model training and adaptive binary quantization in a noisy environment. Background Art

[0002] In the era of big data, the quantity of visual data (such as images and videos) has increased sharply. Therefore, people's interest in technologies that can effectively save storage space and retrieval time has gradually grown. In existing approximate nearest neighbor retrieval technologies, hash learning maps high-dimensional data into a binary Hamming space through a machine learning mechanism, thereby significantly reducing the storage and communication overhead of data. In recent years, with the remarkable performance of deep neural networks in various tasks, researchers have attempted to introduce this technology into the field of data retrieval and developed hash learning algorithms based on deep neural networks.

[0003] Currently, supervised deep hash learning usually uses large-scale datasets for optimization training. In these datasets, most sample labels are manually annotated and verified. However, in real-world scenarios, accurately labeled datasets generally do not exist. Inevitably, many noisy labels will be introduced during the annotation process of sample labels, and deep neural networks have the ability to remember any label of data, resulting in the widespread noisy labels in the dataset seriously increasing the generalization error of the deep hash model, thereby making the robustness of the hash model poor. Currently, the noisy label learning method can only learn continuous representations in the feature space and cannot simultaneously solve the feature quantization loss problem in the deep hash method. Moreover, during the hash learning process, inappropriate binary quantization may introduce huge quantization errors and seriously reduce the accuracy of the hash code values generated by the hash model. Summary of the Invention

[0004] To solve the above problems existing in the related technologies, the present invention provides an image retrieval method for hash model training and adaptive binary quantization in a noisy environment. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0005] The present invention provides a method for training a hash model, including:

[0006] Obtaining a training set including multiple sample images; each sample image corresponds to its own class label;

[0007] Obtaining an initial model: The initial model includes: an initial feature extraction network, an activation function, and an initial classification network; the initial feature extraction network and the initial classification network correspond to initial network parameters;

[0008] At the z-th training, input the multiple sample images into the initial model to obtain the hash code values and the prediction values of the z-th time for each sample image, where the prediction value of each sample image is the probability value that the sample image belongs to each preset category;

[0009] Determine the first similarity retention loss of the z-th time according to the number of the multiple sample images, the prediction value of the z-th time, and the category label of each sample image itself;

[0010] Determine the adaptive regularization loss of the z-th time according to the hyperparameters of the initial model, the number of the multiple sample images, the prediction value of the z-th time, and the prediction values of the (z - 1)-th time for each sample image;

[0011] Adjust the network parameters of the model obtained by the (z - 1)-th training according to the first similarity retention loss of the z-th time and the adaptive regularization loss of the z-th time, and when z + 1 is less than or equal to the preset number of times, continue the (z + 1)-th training until z + a is greater than the preset number of times to obtain the first pre-trained model; a is an integer greater than 1; the first pre-trained model includes: a first pre-trained feature extraction network, the activation function, and a first pre-trained classification network; z is an integer greater than 0; when z = 1, the network parameters of the model obtained by the (z - 1)-th training are the initial network parameters;

[0012] Obtain a first pre-trained hash model according to the first pre-trained feature extraction network and the activation function.

[0013] The present invention also provides an image retrieval method for adaptive binary quantization in a noise environment, including:

[0014] Obtain the image to be retrieved;

[0015] Input the image to be retrieved into the above-mentioned first pre-trained hash model to obtain the hash code value;

[0016] Retrieve the target image corresponding to the image to be retrieved from the preset image library according to the hash code value.

[0017] The present invention also provides a hash model training method, including:

[0018] Obtain a training set including multiple sample images and an initial model; the initial model includes: an initial feature extraction network, an activation function, and an initial classification network; the initial feature extraction network and the initial classification network correspond to initial network parameters;

[0019] When each sample image in the training set has its own class label, at the z-th training, input the sample image into the initial model to obtain the z-th hash code value and the z-th prediction value of each sample image. The prediction value of each sample image is the probability value that the sample image belongs to each preset class;

[0020] Based on the number of the multiple sample images, the z-th prediction value, the class label of each sample image, the hyperparameters of the initial model, and the (z - 1)-th prediction value of each sample image, determine the first similarity retention loss at the z-th time and the adaptive regularization loss at the z-th time. According to the first similarity retention loss and the adaptive regularization loss, adjust the network parameters of the model obtained by the (z - 1)-th training until the number of training times reaches the preset number of times, and then obtain the first pre-trained model; z is an integer greater than 0; when z = 1, the network parameters of the model obtained by the (z - 1)-th training are the initial network parameters; the first pre-trained model includes: a first pre-trained feature extraction network, the activation function, and a first pre-trained classification network;

[0021] When there is a sample pair class label between any two sample images in the training set, at the z-th training, input the sample image into the initial model to obtain the z-th hash code value of each sample image. According to the sample pair class label of each sample image and the z-th hash code value of each sample image, determine the second similarity retention loss at the z-th time. According to the second similarity retention loss at the z-th time, adjust the network parameters of the model obtained by the (z - 1)-th training until the number of training times reaches the preset number of times, and then obtain the second pre-trained model; the second pre-trained model includes: a second pre-trained feature extraction network, the activation function, and a second pre-trained classification network;

[0022] According to the first pre-trained feature extraction network and the activation function, obtain the first pre-trained hash model, or, according to the second pre-trained feature extraction network and the activation function, obtain the second pre-trained hash model.

[0023] The present invention has the following beneficial technical effects:

[0024] Since noisy labels (erroneous labels) are generally considered as damages at the sample point level, when training the model with multiple sample images that only have their own class labels (i.e., labels based on sample points), by adjusting the model's parameters through the first similarity retention loss and the adaptive regularization loss obtained each time, it is possible to mitigate the model's memory effect on noisy labels through the adaptive regularization term and the similarity retention loss function, effectively alleviating the negative impact of noisy labels on the hash model, enhancing the generalization ability of the hash model under real-world data, so that during the training of the hash model, it can combat noisy labels in the hash learning process, improve the robustness of the hash model, and mitigate the impact of binary quantization on the accuracy of the hash code values generated by the hash model, making the accuracy of the hash code values generated by the trained hash model higher.

[0025] Since the hash code values of the images to be retrieved generated by the trained hash model have higher accuracy, when using the hash code values with high accuracy for retrieving the target images, the retrieval accuracy is higher, improving the hash retrieval performance.

[0026] By selecting different loss functions for training the hash model for sample images with different types of labels, it is possible to unify different types of hash learning schemes, facilitating subsequent data analysis and model training.

[0027] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Description of the Drawings

[0028] Figure 1 It is a flowchart of the hash model training method provided by the embodiment of the present invention;

[0029] Figure 2 It is a flowchart of the image retrieval method with adaptive binary quantization in a noisy environment provided by the embodiment of the present invention;

[0030] Figure 3 It is a framework diagram of the exemplary hash model training provided by the embodiment of the present invention;

[0031] Figure 4A It is an average accuracy rate curve graph of the model trained with a simple cross-entropy loss function on the CIFAR-10 dataset provided by the embodiment of the present invention;

[0032] Figure 4B It is an average accuracy rate curve graph of the model trained with a simple cross-entropy loss function on the CIFAR-100 dataset provided by the embodiment of the present invention. Detailed Embodiments

[0033] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0034] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.

[0035] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0036] Although the present invention has been described herein in connection with various embodiments, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit can implement several functions recited in the claims. Certain measures are recited in mutually different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0037] Figure 1 is an optional flowchart of the hash model training method provided by an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:

[0038] S101. Obtain a training set including multiple sample images; each sample image corresponds to its own class label.

[0039] In the embodiment of the present invention, the training set may include N sample images, and each sample image only has its own corresponding class label (i.e., the label based on the sample point); this class label can represent the class to which this sample image belongs. Exemplarily, this class label can be manually labeled; N is an integer greater than 0, and the value of N can be set according to actual needs.

[0040] S102. Obtain an initial model: The initial model includes: an initial feature extraction network, an activation function, and an initial classification network; the initial feature extraction network and the initial classification network correspond to initial network parameters.

[0041] In an embodiment of the present invention, the activation function may be a sign function.

[0042] In an embodiment of the present invention, the network parameters corresponding to the initial feature extraction network and the initial classification network may be used as the initial network parameters of the initial model; and, the initial model further has hyperparameters λ and β, and λ and β may be preset values and may be set according to actual needs.

[0043] S103. At the z-th training, input multiple sample images into the initial model to obtain the z-th hash code value and the z-th prediction value of each sample image, and the prediction value of each sample image is the probability value that the sample image belongs to each preset category.

[0044] In an embodiment of the present invention, at the z-th training, when inputting N sample images into the initial model, for each sample image, first, the initial feature extraction network generates a feature vector of the sample image, then, the activation function generates the hash code value of the sample image according to the feature vector, and finally, the classification network generates a C-dimensional vector (prediction value) according to the hash code value, and each dimension vector in the C-dimensional vector represents the probability that the image belongs to one of the C preset categories; and, these C preset categories are the categories corresponding to the N sample images, and C is an integer greater than or equal to 2.

[0045] S104. Determine the first similarity retention loss at the z-th time according to the number of multiple sample images, the z-th prediction value, and the category label of each sample image itself.

[0046] In an embodiment of the present invention, for each image, the z-th sub-loss of the image may be calculated according to the z-th prediction value of the image and the category label of the image itself, and based on the number N of the sample images input at the z-th time and the z-th sub-loss of each of the N sample images, the first similarity retention loss at the z-th time is determined.

[0047] Exemplarily, the first similarity retention loss at the z-th time is calculated by the following formula:

[0048]

[0049] Wherein, represents the first similarity retention loss at the z-th time, N represents the number of sample images input at the z-th time, i represents the i-th image among the N images (i = 1, 2,..., N), represents the sub-loss of the i-th image, C represents the number of preset categories, c represents the c-th category among the C categories (c = 1, 2,..., C), represents the value corresponding to the c-th category of the class label of the i-th image itself, represents the probability value of the c-th category in the z-th prediction value of the i-th image. Here, is 0 or 1, where when the class label of the i-th image itself is the c-th category it is 1, and when the class label of the i-th image itself is not the c-th category it is 0.

[0050] S105. Determine the z-th adaptive regularization loss according to the hyperparameters of the initial model, the number of multiple sample images, the z-th prediction value, and the (z - 1)-th prediction value of each sample image.

[0051] In the embodiments of the present invention, the z-th prediction value and the (z - 1)-th prediction value of each sample image are both C-dimensional vectors. Correspondingly, the (z - 1)-th prediction value of each sample image is the prediction value of this sample image generated during the (z - 1)-th training. For each image, the inner product value of the z-th time of this image can be obtained according to the vector inner product between the z-th prediction value and the (z - 1)-th prediction value of this image. Based on the inner product values of each of the N sample images input at the z-th time, the logarithmic total value is determined. Finally, according to the hyperparameters of the initial model, the number N of sample images input at the z-th time, and the obtained logarithmic total value, the z-th adaptive regularization loss is determined.

[0052] Exemplarily, the z-th adaptive regularization loss is calculated by the following formula:

[0053]

[0054] where represents the z-th adaptive regularization loss, N represents the number of sample images input at the z-th time, i represents the i-th image among the N images (i = 1, 2,..., N), <p i ,t i > represents the inner product value of the i-th image at the z-th time, p i represents the z-th prediction value of the i-th image, t i represents the (z - 1)-th prediction value of the i-th image, λ represents the hyperparameter, and log(.) represents the logarithmic function.

[0055] S106. Retain the loss of the first similarity and the adaptive regularization loss of the z-th time, adjust the network parameters of the model obtained by the (z - 1)-th training, and when z + 1 is less than or equal to the preset number of times, continue the (z + 1)-th training until z + a is greater than the preset number of times to obtain the first pre-trained model; a is an integer greater than 1; the first pre-trained model includes: the first pre-trained feature extraction network, the activation function, and the first pre-trained classification network; z is an integer greater than 0; when z = 1, the network parameters of the model obtained by the (z - 1)-th training are the initial network parameters.

[0056] In the embodiment of the present invention, the sum of the similarity retention loss of the z-th time and the adaptive regularization loss of the z-th time can be used as the total loss of the z-th time. Adjust the network parameters of the model obtained by the (z - 1)-th training according to the total loss of the z-th time to obtain the network parameters adjusted in the z-th time, and obtain the model obtained by the z-th training according to the network parameters adjusted in the z-th time; then, determine whether z + 1 is less than or equal to the preset number of times, and when z + 1 is less than or equal to the preset number of times, perform the (z + 1)-th training of the model according to the model obtained by the z-th training until z + a is greater than the preset number of times to obtain the second pre-trained model.

[0057] Here, when z = 1, it means that the first iterative training is performed on the initial model at this time, so the network parameters of the model obtained by the (z - 1)-th training are the initial network parameters of the initial model; for example, when z = 3, it means that the third iterative training is performed on the initial model at this time, so the network parameters of the model obtained by the (z - 1)-th training are the network parameters obtained after the second iterative adjustment of the initial network parameters.

[0058] Here, the preset number of times can be set according to actual needs, and the embodiment of the present invention does not limit this.

[0059] Here, when the first pre-trained model is obtained, the first pre-trained feature extraction network, the activation function, and the first pre-trained classification network are obtained.

[0060] S107. Obtain the first pre-trained hash model according to the first pre-trained feature extraction network and the activation function.

[0061] In the embodiments of the present invention, since noise labels (erroneous labels) are generally considered as damages in units of sample points, when training a model with multiple sample images each having only its own class label, by adjusting the parameters of the model through the first similarity retention loss and the adaptive regularization loss obtained each time, it is possible to alleviate the memory effect of the model on noise labels through the adaptive regularization term and the similarity retention loss function, effectively alleviate the negative impact of noise labels on the hash model, enhance the generalization ability of the hash model under real-world data, thereby being able to combat noise labels in the process of hash learning when training the hash model, improve the robustness of the hash model, and alleviate the impact of binary quantization on the accuracy of the hash code values generated by the hash model, so that the accuracy of the hash code values generated by the trained hash model is higher.

[0062] The embodiments of the present invention also provide an image retrieval method with adaptive binary quantization in a noisy environment, as Figure 2 shown, the method includes:

[0063] S201. Obtain the image to be retrieved.

[0064] S202. Input the image to be retrieved into the above-mentioned first pre-trained hash model to obtain a hash code value.

[0065] S203. According to the hash code value, retrieve the target image corresponding to the image to be retrieved from the preset image library.

[0066] In the embodiments of the present invention, since the accuracy of the hash code value of the image to be retrieved generated by the trained hash model is higher, when using the hash code value with high accuracy to retrieve the target image, the retrieval accuracy is higher, improving the hash retrieval performance.

[0067] The embodiments of the present invention also provide a hash model training method, the method includes:

[0068] S301. Obtain a training set including multiple sample images and an initial model; the initial model includes: an initial feature extraction network, an activation function, and an initial classification network; the initial feature extraction network and the initial classification network correspond to initial network parameters.

[0069] S302. When each sample image in the training set has its own class label, at the z-th training, input the sample image into the initial model to obtain the z-th hash code value and the z-th prediction value of each sample image, and the prediction value of each sample image is the probability value that the sample image belongs to each preset class.

[0070] S303. Based on the number of multiple sample images, the prediction value at the z-th time, the class label of each sample image, the hyperparameters of the initial model, and the prediction value of each sample image at the (z - 1)-th time, determine the first similarity retention loss and the adaptive regularization loss at the z-th time. According to the first similarity retention loss and the adaptive regularization loss, adjust the network parameters of the model obtained by the (z - 1)-th training until the number of training times reaches the preset number of times, and obtain the first pre-trained model; z is an integer greater than 0; when z = 1, the network parameters of the model obtained by the (z - 1)-th training are the initial network parameters; the first pre-trained model includes: the first pre-trained feature extraction network, the activation function, and the first pre-trained classification network.

[0071] Here, the training methods of S302 - S303 are the same as those of S102 - S106 above.

[0072] S304. When there are sample pair class labels between any two sample images in the training set, at the z-th training, input the sample images into the initial model to obtain the hash code values of each sample image at the z-th time. According to the sample pair class labels of each sample image and the hash code values of each sample image at the z-th time, determine the second similarity retention loss at the z-th time. According to the second similarity retention loss at the z-th time, adjust the network parameters of the model obtained by the (z - 1)-th training until the number of training times reaches the preset number of times, and obtain the second pre-trained model; the second pre-trained model includes: the second pre-trained feature extraction network, the activation function, and the second pre-trained classification network.

[0073] In the embodiments of the present invention, when there are sample pair class labels between any two sample images in the training set (i.e., this label can be called a sample pair-based label), the network parameters of the model can be adjusted only through the second similarity retention loss. The sample pair class label corresponding to each pair of sample images is used to represent whether the categories of this pair of sample images are the same category.

[0074] In the embodiments of the present invention, the second similarity retention loss at the z-th time can be used as the loss at the z-th time of the model to adjust the network parameters of the model obtained by the (z - 1)-th training, and obtain the model obtained by the z-th training. Then, determine whether z + 1 is less than or equal to the preset number of times, and when z + 1 is less than or equal to the preset number of times, perform the (z + 1)-th training according to the model obtained by the z-th training until z + b is greater than the preset number of times, and obtain the first pre-trained model; b is an integer greater than 1.

[0075] Here, the preset number of times can be set according to actual needs, and the embodiments of the present invention do not limit this.

[0076] Exemplarily, the second similarity retention loss at the z-th time is calculated by the following formula:

[0077]

[0078]

[0079]

[0080]

[0081] Among them, θ represents the network parameters of the model obtained from the (z - 1)-th training, represents the second similarity retention loss of the z-th time, N represents the number of sample images input at the z-th time, i represents the i-th image among the N images (i = 1, 2,..., N), represents the sub-loss of the i-th image at the z-th time, K represents the number of sample pair category labels between the i-th image and the corresponding sample images in the sample images input at the z-th time, which represent that the i-th image and the corresponding sample images belong to the same category, L represents the number of sample pair category labels between the i-th image and the corresponding sample images in the sample images input at the z-th time, which represent that the i-th image and the corresponding sample images belong to different categories, x i represents the hash coding value of the i-th image at the z-th time, x j represents the hash coding value of the j-th sample image (j = 1, 2,..., K) at the z-th time among the K sample images belonging to the same category as the i-th image, x y represents the hash coding value of the y-th sample image (y = 1, 2,..., L) at the z-th time among the L sample images belonging to different categories from the i-th image, represents the j-th similarity score among the K similarity scores, represents the y-th dissimilarity score among the L dissimilarity scores, H dist (.,.) represents a metric function for measuring the Hamming distance between two hash coding values, and log[.] represents the logarithmic function.

[0082] S305. Obtain the first pre-trained hash model according to the first pre-trained feature extraction network and the activation function, or obtain the second pre-trained hash model according to the second pre-trained feature extraction network and the activation function.

[0083] In the embodiments of the present invention, by selecting different loss functions for training the hash model for sample images with different types of labels, different types of hash learning schemes can be unified.

[0084] Exemplarily, Figure 3 is a framework diagram of hash model training; as Figure 3As shown, when inputting an image (sample image) into a network (initial model) during each training, a prediction value for this time can be obtained. When the image only has its own class label (true value), based on the prediction value output by the model this time and the true value of the image, the first similarity retention loss for this time can be determined, and based on the prediction value output by the model last time, the prediction value output this time, and the temporal integration parameter, the target value of the image can be obtained. Based on the target value of the image and the true value of the image, the adaptive regularization loss (adaptive regularization term) for this time can be determined. Thus, based on the first similarity retention loss for this time and the adaptive regularization loss for this time, the network parameters of the model are adjusted; while when the image has the class label (true value) of a sample pair, based on the prediction value output by the model this time and the true value of the image, the second similarity retention loss for this time can be determined. Thus, based on the second similarity retention loss for this time, the network parameters of the model are adjusted until a trained model is obtained. It should be noted that for ease of understanding, Figure 3 the first similarity retention loss and the second similarity retention loss are uniformly expressed as the similarity retention loss.

[0085] The following further explains the effect of the above hash model training method through a derivation process.

[0086] First, a data set containing N samples is represented as where x j ∈X represents the j-th sample, and y j is the corresponding class label. In the real-world scenario, the clean (correctly labeled) label y j may be randomly corrupted into a noisy label before being observed. Therefore, we assume that during training, the data set we can obtain is a data set containing noisy labels. To accurately evaluate the performance of the proposed method, we assume that there is a clean test data set D te .

[0087] Definition 1 (Robust Hamming Space Retrieval): To achieve effective approximate nearest neighbor retrieval, we map different inputs into the general Hamming space B through a hash function h: X → B, and this function can be implemented by a deep neural network with θ as a parameter. Then, given an input data x j , we can obtain its corresponding binary code, that is:

[0088] b j = h(x j , θ) (1);

[0089] For robust Hamming space retrieval, we hope to use a noisy training data set Learn the hash function h and make the hash codes calculated from formula (1) perform well on the clean test dataset D te and show good performance effects. In the following, we will first elaborate a general hash learning framework to maximize the semantic similarity of samples in the Hamming space. Then, we study the problems reflected by this framework when dealing with noisy label data. Finally, we propose an adaptive regularization term to mitigate the impact of noisy label data on the hash model.

[0090] This solution proposes a unified hash learning framework to simultaneously apply to class labels based on sample points and sample pairs. Given the training dataset We first generate a set of hash codes where b i represents the hash code of sample x i . Then, for the given binary code b, we assume that there are K similarity scores and L dissimilarity scores respectively. To preserve the above similarity relationships in the Hamming space, we propose to maximize the similarity scores and minimize the dissimilarity scores through the following constraint functions:

[0091]

[0092] For the labels based on sample pairs, we can calculate the similarity scores between the binary code b of sample x and other samples in the mini-batch. Specifically, we can calculate the similarity scores through (x i is the i-th sample semantically similar to sample x) and calculate the dissimilarity scores through (x j is the j-th sample semantically dissimilar to sample x), where H dist (.,.) is a metric function used to measure the Hamming distance between the binary codes (hash code values) of two samples.

[0093] For the labels based on sample points, we can implement formula (2) through the cross-entropy loss, and use the sample hash code b and the classifier weights w i (i = 1, 2,..., C) to calculate the similarity scores. Specifically, we can obtain C - 1 inter-class similarity scores through (w j represents the j-th non-target weight vector) and obtain the intra-class similarity scores through (w yObtaining an intra-class similarity score by (denoting the target weight vector). Under this condition, if we regard the hash code as a general continuous representation, formula (2) can be degenerated into a cross-entropy loss function. Since noise labels are generally considered to be corruptions at the sample point level, in this paper, regarding noise labels, we focus on labels based on sample points.

[0094] Existing schemes show that the loss function in formula (2) can achieve excellent performance on a clean (correctly labeled) training dataset. However, for a training dataset with noise labels, this loss function may have serious overfitting problems, leading to the degradation of the final retrieval performance. To better analyze the impact of noise labels on hash learning, we conducted experimental experiments. Specifically, we used ResNet-34 as the network backbone and trained a deep hash model on the CIFAR-10 and CIFAR-100 datasets with synthetic noise labels at different levels using the loss function in formula (2). Figure 4A and Figure 4B show the average accuracy curves of the two datasets at different training epochs. Figure 4A is the average accuracy curve graph of the trained deep hash model on the CIFAR-10 dataset, Figure 4B is the average accuracy curve graph of the trained deep hash model on the CIFAR-100 dataset; among them, the lengths of the hash codes are 32 bits and 64 bits respectively, and the noise labels are generated using symmetric noise with flipping rates of 0.4 and 0.8. From Figure 4A and Figure 4B it can be seen that the average accuracies of the two datasets both show an upward trend in the initial stage of training, and then start to decline rapidly after a relatively small number of training epochs. A similar phenomenon also occurs in image classification tasks, that is, deep neural networks will first tend to memorize and adapt to the majority (clean) patterns, and then gradually transition to adapt to the minority (noise) patterns. The above phenomenon is called the memory effect of deep neural networks.

[0095] To theoretically analyze the memory effect in the hash learning process, we first analyze formula (2). For the sample point labels, formula (2) can be rewritten as

[0096]

[0097] Assume:

[0098]

[0099] where p c (c = 1, 2,..., C) represents the conditional probability of each class, and 1[condition] is an indicator function that outputs 1 if the condition condition is satisfied, otherwise outputs 0.

[0100] Combined with formula (4), the gradient of formula (3) is as follows:

[0101]

[0102] where is the Jacobian matrix of the neural network encoding with respect to the parameter θ. It can be seen from the gradient analysis that the influence of noisy labels on the deep neural network is mainly affected by the conditional probability p and the difference between the labels Assuming c is the correct class and c′ is the wrong class, since noise will make the correct label and the wrong label resulting in the corresponding gradient being flipped. Therefore, performing stochastic gradient descent will ultimately produce a memory effect.

[0103] Based on the early learning phenomenon, the deep neural network will preferentially fit clean samples in the early learning stage and then continuously memorize mislabeled samples. Therefore, we propose to make the model adaptively learn network parameters by using the moving average method, thereby weakening the negative impact of noisy labels on the neural network in the later stage of training. Referring to the temporal ensembling technique in semi-supervised learning, we propose to use the output of the model at past times to estimate the target conditional probability t i of the i-th sample at time k. Let

[0104] t i (k) = βt i (k - 1) + (1 - β)p i (k) (6).

[0105] where t i (k) represents the target vector, and p i (k) represents the model output of the i-th sample at the k-th training epoch. 0 ≤ β < 1 is the temporal ensembling parameter (a hyperparameter of the model). We propose an adaptive regularization term by utilizing the early learning phenomenon and combine it with the hash learning model to improve the robustness of the learned hash codes. We combine the adaptive regularization term with the similarity retention loss function of formula (2) above to alleviate the memory effect, and the combined loss function is:

[0106]

[0107] where is the similarity retention loss of the i-th sample (the first similarity retention loss), and <.,.> represents the inner product between vectors. Then the gradient of the loss function in formula (7) is:

[0108]

[0109] where g i ∈R C is given by:

[0110]

[0111] As can be seen from Equation (9), the sign of g i is determined by the weighted combination of the difference between and other terms in the target value. Assume that c is the correct class. Since the deep neural network will preferentially fit clean samples in the early training stage, the c-th term of t i will dominate at this time, making the c-th term of g i negative. For clean samples, since the conditional probability p i output by the model will gradually equal the label Therefore will gradually converge to 0, making the noisy samples tend to dominate. In this case, adding the negative term can maintain the gradient of clean samples at a relatively large level to offset the negative gradient impact caused by mislabeled samples in the later learning stage. Therefore, adding an adaptive regularization term can help alleviate the memory effect phenomenon.

[0112] The effectiveness of the above hash model training method is further illustrated by experimental data below.

[0113] The proposed training method is evaluated on two widely used benchmark datasets CIFAR-10 and CIFAR-100 with synthetic noisy labels, and a dataset Clothing1M with real-world noise. We conducted extensive experiments on the proposed method and compared the experimental results with the current state-of-the-art methods to demonstrate the robustness of our method to noisy labels.

[0114] I. Datasets

[0115] Normalized data augmentation techniques (random cropping and horizontal flipping) are applied to all training datasets. For the CIFAR dataset, the size of the random cropping is set to 32×32, and for the Clothing1M dataset, the size of the random cropping is set to 224×224.

[0116] CIFAR-10: The CIFAR-10 dataset contains 60,000 color images of size 32×32, and the total number of image classes is 10. The CIFAR-10 dataset consists of 50,000 training data and 10,000 test data. In this experiment, we use the entire training set to train the network and divide the test set into a retrieval set and a query set at a ratio of 8:2.

[0117] CIFAR-100: The CIFAR-100 dataset contains 60,000 color images of size 32×32, and the total number of image categories is 100. The CIFAR-100 dataset consists of 50,000 training data and 10,000 test data. In this experiment, we use the entire training set to train the network and divide the test set into a retrieval set and a query set at a ratio of 8:2.

[0118] Clothing1M: Clothing1M is a large-scale dataset. The relevant data is collected from online shopping websites, and the labels in the dataset are generated from the text around the images. The noise rate of Clothing1M is approximately 38.5%. The overall dataset contains more than 1 million images, and the total number of categories is 14. The Clothing1M dataset consists of 1 million training data, 14,000 validation data, and 10,000 test data. We use the training set to train the network and the test set to retrieve the validation set.

[0119] II. Evaluation Metrics

[0120] In the experiment, we adopted three evaluation criteria, namely mean average precision, accuracy, and recall rate, to evaluate the performance of the method proposed in the present invention. Given a sample and a list containing R retrieval results, the average precision (AP) is defined as:

[0121]

[0122] where N is the number of samples in the retrieval set that are truly relevant to the query sample, and P(r) represents the precision of the first r retrieved samples. When the r-th retrieved sample is relevant to the query sample, δ(r) = 1; otherwise, δ(r) = 0. The mean average precision is defined as the mean of the average precisions of all query samples. The truly relevant samples of the query sample are defined as the instances that have at least one same label as the query sample. In this experiment, R is set to 5000.

[0123] III. Experimental Results

[0124] The method proposed in the present invention is compared with the currently most advanced hashing algorithm, and the robustness of the method proposed in the present invention is verified. The specific comparison results are shown in Tables 1, 2, and 3 below. Table 1 shows the average precision means of the method proposed in the present invention and the currently most advanced hashing algorithm under different settings on the CIFAR-10 dataset; Table 2 shows the average precision means of the method proposed in the present invention and the currently most advanced hashing algorithm under different settings on the CIFAR-100 dataset; Table 3 shows the average precision means of the method proposed in the present invention and the currently most advanced hashing algorithm on the Clothing1M dataset. Obviously, compared with the currently most advanced method, the model trained by the training method proposed in the present invention has high robustness.

[0125]

[0126] Table 1

[0127]

[0128] Table 2

[0129]

[0130] Table 3

[0131] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for training a hash model, characterized in that, Comprising: Obtain a training set including multiple sample images; Each sample image corresponds to its own class label; Obtain an initial model: The initial model includes: an initial feature extraction network, an activation function, and an initial classification network; the initial feature extraction network and the initial classification network correspond to initial network parameters; At the th training, input the multiple sample images into the initial model to obtain the hash code values of each sample image at the th time and the prediction values at the th time. The prediction value of each sample image is the probability value that the sample image belongs to each preset category; Based on the quantity of the multiple sample images, the prediction value of the -th time, and the class label of each sample image itself, determine the first similarity retention loss of the -th time; Based on the hyperparameters of the initial model, the number of the multiple sample images, the prediction value of the -th time, and the prediction values of each sample image at the -th time, determine the adaptive regularization loss of the -th time; According to the first similarity retention loss of the -th time and the adaptive regularization loss of the -th time, adjust the network parameters of the model obtained by the -th training, and when it is less than or equal to the preset number of times, continue the -th training until is greater than the preset number of times to obtain the first pre-trained model; is an integer greater than ; the first pre-trained model includes: a first pre-trained feature extraction network, the activation function, and a first pre-trained classification network; is an integer greater than ; when , the network parameters of the model obtained by the -th training are the initial network parameters; According to the first pre-trained feature extraction network and the activation function, obtain a first pre-trained hash model.

2. The method for training a hash model according to claim 1, characterized in that, The predicted value of the -th time and the predicted value of the -th time are both -dimensional vectors; represents the number of the preset categories; is an integer greater than or equal to ; Determining the adaptive regularization loss for the -th time according to the hyperparameters of the initial model, the number of the multiple sample images, the prediction value of the -th time, and the prediction values of the -th time for each sample image, includes: For each image, based on the vector inner product between the predicted value of the -th time and the predicted value of the -th time of this image, the inner product value of the -th time of this image is obtained; Based on the inner product values of each image, determine the logarithmic total value; Determine the adaptive regularization loss for the th time according to the hyperparameters, the number of the multiple sample images, and the total logarithm value.

3. The method for training a hash model according to claim 2, characterized in that, The adaptive regularization loss for the ; Among them, represents the adaptive regularization loss for the number of the multiple sample images, represents the th image among the images, represents the inner product value for the th image at the th time, represents the predicted value for the th image at the th time, represents the hyperparameter, represents the logarithmic function.

4. The method for training a hash model according to claim 1, characterized in that, Determining the first similarity retention loss for the th time according to the number of the multiple sample images, the prediction value for the th time, and the class label of each sample image itself, includes: For each image, according to the predicted value of the -th time of this image and the class label of this image itself, calculate the sub-loss of the -th time of this image; Based on the number of the multiple sample images and the sub-losses of each image at the th time, determine the first similarity retention loss at the th time.

5. The method for training a hash model according to claim 4, characterized in that, The first similarity retention loss of the th time is calculated by the following formula: ; Among them, represents the first similarity retention loss of the th time, represents the number of the multiple sample images, represents the th image among the images, represents the sub-loss of the th image, represents the number of preset categories, represents the th category among the categories, represents the value corresponding to the category label of the th image itself in the th category, represents the probability value of the th prediction value of the th image in the th category.

6. The method for training a hash model according to claim 1, characterized in that, The first similarity retention loss for the th time and the adaptive regularization loss for the th time are used to adjust the network parameters of the model obtained from the th training. When the number of times is less than or equal to the preset number of times, continue the th training until is greater than the preset number of times, and a first pre-trained model is obtained, including: ​ Based on the sum of the similarity retention loss of the th time and the adaptive regularization loss of the th time, it is determined as the total loss of the th time; Adjust the network parameters of the model obtained from the -th total loss, and obtain the model obtained from the -th training to get the model obtained from the -th training; Determine whether it is less than or equal to the preset number of times; When less than or equal to the preset number of times, according to the model obtained from the th training, perform the th training until greater than the preset number of times, obtain the first pre-trained model.

7. An image retrieval method for adaptive binary quantization in a noisy environment, characterized in that, Comprising: Obtain an image to be retrieved; Input the image to be retrieved into the first pre-trained hash model according to any one of claims 1 to 6 above to obtain a hash code value; According to the hash code value, retrieve the target image corresponding to the image to be retrieved from a preset image library.

8. A method for training a hash model, characterized in that, Comprising: Obtain a training set including multiple sample images and an initial model; The initial model includes: an initial feature extraction network, an activation function, and an initial classification network; the initial feature extraction network and the initial classification network correspond to initial network parameters; When each sample image in the training set has its own class label, at the th training, input the sample image into the initial model to obtain the hash code value of each sample image at the th time and the prediction value at the th time. The prediction value of each sample image is the probability value that the sample image belongs to each preset category; Based on the number of the multiple sample images, the prediction value of the -th time, the class label of each sample image, the hyperparameters of the initial model, and the prediction value of each sample image at the -th time, determine the first similarity retention loss of the -th time and the adaptive regularization loss of the -th time. According to the first similarity retention loss and the adaptive regularization loss, adjust the network parameters of the model obtained by the -th training until the number of training times reaches the preset number of times, and obtain the first pre-trained model; is an integer greater than . When , the network parameters of the model obtained by the -th training are the initial network parameters; the first pre-trained model includes: a first pre-trained feature extraction network, the activation function, and a first pre-trained classification network; When there is a sample pair class label between any two sample images in the training set, at the th training, input the sample images into the initial model to obtain the hash code values of each sample image at the th time. According to the sample pair class labels of each sample image and the hash code values of each sample image at the th time, determine the second similarity retention loss at the th time. Adjust the network parameters of the model obtained from the th training according to the second similarity retention loss at the th time until the training times reach the preset times, and obtain a second pre-trained model; the second pre-trained model includes: a second pre-trained feature extraction network, the activation function, and a second pre-trained classification network; According to the first pre-trained feature extraction network and the activation function, obtain a first pre-trained hash model, or, according to the second pre-trained feature extraction network and the activation function, obtain a second pre-trained hash model.

9. The method for training a hash model according to claim 8, characterized in that, The second similarity retention loss of the -th time is calculated by the following formula: ; ; ; ; Among them, represents the network parameters of the model obtained from the th training, represents the second similarity retention loss of the th time, represents the number of sample images input for the th time, represents the th image among the th sub - loss of the th image, represents the number of images in the sample pair category labels between the sample images input for the th time and the th image, where the th image and the corresponding sample image belong to the same category, represents the number of images in the sample pair category labels between the sample images input for the th time and the th image, where the th image and the corresponding sample image belong to different categories, represents the hash - coding value of the th th image, represents the hash - coding value of the th th sample image among the sample images that belong to the same category as the th th time, represents the hash - coding value of the th th sample image among the sample images that belong to different categories from the th th time, represents the th similarity score among represents the th dissimilarity score among represents a metric function that measures the Hamming distance between two hash - coding values, represents the logarithmic function.

10. The method for training a hash model according to claim 8, characterized in that, The second similarity retention loss according to the times is used to adjust the network parameters of the model obtained by the th training until the number of training times reaches the preset number of times, and a second pre-trained model is obtained, including: Using the second similarity retention loss of the th time, adjust the network parameters of the model obtained by the th training to obtain the model obtained by the th training; Determine whether it is less than or equal to the preset number of times; When less than or equal to the preset number of times, according to the model obtained from the th training, the training is carried out until greater than the preset number of times, the second pre-training model is obtained; is an integer greater than .

Citation Information

Patent Citations

  • Cross-modal hash retrieval method based on triple deep networks

    CN108170755A

  • Deep non-relaxation Hash image retrieval method based on point pair similarity

    CN109783682A