Hash retrieval method for solving concept drift phenomenon

Through the deep incremental sampling hashing method, the robust feature learning and knowledge distillation technology of deep neural networks are used to solve the problem of retrieval performance degradation caused by concept drift in dynamic data environments, and a more discernible hash encoding is generated, which improves the adaptability and efficiency of hash retrieval.

CN120234434APending Publication Date: 2025-07-01SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510311372.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing dynamic deep hash retrieval method fails to effectively solve the concept drift phenomenon, resulting in a significant decline in retrieval performance, and online training of hash networks is time-consuming and impractical.

Method used

The deep incremental sampling hashing method is used to utilize the robust feature learning ability of deep neural networks, and the deep hash network is updated online, combining knowledge distillation and multi-loss function training to generate hash encoding to alleviate catastrophic forgetting caused by concept drift.

Benefits of technology

It effectively solves the problem of retrieval performance degradation caused by concept drift, improves the adaptability and efficiency of hash retrieval, reduces old information forgetting, and generates a more discernible hash encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234434A_ABST
    Figure CN120234434A_ABST
Patent Text Reader

Abstract

The invention discloses a Hash retrieval method for solving a concept drift phenomenon. The Hash retrieval method comprises the following steps: acquiring a target image and preprocessing the target image; performing feature extraction by using an online updated deep hash network and calculating a feature vector center of each class of training set image; using the sample set of the previous time step and the training set image of the current batch to train a Hash model; according to the designed loss function, utilizing knowledge distillation to guide deep Hash network training, and generating a Hash code; and calculating the similarity according to the Hash codes of the training set images and the test set images, and sorting. According to the method, the problem of online updating of the deep hash network and the problem of disastrous forgetting occurring along with continuous accumulation of images are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning and image retrieval, and particularly to a hashing retrieval method for solving the concept drift phenomenon. Background Art

[0002] With the explosive growth of Internet information, the efficient management of a large number of images has become increasingly important. One of the most basic tasks is the effective retrieval of large-scale images. However, due to the redundancy and inaccuracy of image tags, retrieval methods relying solely on image tags are usually inefficient. Therefore, image retrieval methods based on visual features have become increasingly popular. In practical applications, the data environment is usually online and non-stationary, and data usually appears in the form of data streams or continuous data blocks, that is, images appear in a continuous form or as a set of multiple images in batches. In this case, the concept drift phenomenon is inevitable, such as the appearance of new classes or the change in the distribution of existing classes of images. As images appear in batches, new classes of images that have not appeared in the previous batches may appear, that is, the phenomenon of new class addition; or the class distribution of images in the new batch may be inconsistent with the class distribution of images in the previous batch, that is, the phenomenon of image distribution change. Therefore, traditional hashing methods trained only based on the initial image set cannot solve the concept drift phenomenon, resulting in a significant decline in retrieval performance. In addition, it will be time-consuming to retrain the hashing network using all the training sets that have appeared at each time step, and it is impractical to reuse the historical training sets in practical applications. In order to make the hashing retrieval method better adapt to the real dynamic data environment, a dynamic deep hashing retrieval method has been proposed.

[0003] Although significant progress has been made in the current research on dynamic deep hashing retrieval, the existing dynamic depth methods do not consider the concept drift phenomenon, and the existing dynamic depth methods perform poorly in the data scenario of concept drift. For this reason, we propose a deep incremental sampling hashing method that simultaneously performs feature extraction and online network training using the robust feature learning ability of a deep neural network. In addition, the deep neural network reduces the difference between the network output and the target binary code by minimizing the binary cross-entropy loss function, maintains the similarity between samples by the similarity-preserving loss function, and uses the hashing code balance loss function to ensure the balance of the hashing code. Summary of the Invention

[0004] The object of the present invention is to overcome the deficiencies of the prior art and propose a hashing retrieval method for solving the concept drift phenomenon, which can solve the problem of online network update and the catastrophic forgetting problem that occurs as images accumulate continuously.

[0005] To achieve the above object, the technical solution provided by the present invention is: a hash retrieval method for solving the concept drift phenomenon, comprising the following steps:

[0006] S1: Obtain a target image and perform preprocessing to obtain a target image of a unified size; divide the target image into multiple batches, and divide them into a training set image and a test set image according to a ratio, which are respectively used for the training phase and the test phase; different batches of target images appear at different time steps, so as to simulate a dynamic data environment and the possible concept drift phenomenon that may follow. The concept drift includes: as the images appear in batches, new categories of images that have never appeared in the previous batches of images will appear, that is, the phenomenon of new class addition; and the category distribution of the images in the new batch is inconsistent with the category distribution of the images in the previous batch, that is, the phenomenon of image distribution change;

[0007] S2: In the training phase, use an online updated deep hash network to extract features from each batch of target images to obtain image feature vectors with richer semantic information, and calculate the feature vector centers of each category of images in the training set images of the current batch;

[0008] S3: According to the feature vector centers of each category of images in the training set images of the current batch, calculate by weighted average to online update the feature vector centers of each category of images in all batches of training set images up to the current moment as the historical feature vector centers of each category of images; calculate the distances between the historical feature vector centers of each category of images and the feature vectors of each training set image, select the M training set images closest to the historical feature vector centers of each category of images as the sample set, and save them; mix the training set images of the current batch with the sample set saved in the previous batch as the training images, and use them together for the training of the deep hash network in step S2;

[0009] S4: Use knowledge distillation to guide the training of the deep hash network in step S2 to alleviate the catastrophic forgetting problem caused by the concept drift phenomenon; among them, use the deep hash network in the previous time step as the teacher hash network, and use the deep hash network in the current time step as the student hash network; first use the teacher hash network to learn the training images, guide the training of the teacher hash network according to the designed loss function, use the probability distribution generated by the teacher hash network as the supervision signal of the student hash network, then use the training images to train the student hash network, guide the update of the network parameters of the student hash network through the backpropagation algorithm, and finally generate the hash code of the training images through the hash layer mapping;

[0010] S5: In the testing stage, use the trained student hash network to directly generate the hash codes of the current batch of test set images and all the training set images up to the current batch. Calculate the similarity between the test set images and the training set images based on their hash codes and sort them. For each test set image, select the top n images with the highest similarity in the training set images as the final query results.

[0011] Furthermore, in step S1, preprocess the target images, including random cropping, random horizontal flipping, color jittering, and normalization operations. Divide the target images into multiple batches according to time steps and proportionally divide them into training set images and test set images. Simulate the concept drift phenomenon by introducing new category images in different batches of target images and making the category distribution of the target images fit a Gaussian distribution.

[0012] Furthermore, in step S2, use a deep hash network based on ResNet-18 to extract the image feature vectors of each batch of target images. Among them, use the network parameters of the deep hash network trained in the previous time step to initialize the deep hash network at the current time step. And by taking the mean value, calculate the feature vector center μ of each category of images in the current batch of training set images. The formula is as follows:

[0013]

[0014] In the formula, x represents an image in a set X of images of a certain category in the current batch of training set images, represents the feature extraction function, and n represents the number of images in the set X of images of this category.

[0015] Furthermore, the specific operation steps of step S3 are as follows:

[0016] S31: According to the feature vector centers of each category of images in the current batch of training set images, calculate the weighted average to online update the feature vector centers of each category of images in all batches of training set images up to the current moment, and use them as the historical feature vector centers of each category of images. In this way, the calculation does not require calling historical images again, thus improving the calculation efficiency of the historical feature vector center μ t The efficiency, and the calculation formula is as follows:

[0017]

[0018] In the formula, α t-1 represents the total number of training set images of this category up to the t-1 moment; β t represents the number of training set images of this category at the t moment; represents the current batch, i.e., the center of the feature vectors of the training set images of this category at time t; μ t-1 represents the historical center of the feature vectors of the training set images of this category up to time t - 1;

[0019] S32: Calculate the distances between the historical centers of the feature vectors of each category of images and the feature vectors of each training set image, and select the M training set images that are closest to the historical centers of the feature vectors of each category of images as the sample set. In this way, it is required that the average feature vector of a certain category of images in the sample set is close to the average feature vector of all the training set images of this category seen so far. Essentially, the constructed sample set is a priority list, where the instances with higher rankings will better fit the average feature vector p of all the training set images of this category k , therefore, iterate through the feature vectors of all the training set images in sequence, divide them by the number of the current category in the sample set to calculate the average, and gradually select the first m images that are closest to the average feature vector center μ and add them to the sample sets of the corresponding categories. The calculation formula is as follows:

[0020]

[0021] In the formula, x represents an image in a set X of a certain category of images in the training set images of the current batch, μ is the average feature vector center of the training set images of the current category, m represents the number of images of a certain category in the sample set, represents the feature extraction function, p r represents the r-th ranked image of the current category in the sample set;

[0022] S33: Mix the training set images of the current batch with the sample set saved in the previous batch as the training images, and use them together for the training of the deep hashing network in step S2. The formula for the output function g(x) of the deep hashing network is as follows:

[0023]

[0024] In the formula, ω t represents the network weight parameters of the deep hashing network at the current time t.

[0025] Furthermore, the specific operation steps of step S4 are as follows:

[0026] S41: According to the principle of knowledge distillation, the deep hashing network in the previous time step is used as the teacher hashing network, and the deep hashing network in the current time step is used as the student hashing network; the teacher hashing network is used to learn the training images obtained after mixing, and the training of the teacher hashing network is guided according to the designed loss function, where the loss function consists of three parts, namely the binary cross-entropy loss function, the similarity preservation loss function, and the hash code balance loss function; the probability distribution output by the teacher hashing network is used as the additional supervised soft label of the student hashing network to adjust the learning of the student hashing network to reduce the forgetting of old knowledge when new categories appear; the formula for the output probability distribution q of the teacher hashing network is as follows:

[0027] q = g′(x)

[0028] In the formula, g′(x) represents the output function of the teacher hashing network;

[0029] S42: According to the additional supervised soft label output by the teacher hashing network and the image label, a binary cross-entropy loss function L1 is constructed; the student hashing network is made to output the correct class symbol for the training images of the newly emerging categories and to output close to the output of its teacher hashing network for the training images of the existing categories; the binary cross-entropy loss function L1 consists of two parts, namely the classification loss and the distillation loss. Among them, when the class label of the training image belongs to the sample set, that is, the existing category, it participates in the calculation of the distillation loss; when the class label of the training image belongs to the new category, it participates in the calculation of the classification loss, and the formula is as follows:

[0030]

[0031] In the formula, P represents the sample set, X t represents the current batch of training set images arriving at time t, x i represents one of the training images obtained by mixing P and X t y i represents the sample label predicted by the student hashing network for the image x i y represents the true label of the image x i m represents the class of the images in the sample set, q i represents the probability distribution output by the teacher hashing network when the image x i belongs to the sample set, g y (x i ) represents the output of the student hashing network, represents when the sample label y i predicted by the student hashing network is consistent with the true label y, represents when the sample label y i predicted by the student hashing network is inconsistent with the true label y, n tDenote the category of the newly emerging image at time t;

[0032] S43: To minimize the Hamming distance between two similar images and maximize the Hamming distance between dissimilar images, a similarity-preserving loss function L2 was constructed;

[0033] First, using the pairwise similarity matrix S that can indicate the similarity relationship between images, construct the posterior check estimate logP(B|S) of the hash code set B. According to Bayes' theorem, the posterior probability is proportional to the product of the likelihood part P(S|B) and the prior part P(B), that is, P(B|S) ∝ P(S|B)P(B); since the denominator P(S) is a constant term and can be ignored during optimization, it can be directly written as a proportional relationship. After taking the logarithm, the formula is expressed as follows:

[0034]

[0035] In the formula, i and j respectively represent a certain image in the current total number of images N, s ij represents the similarity value between image x i and image x j in the pairwise similarity matrix S, ω ij represents the weight assigned to each training image pair (x i , x j , s ij ), the hash code set B = [b1, b2,..., b i ,..., b j ,..., b N , b i , b j respectively represent the hash codes of image x i and image x j ;

[0036] Among them, represents the weighted likelihood function, represents that the prior probability distribution assumes that each hash code b i is independently generated, so it is modeled as a Bernoulli distribution and converted to a summation form after taking the logarithm;

[0037] Secondly, because the number of similar pairs in the training images is less than the number of dissimilar pairs, which may lead to an imbalance problem. Considering the impact brought by this, the weight ω ij of the training image pairs was recalculated, and its formula is expressed as follows:

[0038]

[0039] In the formula, S1 = {s ij ∈S: s ij={1} represents the set of similar training image pairs, |S1| represents the number of image pairs contained in S1, and S0 = {s ij ∈S: s ij = 0} represents the set of dissimilar training image pairs, |S0| represents the number of image pairs contained in S0, and |S| represents the total number of image pairs in the training images;

[0040] Finally, since P(s ij |B) is the conditional probability of the similarity label s ij given the corresponding set of hash codes B, therefore, P(S|B) can be naturally obtained from the Bernoulli distribution because for each similarity label s ij ∈ {0, 1}, the conditional probability is defined as:

[0041]

[0042] In the formula, Θ ij represents the inner product of the hash code b i of the image x i and the hash code b j of the image x j ; where σ(·) is the Sigmoid function;

[0043] Since maximizing the posterior probability is equivalent to minimizing the negative log-likelihood function, for each image pair (i, j), based on the pairwise similarity matrix S, calculate its negative log-likelihood loss, and construct the similarity-preserving loss function L2. After simplification, the formula is expressed as follows:

[0044]

[0045] S44: Considering the discrete probability distribution of the binary hash code, use to calculate the average value of the s-th bit of the hash codes of all training images in the current batch; when minimizing the above formula, each bit in the hash code will be balanced, that is, the probability of each bit taking the values of -1 and 1 is approximately equal, which can maximize the information content of each bit; in addition, the average relationship between the bits of the hash codes of two different training images is also considered at the same time, ensuring the balance of each bit and enhancing the distinguishability between the hash codes of different training images; the formula of the hash code balance loss function L3 is expressed as follows:

[0046]

[0047] In the formula, K represents the number of bits of the hash code, and respectively represent the average values of the s-th bits of the hash codes of two different training images;

[0048] S45: Finally, the formula of the overall loss function L for training the deep hashing network is expressed as follows:

[0049]

[0050] In the formula, α and β respectively represent the weights of the similarity preservation loss function L2 and the hash code balance loss function L3;

[0051] S46: Using the designed loss function L, the update of the network parameters of the deep hashing network is guided by the backpropagation algorithm. Finally, through the mapping of the hash layer, the hash code of the training image is generated.

[0052] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0053] 1. The present invention supports the online update of the deep hashing network and the feature center of the training image, avoiding the overly complex problem when using all the images that have appeared in the online data environment to train the hashing network.

[0054] 2. The present invention generates hash codes in an end-to-end manner. The binary cross-entropy loss function, the hash code balance loss function, and the similarity preservation loss function are used to train the deep hashing network. By optimizing a single loss function, the entire deep hashing network is trained, avoiding the potential mismatch problem when using multiple modules.

[0055] 3. The present invention constructs a sample set with representative image features in a dynamic data environment, learns the deep hashing network based on new images and the images in the sample set, can effectively reduce the forgetting of old information, and provides guidance for the classification of the deep hashing network.

[0056] 4. The present invention uses knowledge distillation to guide the training of the deep hashing network. According to the designed loss function, the training of the teacher hashing network is guided, and the probability distribution generated by the teacher hashing network is used as the supervision signal for the student hashing network. Then the student hashing network is trained, and the hash code of the training image is generated. This effectively solves the catastrophic forgetting problem caused by the concept drift phenomenon.

[0057] In summary, the present invention can utilize the robust feature learning ability of the deep neural network to perform feature extraction and network online update simultaneously. By minimizing the binary cross-entropy loss function, the difference between the network output and the target binary code is reduced. By using the similarity preservation loss function, the similarity between images is maintained. The hash code balance loss function is used to ensure the balance of the hash code. It can solve the problem of network online update and the catastrophic forgetting problem that occurs as data accumulates. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1Schematic diagram of the feature extraction and deep hashing network learning logic process of the present invention.

[0059] Figure 2 Frame diagram of the method of the present invention. Specific implementation manners

[0060] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0061] As Figure 1 and Figure 2 shown, this embodiment discloses a hashing retrieval method for solving the concept drift phenomenon, which uses Resnet-18 as the base network for feature extraction, and includes the following steps:

[0062] 1) Feature learning:

[0063] For image feature extraction, a deep hashing network based on ResNet-18 is used to extract the image feature vectors of each batch of target images. Among them, the network parameters of the deep hashing network trained in the previous time step are used to initialize the deep hashing network of the current time step. And by calculating the mean value, the feature vector center μ of each category of images in the training set images of the current batch is calculated, and the formula is expressed as follows:

[0064]

[0065] In the formula, x represents an image in a certain category image set X in the training set images of the current batch, represents the feature extraction function, and n represents the number of images in the category image set X.

[0066] 2) Deep hashing network training:

[0067] Knowledge distillation is used to guide network training to alleviate the catastrophic forgetting problem caused by the concept drift phenomenon. The deep hashing network in the previous time step is used as the teacher hashing network, and the deep hashing network in the current time step is used as the student hashing network. The teacher hashing network is used to learn the mixed training images. The probability distribution output by the teacher hashing network is used as the additional supervised soft label of the student network to adjust the learning of the student hashing network to reduce the forgetting of old knowledge when new categories appear. The formula for the probability distribution q output by the teacher hashing network is expressed as follows:

[0068] q = g′(x)

[0069] In the formula, g′(x) represents the output function of the teacher hashing network.

[0070] 3) Online update of the feature center and the sample set

[0071] Use the online-updated deep hashing network to extract features from each batch of target images, obtaining image feature vectors with richer semantic information. According to the feature vector centers of images of each category in the training set images of the current batch, through weighted average calculation, online update the feature vector centers of images of each category in all batches of training set images up to the current moment, and use them as the historical feature vector centers of images of each category; calculate the distances between the historical feature vector centers of images of each category and the feature vectors of each training set image, select the M training set images with the closest distances to the historical feature vector centers of images of each category as the sample set, and save them; mix the training set images of the current batch with the sample set saved in the previous batch as the training images, and use them together for the training of the deep hashing network, which includes the following steps:

[0072] 3.1) According to the feature vector centers of images of each category in the training set images of the current batch, through weighted average calculation, online update the feature vector centers of images of each category in all batches of training set images up to the current moment, and use them as the historical feature vector centers of images of each category. In this way, the calculation does not require calling historical images again, thereby improving the efficiency of calculating the historical feature vector center μ t The efficiency is expressed by the following calculation formula:

[0073]

[0074] In the formula, α t-1 represents the total number of training set images of this category up to the (t - 1)th moment, β t represents the number of training set images of this category at the tth moment, represents the current batch, that is, the feature vector center of the training set images of this category at the tth moment, and μ t-1 represents the historical feature vector center of the training set images of this category up to the (t - 1)th moment.

[0075] 3.2) Calculate the distances between the historical feature vector centers of images of each category and the feature vectors of each training set image, and select the M training set images with the closest distances to the historical feature vector centers of images of each category as the sample set. In this way, it is required that the average feature vector of a certain category of images in the sample set is as close as possible to the average feature vector of all training set images of this category seen so far. Essentially, the constructed sample set is a priority list, where instances with higher rankings are more fitting to the average feature vector p k of all training set images of this category. Therefore, iterate through the feature vectors of all training set images in turn, divide them by the number of the current category in the sample set to calculate the average value, and gradually select the first m images closest to the average feature vector center μ to add to the sample set of the corresponding category. The calculation formula is expressed as follows:

[0076]

[0077] wherein, x represents an image in a set X of images of a certain category in the training set images of the current batch, μ is the average feature vector center of the training set images of the current category, m represents the number of images of a certain category in the sample set, represents a feature extraction function, p r represents the r-th ranked image of the current category in the sample set.

[0078] 3.3) Mix the training set images of the current batch with the sample set saved in the previous batch as training images and use them together for the training of the deep hashing network in step 2). The formula of the output function g(x) of the deep hashing network is expressed as follows:

[0079]

[0080] wherein, ω t represents the network weight parameters of the deep hashing network at the current moment t.

[0081] 4) As Figure 2 shown, the present invention uses knowledge distillation to guide the training of the deep hashing network and alleviates the catastrophic forgetting problem caused by the concept drift phenomenon. First, use the teacher hashing network to learn the training images obtained after mixing, and guide the training of the teacher hashing network according to the designed loss function. The loss function includes three parts, namely the binary cross-entropy loss function, the similarity preservation loss function, and the hash code balance loss function. The probability distribution output by the teacher hashing network is used as the additional supervised soft label of the student network to adjust the learning of the student hashing network. Finally, through the mapping of the hash layer, the corresponding hash code is obtained. Through the design of the sample set, knowledge distillation, and various loss functions, the problems of online network update and catastrophic forgetting occurring with the continuous accumulation of data are solved. The specific operation steps are as follows:

[0082] 4.1) Construct a binary cross-entropy loss function L1 according to the additional supervised soft label output by the teacher hashing network and the image label. Make the student hashing network output the correct class symbol for the training images of the newly emerging categories and make the training images of the existing categories output close to the output of its teacher hashing network. The binary cross-entropy loss function L1 consists of two parts, namely the classification loss and the distillation loss. Among them, when the class label of the training image belongs to the sample set, that is, the existing category, it participates in the calculation of the distillation loss; when the class label of the training image belongs to the new category, it participates in the calculation of the classification loss. The formula is expressed as follows:

[0083]

[0084] wherein, P represents the sample set, Xt Denote the current batch of training set images arriving at time t, x i Denote P and X t One of the training images obtained by mixing, y i Denote the image x i The sample label predicted by the student hash network, y denotes the image x i The true label, m denotes the category of the images in the sample set, q i Denote when the image x i Belongs to the sample set, the probability distribution output by the teacher hash network, g y (x i ) denotes the output of the student hash network, Denote when the sample label y predicted by the student hash network i Is consistent with the true label y, Denote when the sample label y predicted by the student hash network i Is inconsistent with the true label y, n t Denote the category of the newly emerged image at time t.

[0085] 4.2) To minimize the Hamming distance between two similar images and maximize the Hamming distance between dissimilar images, a similarity-preserving loss function L2 is constructed.

[0086] First, using the pairwise similarity matrix S that can indicate the similarity relationship between images, construct the posterior check estimate logP(B|S) of the hash code set B. According to Bayes' theorem, the posterior probability is proportional to the product of the likelihood part P(S|B) and the prior part P(B), that is, P(B|S) ∝ P(S|B)P(B). Since the denominator P(S) is a constant term and can be ignored during optimization, it can be directly written as a proportional relationship. After taking the logarithm, the formula is expressed as follows:

[0087]

[0088] In the formula, i and j respectively represent a certain image in the current total number of images N, s ij Denote the similarity value between the image x i And the image x j In the pairwise similarity matrix S, ω ij Denote the weight assigned to each training image pair (x i , x j , s ij ), the hash code set B = [b1, b2,..., b i ,..., b j ,..., b N , b i , b j Respectively the image xi and the hash code of image x j .

[0089] Among them, represents the weighted likelihood function, indicating that the prior probability distribution assumes that each hash code b i is generated independently, so it is modeled as a Bernoulli distribution and transformed into a summation form after taking the logarithm,

[0090] Secondly, because the number of similar pairs in the training images is usually much smaller than the number of dissimilar pairs, which may lead to an imbalance problem. Considering the impact brought by this, the weights ω of the training image pairs are recalculated ij , and its formula is as follows:

[0091]

[0092] In the formula, S1 = {s ij ∈S: s ij = 1} represents the set of similar training image pairs, |S1| represents the number of image pairs included in S1, S0 = {s ij ∈S: s ij = 0} represents the set of dissimilar training image pairs, |S0| represents the number of image pairs included in S0, and |S| represents the total number of image pairs in the training images.

[0093] Finally, since P(s ij |B) is the conditional probability of the similarity label s ij given the corresponding set of hash codes B. Therefore, P(S|B) can be naturally obtained from the Bernoulli distribution. Because for each similarity label s ij ∈{0,1}, the conditional probability is defined as:

[0094]

[0095] In the formula, Θ ij represents the inner product of the hash code b i of image x i and the hash code b j of image x j , where σ(·) is the Sigmoid function.

[0096] Since maximizing the posterior probability is equivalent to minimizing the negative log-likelihood function. For each image pair (i,j), based on the pairwise similarity matrix S, its negative log-likelihood loss is calculated, and a similarity-preserving loss function L2 is constructed. After simplification, the formula is as follows:

[0097]

[0098] 4.3) Considering the discrete probability distribution of binary hash codes, use to calculate the average value of the s-th bit of the hash codes of all training images in the current batch. When minimizing the above formula, each bit in the hash code will be balanced, that is, the probabilities of the values -1 and 1 for each bit in the code are approximately equal, which can maximize the information content of each bit. Additionally, the average relationship between the bits of the hash codes of two different training images is also considered simultaneously, ensuring the balance of each bit and enhancing the distinguishability between the hash codes of different training images. The formula for the hash code balance loss function L3 is expressed as follows:

[0099]

[0100] In the formula, K represents the number of bits of the hash code, and respectively represent the average values of the s-th bits of the hash codes of two different training images.

[0101] 4.4) Finally, the formula for the overall loss function L of the deep hash network training is expressed as follows:

[0102]

[0103] In the formula, α and β respectively represent the weights of the loss functions L2 and L3.

[0104] 4.5) At time t, use the designed loss function L to guide the update of the network parameters of the student hash network through the backpropagation algorithm, and finally generate the hash codes of the training images through the mapping of the hash layer.

[0105] 5) In the test stage, use the trained student hash network to directly generate the hash codes of the test set images Q t in the current batch and all the training set images up to the current batch, calculate the similarity between the test set images and the training set images based on their hash codes, and perform sorting; for each test set image, select the top n images with the highest similarity ranking among all the training set images as the final query results.

[0106] The MNIST dataset consists of 70,000 handwritten digit images from 0 to 9, with corresponding labels. Each image is a 28×28 pixel grayscale image of a handwritten digit. The CIFAR10 dataset consists of 60,000 images belonging to 10 classes of real-world objects. Each image is an RGB image with 32×32 pixels and 3 channels. The CIFAR100 dataset consists of 60,000 images belonging to 20 superclasses, and each superclass consists of 5 subclasses. Each image corresponds to a superclass and a subclass. Each subclass includes 600 images. Each image is an RGB image with 32×32 pixels and 3 channels.

[0107] To verify the effectiveness of the experiment, the data environment was further simulated, and the simulation results are shown in Tables 1, 2, and 3.

[0108] Table 1 Data scenario settings for the appearance of new classes

[0109]

[0110] Table 2 Data scenario settings for distribution drift

[0111]

[0112] Table 3 Data scenario settings for combined non-stationarity

[0113]

[0114] Hamming ranking is a commonly used method for evaluating the performance of cross-modal hashing algorithms. In this experiment, the mean Average Precision (MAP) was used as the evaluation criterion.

[0115] To verify the effectiveness of the present invention, using MAP as the evaluation criterion, on three commonly used datasets MNIST, CIFAR-10, and CIFAR-100, when the hash code length is 64 bits, a comparative analysis was carried out with the methods of OSH, ICH, IBL, CIHR, CPH, HMOH, ITQ, REPH, SCPH, OselH, AQOH, and FCOH. The experimental results are shown in Tables 4 and 5.

[0116] Table 4 Mean precision results for the appearance of new class scenarios

[0117]

[0118] Table 5 Mean precision results for the CIFAR100 simulation scenarios

[0119]

[0120] The experimental results can fully prove the effectiveness of the method of the present invention. The present invention retains the previous knowledge in a non-stationary environment by means of knowledge distillation and using a representative sample set, so as to generate more effective and more discriminative hash codes in an end-to-end manner, thereby achieving an improvement in retrieval performance.

[0121] Experimental conclusion: Aiming at the concept drift phenomenon existing in the dynamic data environment of the existing algorithms, the present invention proposes a hash retrieval method for coping with the concept drift phenomenon based on deep learning. Experimental evaluations on three publicly available multi-modal datasets MNIST, CIFAR10, and CIFAR100 show that the use of the method of the present invention has a certain improvement in retrieval accuracy and is superior to the existing methods. In the following research, the problem of how to apply incremental learning in multi-modal retrieval tasks will be explored, which has good application prospects and is worthy of promotion.

[0122] The above-described embodiments are only the preferred embodiments of the present invention, and do not limit the scope of implementation of the present invention. Therefore, all changes made according to the shape and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A hash retrieval method for solving the concept drift phenomenon, characterized in that: The following steps are involved: S1: Acquire a target image and perform preprocessing to obtain a target image of uniform size; divide the target image into multiple batches, and divide them into training set images and test set images in proportion, which are used in the training phase and the test phase respectively; different batches of target images appear at different time steps, so as to simulate a dynamic data environment and the concept drift phenomenon that may occur therewith, wherein the concept drift includes: as images appear in batches, new categories of images that have never appeared in the images of the previous batches appear, i.e., the phenomenon of new category addition; and the category distribution of images in the new batch is inconsistent with the category distribution of images in the previous batch, i.e., the phenomenon of image distribution change; S2: In the training phase, the online updated deep hashing network is used to extract features from each batch of target images to obtain image feature vectors with richer semantic information, and the feature vector centers of each category of images in the current batch of training set images are calculated; S3: According to the feature vector center of each category image in the training set images of the current batch, the feature vector center of each category image in all batches of training set images up to the current moment is updated online through weighted average calculation as the historical feature vector center of each category image; the distance between the historical feature vector center of each category image and the feature vector of each training set image is calculated, and the M training set images with the closest distance to the historical feature vector center of each category image are selected as the sample set, and saved; the training set images of the current batch are mixed with the sample set saved in the previous batch as training images, and used together for the training of the deep hash network in step S2; S4: Use knowledge distillation to guide the deep hash network training in step S2 to alleviate the catastrophic forgetting problem caused by the concept drift phenomenon; wherein the deep hash network in the previous time step is used as the teacher hash network, and the deep hash network in the current time step is used as the student hash network; first use the teacher hash network to learn the training image, guide the training of the teacher hash network according to the designed loss function, use the probability distribution generated by the teacher hash network as the supervision signal of the student hash network, then use the training image to train the student hash network, guide the update of the network parameters of the student hash network through the back propagation algorithm, and finally generate the hash code of the training image through hash layer mapping; S5: In the testing phase, the trained student hash network is used to directly generate hash codes for the test set images of the current batch and all the training set images up to the current batch. The similarity between the test set images and the training set images is calculated based on their hash codes, and the two are sorted. For each test set image, the top n images in terms of similarity among all the training set images are selected as the final query results.

2. A hash retrieval method for solving the concept drift phenomenon according to claim 1, characterized in that: In step S1, the target image is preprocessed, including random cropping, random horizontal flipping, color jittering and standardization operations; the target image is divided into multiple batches according to the time step, and divided into training set images and test set images in proportion; the concept drift phenomenon is simulated by allowing new categories of images to appear in target images of different batches and fitting the target image category distribution with a Gaussian distribution.

3. A hash retrieval method for solving the concept drift phenomenon according to claim 1, characterized in that: In step S2, a deep hash network based on ResNet-18 is used to extract image feature vectors of each batch of target images; the network parameters of the deep hash network trained in the previous time step are used to initialize the deep hash network of the current time step, and the feature vector center μ of each category image in the current batch of training set images is calculated by averaging, and the formula is expressed as follows: In the formula, x represents an image in a certain category image set X in the current batch training set images, Represents the feature extraction function, and n represents the number of images in the image set X of this category.

4. A hash retrieval method for solving the concept drift phenomenon according to claim 1, characterized in that: The specific operation steps of step S3 are as follows: S31: According to the feature vector center of each category image in the current batch of training set images, the feature vector center of each category image in all batches of training set images up to the current moment is updated online through weighted average calculation, and it is used as the historical feature vector center of each category image. In this way, the calculation does not need to call the historical image again, thereby improving the calculation of the historical feature vector center μ t The efficiency is calculated as follows: In the formula, α t-1 represents the total number of training set images of this category up to time t-1; β t Represents the number of training set images of this category at time t; Represents the current batch, that is, the feature vector center of the training set image of this category at time t; μ t-1 Represents the historical feature vector center of the training set image of this category up to time t-1; S32: Calculate the distance between the center of the historical feature vector of each category image and the feature vector of each training set image, and select the M training set images with the closest distance to the center of the historical feature vector of each category image as the sample set. This requires that the average feature vector of a category image in the sample set is close to the average feature vector of all training set images of the category seen so far. In essence, the constructed sample set is a priority list, in which the high-ranking instances will better fit the average feature vector p of all training set images of the category. k Therefore, the feature vectors of all training set images are iterated in turn, and the average value is calculated by dividing them by the number of current categories in the sample set. The first m images closest to the average feature vector center μ are gradually selected and added to the sample set of the corresponding category. The calculation formula is as follows: In the formula, x represents an image in a certain category of image set X in the current batch of training set images, μ is the average feature vector center of the training set images of the current category, and m represents the number of images of a certain category in the sample set. represents the feature extraction function, p r Represents the rth ranked image of the current category in the sample set; S33: The training set images of the current batch are mixed with the sample set saved in the previous batch as training images, and used together for the training of the deep hash network in step S2. The formula of the output function g(x) of the deep hash network is expressed as follows: In the formula, ω t Represents the network weight parameters of the deep hashing network at the current time t.

5. A hash retrieval method for solving the concept drift phenomenon according to claim 1, characterized in that: The specific operation steps of step S4 are as follows: S41: According to the principle of knowledge distillation, the deep hash network in the previous time step is used as the teacher hash network, and the deep hash network in the current time step is used as the student hash network; the teacher hash network is used to learn the mixed training image, and the training of the teacher hash network is guided by the designed loss function, where the loss function contains three parts, namely, the binary cross entropy loss function, the similarity preservation loss function and the hash coding balance loss function; the probability distribution output by the teacher hash network is used as an additional supervised soft label of the student hash network to adjust the learning of the student hash network to reduce the forgetting of old knowledge when new categories appear; the formula of the output probability distribution q of the teacher hash network is expressed as follows: q=g′(x) Where g′(x) represents the output function of the teacher hashing network; S42: Based on the additional supervised soft labels and image labels output by the teacher hash network, a binary cross entropy loss function L1 is constructed; the student hash network is made to output the correct category symbol for the training images of the newly appeared category, and the output of the training images of the existing category is made close to the output of its teacher hash network; the binary cross entropy loss function L1 consists of two parts: classification loss and distillation loss. When the category label of the training image belongs to the sample set, that is, the category that has appeared, it participates in the calculation of the distillation loss; when the category label of the training image belongs to the new category, it participates in the calculation of the classification loss. The formula is as follows: In the formula, P represents the sample set, X t represents the current batch of training set images at time t, x i Represents P and X t One of the mixed training images, y i Represents image x i The sample label predicted by the student hashing network, y represents the image x i The true label of m represents the category of the image in the sample set, q i It means that when the image x i The probability distribution of the teacher hash network output when it belongs to the sample set, g y (x i ) represents the output of the student hashing network, Represents the sample label y predicted by the student hashing network i When it is consistent with the true label y, Represents the sample label y predicted by the student hashing network i When it is inconsistent with the true label y, n t Indicates the category of the new image that appears at time t; S43: In order to minimize the Hamming distance between two similar images and maximize the Hamming distance between dissimilar images, a similarity-preserving loss function L2 is constructed; First, using the pairwise similarity matrix S that can indicate the similarity relationship between images, the posterior check estimate logP(B|S) of the hash code set B is constructed. According to Bayes' theorem, the posterior probability is proportional to the product of the likelihood part P(S|B) and the prior part P(B), that is, P(B|S)∝P(S|B)P(B); since the denominator P(S) is a constant term and can be ignored during optimization, it can be directly written as a proportional relationship. After taking the logarithm, the formula is expressed as follows: In the formula, i and j represent an image in the total number of current images N, s ij Represents the pairwise similarity matrix S of image x i and image x j The similarity value between ij Denotes the value assigned to each training image pair (x i ,x j ,s ij ), the hash code set B = [b1, b2, ..., b i ,...,b j ,...,b N ], b i 、b j Represents the image x i and image x j The hash code of in, represents the weighted likelihood function, The prior probability distribution is assumed to be that each hash code b i They are generated independently, so they are modeled as Bernoulli distributions, and converted into sum form after taking the logarithm; Secondly, because the number of similar pairs in the training images is less than the number of dissimilar pairs, this may cause an imbalance problem. Considering this impact, the weights ω of the training image pairs are recalculated. ij , the formula is as follows: Where, S1 = {s ij ∈S:s ij =1} represents a set of similar training image pairs, |S1| represents the number of image pairs contained in S1, S0 = {s ij ∈S:s ij =0} represents the set of dissimilar training image pairs, |S0| represents the number of image pairs contained in S0, and |S| represents the total number of image pairs in the training images; Finally, due to P(s ij |B) is the similarity label s given the corresponding hash code set B ij Therefore, P(S|B) can be naturally derived from the Bernoulli distribution, because for each similarity label s ij ∈{0,1}, the conditional probability is defined as: In the formula, Θ ij Represents image x i The hash code of b i and image x j The hash code of b j The inner product of ij =b i T b j ; Where σ(·) is the Sigmoid function; Since maximizing the posterior probability is equivalent to minimizing the negative log-likelihood function, for each image pair (i, j), based on the pairwise similarity matrix S, its negative log-likelihood loss is calculated, and the similarity preservation loss function L2 is constructed. After simplification, the formula is expressed as follows: S44: Considering the discrete probability distribution of binary hash codes, use Calculate the average value of the sth bit of the hash code of all training images in the current batch; when the above formula is minimized, each bit in the hash code will be balanced, that is, the probability of each bit in the code taking the value of -1 and 1 is roughly equal, which can maximize the amount of information in each bit; in addition, the average relationship between the bits of the hash codes of two different training images is also considered at the same time, ensuring the balance of each bit and enhancing the distinguishability between the hash codes of different training images; the formula of the hash code balance loss function L3 is expressed as follows: In the formula, K represents the number of bits of the hash code, and Respectively represent the average of the sth bit of the hash code of two different training images; S45: Finally, the formula of the overall loss function L of the deep hash network training is as follows: Where α and β represent the weights of the similarity preservation loss function L2 and the hash coding balance loss function L3 respectively; S46: Use the designed loss function L to guide the update of the network parameters of the deep hash network through the back-propagation algorithm, and finally generate the hash code of the training image through hash layer mapping.

Citation Information

Cited By

  • An incremental image hash retrieval method based on feature enhancement and knowledge distillation

    CN122527360B