Image retrieval method, device and equipment based on dynamic hash code and storage medium
Through the image retrieval method based on dynamic hash code, the pre-learning hash projection matrix and Manhattan distance are used to solve the problem of large amount of image retrieval and low efficiency in the prior art, and efficient image retrieval is achieved.
Patent Information
- Application Number
- CN202510067079.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has a large amount of calculation in image retrieval, resulting in low retrieval efficiency.
The image retrieval method based on dynamic hash code is used to calculate the real-value matrix of the image to be query and the candidate image in the intermediate space through the pre-learning hash projection matrix, and binarized and representative values are calculated, and the matching image is determined using the Manhattan distance.
The calculation amount of image retrieval is reduced, the retrieval efficiency is improved, and images can be effectively retrieved in a dynamic data environment.
Smart Images

Figure CN120045732A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technologies, and in particular, to an image retrieval method, apparatus, device, and storage medium based on dynamic hash codes. Background Art
[0002] Due to the continuous growth of the number of images on the Internet, researching efficient methods to retrieve images similar to a query image has become a severe challenge. The earliest method was linear comparison, which directly calculated the Euclidean distance between the query image and candidate images using a brute-force method. However, since real-world data often has thousands of dimensions, this method is prone to the curse of dimensionality. Therefore, due to its extremely low computational and time complexity, the hash method for finding the nearest neighbor has attracted increasing attention. The hash method uses a set of hash functions to map images from a high-dimensional space to a low-dimensional Hamming space and preserves the similarity between data. Each image is converted into a compact binary code. The similarity between images is calculated by the Hamming distance, and a lower Hamming distance represents a higher similarity, and so on.
[0003] However, the computational amount of the existing technology is large, reducing the efficiency of image retrieval. Summary of the Invention
[0004] The main objective of the embodiments of the present application is to propose an image retrieval method, apparatus, device, and storage medium based on dynamic hash codes to reduce the computational amount of image retrieval and improve the retrieval efficiency.
[0005] To achieve the above objective, on the one hand, an embodiment of the present application proposes an image retrieval method based on dynamic hash codes, and the method includes the following steps:
[0006] Calculating a first real-valued matrix of a query image in an intermediate space based on a pre-learned hash projection matrix;
[0007] Calculating a second real-valued matrix of each candidate image in the intermediate space based on the pre-learned hash projection matrix, and then respectively binarizing each of the second real-valued matrices to obtain respective binarized matrices;
[0008] Calculating a representative value of each of the second real-valued matrices in the intermediate space according to the binarized matrix;
[0009] Calculating the Manhattan distance between the first real-valued matrix and each of the representative values respectively;
[0010] Determining a matching image of the query image among each of the candidate images according to each of the Manhattan distances.
[0011] In some embodiments, calculating the representative values of each of the second real-value matrices in the intermediate space according to the binarization matrix includes the following steps:
[0012] Under each dimension, divide the second real-value matrices with a hash code of 0 into a first subset according to each binarization matrix, and divide the second real-value matrices with a hash code of 1 into a second subset according to each binarization matrix;
[0013] Calculate a first representative value of the first subset under each dimension, and calculate a second representative value of the second subset under each dimension.
[0014] In some embodiments, calculating the first representative value of the first subset under each dimension includes the following steps:
[0015] Calculate a first Gaussian kernel function of each of the second real-value matrices in the first subset under each dimension;
[0016] Calculate the first representative value of the first subset according to the first Gaussian kernel function under each dimension;
[0017] Calculating the second representative value of the second subset under each dimension includes the following steps:
[0018] Calculate a second Gaussian kernel function of each of the second real-value matrices in the second subset under each dimension;
[0019] Calculate the second representative value of the second subset according to the second Gaussian kernel function under each dimension.
[0020] In some embodiments, calculating the first representative value of the first subset according to the first Gaussian kernel function under each dimension includes the following steps:
[0021] Traverse a first value set, use the value currently traversed in the first value set as a first current value, and divide each value in the first value set into a first sample and a second sample with the first current value as a boundary; wherein, the first value set is all values between the minimum value and the maximum value of each of the second real-value matrices in the first subset;
[0022] Calculate a first classification probability of the first sample and a second classification probability of the second sample according to the first Gaussian kernel function;
[0023] Calculate a first mean of the first sample according to the first classification probability and the first Gaussian kernel function, and calculate a second mean of the second sample according to the second classification probability and the first Gaussian kernel function;
[0024] Calculate the first between-class variance of the first current value based on the first classification probability, the second classification probability, the first mean, and the second mean, and then obtain the first between-class variances corresponding to all the first current values after traversing the first value set;
[0025] Determine the maximum value among the first between-class variances as the first representative value;
[0026] Calculating the second representative value of the second subset according to the second Gaussian kernel function in each dimension includes the following steps:
[0027] Traverse the second value set, use the value currently traversed in the second value set as the second current value, and divide each value in the second value set into a third sample and a fourth sample with the second current value as the boundary; wherein, the second value set is all values between the minimum value and the maximum value of each second real value matrix in the second subset;
[0028] Calculate the third classification probability of the third sample and the fourth classification probability of the fourth sample according to the second Gaussian kernel function;
[0029] Calculate the third mean of the third sample according to the third classification probability and the second Gaussian kernel function, and calculate the fourth mean of the fourth sample according to the fourth classification probability and the second Gaussian kernel function;
[0030] Calculate the second between-class variance of the second current value based on the third classification probability, the fourth classification probability, the third mean, and the fourth mean, and then obtain the second between-class variances corresponding to all the second current values after traversing the second value set;
[0031] Determine the maximum value among the second between-class variances as the second representative value.
[0032] In some embodiments, the method further includes the following steps:
[0033] Obtain the pre-stored first kernel density sum and the first classification probability, the second classification probability, the first mean, and the second mean at the previous moment of the current moment; wherein, the first kernel density sum is the sum of the kernel densities corresponding to all values in the first value set at all moments before the current moment;
[0034] Calculate the second kernel density sum corresponding to all values in the first value set at the current moment;
[0035] Calculate the first classification probability at the current moment based on the sum of the first kernel densities, the sum of the second kernel densities, the first classification probability at the previous moment of the current moment, the first Gaussian kernel function, and the first set of values;
[0036] Subtract the first classification probability at the current moment from the sum of the second kernel densities to obtain the second classification probability at the current moment;
[0037] Calculate the first mean at the current moment based on the sum of the first kernel densities, the sum of the second kernel densities, the first classification probability at the current moment, the first classification probability at the previous moment, and the first mean, the first Gaussian kernel function, and the first set of values;
[0038] Calculate the second mean at the current moment based on the sum of the first kernel densities, the sum of the second kernel densities, the second classification probability at the current moment, the second classification probability at the previous moment, and the second mean, the first Gaussian kernel function, and the first set of values;
[0039] Calculate the between-class variance of the first class at the current moment based on the first classification probability, the second classification probability, the first mean, and the second mean at the current moment, and then obtain all the between-class variances of the first class corresponding to the current moment after traversing the first set of values;
[0040] Determine the maximum value among all the between-class variances of the first class at the current moment as the first representative value at the current moment.
[0041] In some embodiments, the calculating the Manhattan distances between the first real-value matrix and each of the representative values respectively includes the following steps:
[0042] Calculate the first Manhattan distance between the first real-value matrix and the first representative value in each dimension, and calculate the second Manhattan distance between the first real-value matrix and the second representative value in each dimension;
[0043] Sum up the first Manhattan distances in each dimension to obtain the first Manhattan distance matrix, and sum up the second Manhattan distances in each dimension to obtain the second Manhattan distance matrix;
[0044] Determine the asymmetric distance matrix between the image to be queried and each of the candidate images according to the binarization matrix, the first Manhattan distance matrix, and the second Manhattan distance matrix;
[0045] The expression of the asymmetric distance matrix is:
[0046]
[0047] Among them, D A represents the asymmetric distance matrix, represents the binarization matrix, M 0 represents the first Manhattan distance matrix, M 1 represents the second Manhattan distance matrix.
[0048] In some embodiments, determining the matching image of the image to be queried in each of the candidate images according to each of the Manhattan distances includes the following steps:
[0049] Determining the candidate image corresponding to the minimum value among each of the Manhattan distances as the matching image of the image to be queried.
[0050] To achieve the above object, on the other hand, an image retrieval device based on dynamic hash codes according to an embodiment of the present application is proposed. The device includes:
[0051] A to-be-query image projection unit, configured to calculate a first real-value matrix of the to-be-query image in an intermediate space based on a pre-learned hash projection matrix;
[0052] A candidate image processing unit, configured to calculate a second real-value matrix of each candidate image in the intermediate space based on the pre-learned hash projection matrix, and then binarize each of the second real-value matrices to obtain each binarization matrix;
[0053] A representative value calculation unit, configured to calculate a representative value of each of the second real-value matrices in the intermediate space according to the binarization matrix;
[0054] A distance calculation unit, configured to calculate the Manhattan distance between the first real-value matrix and each of the representative values respectively;
[0055] An image matching unit, configured to determine the matching image of the to-be-query image in each of the candidate images according to each of the Manhattan distances.
[0056] To achieve the above object, on the other hand, an electronic device according to an embodiment of the present application is proposed. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned image retrieval method based on dynamic hash codes is implemented.
[0057] To achieve the above object, on the other hand, a computer-readable storage medium according to an embodiment of the present application is proposed. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned image retrieval method based on dynamic hash codes is implemented.
[0058] The embodiments of the present application at least include the following beneficial effects:
[0059] The present application can calculate a first real-value matrix of the image to be queried in the intermediate space based on a pre-learned hash projection matrix; calculate second real-value matrices of each candidate image in the intermediate space based on the pre-learned hash projection matrix, and then binarize each second real-value matrix to obtain each binarized matrix; calculate a representative value of each second real-value matrix in the intermediate space according to the binarized matrix; calculate the Manhattan distance between the first real-value matrix and each representative value respectively; and determine a matching image of the image to be queried among each candidate image according to each Manhattan distance. The present application calculates a representative value of a candidate image in the intermediate space based on a pre-learned hash projection matrix, and then retrieves the image most relevant to the image to be queried among each candidate image by using the Manhattan distance between the representative value and the first real-value matrix. Using the representative value reduces the retrieval calculation amount and improves the efficiency of image retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0061] Figure 1 It is a schematic flowchart of an image retrieval method based on a dynamic hash code provided by an embodiment of the present application;
[0062] Figure 2 It is an example flowchart of an image retrieval method based on a dynamic hash code provided by an embodiment of the present application;
[0063] Figure 3 It is a schematic diagram of two subsets obtained by dividing all candidate images according to hash codes in the intermediate space provided by an embodiment of the present application;
[0064] Figure 4 provided by an embodiment of the present application The mean value and representative value α in 0 distribution schematic diagram;
[0065] Figure 5 It is a schematic structural diagram of an image retrieval device based on a dynamic hash code provided by an embodiment of the present application;
[0066] Figure 6 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] In order to make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of this application. They are merely examples of devices and methods that are consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0068] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information. Similarly, the second information may also be referred to as the first information. Depending on the context, the words "if", "when" as used herein may be interpreted as "when...", "while...", or "in response to determining".
[0069] The terms "at least one", "multiple", "each", "any one", etc. used in this application, at least one includes one, two, or more than two, multiple includes two or more than two, each refers to each one in the corresponding multiple, and any one refers to any one in the multiple.
[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0071] Before elaborating in detail on the embodiments of this application, first, some related technologies that may be involved in the embodiments of this application are described as follows:
[0072] Existing hashing methods have achieved good retrieval effects. However, most methods are designed for a stable data environment. When new data appears or the distribution of old data changes, existing static methods need to re-learn all hash functions, which will result in extremely high computational complexity and cannot meet the needs of practical applications. To solve this problem, in recent years, many non-stationary hashing methods in non-stationary data environments have been proposed and have achieved good results.
[0073] Non-stationary hashing methods, also known as dynamic hashing methods, have attracted increasing attention. OKH is the first method to update the hash function using newly emerging data. FOH only updates a small part of the binary code and creates a similarity matrix using multi-label supervision information. SDOH-HC uses two modules based on replay and knowledge distillation to learn hash codes. AQOH learns compact hash codes by minimizing the quantization error between the cosine distance of features and the Hamming similarity of hash codes. MIHash measures whether data distributions overlap and avoids unnecessary hash table updates. SPLH learns a latent intermediate space to capture the relationships between different concepts. HMOH uses the Hadamard matrix as the target code and regards the learning process of the hash function as a set of binary classification tasks. OSelH independently learns a large number of hash functions and then selects a better part from the hash function pool. SSH learns hash functions in a semi-supervised manner using the distribution information of all images and the semantic information of labeled images. AHFSS proposes a learning method based on stochastic gradient descent. FROSH and DFROSH are based on the idea of sketches and use the sketches of the dataset to retain most of the features of the original data.
[0074] In addition, most existing hash methods use the Hamming distance of binary hash codes to evaluate the similarity between images. However, due to the binarization of the hash mapping value of the query image, this will result in the loss of location information. In fact, if binarization is not performed, the hash mapping value in the real-valued form of the query image more precisely retains the location information of the query image. Therefore, many asymmetric distance calculation methods that do not binarize the query image have been proposed. ADBE proposes an asymmetric distance calculation method based on the mean, using the mean of images with hash codes of 0 or 1 as the representative value to calculate the Euclidean distance. ASPH proposes an asymmetric distance calculation method on the hypersphere and uses an adjustment method to calculate the representative value of candidate images. WoRank and WsRank calculate bit-level weights based on the asymmetric distance according to the discriminative ability of each bit.
[0075] Most existing non-stationary hashing methods evaluate the similarity between images based on the Hamming distance. They binarize all candidate images and the query image, which leads to the following problems: First, binarizing the query image doesn't have much effect but only aims to improve the efficiency of distance calculation between images. Although it improves the calculation efficiency to some extent, the cost is a great loss of the position information of the query image. Second, for a given query image, there are many candidate images with the same Hamming distance to it. This results in a decrease in the discriminability of the method for candidate images, making it more difficult to find the most similar image, thus reducing the retrieval accuracy. Many asymmetric distance calculation methods have been proposed, and most of them use non-binarized query images to avoid the loss of position information. However, in the real-world non-stationary data environment, concept drift is an unavoidable problem. Concept drift refers to the change in data distribution or the emergence of new category images. When concept drift occurs, the data environment becomes more complex, leading to a further decrease in accuracy. Additionally, existing asymmetric distance calculation methods are all static. For non-stationary data environments with concept drift, there is currently no corresponding asymmetric distance calculation method.
[0076] Moreover, existing asymmetric distance calculation methods basically use the mean as the representative value of the image distribution. However, real-world data is often non-uniformly distributed. Using the mean to represent the distribution of most samples is inaccurate, which leads to a decrease in the retrieval accuracy. Additionally, all existing asymmetric distance calculation methods are based on stationary data environments. When new data appears, existing methods need to obtain all the old data and frequently recalculate the asymmetric distance according to the old data, which does not meet the retrieval requirements of dynamic hashing.
[0077] To solve at least one problem of the existing technology, embodiments of the present application provide an image retrieval method, device, equipment, and storage medium based on dynamic hash codes. The technical solution of the present application includes: calculating a first real-value matrix of the image to be queried in the intermediate space based on a pre-learned hash projection matrix; calculating a second real-value matrix of each candidate image in the intermediate space based on the pre-learned hash projection matrix, and then binarizing each second real-value matrix to obtain each binarized matrix; calculating the representative value of each second real-value matrix in the intermediate space according to the binarized matrix; calculating the Manhattan distance between the first real-value matrix and each representative value; and determining the matching image of the image to be queried among each candidate image according to each Manhattan distance. The present application calculates the representative value of the candidate image in the intermediate space based on the pre-learned hash projection matrix, and then uses the Manhattan distance between the representative value and the first real-value matrix to retrieve the image most relevant to the image to be queried among each candidate image. Using the representative value reduces the retrieval calculation amount and improves the efficiency of image retrieval.
[0078] The embodiments of the present application provide an image retrieval method, device, equipment and storage medium based on a dynamic hash code, which relates to the technical field of image processing. The image retrieval method, device, equipment and storage medium based on a dynamic hash code provided by the embodiments of the present application can be applied to a terminal, or can be applied to a server, or can also be software running on a terminal or a server. In some embodiments, the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side may be configured as an independent physical server, or may be configured as a server cluster or a distributed system composed of multiple physical servers, or may also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server may also be a node server in a blockchain network; the software may be an application implementing a knowledge extraction method, etc., but is not limited to the above forms.
[0079] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0080] Referring to Figure 1 , the embodiments of the present application provide an image retrieval method based on a dynamic hash code. The method may include but is not limited to S100 to S140, specifically as follows:
[0081] S100: Calculate a first real-value matrix of the image to be queried in the intermediate space based on a pre-learned hash projection matrix.
[0082] S110: Calculate a second real-value matrix of each candidate image in the intermediate space based on the pre-learned hash projection matrix, and then binarize each of the second real-value matrices to obtain respective binarized matrices.
[0083] S120: Calculate the representative values of each of the second real-value matrices in the intermediate space according to the binarization matrix.
[0084] Further, S120 may include the following steps S121 to S122:
[0085] S121: Divide the second real-value matrices with hash code 0 into a first subset according to each of the binarization matrices in each dimension, and divide the second real-value matrices with hash code 1 into a second subset according to each of the binarization matrices in each dimension;
[0086] S122: Calculate the first representative value of the first subset in each dimension, and calculate the second representative value of the second subset in each dimension.
[0087] Even further, calculating the first representative value of the first subset in each dimension in S122 includes the following steps S1221 to S1222:
[0088] S1221: Calculate the first Gaussian kernel function of each of the second real-value matrices in the first subset in each dimension;
[0089] S1222: Calculate the first representative value of the first subset according to the first Gaussian kernel function in each dimension.
[0090] As a more specific implementation manner, S1222 may include the following steps:
[0091] Traverse the first value set, use the value currently traversed in the first value set as the first current value, and divide each value in the first value set into a first sample and a second sample with the first current value as the boundary; wherein, the first value set is all values between the minimum value and the maximum value of each of the second real-value matrices in the first subset;
[0092] Calculate the first classification probability of the first sample and the second classification probability of the second sample according to the first Gaussian kernel function;
[0093] Calculate the first mean of the first sample according to the first classification probability and the first Gaussian kernel function, and calculate the second mean of the second sample according to the second classification probability and the first Gaussian kernel function;
[0094] Calculate the first between-class variance of the first current value according to the first classification probability, the second classification probability, the first mean, and the second mean, and then obtain the first between-class variances corresponding to all the first current values after traversing the first value set;
[0095] Determine the maximum value among each of the first-class between-class variances as the first representative value.
[0096] Furthermore, calculating the second representative value of the second subset under each dimension in S122 includes the following steps S1223 to S1224:
[0097] S1223: Calculate the second Gaussian kernel function of each of the second real-value matrices in the second subset under each dimension;
[0098] S1224: Calculate the second representative value of the second subset according to the second Gaussian kernel function under each dimension.
[0099] As a more specific implementation manner, S1224 may include the following steps:
[0100] Traverse the second value set, take the value currently traversed in the second value set as the second current value, and divide each value in the second value set into a third sample and a fourth sample with the second current value as the boundary; wherein, the second value set is all values between the minimum value and the maximum value of each of the second real-value matrices in the second subset;
[0101] Calculate the third classification probability of the third sample and the fourth classification probability of the fourth sample according to the second Gaussian kernel function;
[0102] Calculate the third mean of the third sample according to the third classification probability and the second Gaussian kernel function, and calculate the fourth mean of the fourth sample according to the fourth classification probability and the second Gaussian kernel function;
[0103] Calculate the second-class between-class variance of the second current value according to the third classification probability, the fourth classification probability, the third mean, and the fourth mean, and then obtain the second-class between-class variances corresponding to all the second current values after traversing the second value set;
[0104] Determine the maximum value among each of the second-class between-class variances as the second representative value.
[0105] Considering that if there is a newly added candidate image, recalculating the representative value requires a large amount of computation, so this embodiment can also directly update the representative value according to the calculation results at the historical moment without recalculation, further reducing the amount of computation.
[0106] The following takes the update of the first representative value as an example for illustration. The update of the second representative value can be implemented by referring to the update method of the first representative value, and will not be elaborated in this embodiment. It should be clear that even though the update method of the second representative value is not directly written in this embodiment, it can still be deduced without doubt by referring to the update method of the first representative value.
[0107] Specifically, the embodiment of the present application may further include the following steps:
[0108] Obtain the pre-stored first kernel density sum and the pre-stored first classification probability, second classification probability, first mean, and second mean at the previous moment of the current moment; wherein, the first kernel density sum is the sum of the kernel densities corresponding to all values in the first value set at all moments before the current moment;
[0109] Calculate the second kernel density sum corresponding to all values in the first value set at the current moment;
[0110] Calculate the first classification probability at the current moment according to the first kernel density sum, the second kernel density sum, the first classification probability at the previous moment of the current moment, the first Gaussian kernel function, and the first value set;
[0111] Subtract the first classification probability at the current moment from the second kernel density sum to obtain the second classification probability at the current moment;
[0112] Calculate the first mean at the current moment according to the first kernel density sum, the second kernel density sum, the first classification probability at the current moment, the first classification probability at the previous moment, the first mean, the first Gaussian kernel function, and the first value set;
[0113] Calculate the second mean at the current moment according to the first kernel density sum, the second kernel density sum, the second classification probability at the current moment, the second classification probability at the previous moment, the second mean, the first Gaussian kernel function, and the first value set;
[0114] Calculate the first between-class variance at the current moment according to the first classification probability, second classification probability, first mean, and second mean at the current moment, and then obtain all the first between-class variances corresponding to the current moment after traversing the first value set;
[0115] Determine the maximum value among all the first between-class variances at the current moment as the first representative value at the current moment.
[0116] S130: Calculate the Manhattan distances between the first real - valued matrix and each of the representative values respectively.
[0117] Further, S130 may include the following steps S131 - S133:
[0118] S131: Calculate the first Manhattan distance between the first real - valued matrix and the first representative value in each dimension, and calculate the second Manhattan distance between the first real - valued matrix and the second representative value in each dimension;
[0119] S132: Aggregate the first Manhattan distances in each dimension to obtain a first Manhattan distance matrix, and aggregate the second Manhattan distances in each dimension to obtain a second Manhattan distance matrix;
[0120] S133: Determine an asymmetric distance matrix between the image to be queried and each of the candidate images according to the binarization matrix, the first Manhattan distance matrix, and the second Manhattan distance matrix;
[0121] The expression of the asymmetric distance matrix is:
[0122]
[0123] where D A represents the asymmetric distance matrix, represents the binarization matrix, M 0 represents the first Manhattan distance matrix, M 1 represents the second Manhattan distance matrix.
[0124] S140: Determine the matching image of the image to be queried among each of the candidate images according to each of the Manhattan distances.
[0125] Further, S140 may include step S141:
[0126] S141: Determine the candidate image corresponding to the minimum value among each of the Manhattan distances as the matching image of the image to be queried.
[0127] Next, specific application examples will be combined to introduce and illustrate the solution of the embodiment of the present application in detail:
[0128] This embodiment presents a dynamic data environment in which the concept drift phenomenon occurs. This embodiment proposes an online asymmetric distance calculation method (abbreviated as ICHAD) based on GKDE and representative value calculation. First, GKDE is used to estimate the distribution of images, and then the method calculates the representative values of images corresponding to 0 and 1 for each dimension of the hash code. In each batch, the intermediate variables involved in the representative value calculation process are recorded and used for the calculation of representative values in the new batch. In this way, ICHAD does not need to access old data to update the representative values, which is very important for asymmetric distance calculation. Through the above process, this embodiment proposes the first method for online calculating asymmetric distance in a non-stationary data environment with concept drift.
[0129] Then, the parameter variables involved in this embodiment are described as follows:
[0130] At time T, represents the set containing all candidate images (or candidate data set), X represents the newly added candidate image, and X q represents the image to be queried (or query data set). The hash function projection matrix is denoted by W, I represents the real-value matrix of the query data set in the intermediate space, respectively represent the real-value matrices of all data in the intermediate space and their hash codes, sgn(·) represents the sign function, and x i and x q respectively represent a sample in X and X q , i represents the i-th sample in the data set, and k represents the k-th dimension of the data set. contains the indices i of candidate images that satisfy , contains the indices i of candidate images that satisfy . and respectively record the real-value vectors of candidate images with hash codes 0 or 1 in a certain dimension. GKD is the abbreviation of Gaussian Kernel Density function, GKDE is the abbreviation of Gaussian Kernel Density Estimation, and the OTSU method is the abbreviation of the Otsu method. and represent and 's Gaussian kernel density functions, h represents the bandwidth of the Gaussian kernel function, and σ represents the standard deviation of candidate images. C 1 and C 2 represent dividing or 's Gaussian kernel function into two categories with the currently iterated representative value point as the dividing line. g minand g max represents or The minimum and maximum values within, at an interval of bandwidth h, and all the values between them form a matrix denoted as L 0 is represented by p 1 and p 2 represent the probabilities that the image is of class C 1 or class C 2 The mean values of class C 1 and m 2 represent the means of class C 1 class and class C 2 The between-class variance between two classes C 2 is represented by δ 1 and C 2 between. α 0 or α 1 respectively represent or The representative value of the Gaussian kernel function in a certain dimension. All the α 0 and α 1 constitute a matrix N 0 ∈R 1×k and N 1 ∈R 1×k . M 0 and M 1 represent the Manhattan distance matrix between the Gaussian kernel function representative value points of the image to be queried and the hash code 0 or 1 in the intermediate space. D A represents the asymmetric distance matrix.
[0131] Specifically, the ICHAD method of this embodiment calculates the asymmetric distance using the hash code of the candidate image and the real-value vector of the image to be queried, and evaluates the similarity between images using the asymmetric distance. The calculation method of the asymmetric distance is divided into two parts: the calculation of the representative value and the calculation of the query-related value. The representative value is calculated based on the GKD function, and the query-related values are used in the final retrieval process. For each query image, the candidate image with the smallest asymmetric distance from it is regarded as the most similar to it. When new data appears, the representative value will be updated online, which avoids frequent access and reading of old data. The proposed asymmetric distance calculation method in this embodiment only needs to access the newly emerging data, which greatly reduces the computational complexity and retrieval time, and realizes the online calculation of the asymmetric distance. Since this embodiment does not involve the learning of the hash function, the hash function is learned by SDOH-HC.
[0132] This embodiment may include the following solutions:
[0133] 1. Asymmetric distance calculation.
[0134] When the time step \(t = 0\), based on the learned hash projection matrix \(W\), ICHAD first maps the query image to the intermediate space as follows:
[0135] \(I = W\) T X q (1)
[0136] where \(I\) represents the real - valued matrix of the query dataset in the intermediate space. The candidate dataset is processed in a similar way, with the difference that the candidate dataset needs to be binarized as follows:
[0137]
[0138] where \(sgn(\cdot)\) represents the sign function. When the time step \(t = T\), ICHAD first calculates the representative value for each dimension \(k\) based on all data Since all data are binarized, in the \(k\) - th dimension, the hash code can be used to divide all data into two subsets, namely and where \(k\) represents the \(k\) - th dimension. contains the indices of candidate images that satisfy while contains the indices of candidate images that satisfy
[0139] Considering that all data must be mapped to the intermediate space before binarization, and the real - valued matrices of all candidate images are available. Based on these real - valued matrices, the data with hash code 0 or 1 can be divided into two subsets in the intermediate space. and respectively record the real - valued vectors of candidate images with hash code 0 or 1, which will be initialized as all - zero matrices. Next, these two matrices will be defined as follows:
[0140]
[0141] Then, analyze the distribution of real - valued vectors in the intermediate space under a specific dimension. For the sake of simplicity in representation, the parameter \(k\) is ignored. The method only uses the non - zero values in and Similar to the calculation method of Hamming distance, for a given query image \(x\) q , ICHAD needs to calculate its Manhattan distance in the intermediate space from the candidate images in and two sets. Using this distance, the method can calculate which class of candidate samples the query sample is closer to.
[0142] However, in this case, the method of this embodiment needs to calculate the Manhattan distance between the image to be queried and all candidate images, which results in a very high computational complexity. Therefore, this embodiment considers using representative values to represent and the distribution of images in the intermediate space. The selection of representative values greatly affects the calculation of the asymmetric distance and the final retrieval effect. The Gaussian kernel functions of the real-valued vectors of candidate images in the two sets are obtained by using the GKDE method respectively. and represent and the Gaussian kernel functions. If there are n images in , then:[[]]
[0143]
[0144] can be calculated in a similar way. After Gaussian kernel density estimation, the frequency of the distribution of candidate images in the intermediate space can be estimated. Then, this embodiment uses this frequency to calculate the representative values of candidate samples. These values need to be able to represent the distribution of the vast majority of samples. If the method uses the average value or the median value as the representative value, it will be easily affected by outliers. The existing calculation methods do not exclude the influence of outliers, which results in a very small number of abnormal images in the images affecting the calculation of the representative value. OTSU is a commonly used threshold setting method in the field of image segmentation. Since it can calculate the point that maximizes the between-class variance, it ensures that the representative value can represent the distribution of the vast majority of images. Therefore, OTSU is very suitable for calculating the representative value. Based on the OTSU method, this embodiment proposes an effective way to calculate the representative value. This embodiment traverses the values between the minimum value g and the maximum value g min and the maximum value g max . The set of all values between g min and g max is represented by L 0 . Taking the currently traversed value as the boundary, the samples are divided into two categories, namely U 1 and U 2 . The probabilities of the samples being classified into these two categories are p 1 and p 2 respectively, and the calculation methods are as follows:
[0145]
[0146] The means of U 1 and U 2 are represented by m 1 and m 2 respectively, and the calculation method of the between-class variance δ 2 is as follows:
[0147] δ 2 = p 1 p 2 (m 1 - m 2 ) 2
[0148]
[0149] After traversing all values, the point that can maximize the between-class variance of U 1 and U 2 is selected as the representative value α 0 . In the same way, the representative value α 1 can also be calculated.
[0150] After calculating the representative values in each dimension, all α 0 and α 1 can form matrix N 0 ∈ R 1×k and N 1 ∈ R 1×k . Subsequently, for a given query image x q , ICHAD calculates the value related to the query to measure the distance between the query image and other candidate images, and the calculation method is as follows:
[0151]
[0152] where d(·) represents the Manhattan distance, M 0 and M 1 represent the Manhattan distance matrices of the query image and all representative values with hash codes of 0 or 1 in the intermediate space. After the above operations in each dimension, the results of all dimensions will be aggregated into the final asymmetric distance, and the calculation method is as follows:
[0153]
[0154] where D A represents the asymmetric distance matrix. Through the above method, the asymmetric distance between the query image and the candidate images is divided into the Manhattan distances between the query image and the representative values in each dimension. The method of this embodiment no longer requires binarizing the query image, but only needs to map the query image to the intermediate space. Based on this, ICHAD makes better use of the position information of the query image. In addition, since the calculation process is divided into each dimension, the representative values can more accurately retain the similarity, and the calculation complexity of the Manhattan distance can also be greatly reduced.
[0155] 2. Online calculation of asymmetric distance.
[0156] According to formula (6), calculating the between-class variance requires four intermediate variables, namely m 1 , m 2 , p 1 , and p 2 . These four represent the distribution of all candidate images. New data starts to appear at t > 0. At this time, the main problem becomes how to calculate the values of these four variables using only the newly emerging images without accessing the old images and make them represent all candidate images. Assume that at this time it is time t = T. At the previous time t = T - 1, the means of U 1 and U 2 are and respectively. The probabilities that candidate images are classified as U 1 or U 2 are and respectively. Assume that φ 1 represents the sum of the kernel densities corresponding to all values in L 0 at all times before time T, and φ 2 represents the sum of the kernel densities corresponding to all values in L 0 at time t = T. For the sake of simplicity of representation, ignoring the parameter k, p 1 and p 2 in each dimension can be updated in the following way:
[0157]
[0158] After the calculation is completed at each moment, the only two items related to the old data in the above formula, φ 1 and will be directly stored in the memory and no longer need to be recalculated at the next moment. This makes ICHAD not need to access the old data. The mean of the kernel density of candidate images can also be calculated in a similar way, and the calculation method is as follows:
[0159]
[0160] The between-class variance will be calculated using the mean and probability after the online update of the above parameters is completed, and the calculation method is as follows:
[0161]
[0162] Each value in L 0 will respectively correspond to two means and two probabilities. When all the points in L 0 have been traversed, the representative values and of candidate samples at the current time stepWill be updated. M 0 and M 1 are values related to the query and do not require accessing old data during the calculation process. Therefore, this embodiment updates them according to the ways of formula (7) and formula (8). Through the above online update method, at a new moment, when new data appears, ICHAD only needs to record the mean and probability obtained from each calculation and store them in memory to achieve the online calculation of the asymmetric distance. Such an approach only incurs a minimal storage cost but can greatly reduce the computational complexity and calculation time.
[0163] Through this online calculation method of the asymmetric distance, this embodiment no longer needs to access all old data for similarity evaluation. Using this online updated asymmetric distance, ICHAD can reduce the computational complexity to a level similar to that of the Hamming distance while utilizing more accurate data position information, which not only reduces the retrieval time but also improves the retrieval accuracy.
[0164] The method of this embodiment can be implemented through the following algorithm:
[0165] ICHAD algorithm at time T:
[0166] Input: The newly added candidate image X, all candidate images The image X to be queried q , the intermediate variables related to the representative value calculation at the previous moment and
[0167] Output: Asymmetric distance matrix D A .
[0168] 1. Use SDOH-HC to obtain the hash codes H of X and and
[0169] 2. Select the image index matrices C with hash codes 0 or 1 respectively 0 and C 1 .
[0170] 3. According to formula (3), use C 0 and C 1 to obtain G 0 and G 1 .
[0171] 4. According to formula (4), calculate the Gaussian kernel function and
[0172] 5. Traverse all values from g min to g max to calculate the intermediate variables related to the representative system.
[0173] 6. According to formulas (10), (11), (12), and (13), using and update and
[0174] 7. According to formula (14), calculate the new representative value by minimizing the between-class variance δ 2
[0175] 8. According to formulas (7) and (8), calculate the values M 0 and M 1 .
[0176] 9. According to formula (9), calculate the asymmetric distance matrix D A .
[0177] 10. Return D A .
[0178] In summary, the technical features of this embodiment include:
[0179] 1. The proposed ICHAD method is a hash-based method. In this embodiment, SDOH-HC is used for hash code learning, which can cope with the emergence of concept drift.
[0180] 2. The ICHAD method proposes a brand-new asymmetric distance calculation method. Based on the OTSU method, in this embodiment, the between-class variance is maximized to select representative values, which can make the representative values better represent the distribution of most samples and utilize more accurate sample position information, thereby improving the accuracy.
[0181] 3. The ICHAD method proposes an online calculation method for asymmetric distance, which can calculate the asymmetric distance without accessing old data, greatly reducing the retrieval time.
[0182] The beneficial effects of this embodiment include:
[0183] This embodiment (ICHAD) overcomes the disadvantages of the existing hash methods based on Hamming distance that cannot utilize accurate position information and cannot cope with concept drift by the proposed online calculation method for asymmetric distance, and solves the problem that the existing asymmetric calculation method cannot calculate the asymmetric distance online, resulting in extremely high computational complexity. This embodiment can efficiently perform retrieval in a dynamic data environment with concept drift, reduce the complexity of the asymmetric distance to be close to that of the Hamming distance, and at the same time achieve excellent retrieval performance.
[0184] This application also provides more specific implementation manners as follows:
[0185] The specific implementation plan of the proposed ICHAD for nearest neighbor image retrieval is as follows Figure 2 shown (in this plan, the maximum number of iterations is 10, the length of the hash code is set to 64, and the bandwidth h is set to 1e-4. Since the hash code uses SDOH-HC learning, all relevant parameters are the same as those set in the original SDOH-HC paper). This implementation plan realizes the function of nearest neighbor image retrieval, and it is required to retrieve 100 samples in the database that are most relevant to the current query sample. At time T, ICHAD first uses the SDOH-HC method to map and binarize the candidate samples into hash codes, and maps the query sample to the intermediate space. Then, all the mapped results are stored in the database, and the asymmetric distance between the samples is calculated using the data stored in the database and the above method, and the similarity between the samples is evaluated based on this. Finally, 100 candidate samples with the smallest asymmetric distance from the query sample are returned as the retrieval results.
[0186] Figure 3 shows a schematic diagram of dividing the candidate image dataset into two subsets according to whether the hash code is 0 or 1, and represent the images with hash codes 0 and 1 respectively in a certain dimension.
[0187] Figure 4 shows the difference between the sample kernel density mean and the representative points obtained by the method of maximizing the between-class variance in this work during the calculation of the representative value. The representative points should be able to represent most of the samples as much as possible without being interfered by a small number of outliers. Obviously, the mean is affected by the outliers with lower density in the kernel function, so it is located to the left of the ICHAD representative points in the figure and cannot more accurately represent the distribution of most of the samples.
[0188] Referring to Figure 5 , the embodiment of this application also provides an image retrieval device based on dynamic hash codes, which can implement the above-mentioned image retrieval method based on dynamic hash codes. The device includes:
[0189] A query image projection unit for calculating the first real-valued matrix of the query image in the intermediate space based on a pre-learned hash projection matrix;
[0190] A candidate image processing unit for calculating the second real-valued matrix of each candidate image in the intermediate space based on the pre-learned hash projection matrix, and then binarizing each of the second real-valued matrices to obtain each binarized matrix;
[0191] A representative value calculation unit for calculating the representative value of each of the second real-valued matrices in the intermediate space according to the binarized matrix;
[0192] A distance calculation unit for calculating the Manhattan distance between the first real-value matrix and each of the representative values respectively;
[0193] An image matching unit for determining a matching image of the image to be queried among the candidate images according to each of the Manhattan distances.
[0194] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented in the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0195] An embodiment of the present application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned image retrieval method based on dynamic hash codes. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0196] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented in the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0197] Please refer to Figure 6 , Figure 6 which shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0198] A processor 601, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., for executing relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0199] A memory 602, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 602 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 602 and are called by the processor 601 to execute the image retrieval method based on dynamic hash codes of the embodiments of the present application;
[0200] An input / output interface 603 for implementing information input and output;
[0201] A communication interface 604 for implementing communication interaction between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0202] A bus 605 for transmitting information between various components of the device (such as a processor 601, a memory 602, an input / output interface 603, and a communication interface 604);
[0203] Among them, the processor 601, the memory 602, the input / output interface 603, and the communication interface 604 achieve communication connections with each other inside the device through the bus 605.
[0204] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned image retrieval method based on a dynamic hash code.
[0205] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.
[0206] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0207] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0208] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0209] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0210] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0211] As used in the specification of this application and the above drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0212] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0213] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0214] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0215] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0216] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.
[0217] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. An image retrieval method based on dynamic hash code, characterized in that: The method comprises the following steps: Calculate a first real-valued matrix of the query image in the intermediate space based on the pre-learned hash projection matrix; Calculate the second real-valued matrix of each candidate image in the intermediate space based on the pre-learned hash projection matrix, and then binarize each of the second real-valued matrices to obtain each binarized matrix; Calculate the representative value of each of the second real-valued matrices in the intermediate space according to the binarized matrix; Calculate the Manhattan distance between the first real-valued matrix and each representative value respectively; A matching image of the query image is determined in each of the candidate images according to each of the Manhattan distances.
2. The image retrieval method based on dynamic hash code according to claim 1 is characterized in that: The step of calculating the representative value of each of the second real-valued matrices in the intermediate space according to the binarized matrix comprises the following steps: In each dimension, the second real-valued matrix having a hash code of 0 is divided into a first subset according to each of the binarized matrices, and in each dimension, the second real-valued matrix having a hash code of 1 is divided into a second subset according to each of the binarized matrices; A first representative value of the first subset is calculated in each dimension, and a second representative value of the second subset is calculated in each dimension.
3. The image retrieval method based on dynamic hash code according to claim 2 is characterized in that: The step of calculating the first representative value of the first subset in each dimension comprises the following steps: Calculate the first Gaussian kernel function of each of the second real-valued matrices in the first subset in each dimension; Calculating the first representative value of the first subset according to the first Gaussian kernel function in each dimension; The step of calculating the second representative value of the second subset in each dimension comprises the following steps: Calculate the second Gaussian kernel function of each of the second real-valued matrices in the second subset in each dimension; The second representative value of the second subset is calculated in each dimension according to the second Gaussian kernel function.
4. The image retrieval method based on dynamic hash code according to claim 3 is characterized in that: The step of calculating the first representative value of the first subset according to the first Gaussian kernel function in each dimension comprises the following steps: Traversing the first value set, taking the currently traversed value in the first value set as the first current value, and dividing each value in the first value set into a first sample and a second sample with the first current value as a boundary; wherein the first value set is all values between the minimum value and the maximum value of each of the second real-value matrices in the first subset; Calculate a first classification probability of the first sample and a second classification probability of the second sample according to the first Gaussian kernel function; Calculate a first mean of the first sample according to the first classification probability and the first Gaussian kernel function, and calculate a second mean of the second sample according to the second classification probability and the first Gaussian kernel function; Calculate a first inter-class variance of the first current value according to the first classification probability, the second classification probability, the first mean, and the second mean, and then obtain the first inter-class variances corresponding to all the first current values after traversing the first value set; determining the maximum value among the first inter-class variances as the first representative value; The step of calculating the second representative value of the second subset according to the second Gaussian kernel function in each dimension comprises the following steps: The second value set is traversed, and the currently traversed value in the second value set is used as the second current value, and each value in the second value set is divided into a third sample and a fourth sample with the second current value as a boundary; wherein the second value set is all values between the minimum value and the maximum value of each second real-value matrix in the second subset; Calculating a third classification probability of the third sample and a fourth classification probability of the fourth sample according to the second Gaussian kernel function; Calculating a third mean of the third sample according to the third classification probability and the second Gaussian kernel function, and calculating a fourth mean of the fourth sample according to the fourth classification probability and the second Gaussian kernel function; Calculate the second inter-class variance of the second current value according to the third classification probability, the fourth classification probability, the third mean, and the fourth mean, and then traverse the second value set to obtain the second inter-class variances corresponding to all the second current values; The maximum value among the second between-class variances is determined as the second representative value.
5. The image retrieval method based on dynamic hash code according to claim 4 is characterized in that: The method further comprises the following steps: Obtaining a pre-stored first kernel density sum and the pre-stored first classification probability, the second classification probability, the first mean, and the second mean at a moment before the current moment; wherein the first kernel density sum is the sum of the kernel densities corresponding to all values in the first value set at all moments before the current moment; Calculate the second kernel density sum corresponding to all values in the first value set at the current moment; Calculating the first classification probability at the current moment according to the first kernel density sum, the second kernel density sum, the first classification probability at a moment before the current moment, the first Gaussian kernel function, and the first value set; Subtracting the first classification probability at the current moment from the second kernel density to obtain the second classification probability at the current moment; Calculate the first mean at the current moment according to the first kernel density sum, the second kernel density sum, the first classification probability at the current moment, the first classification probability at the previous moment and the first mean, the first Gaussian kernel function, and the first value set; Calculating the second mean at the current moment according to the first kernel density sum, the second kernel density sum, the second classification probability at the current moment, the second classification probability at the previous moment and the second mean, the first Gaussian kernel function, and the first value set; Calculating the first inter-class variance at the current moment according to the first classification probability, the second classification probability, the first mean, and the second mean at the current moment, and then traversing the first value set to obtain all the first inter-class variances corresponding to the current moment; The maximum value of all the first between-class variances at the current moment is determined as the first representative value at the current moment.
6. The image retrieval method based on dynamic hash code according to claim 2 is characterized in that: The step of respectively calculating the Manhattan distance between the first real-valued matrix and each of the representative values comprises the following steps: Calculate a first Manhattan distance between the first real-valued matrix and the first representative value in each dimension, and calculate a second Manhattan distance between the first real-valued matrix and the second representative value in each dimension; The first Manhattan distances in each dimension are summarized to obtain a first Manhattan distance matrix, and the second Manhattan distances in each dimension are summarized to obtain a second Manhattan distance matrix; Determine an asymmetric distance matrix between the query image and each of the candidate images according to the binarization matrix, the first Manhattan distance matrix, and the second Manhattan distance matrix; The expression of the asymmetric distance matrix is: Among them, D A represents the asymmetric distance matrix, represents the binarization matrix, M 0 Denotes the first Manhattan distance matrix, M 1 represents the second Manhattan distance matrix.
7. The image retrieval method based on dynamic hash code according to any one of claims 1 to 6, characterized in that: Determining a matching image of the image to be queried in each of the candidate images according to each of the Manhattan distances comprises the following steps: The candidate image corresponding to the minimum value of each of the Manhattan distances is determined as the matching image of the query image.
8. An image retrieval device based on dynamic hash code, characterized in that: The device comprises: A to-be-queried image projection unit, used for calculating a first real-valued matrix of the to-be-queried image in the intermediate space based on a pre-learned hash projection matrix; A candidate image processing unit, configured to calculate a second real-valued matrix of each candidate image in the intermediate space based on the pre-learned hash projection matrix, and then binarize each of the second real-valued matrices to obtain each binarized matrix; A representative value calculation unit, used for calculating the representative value of each of the second real-valued matrices in the intermediate space according to the binarized matrix; a distance calculation unit, used for respectively calculating the Manhattan distance between the first real-valued matrix and each of the representative values; An image matching unit is used to determine a matching image of the query image in each of the candidate images according to each of the Manhattan distances.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the image retrieval method based on dynamic hash code as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image retrieval method based on dynamic hash codes according to any one of claims 1 to 7 is implemented.