Vector retrieval method and system based on low codebook quantization

By using a vector search method based on low codebook quantization in high-dimensional data retrieval, and using cumulative distribution function and quantile function to generate encoding, the problem of high computational complexity and inability to adapt to dynamic data changes in traditional methods is solved, and efficient codebook generation and data update are achieved.

CN120067105AInactive Publication Date: 2025-05-30HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510102075.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional product quantization (PQ) method has high computational complexity in high-dimensional data retrieval and cannot effectively deal with dynamic changes in data.

Method used

A vector search method based on low codebook quantization is adopted to construct the subspace of the feature vectors, and the cumulative distribution function and quantile function are used to generate encodings to avoid iterative calculations, and to achieve rapid codebook generation and data distribution updates.

Benefits of technology

It reduces the overall computing overhead, improves the efficiency of codebook generation, and can effectively deal with dynamic data changes. It is suitable for vector retrieval scenarios with small memory and high speed requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067105A_ABST
    Figure CN120067105A_ABST
Patent Text Reader

Abstract

The invention discloses a vector retrieval method and system based on low codebook quantization, and the method comprises the steps: confirming a cumulative distribution function and a quantile function according to the distribution type of data in a subspace divided by a feature vector, and obtaining a coding value and a code word of the feature vector through the cumulative distribution function and the quantile function; judging the similarity between the code value of the feature vector and the query vector according to the distance between the code value of the feature vector and the query vector to finish vector retrieval; compared with the traditional vector retrieval method for searching the optimal clustering center through iteration, the method has the advantages that the iterative computation is avoided, the problems of large computation amount and long processing time are solved, the overall computation overhead is reduced, and the codebook generation efficiency is improved; meanwhile, in the set updating process of the feature vectors, data distribution is updated only through incremental calculation of the mean value and the standard deviation, the problems that in a traditional retrieval method, the cost of retraining a model is high, and the method cannot adapt to dynamic changes of data are solved, and the whole process is more efficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data retrieval, and particularly relates to a vector retrieval method and system based on low-codebook quantization. Background Art

[0002] With the development of modern computer technology, the scale and dimension of data sets have increased significantly. The efficiency of traditional exact search has been significantly reduced, and the computing time increases exponentially with the increase of data scale and dimension. Therefore, approximate nearest neighbor search (ANNS) has become the core technology in large-scale high-dimensional data search; Product Quantization (PQ) has been widely used in high-dimensional data compression and approximate nearest neighbor search due to its high compression ratio, high query efficiency, and excellent query accuracy; the basic idea of product quantization is to decompose the high-dimensional vector space into multiple low-dimensional subspaces, and each subspace is quantized separately. After each subspace is quantized, each part of the vector will be mapped to a finite set of codewords, thereby greatly reducing the storage and computational complexity; in the traditional PQ approach, when quantizing each subspace, the k-means clustering algorithm is usually used to obtain the codebook of each subspace (Codebook), and each codebook contains several codewords (i.e., cluster centers); since the k-means clustering algorithm is an iterative optimization algorithm, it usually needs to calculate the distance from data points to the cluster centers multiple times, reassign data points to clusters, and update the cluster centers. Therefore, the computational complexity of the k-means clustering algorithm is very high; and k-means clustering is a batch processing method that needs to wait for all data points to be loaded and run. If new data needs to be added, k-means needs to run again. Therefore, k-means cannot efficiently handle the dynamic changes of data. Summary of the Invention

[0003] The purpose of the present invention is to provide a vector retrieval method and system based on low-codebook quantization.

[0004] In the first aspect, the present invention provides a vector retrieval method based on low-codebook quantization, which includes the following steps: Step 1: Construct a set of feature vectors to be queried, and divide the data belonging to the same dimension in all feature vectors into a subspace; Step 2: Allocate the number of bits for each subspace, and obtain the size of the codebook according to the number of bits divided for each subspace; Step 3: Obtain the corresponding clustering center points according to the distribution parameters of the data in each subspace and the size of the codebook; Step 4: Obtain the corresponding codewords according to the clustering center points of the feature vectors; calculate the distances between the codewords of the feature vectors and the query vector in different subspaces respectively, and construct a distance table; Step 5: Traverse the set of feature vectors, and retrieve a preset number of vectors with the highest similarity to the query vector from the set of feature vectors based on the distance table as the retrieval result, thus completing the retrieval of vectors.

[0005] Preferably, in step 3, the method for obtaining the clustering center points is as follows: Respectively obtain the encoding value α corresponding to each value in the subspace, and its expression is: where is the cumulative distribution function; x is the value of each feature vector in the subspace; c is the size of the codebook; floor(·) is the floor symbol; Use different encoding values α in a subspace as the clustering center point a of this subspace.

[0006] Preferably, in step 3, use the quantile function of the clustering center point in the subspace as the codeword of this clustering center point.

[0007] Preferably, both the cumulative distribution function and the quantile function are obtained according to the distribution type of the data in the subspace.

[0008] Preferably, in step 5, the method for determining the similarity between the feature vector and the query vector is as follows: Traverse all the feature vectors, and obtain the complete distance D between the feature vector and the query vector according to the distance table, and its expression is: where is the complete distance between the i-th feature vector and the query vector; is the distance between the i-th feature vector and the query vector in the j-th subspace; Compare the complete distances between all feature vectors and the query vector. The smaller the complete distance between the feature vector and the query vector, the more similar the feature vector is to the query vector.

[0009] Preferably, in step 4, if the set of feature vectors is updated before performing vector retrieval on the query vector, then re-execute step 3 to update the clustering center points in each subspace.

[0010] Preferably, in step 2, the number of bits allocated in each subspace is the same; the size c of the codebook of the subspace is ; where b is the number of bits allocated to the subspace.

[0011] In a second aspect, the present invention provides a vector retrieval system based on low codebook quantization, which is used to execute the above-mentioned vector retrieval method based on low codebook quantization; the vector retrieval system includes a vector encoding module, a distance detection module, a query module, and a storage module; the vector encoding module is used to encode the feature vector to obtain the corresponding encoding value; the distance detection module is used to obtain the distance table of each subspace; the query module is used to retrieve the similar vectors of the query vector according to the distance table; the storage module is used to store the encoding values and distance tables in each subspace.

[0012] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, the memory stores the computer program; the processor executes the above-mentioned vector retrieval method.

[0013] In a fourth aspect, the present invention provides a readable storage medium storing a computer program; when the computer program is executed by a processor, it is used to implement the above-mentioned vector retrieval method.

[0014] The beneficial effects of the present invention are as follows: 1. The present invention generates codes through the cumulative distribution function and uses the quantile function for decompression. Compared with finding the optimal clustering center through iteration in traditional image retrieval, the present invention avoids iterative calculation, solves the problems of large computational amount and long processing time, greatly reduces the overall computational overhead, and improves the efficiency of codebook generation.

[0015] 2. During the update process of the image database, the present invention only needs to update the data distribution by incrementally calculating the mean and standard deviation, without re-fitting the data, solves the problems of high cost of re-training the model and inability to adapt to dynamic changes of data in traditional retrieval methods, makes the whole process more efficient, and is applicable to vector retrieval scenarios with small memory, high speed requirements, or frequent data updates. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is the overall flowchart of the present invention.

[0017] Figure 2 is the flowchart of the vector retrieval method in the present invention.

[0018] Figure 3 is the flowchart of constructing the distance table in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The present invention will be further described below with reference to the accompanying drawings.

[0020] As Figure 1 and Figure 2 shown, a vector retrieval method based on low codebook quantization includes the following steps: Step 1: Obtain the image database to be queried, extract the features of the images in the image database using the Resnet-18 model, and convert the images into 512-dimensional feature vectors; construct a set of feature vectors based on the feature vectors V of each image ; where is the i-th feature vector, ; N is the number of feature vectors; is the feature vector the data of the j-th dimension in; is the feature vector ; the dimension of. Divide the data belonging to the same dimension in all feature vectors V into a subspace, and obtain subspaces.

[0021] Step 2: Allocate the number of bits Allocate the number of bits for quantization to each subspace. The bit allocation determines the quantization accuracy of each subspace; if the number of bits allocated to a subspace is b, then the number of cluster centers of the subspace, i.e., the codebook size c, is .

[0022] In some embodiments, the number of bits allocated to each subspace is the same.

[0023] Step 3: Data encoding According to the distribution type of the data in each subspace, obtain the cumulative distribution function and the quantile function of each subspace respectively. According to the cumulative distribution function and the quantile function, obtain the clustering center points in each subspace. Since the one-dimensional data of most data sets conforms to a one-dimensional normal distribution, in this embodiment, the distribution type of the data in all subspaces is a normal distribution. The specific process of obtaining the clustering center points is as follows:

[0024] Obtain the standard deviation stdDev of all the data in each subspace respectively j , according to the standard deviation stdDev of each subspace j obtain the cumulative distribution function of each feature vector in different subspaces , and its expression is: where is the average value of the j-th subspace; is the sum of the data in the j-th subspace; is the number of data in the j-th subspace; erf is the error function.

[0025] Traverse each subspace of each feature vector in the set of feature vectors, according to the cumulative distribution function Obtain the encoded value α corresponding to each value in the subspace, and its expression is: where floor(·) is the floor function.

[0026] Use different encoded values α as the clustering center points a of the subspace, and the value range of the clustering center point a is [0, c). If an image needs to be added to the image dataset, then update the data sum , the number of data in the subspace and the average value , and the update process is as follows:

[0027] where is the updated data sum; is the updated number of data in the subspace; is the updated average value.

[0028] Recalculate the cumulative distribution function based on the updated data, and encode the set of feature vectors to obtain the updated clustering center points.

[0029] Step 4. Construct a distance table As Figure 3 shown, use the Resnet-18 model to extract the features of the query image to obtain the query vector. Use the quantile function to obtain the codeword corresponding to each clustering center point; when the distribution type of the data in all subspaces is a normal distribution, the expression of the quantile function is:

[0030] where is an intermediate variable, .

[0031] Use the set of codewords corresponding to the clustering center points as the codebook. Therefore, there is no need to record the codebook, and the codewords can be obtained by performing a calculation through the quantile function when needed. Moreover, when the one-dimensional data is closer to the normal distribution, the algorithm accuracy is higher and the effect is better. Use the clustering center points as the codeword indices corresponding to the codewords, and the codewords can be quickly located through the codeword indices.

[0032] Obtain the distances d between the query vector and different codewords in the subspace respectively. The set of all distances in a subspace is the distance table of that subspace.

[0033] Step 5: Use the feature vectors in the set of feature vectors as the encoding vectors; traverse all the encoding vectors, and obtain the complete distance D between the encoding vector and the query vector according to the distance table. Its expression is: where, is the complete distance between the i-th encoding vector and the query vector; is the distance between the i-th encoding vector and the query vector in the j-th subspace.

[0034] Compare the complete distances between all encoding vectors and the query vector. The smaller the complete distance between the encoding vector and the query vector, the more similar the image corresponding to the encoding vector is to the query image; retrieve a preset number of images with the highest similarity to the query image in the image database based on the complete distance as the retrieval result, and complete the retrieval of the image.

[0035] The retrieval system that executes the above retrieval method includes a feature extraction module, a vector encoding module, a distance detection module, a query module, and a storage module; the feature extraction module is used to extract the feature vectors of the image; the vector encoding module is used to encode the feature vectors; the distance detection module is used to obtain the distance table of each subspace; the query module is used to retrieve the similar images of the query image according to the distance table; the storage module is used to store the encoding values and distance tables of each subspace.

Claims

1. A vector retrieval method based on low codebook quantization, characterized in that: The following steps are involved: Step 1: Construct a set of queried feature vectors and divide the data belonging to the same dimension in all feature vectors into a subspace; Step 2: Allocate the number of bits to each subspace, and obtain the size of the codebook according to the number of bits divided into each subspace; Step 3: Obtain the corresponding cluster center point according to the distribution parameters of the data in each subspace and the size of the codebook; Step 4: Obtain the corresponding codeword according to the cluster center point of the feature vector; calculate the distance between the codeword of the feature vector and the query vector in different subspaces respectively, and construct a distance table; Step 5: traverse the set of feature vectors, and retrieve a preset number of vectors with the highest similarity to the query vector from the set of feature vectors based on the distance table as retrieval results, thereby completing the vector retrieval.

2. The vector retrieval method based on low codebook quantization according to claim 1, characterized in that: In step 3, the method for obtaining the cluster center point is as follows: Get the encoding value α corresponding to each value in the subspace respectively, and its expression is: in, is the cumulative distribution function; x is the value of each eigenvector in the subspace; c is the size of the codebook; floor(·) is the rounding symbol; Different encoding values ​​α in a subspace are used as the clustering center points of the subspace.

3. The vector retrieval method based on low codebook quantization according to claim 2 is characterized in that: In the step three, the quantile function of the cluster center point in the subspace is used as the codeword of the cluster center point.

4. The vector retrieval method based on low codebook quantization according to claim 3 is characterized in that: The cumulative distribution function and the quantile function are both obtained according to the distribution type of the data in the subspace.

5. The vector retrieval method based on low codebook quantization according to claim 1, characterized in that: In step 5, the method for determining the similarity between the feature vector and the query vector is as follows: Traverse all feature vectors and obtain the complete distance D between the feature vector and the query vector according to the distance table. The expression is: in, is the complete distance between the ith feature vector and the query vector; is the distance between the i-th feature vector and the query vector in the j-th subspace; Compare the complete distances between all feature vectors and the query vector. The smaller the complete distance between the feature vector and the query vector, the more similar the feature vector is to the query vector.

6. The vector retrieval method based on low codebook quantization according to claim 1, characterized in that: In the step 4, if the set of feature vectors is updated before the query vector is searched, the step 3 is re-executed to update the cluster center point in each subspace.

7. The vector retrieval method based on low codebook quantization according to claim 1, characterized in that: In the step 2, the number of bits allocated in each subspace is the same; the codebook size c of the subspace is ; Where b is the number of bits allocated to the subspace.

8. A vector retrieval system based on low codebook quantization, characterized in that: Used to execute a vector retrieval method based on low codebook quantization as described in claim 1; the vector retrieval system includes a vector encoding module, a distance detection module, a query module and a storage module; the vector encoding module is used to encode the feature vector and obtain the corresponding encoding value; the distance detection module is used to obtain the distance table of each subspace; the query module is used to retrieve similar vectors of the query vector according to the distance table; the storage module is used to store the encoding value and distance table in each subspace.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The memory stores a computer program; the processor executes the vector retrieval method according to any one of claims 1 to 7.

10. A readable storage medium storing a computer program; characterized in that: When the computer program is executed by a processor, it is used to implement the vector retrieval method as described in any one of claims 1 to 7.