Approximate nearest neighbor search method and device based on vector database and related equipment
By constructing a quantization codebook and determining the quantization vector, the trade-off between storage efficiency and search accuracy in approximate nearest neighbor search is solved, and efficient approximate nearest neighbor search is achieved.
Patent Information
- Application Number
- CN202510742517.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-09
AI Technical Summary
In the prior art, approximate nearest neighbor search technology cannot achieve high storage efficiency and high search accuracy at the same time.
By constructing a quantization codebook, the quantization vector corresponding to each first vector is determined, and a query result is generated based on the similarity, and an approximate nearest neighbor search is performed using the structural characteristics of the quantization vector.
It achieves high storage efficiency while improving search accuracy, reducing computational complexity and improving query efficiency.
Smart Images

Figure CN120610979A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vector technology, and in particular to an approximate nearest neighbor search method, device and related equipment based on a vector database. Background Art
[0002] Vector databases are used to handle the storage, indexing, and searching of vector data, and they play an important role in a variety of application scenarios. In the fields of database systems, information retrieval, machine learning, etc., approximate nearest neighbor (ANN) search is a crucial operation. With the continuous growth of data volume, how to efficiently store and search vectors in high-dimensional space has become the key to improving system performance. In the prior art, when performing approximate nearest neighbor search, methods such as product quantization (PQ) and scalar quantization (SQ) are usually used. However, this method often faces a trade-off between efficiency and accuracy when processing high-dimensional data. It can usually only meet high storage efficiency or high search accuracy, and cannot improve search accuracy while improving storage efficiency. Therefore, the approximate nearest neighbor search technology in the prior art has the problem of not being able to achieve high storage efficiency and high search accuracy at the same time. Summary of the Invention
[0003] The present invention provides an approximate nearest neighbor search method, apparatus and related equipment based on a vector database, which solves the problem in the prior art that the approximate nearest neighbor search technology cannot achieve high storage efficiency and high search accuracy at the same time.
[0004] To solve the above problems, the present invention is achieved as follows:
[0005] In a first aspect, an embodiment of the present application provides an approximate nearest neighbor search method based on a vector database, the method comprising:
[0006] Constructing a quantization codebook based on a vector database, wherein the vector database includes a plurality of first vectors, and the quantization codebook includes a plurality of second vectors, wherein the plurality of second vectors are vectors obtained by performing data compression on the plurality of first vectors;
[0007] Determining a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors from the plurality of second vectors, wherein the quantized vector is a second vector having the smallest distance from the corresponding first vector from the plurality of second vectors;
[0008] In the case of performing an approximate nearest neighbor search on a query vector based on the vector database, a query result is generated based on the similarity between the query vector and the multiple quantization vectors, the query result including at least one target quantization vector, and the at least one target quantization vector including: all quantization vectors among the multiple quantization vectors whose similarity with the query vector is greater than or equal to a preset threshold.
[0009] Optionally, constructing a quantization codebook based on a vector database includes:
[0010] Determine a vector dimension of each of the first vectors to obtain a plurality of dimension values, where the plurality of dimension values correspond one-to-one to the plurality of first vectors, and the dimension values are unsigned binary numbers;
[0011] Determining a plurality of compressed dimension values corresponding one-to-one to each of the plurality of dimension values, wherein the compressed dimension value is expressed as ±1 / √d, where d is the number of dimensions corresponding to the dimension value;
[0012] Determine a plurality of third vectors corresponding one-to-one to the plurality of compressed dimension values, wherein a vector length of the third vector is a length indicated by the corresponding compressed dimension value;
[0013] Normalizing the plurality of third vectors to obtain a plurality of fourth vectors, where the plurality of fourth vectors correspond one-to-one to the plurality of third vectors;
[0014] Random rotation processing is performed on the multiple fourth vectors to obtain the multiple second vectors.
[0015] Optionally, the determining, from the plurality of second vectors, a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors includes:
[0016] Based on a target formula, calculating a second vector with a closest target distance in the quantization codebook for each of the first vectors to obtain the multiple quantization vectors;
[0017] The target formula is y'=argmin y∈C ||yo|| 2 , where y' is the quantization vector, y is the second vector, C is the quantization codebook, and o is the first vector.
[0018] Optionally, in the case of performing an approximate nearest neighbor search on the query vector based on the vector database, generating a query result based on similarities between the query vector and the multiple quantized vectors includes:
[0019] In a case where an approximate nearest neighbor search is performed on a query vector based on the vector database, calculating an inner product between the query vector and a plurality of quantized vectors to obtain a plurality of inner product values, wherein the plurality of inner product values correspond one-to-one to the plurality of quantized vectors;
[0020] Calculating the Euclidean distances between the query vector and the plurality of quantized vectors to obtain a plurality of Euclidean distance values, where the plurality of Euclidean distance values correspond one-to-one to the plurality of quantized vectors;
[0021] Based on the inner product value and the Euclidean distance value, at least one target quantization vector is determined from the multiple quantization vectors, wherein the at least one target quantization vector includes all quantization vectors from the multiple quantization vectors whose corresponding inner product values are greater than or equal to a preset inner product value and whose corresponding Euclidean distance values are greater than or equal to a preset Euclidean distance value, and the preset threshold includes the preset inner product value and the preset Euclidean distance value.
[0022] Optionally, the quantized vector includes a most significant bit and remaining bits, and the calculating the Euclidean distance between the query vector and the plurality of quantized vectors to obtain a plurality of Euclidean distance values includes:
[0023] Based on the most significant bit of the first quantization vector and the remaining bits of the first quantization vector, the Euclidean distance between the query vector and the first quantization vector is calculated to obtain the Euclidean distance value corresponding to the first quantization vector, wherein the first quantization vector is any one of the multiple quantization vectors.
[0024] Optionally, the calculating the Euclidean distance between the query vector and the first quantized vector based on the most significant bit of the first quantized vector and the remaining bits of the first quantized vector to obtain the Euclidean distance value corresponding to the first quantized vector includes:
[0025] Determining precision requirement information of the vector to be queried, the precision requirement information including search precision of the vector to be queried in the vector database;
[0026] When the precision requirement indicated by the precision requirement information is a first precision value, calculating, based on the most significant bit of the first quantized vector, a Euclidean distance between the query vector and the first quantized vector to obtain the Euclidean distance value corresponding to the first quantized vector;
[0027] When the precision requirement indicated by the precision requirement information is a second precision value, calculating the Euclidean distance between the query vector and the first quantized vector based on the most significant bit of the first quantized vector and the remaining bits of the first quantized vector, to obtain the Euclidean distance value corresponding to the first quantized vector;
[0028] The first precision value is smaller than the second precision value.
[0029] In a second aspect, an embodiment of the present application provides an approximate nearest neighbor search device based on a vector database, the device comprising:
[0030] A construction module, configured to construct a quantization codebook based on a vector database, wherein the vector database includes a plurality of first vectors, and the quantization codebook includes a plurality of second vectors, wherein the plurality of second vectors are vectors obtained by performing data compression on the plurality of first vectors;
[0031] a determining module, configured to determine, from the plurality of second vectors, a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors, wherein the quantized vector is a second vector from the plurality of second vectors having the smallest distance to the corresponding first vector;
[0032] A generation module is used to generate a query result based on the similarity between the query vector and the multiple quantization vectors when performing an approximate nearest neighbor search on the query vector based on the vector database, wherein the query result includes at least one target quantization vector, and the at least one target quantization vector includes: all quantization vectors among the multiple quantization vectors whose similarity with the query vector is greater than or equal to a preset threshold.
[0033] In a third aspect, the present application also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described in the first aspect above.
[0034] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the method described in the first aspect above are implemented.
[0035] In a fifth aspect, the present application also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the steps in the method described in the first aspect above.
[0036] The present application provides an approximate nearest neighbor search method, apparatus, and related equipment based on a vector database. The method comprises: constructing a quantization codebook based on a vector database, wherein the vector database comprises a plurality of first vectors, the quantization codebook comprises a plurality of second vectors, and the plurality of second vectors are vectors obtained after data compression of the plurality of first vectors; determining a plurality of quantization vectors corresponding one-to-one to the plurality of first vectors from the plurality of second vectors, wherein the quantization vector is a second vector with the smallest distance from the corresponding first vector among the plurality of second vectors; and performing an approximate nearest neighbor search on a query vector based on the vector database, determining a quantization codebook based on the distance between the query vector and the plurality of first vectors. The similarity between quantized vectors is measured to generate a query result, wherein the query result includes at least one target quantized vector, and the at least one target quantized vector includes: all quantized vectors among the multiple quantized vectors whose similarity with the vector to be queried is greater than or equal to a preset threshold; the present application constructs a quantization codebook based on a vector database, and determines the quantization vector corresponding to each first vector in the vector database based on the quantization codebook, thereby calculating the similarity between the vector to be queried and the quantized vector when performing an approximate nearest neighbor search on the query vector based on the vector database, and obtaining a query result, thereby achieving high storage efficiency by storing the quantized vectors, and achieving high search accuracy by using the structural characteristics of the quantized vectors. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for the description of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0038] Figure 1 A schematic diagram of a flow chart of an approximate nearest neighbor search method based on a vector database provided in an embodiment of the present application;
[0039] Figure 2 A schematic diagram of the query process provided in the embodiment of the present application;
[0040] Figure 3 A schematic structural diagram of an approximate nearest neighbor search device based on a vector database provided in an embodiment of the present application;
[0041] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0043] The terms "first", "second" etc. in the embodiments of the present application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. In addition, the terms "comprise" and "have" and any deformation thereof are intended to cover non-exclusive inclusions, such as, the process, method, system, product or equipment comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or that are intrinsic to these processes, methods, products or equipment. In addition, "and / or" is used in the present application to represent at least one of connected objects, such as A and / or B and / or C, and represents comprising independent A, independent B, independent C, and A and B all exist, B and C all exist, A and C all exist, and 7 situations that A, B and C all exist.
[0044] See also Figure 1 , Figure 1 Schematic diagram of the flow of the approximate nearest neighbor search method based on the vector database provided in the embodiment of the present application. Figure 1 As shown, the approximate nearest neighbor search method based on the vector database may include the following steps:
[0045] Step 101: construct a quantization codebook based on a vector database, where the vector database includes multiple first vectors, and the quantization codebook includes multiple second vectors, where the multiple second vectors are vectors obtained by performing data compression on the multiple first vectors.
[0046] In this embodiment, the vector database is a database specifically used to store and retrieve high-dimensional vector data, including multiple first vectors. A quantization codebook is a structure used in the data compression and quantization process to convert the representation of high-dimensional data or signals into a low-dimensional or discrete representation. The core concept of the quantization codebook is to approximate the original data using multiple representative vectors (also called "codewords"), thereby reducing storage and computing requirements.
[0047] Thus, in this embodiment, multiple first vectors are converted into multiple second vectors. Specifically, in this embodiment, a codebook can be constructed by shifting, normalizing, and randomly rotating vectors of unsigned integers. In this way, compared with the prior art method of directly using randomly rotated hypercube vertices, the method provided by this application allows the randomness and uniformity of the codebook to be maintained at a higher compression rate, thereby supporting more efficient quantization and distance estimation.
[0048] Step 102: Determine a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors from among the plurality of second vectors, wherein the quantized vector is a second vector having the smallest distance from the corresponding first vector among the plurality of second vectors.
[0049] In this embodiment, among multiple second vectors, the second vector corresponding to each first vector is determined, and the second vector is used as the quantization vector, that is, the closest second vector in the codebook is found among the multiple first vectors as its quantization vector, thereby optimizing the vector quantization process and enabling the vector to achieve more accurate vector representation while maintaining low computational complexity.
[0050] It should be noted that the distance between the first vector and the second vector can be calculated using distance calculation methods such as Euclidean distance, Manhattan distance, cosine similarity, etc., which is not specifically limited in this embodiment.
[0051] The quantization strategy for generating quantized vectors theoretically minimizes search error. This asymptotic optimality is achieved through a carefully designed quantization codebook and quantization code, ensuring higher search accuracy at the same storage cost. Furthermore, high-accuracy search results are provided even at high compression rates, without the need to store the original vectors for reordering, thus saving storage space and improving efficiency.
[0052] Step 103: When performing an approximate nearest neighbor search on the query vector based on the vector database, generate a query result based on the similarity between the query vector and the multiple quantization vectors, wherein the query result includes at least one target quantization vector, and the at least one target quantization vector includes: all quantization vectors among the multiple quantization vectors whose similarity with the query vector is greater than or equal to a preset threshold.
[0053] In this embodiment, the query vector is a user-entered vector that is used in an approximate nearest neighbor search (ANNS) within a vector database. Approximate nearest neighbor search is an algorithm that quickly finds data points similar to or "closest" to a given query point in a large dataset, improving search efficiency while maintaining acceptable accuracy. The ANNS algorithm efficiently navigates the search space through intelligent shortcuts and data structures, finding sufficiently close matching points in most practical scenarios without requiring a fully accurate calculation of the nearest neighbor vector, thus reducing computation time and resources.
[0054] In the process of searching for a query vector, a query result is generated by calculating the similarity between the query vector and the quantized vectors in the quantized codebook. Specifically, the similarity calculation can be implemented through efficient arithmetic operations, such as bit operations.
[0055] The query result includes at least one target quantization vector, where the target quantization vector is all quantization vectors whose similarity to the query vector is greater than or equal to a preset threshold. It should be noted that the preset threshold can be adaptively set based on actual conditions and is not specifically limited in this embodiment.
[0056] The present application constructs a quantization codebook based on a vector database, and determines the quantization vector corresponding to each first vector in the vector database based on the quantization codebook, thereby calculating the similarity between the query vector and the quantized vector when performing an approximate nearest neighbor search for the query vector based on the vector database, and obtaining a query result. This achieves high storage efficiency by storing the quantized vectors, and achieves high search accuracy by utilizing the structural characteristics of the quantized vectors.
[0057] In some feasible implementations, optionally, constructing a quantization codebook based on a vector database includes:
[0058] Determine a vector dimension of each of the first vectors to obtain a plurality of dimension values, where the plurality of dimension values correspond one-to-one to the plurality of first vectors, and the dimension values are unsigned binary numbers;
[0059] Determining a plurality of compressed dimension values corresponding one-to-one to each of the plurality of dimension values, wherein the compressed dimension value is expressed as ±1 / √d, where d is the number of dimensions corresponding to the dimension value;
[0060] Determine a plurality of third vectors corresponding one-to-one to the plurality of compressed dimension values, wherein a vector length of the third vector is a length indicated by the corresponding compressed dimension value;
[0061] Normalizing the plurality of third vectors to obtain a plurality of fourth vectors, where the plurality of fourth vectors correspond one-to-one to the plurality of third vectors;
[0062] Random rotation processing is performed on the multiple fourth vectors to obtain the multiple second vectors.
[0063] In this embodiment, the quantization codebook is a predefined set of vectors that are carefully selected or calculated to represent points in the original high-dimensional space. By determining the vector dimension of each first vector in a plurality of first vectors, a plurality of dimension values are obtained, and a plurality of compressed dimension values corresponding to each of the plurality of dimension values are determined, and a plurality of third vectors corresponding to the plurality of compressed dimension values are determined. The dimension value refers to the number of dimensions included in the corresponding vector, and the compressed dimension value is an expression of the number of dimensions, that is, each third vector can be represented by ±1 / √d, where d is the number of dimensions corresponding to the dimension value. Thus, the plurality of first vectors can be represented by d-bit unsigned binary numbers (1 represents 1 / √d, 0 represents -1 / √d), ensuring low space occupancy of the vectors in the codebook. The vector lengths of the plurality of third vectors are the lengths indicated by the corresponding compressed dimension values, and the vector directions are random directions without considering the numerical values of the plurality of dimensions in the vector.
[0064] After obtaining the plurality of third vectors, the plurality of third vectors are normalized so that the vector lengths of the plurality of third vectors are normalized to 1, thereby ensuring that they are distributed on a hypersphere with a radius of 1 and the origin as the center in space, thereby obtaining a plurality of fourth vectors.
[0065] Then, multiple fourth vectors are randomly rotated. This step is to make the multiple fourth vectors in the codebook more evenly distributed and reduce the loss of precision during compression. Specifically, a random orthogonal matrix of the fourth vector is first generated to ensure that the vector length remains unchanged while the basis fourth vector is randomly rotated. Among them, the random orthogonal matrix can be generated by a variety of methods. For example, a matrix can be randomly generated using a Gaussian distribution, and then an orthogonal matrix can be obtained by QR decomposition. QR decomposition is a mathematical method that can decompose any matrix into the product of an orthogonal matrix and an upper triangular matrix. The generated random orthogonal matrix Q is then applied to each fourth vector y, Qy is calculated to obtain the rotated vector, and the calculated vector is added to the codebook to obtain multiple second vectors.
[0066] Through the approach in this embodiment, low space occupancy of vectors in the quantization codebook is ensured, vectors in the quantization codebook are distributed more evenly, and precision loss is reduced during compression.
[0067] Optionally, the determining, from the plurality of second vectors, a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors includes:
[0068] Based on a target formula, calculating a second vector with a closest target distance in the quantization codebook for each of the first vectors to obtain the multiple quantization vectors;
[0069] The target formula is y'=argmin y∈C ||yo|| 2 , where y' is the quantization vector, y is the second vector, C is the quantization codebook, and o is the first vector.
[0070] In this embodiment, when determining quantization vectors corresponding to multiple first vectors from multiple second vectors, the second vector with the closest target distance to each first vector in the quantization codebook is calculated according to the target formula, thereby obtaining multiple quantization vectors.
[0071] Specifically, for a plurality of first vectors that have been normalized, for each first vector o, find the vector y' closest to it in the codebook. The target formula is y'=argmin y∈C ||yo|| 2 , where y' is the quantization vector, y is the second vector, C is the quantization codebook, and o is the first vector.
[0072] The binary representation of y' is stored as a compressed form of the first vector o, and the first vector length ||o|| may also be stored in a compressed form, thereby further reducing the precision loss during the storage process.
[0073] Optionally, in the case of performing an approximate nearest neighbor search on the query vector based on the vector database, generating a query result based on similarities between the query vector and the multiple quantized vectors includes:
[0074] In a case where an approximate nearest neighbor search is performed on a query vector based on the vector database, calculating an inner product between the query vector and a plurality of quantized vectors to obtain a plurality of inner product values, wherein the plurality of inner product values correspond one-to-one to the plurality of quantized vectors;
[0075] Calculating the Euclidean distances between the query vector and the plurality of quantized vectors to obtain a plurality of Euclidean distance values, where the plurality of Euclidean distance values correspond one-to-one to the plurality of quantized vectors;
[0076] Based on the inner product value and the Euclidean distance value, at least one target quantization vector is determined from the multiple quantization vectors, wherein the at least one target quantization vector includes all quantization vectors from the multiple quantization vectors whose corresponding inner product values are greater than or equal to a preset inner product value and whose corresponding Euclidean distance values are greater than or equal to a preset Euclidean distance value, and the preset threshold includes the preset inner product value and the preset Euclidean distance value.
[0077] In this embodiment, an approximate nearest neighbor search is performed on a query vector using inner product values and Euclidean distance values. Specifically, inner products are first calculated between the query vector and multiple quantized vectors to obtain multiple inner product values. Then, Euclidean distances are calculated between the query vector and the multiple quantized vectors to obtain multiple Euclidean distance values. Based on the multiple inner product values and the multiple Euclidean distance values, at least one target quantized vector is determined from the multiple quantized vectors.
[0078] Specifically, the two quantized vectors and (corresponding to the quantized representation of the quantized vector o and the quantized representation of the query vector q respectively), their inner product They can be directly calculated using their binary codes.
[0079] For the Euclidean distance, the query vector q is quantized and the quantized vector o and its quantized representation The Euclidean distance between them It can be approximately calculated by the following formula:
[0080]
[0081] here and can be precomputed and stored because they are not relevant to a particular query, whereas It needs to be calculated dynamically for each query.
[0082] Among them, the inner product value corresponding to the target quantization vector is greater than or equal to the preset inner product value, and the corresponding Euclidean distance value is greater than or equal to the preset Euclidean distance value. The preset inner product value and the preset Euclidean distance value can be set according to actual conditions and are not specifically limited in this embodiment.
[0083] The calculation method in this embodiment utilizes the structural characteristics of the codebook, allowing distance calculation to be performed through efficient arithmetic operations, such as bitwise operations. This not only improves the calculation speed but also makes the entire search process more suitable for hardware acceleration, such as parallel processing using the SIMD instruction set.
[0084] Optionally, the quantized vector includes a most significant bit and remaining bits, and the calculating the Euclidean distance between the query vector and the plurality of quantized vectors to obtain a plurality of Euclidean distance values includes:
[0085] Based on the most significant bit of the first quantization vector and the remaining bits of the first quantization vector, the Euclidean distance between the query vector and the first quantization vector is calculated to obtain the Euclidean distance value corresponding to the first quantization vector, wherein the first quantization vector is any one of the multiple quantization vectors.
[0086] In this embodiment, each quantization vector is composed of the most significant bits and the remaining bits, such as Figure 2 As shown, Figure 2 This is a diagram of the query flow in this embodiment. When calculating the Euclidean distance, the calculation is performed based on the most significant bit and the remaining bits. For example, the specific calculation method can first estimate the distance. If the accuracy is insufficient, the remaining bits are accessed to incrementally estimate the distance with higher accuracy. This approach improves the efficiency and accuracy of distance estimation, especially when processing a large number of candidate vectors, and can achieve better results under different performance and recall requirements.
[0087] Optionally, the calculating the Euclidean distance between the query vector and the first quantized vector based on the most significant bit of the first quantized vector and the remaining bits of the first quantized vector to obtain the Euclidean distance value corresponding to the first quantized vector includes:
[0088] Determining precision requirement information of the vector to be queried, the precision requirement information including search precision of the vector to be queried in the vector database;
[0089] When the precision requirement indicated by the precision requirement information is a first precision value, calculating, based on the most significant bit of the first quantized vector, a Euclidean distance between the query vector and the first quantized vector to obtain the Euclidean distance value corresponding to the first quantized vector;
[0090] When the precision requirement indicated by the precision requirement information is a second precision value, calculating the Euclidean distance between the query vector and the first quantized vector based on the most significant bit of the first quantized vector and the remaining bits of the first quantized vector, to obtain the Euclidean distance value corresponding to the first quantized vector;
[0091] The first precision value is smaller than the second precision value.
[0092] In this embodiment, when a query vector is obtained, distance calculations with different precision requirements are performed based on the precision requirement information of the query vector. Specifically, when the precision requirement indicated by the precision requirement information is a first precision value, only the most significant bit is used for the distance calculation. This allows the system to quickly calculate the distance estimate between the query vector and the data vector. This fast estimate is often sufficient to eliminate candidate vectors that are clearly not nearest neighbors, thereby reducing the number of vectors that require further processing.
[0093] When the precision requirement information indicates the second precision value, the distance calculation is performed using the most significant bit and the remaining bits. Specifically, the number of bits used may be gradually increased to gradually improve the accuracy of the distance estimate. By performing higher-precision calculations only on those candidate vectors that remain possible after preliminary screening, the system can more efficiently utilize computing resources and avoid unnecessary computational overhead.
[0094] Finally, based on the combined information from the most significant bit and the remaining bits, the system calculates a final distance estimate between each quantized vector and the query vector. The system selects the candidate vector with the smallest estimated distance as the nearest neighbor of the query and outputs this result, ensuring the accuracy and reliability of the query results.
[0095] As shown in Tables 1 and 2, experimental results on the space-accuracy trade-off demonstrate that our method consistently achieves better accuracy than all baseline methods when using the same number of bits across all test datasets. In particular, when the number of bits is greater than 6, the average relative error of our method is 1.3 to 3.1 times smaller than that of Learning Vector Quantization (LVQ) and SQ. For PQ and Optimized Product Quantization (OPQ), their errors do not decrease as quickly as our method, SQ, or LVQ when the number of bits increases.
[0096]
[0097] Table 1: Comparison of distance estimation errors (average relative error, Lower is Better)
[0098]
[0099] Table 2: ANN search efficiency comparison (Recall@100, Higher is Better)
[0100] Experimental results on the time-accuracy tradeoff of ANNS show that it performs better than LVQ in terms of time-accuracy tradeoff at the same bit size (4 and 8 bits). In particular, the proposed method can improve accuracy while maintaining efficiency, especially in the small bit size setting.
[0101] Experimental results on recall show that using 4-bit, 5-bit, and 7-bit quantization can stably produce recalls exceeding 90%, 95%, and 99% respectively on all test datasets without the need for re-ranking.
[0102] Experimental results on average distance ratios show that our method using 5-bit quantization produces near-perfect average distance ratios on all datasets, although it typically only produces 97% recall.
[0103] The distance calculation method in this embodiment provides different levels of accuracy as needed while maintaining query efficiency, and is particularly suitable for applications that are sensitive to response time. Secondly, the system can flexibly adjust the calculation strategy based on the specific requirements of the query and resource conditions to adapt to different search scenarios. Finally, this method is applicable to a variety of different query scenarios, including those with special requirements for response time or query accuracy. Through this progressive query processing strategy, the ANNS method proposed in the paper can provide high-precision search results as needed while maintaining high efficiency, thereby achieving better performance and user experience in actual applications.
[0104] The present application constructs a quantization codebook based on a vector database, and determines the quantization vector corresponding to each first vector in the vector database based on the quantization codebook, thereby calculating the similarity between the query vector and the quantized vector when performing an approximate nearest neighbor search for the query vector based on the vector database, and obtaining a query result. This achieves high storage efficiency by storing the quantized vectors, and achieves high search accuracy by utilizing the structural characteristics of the quantized vectors.
[0105] See also Figure 3 , Figure 3 : is a structural diagram of an approximate nearest neighbor search device based on a vector database provided by an embodiment of the present application. Figure 3 As shown, the approximate nearest neighbor search device 500 based on the vector database includes:
[0106] A construction module 510 is configured to construct a quantization codebook based on a vector database, wherein the vector database includes a plurality of first vectors, and the quantization codebook includes a plurality of second vectors, wherein the plurality of second vectors are vectors obtained by performing data compression on the plurality of first vectors;
[0107] a determination module 520 configured to determine, from the plurality of second vectors, a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors, wherein the quantized vector is a second vector from the plurality of second vectors having the smallest distance to the corresponding first vector;
[0108] A generation module 530 is configured to generate a query result based on the similarity between the query vector and the multiple quantization vectors when performing an approximate nearest neighbor search on the query vector based on the vector database, wherein the query result includes at least one target quantization vector, and the at least one target quantization vector includes: all quantization vectors among the multiple quantization vectors whose similarity with the query vector is greater than or equal to a preset threshold.
[0109] Optionally, the building block 510 includes:
[0110] A first determining submodule, configured to determine a vector dimension of each of the first vectors to obtain a plurality of dimension values, wherein the plurality of dimension values correspond one-to-one to the plurality of first vectors, and the dimension values are unsigned binary numbers;
[0111] A second determining submodule is configured to determine a plurality of compressed dimension values corresponding one-to-one to each of the plurality of dimension values, wherein the compressed dimension value is expressed as ±1 / √d, where d is the number of dimensions corresponding to the dimension value;
[0112] A third determining submodule, configured to determine a plurality of third vectors corresponding one-to-one to the plurality of compressed dimension values, wherein a vector length of the third vector is a length indicated by the corresponding compressed dimension value;
[0113] a first processing submodule, configured to perform normalization processing on the plurality of third vectors to obtain a plurality of fourth vectors, wherein the plurality of fourth vectors correspond one-to-one to the plurality of third vectors;
[0114] The second processing submodule is configured to perform random rotation processing on the plurality of fourth vectors to obtain the plurality of second vectors.
[0115] Optionally, the determining module 520 includes:
[0116] a calculation submodule, configured to calculate, based on a target formula, a second vector with a closest target distance in the quantization codebook for each of the first vectors, to obtain the multiple quantization vectors;
[0117] The target formula is y'=argmin y∈C ||yo|| 2 , where y' is the quantization vector, y is the second vector, C is the quantization codebook, and o is the first vector.
[0118] Optionally, the generating module 530 includes:
[0119] a first calculation submodule, configured to calculate, when performing an approximate nearest neighbor search on a query vector based on the vector database, an inner product between the query vector and a plurality of quantized vectors, to obtain a plurality of inner product values, wherein the plurality of inner product values correspond one-to-one to the plurality of quantized vectors;
[0120] a second calculation submodule, configured to calculate the Euclidean distances between the query vector and the plurality of quantized vectors to obtain a plurality of Euclidean distance values, wherein the plurality of Euclidean distance values correspond one-to-one to the plurality of quantized vectors;
[0121] A fourth determination submodule is used to determine at least one target quantization vector from the multiple quantization vectors based on the inner product value and the Euclidean distance value, wherein the at least one target quantization vector includes all quantization vectors from the multiple quantization vectors whose corresponding inner product values are greater than or equal to a preset inner product value and whose corresponding Euclidean distance values are greater than or equal to a preset Euclidean distance value, and the preset threshold includes the preset inner product value and the preset Euclidean distance value.
[0122] Optionally, the second calculation submodule includes:
[0123] A calculation unit is used to calculate the Euclidean distance between the query vector and the first quantization vector based on the most significant bit of the first quantization vector and the remaining bits of the first quantization vector, so as to obtain the Euclidean distance value corresponding to the first quantization vector, wherein the first quantization vector is any one of the multiple quantization vectors.
[0124] Optionally, the computing unit includes:
[0125] a determination subunit, configured to determine precision requirement information of the vector to be queried, wherein the precision requirement information includes a search precision of the vector to be queried in the vector database;
[0126] a first calculation subunit, configured to calculate, when the precision requirement indicated by the precision requirement information is a first precision value, a Euclidean distance between the query vector and the first quantized vector based on the most significant bit of the first quantized vector, to obtain the Euclidean distance value corresponding to the first quantized vector;
[0127] a second calculation subunit, configured to, when the precision requirement indicated by the precision requirement information is a second precision value, calculate the Euclidean distance between the query vector and the first quantized vector based on the most significant bit of the first quantized vector and the remaining bits of the first quantized vector, to obtain the Euclidean distance value corresponding to the first quantized vector;
[0128] The first precision value is smaller than the second precision value.
[0129] The present application constructs a quantization codebook based on a vector database, and determines the quantization vector corresponding to each first vector in the vector database based on the quantization codebook, thereby calculating the similarity between the query vector and the quantized vector when performing an approximate nearest neighbor search for the query vector based on the vector database, and obtaining a query result. This achieves high storage efficiency by storing the quantized vectors, and achieves high search accuracy by utilizing the structural characteristics of the quantized vectors.
[0130] The present application also provides an electronic device. Figure 4 , the electronic device may include a processor 601, a memory 602, and a program 6021 stored in the memory 602 and executable on the processor 601.
[0131] When the program 6021 is executed by the processor 601, it can achieve Figure 1 Any step in the corresponding method embodiment:
[0132] Constructing a quantization codebook based on a vector database, wherein the vector database includes a plurality of first vectors, and the quantization codebook includes a plurality of second vectors, wherein the plurality of second vectors are vectors obtained by performing data compression on the plurality of first vectors;
[0133] Determining a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors from the plurality of second vectors, wherein the quantized vector is a second vector having the smallest distance from the corresponding first vector from the plurality of second vectors;
[0134] In the case of performing an approximate nearest neighbor search on a query vector based on the vector database, a query result is generated based on the similarity between the query vector and the multiple quantization vectors, the query result including at least one target quantization vector, and the at least one target quantization vector including: all quantization vectors among the multiple quantization vectors whose similarity with the query vector is greater than or equal to a preset threshold.
[0135] Optionally, constructing a quantization codebook based on a vector database includes:
[0136] Determine a vector dimension of each of the first vectors to obtain a plurality of dimension values, where the plurality of dimension values correspond one-to-one to the plurality of first vectors, and the dimension values are unsigned binary numbers;
[0137] Determining a plurality of compressed dimension values corresponding one-to-one to each of the plurality of dimension values, wherein the compressed dimension value is expressed as ±1 / √d, where d is the number of dimensions corresponding to the dimension value;
[0138] Determine a plurality of third vectors corresponding one-to-one to the plurality of compressed dimension values, wherein a vector length of the third vector is a length indicated by the corresponding compressed dimension value;
[0139] Normalizing the plurality of third vectors to obtain a plurality of fourth vectors, where the plurality of fourth vectors correspond one-to-one to the plurality of third vectors;
[0140] Random rotation processing is performed on the multiple fourth vectors to obtain the multiple second vectors.
[0141] Optionally, the determining, from the plurality of second vectors, a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors includes:
[0142] Based on a target formula, calculating a second vector with a closest target distance in the quantization codebook for each of the first vectors to obtain the multiple quantization vectors;
[0143] The target formula is y'=argmin y∈C ||yo|| 2 , where y' is the quantization vector, y is the second vector, C is the quantization codebook, and o is the first vector.
[0144] Optionally, in the case of performing an approximate nearest neighbor search on the query vector based on the vector database, generating a query result based on similarities between the query vector and the multiple quantized vectors includes:
[0145] In a case where an approximate nearest neighbor search is performed on a query vector based on the vector database, calculating an inner product between the query vector and a plurality of quantized vectors to obtain a plurality of inner product values, wherein the plurality of inner product values correspond one-to-one to the plurality of quantized vectors;
[0146] Calculating the Euclidean distances between the query vector and the plurality of quantized vectors to obtain a plurality of Euclidean distance values, where the plurality of Euclidean distance values correspond one-to-one to the plurality of quantized vectors;
[0147] Based on the inner product value and the Euclidean distance value, at least one target quantization vector is determined from the multiple quantization vectors, wherein the at least one target quantization vector includes all quantization vectors from the multiple quantization vectors whose corresponding inner product values are greater than or equal to a preset inner product value and whose corresponding Euclidean distance values are greater than or equal to a preset Euclidean distance value, and the preset threshold includes the preset inner product value and the preset Euclidean distance value.
[0148] Optionally, the quantized vector includes a most significant bit and remaining bits, and the calculating the Euclidean distance between the query vector and the plurality of quantized vectors to obtain a plurality of Euclidean distance values includes:
[0149] Based on the most significant bit of the first quantization vector and the remaining bits of the first quantization vector, the Euclidean distance between the query vector and the first quantization vector is calculated to obtain the Euclidean distance value corresponding to the first quantization vector, wherein the first quantization vector is any one of the multiple quantization vectors.
[0150] Optionally, the calculating the Euclidean distance between the query vector and the first quantized vector based on the most significant bit of the first quantized vector and the remaining bits of the first quantized vector to obtain the Euclidean distance value corresponding to the first quantized vector includes:
[0151] Determining precision requirement information of the vector to be queried, the precision requirement information including search precision of the vector to be queried in the vector database;
[0152] When the precision requirement indicated by the precision requirement information is a first precision value, calculating, based on the most significant bit of the first quantized vector, a Euclidean distance between the query vector and the first quantized vector to obtain the Euclidean distance value corresponding to the first quantized vector;
[0153] When the precision requirement indicated by the precision requirement information is a second precision value, calculating the Euclidean distance between the query vector and the first quantized vector based on the most significant bit of the first quantized vector and the remaining bits of the first quantized vector, to obtain the Euclidean distance value corresponding to the first quantized vector;
[0154] The first precision value is smaller than the second precision value.
[0155] The present application constructs a quantization codebook based on a vector database, and determines the quantization vector corresponding to each first vector in the vector database based on the quantization codebook, thereby calculating the similarity between the query vector and the quantized vector when performing an approximate nearest neighbor search for the query vector based on the vector database, and obtaining a query result. This achieves high storage efficiency by storing the quantized vectors, and achieves high search accuracy by utilizing the structural characteristics of the quantized vectors.
[0156] The present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the various processes of the above-mentioned embodiment of the approximate nearest neighbor search method based on a vector database are implemented, and the same technical effects are achieved. To avoid repetition, the details are not described here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0157] An embodiment of the present application further provides a computer program product, which is stored in a storage medium. The computer program product is executed by at least one processor to implement the various processes of the above-mentioned embodiment of the approximate nearest neighbor search method based on a vector database, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0158] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0159] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0160] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. An approximate nearest neighbor search method based on a vector database, characterized in that: The method comprises: Constructing a quantization codebook based on a vector database, wherein the vector database includes a plurality of first vectors, and the quantization codebook includes a plurality of second vectors, wherein the plurality of second vectors are vectors obtained by performing data compression on the plurality of first vectors; Determining a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors from the plurality of second vectors, wherein the quantized vector is a second vector having the smallest distance from the corresponding first vector from the plurality of second vectors; In the case of performing an approximate nearest neighbor search on a query vector based on the vector database, a query result is generated based on the similarity between the query vector and the multiple quantization vectors, the query result including at least one target quantization vector, and the at least one target quantization vector including: all quantization vectors among the multiple quantization vectors whose similarity with the query vector is greater than or equal to a preset threshold.
2. The method according to claim 1, characterized in that The constructing of a quantization codebook based on a vector database includes: Determine a vector dimension of each of the first vectors to obtain a plurality of dimension values, where the plurality of dimension values correspond one-to-one to the plurality of first vectors, and the dimension values are unsigned binary numbers; Determining a plurality of compressed dimension values corresponding one-to-one to each of the plurality of dimension values, wherein the compressed dimension value is expressed as ±1 / √d, where d is the number of dimensions corresponding to the dimension value; Determine a plurality of third vectors corresponding one-to-one to the plurality of compressed dimension values, wherein a vector length of the third vector is a length indicated by the corresponding compressed dimension value; Normalizing the plurality of third vectors to obtain a plurality of fourth vectors, where the plurality of fourth vectors correspond one-to-one to the plurality of third vectors; Random rotation processing is performed on the plurality of fourth vectors to obtain the plurality of second vectors.
3. The method according to claim 1, characterized in that The determining, from the plurality of second vectors, a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors comprises: Based on a target formula, calculating a second vector with a closest target distance in the quantization codebook for each of the first vectors to obtain the multiple quantization vectors; The target formula is y'=argmin y∈C ||yo|| 2 , where y' is the quantization vector, y is the second vector, C is the quantization codebook, and o is the first vector.
4. The method according to claim 1, wherein The step of generating a query result based on similarities between the query vector and the plurality of quantized vectors when performing an approximate nearest neighbor search on the query vector based on the vector database includes: In a case where an approximate nearest neighbor search is performed on a query vector based on the vector database, calculating an inner product between the query vector and a plurality of quantized vectors to obtain a plurality of inner product values, wherein the plurality of inner product values correspond one-to-one to the plurality of quantized vectors; Calculating the Euclidean distances between the query vector and the plurality of quantized vectors to obtain a plurality of Euclidean distance values, where the plurality of Euclidean distance values correspond one-to-one to the plurality of quantized vectors; Based on the inner product value and the Euclidean distance value, at least one target quantization vector is determined from the multiple quantization vectors, wherein the at least one target quantization vector includes all quantization vectors from the multiple quantization vectors whose corresponding inner product values are greater than or equal to a preset inner product value and whose corresponding Euclidean distance values are greater than or equal to a preset Euclidean distance value, and the preset threshold includes the preset inner product value and the preset Euclidean distance value.
5. The method according to claim 4, characterized in that The quantized vector includes a most significant bit and remaining bits, and the calculating of the Euclidean distance between the query vector and the plurality of quantized vectors to obtain a plurality of Euclidean distance values includes: Based on the most significant bit of the first quantization vector and the remaining bits of the first quantization vector, the Euclidean distance between the query vector and the first quantization vector is calculated to obtain the Euclidean distance value corresponding to the first quantization vector, wherein the first quantization vector is any one of the multiple quantization vectors.
6. The method according to claim 5, characterized in that The calculating, based on the most significant bit of the first quantized vector and the remaining bits of the first quantized vector, the Euclidean distance between the query vector and the first quantized vector to obtain the Euclidean distance value corresponding to the first quantized vector includes: Determining precision requirement information of the vector to be queried, the precision requirement information including search precision of the vector to be queried in the vector database; When the precision requirement indicated by the precision requirement information is a first precision value, calculating, based on the most significant bit of the first quantized vector, a Euclidean distance between the query vector and the first quantized vector to obtain the Euclidean distance value corresponding to the first quantized vector; When the precision requirement indicated by the precision requirement information is a second precision value, calculating the Euclidean distance between the query vector and the first quantized vector based on the most significant bit of the first quantized vector and the remaining bits of the first quantized vector, to obtain the Euclidean distance value corresponding to the first quantized vector; The first precision value is smaller than the second precision value.
7. An approximate nearest neighbor search device based on a vector database, characterized in that: The device comprises: A construction module, configured to construct a quantization codebook based on a vector database, wherein the vector database includes a plurality of first vectors, and the quantization codebook includes a plurality of second vectors, wherein the plurality of second vectors are vectors obtained by performing data compression on the plurality of first vectors; a determining module, configured to determine, from the plurality of second vectors, a plurality of quantized vectors corresponding one-to-one to the plurality of first vectors, wherein the quantized vector is a second vector from the plurality of second vectors having the smallest distance to the corresponding first vector; A generation module is used to generate a query result based on the similarity between the query vector and the multiple quantization vectors when performing an approximate nearest neighbor search on the query vector based on the vector database, wherein the query result includes at least one target quantization vector, and the at least one target quantization vector includes: all quantization vectors among the multiple quantization vectors whose similarity with the query vector is greater than or equal to a preset threshold.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which implements the steps of the method according to any one of claims 1 to 6 when executed by a processor.
10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Approximate nearest neighbor search method based on local vector quantization
CN118013085A