Multi-region crossing vector approximate retrieval method and system based on division
By adopting the partition-based multi-zone traversal vector approximation search method in graph algorithm search, the problem of invalid search extension is solved, and more efficient query performance and lower latency are achieved.
Patent Information
- Application Number
- CN202411994903.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-16
AI Technical Summary
The existing technology has invalid search extensions in graph algorithm search, resulting in poor query performance and difficult to meet the business scenario needs of low latency and high precision.
Using the multi-zone traversal vector approximation search method based on division, we randomly divide the data set and build sparse approximation nearest neighbor maps to reduce invalid searches, and design a two-stage search strategy to optimize the query process.
It effectively reduces invalid search extensions, improves query performance, and can significantly accelerate the query process while maintaining the same query accuracy.
Smart Images

Figure CN120011598A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of graph algorithm search technology, and in particular to a method and system for approximate retrieval of multi-region crossing vectors based on division. Background Art
[0002] Given a dataset Represents a data set containing n vectors, where v i Indicates that one of D-dimensional vector in Euclidean space. The distance between any two vectors p and q can be expressed as δ(p,q). The problem of this study is defined as given a query q, returning the set of k vectors closest to q within a certain allowable error range.
[0003] For low-latency and high-precision business scenarios, existing systems usually adopt solutions based on approximate proximity graphs (PG). That is, a single large approximate proximity graph is constructed on the data set, using nodes to represent a vector in the data set, and using edges to represent the neighbor relationship between vectors. When a query is given, the system uses a greedy beam search strategy, that is: the algorithm starts from a point on the graph, maintains a candidate set with a limited length, and continuously expands to neighbor nodes closer to the query vector until there are no vectors closer to the query vector in the candidate set. This process is like Figure 1 As shown in Figure a of the data set is defined in On the vector dataset, represents the query vector, and the parallelogram represents the area close to the query during the search process. The existing solution builds a single large approximate neighbor graph as shown in Figure b. The search starts from v1 and continues to Greed approaches.
[0004] With the rapid development of large models and retrieval enhancement generation technology, vector approximate search has more and more extensive application value in large models, recommendation systems, and information retrieval. In order to cope with the increasingly high end-to-end latency and accuracy requirements of various business scenarios, it is urgent to provide unified optimization for different graph algorithms to reduce invalid searches. Summary of the invention
[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art that lead to many unnecessary expansions and to provide a partition-based multi-zone crossing vector approximate retrieval method and system that reduces invalid searches and provides a unified optimization for different graph algorithms.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] A multi-zone crossing vector approximate retrieval method based on division includes the following steps:
[0008] Obtain the data set to be searched, sample a part of the vectors of the data set with a sampling ratio λ as the routing vector, randomly divide the unsampled vectors in the data set into m partitions, and use the routing vector as the shared vector of each partition; construct a sparse approximate neighbor graph for each of the m partitions;
[0009] Receive query information, perform a first-stage search with the first candidate set queue length, and approach an area near the query information;
[0010] A second-stage search is performed in the approaching area with the second candidate set queue length. If the search is extended to the routing vector, the copies of the routing vector in all partitions are added to the candidate set queue, and a dynamic search across partitions is performed to finally obtain the search result; the second candidate set queue length is greater than the first candidate set queue length.
[0011] Furthermore, the first stage search only searches a sparse approximate neighbor graph.
[0012] Furthermore, the update operation of the sparse approximate neighbor graph includes vector insertion and vector deletion;
[0013] The vector insertion specifically includes: randomly inserting the vector to be inserted into m sparse approximate neighbor graphs, using the probability determined by the sampling ratio λ to determine whether the vector to be inserted is selected as a routing vector, and if it is selected as a routing vector, inserting the vector into each sparse approximate neighbor graph.
[0014] Furthermore, the vector deletion specifically includes: performing a vector deletion operation on each sparse approximate neighbor graph containing the vector to be deleted; if the vector to be deleted is a routing vector, deleting the routing vector from each sparse approximate neighbor graph.
[0015] Furthermore, the method is used in vector query-driven AI application scenarios, recommendation systems, e-commerce, computer vision, information retrieval or financial risk management fields.
[0016] The present invention also provides a multi-region crossing vector approximate retrieval system based on division, comprising:
[0017] A random partitioning framework construction module is used to obtain a data set to be searched, sample a portion of vectors in the data set with a sampling ratio λ as routing vectors, randomly divide the unsampled vectors in the data set into m partitions, and use the routing vectors as shared vectors of each partition; and construct sparse approximate neighbor graphs for each of the m partitions.
[0018] A first stage search module, used for receiving query information, performing a first stage search with a first candidate set queue length, and approaching an area near the query information;
[0019] The second stage search module is used to perform a second stage search in the approaching area with a second candidate set queue length. If the search process is extended to the routing vector, the copies of the routing vector in all partitions are added to the candidate set queue, and a dynamic search across partitions is performed to finally obtain the search result; the second candidate set queue length is greater than the first candidate set queue length.
[0020] Furthermore, the first-stage search module searches only a sparse approximate neighbor graph.
[0021] Furthermore, in the random partitioning framework construction module, the updating operation of the sparse approximate neighbor graph includes vector insertion and vector deletion;
[0022] The vector insertion specifically includes: randomly inserting the vector to be inserted into m sparse approximate neighbor graphs, using the probability determined by the sampling ratio λ to determine whether the vector to be inserted is selected as a routing vector, and if it is selected as a routing vector, inserting the vector into each sparse approximate neighbor graph.
[0023] Furthermore, the vector deletion specifically includes: performing a vector deletion operation on each sparse approximate neighbor graph containing the vector to be deleted; if the vector to be deleted is a routing vector, deleting the routing vector from each sparse approximate neighbor graph.
[0024] Furthermore, the system is used in vector query-driven AI application scenarios, recommendation systems, e-commerce, computer vision, information retrieval or financial risk management fields.
[0025] Compared with the prior art, the present invention has the following advantages:
[0026] (1) In order to reduce some meaningless extended calculations in the graph search process and further improve the query performance of graph search, the present invention proposes a general framework: Crossing Sparse Proximity Graph (CSPG). This framework provides a simple and efficient mode to optimize the existing mainstream graph indexing algorithm by randomly partitioning the data set and maintaining a set of routing vectors. Since the divided sub-datasets are smaller and sparser than the original data set, the constructed sparse approximate neighbor graph is also sparser than the approximate neighbor graph constructed on the entire data set, thus allowing a longer compensation to be used in the search process to approximate the query vector faster.
[0027] Due to its high flexibility, the framework can be embedded into almost all existing state-of-the-art graph indexing algorithms.
[0028] (2) For cross-partition multi-partition views, the present invention designs an efficient two-stage search strategy on CSPG, including a fast approximation stage and a cross-partition fine search stage. In the fast approximation stage, the CSPG search algorithm can use fewer extensions and larger compensations to quickly approximate the area near the query; in the cross-partition fine search stage, if the route vector is expanded, its copies in all partitions will be added to the candidate set queue. This operation enables the search process to dynamically traverse between different partitions, maximizing the search space of the vector, and at the same time making the results obtained in the search process taken from different sparse approximate neighbor graphs, thereby maximizing the search refinement near the query and greatly improving the query efficiency overall.
[0029] (3) A series of experiments were conducted on the most widely used public datasets and the most representative mainstream graph indexes. The experimental results showed that CSPG can effectively reduce meaningless graph search expansion and accelerate the query process while maintaining the same query accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A schematic diagram of a flow chart of a method for approximate retrieval of multi-region crossing vectors based on division provided in an embodiment of the present invention;
[0031] Figure 2 An example diagram of an existing approximate neighbor graph and its search process provided in an embodiment of the present invention;
[0032] Figure 3 A schematic diagram of two sparse approximate neighbor graphs and their search process provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0034] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0035] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0036] Example 1
[0037] like Figure 1 As shown, this embodiment provides a method for approximate retrieval of multi-zone crossing vectors based on division, including:
[0038] Construct S1 based on the random partitioning framework: obtain the data set to be searched, sample a part of the vectors in the data set with a sampling ratio λ as the routing vector, randomly divide the unsampled vectors in the data set into m partitions, and use the routing vector as the shared vector of each partition; construct a sparse approximate neighbor graph for each of the m partitions;
[0039] Stage 1: Single-partition fast approximation process S2
[0040] Receive query information, perform a first-stage search with the first candidate set queue length, and approach an area near the query information;
[0041] Phase 2: Fine-grained S3 search across partitions
[0042] The second stage search is performed in the approaching area with the second candidate set queue length. If the search is extended to the routing vector, the copies of the routing vector in all partitions are added to the candidate set queue, and a dynamic search across partitions is performed to finally obtain the search result; the second candidate set queue length is greater than the first candidate set queue length.
[0043] The specific process includes the following:
[0044] 1. Based on random partitioning framework
[0045] When solving the k-ANNS problem, it is hoped that as many vectors as possible can be searched with the minimum cost. A direct approach is to relax the edge selection strategy of the approximate neighbor graph to allow a vector to have both close and distant neighbors. However, in order to ensure the efficiency of search filtering, the node readings are usually not set too high, which makes it more difficult to set long and short edges at the same time. To solve this problem, this embodiment proposes a method to maximize the number of search vectors near the query node, but it will not significantly increase the node degree. This method is implemented through a simple and efficient random data set partitioning method. Specifically, the data set is partitioned into Randomly divide into m groups, which share a part of the vectors. These shared vectors are called route vectors (RV). Then a sparse approximate proximity graph (SPG) is constructed for each group of divided vectors. Since the divided sub-datasets are smaller and sparser than the original data set, the constructed sparse approximate proximity graph is also sparser than the approximate proximity graph constructed on the entire data set, thus allowing a longer offset to be used during the search process to approximate the query vector faster.
[0046] The construction of the framework consists of three main steps: 1) According to the sampling ratio λ, Sampling a portion of the vector These vectors are common to every partition. 2) For For each vector in , the algorithm randomly assigns them to m partitions. 3) Construct sparse approximate neighbor graphs for each of the m partitions.
[0047] by Figure 2 As an example (only two partitions are included), the construction process of the CSPG framework is explained in detail. First, some vectors are sampled by a certain ratio λ These vectors are shared by both partitions and are called routing vectors. For those non-routing vectors that have not been sampled, they are randomly assigned to a partition. The first partition contains Second partition Then, respectively and Construct a sparse approximate neighbor graph on as well as
[0048] Consider the original dataset The time complexity and space complexity of constructing the n vectors on . Assume that the time and space cost of constructing the approximate neighbor graph on the entire dataset is and Obviously, for each partition there is vectors, then the time complexity of CSPG construction is The space complexity is The time and space overhead of this algorithm is of the same order as the original algorithm.
[0049] Since the CSPG framework is built on the mainstream approximate neighbor graph, the existing vector update operations for the underlying sparse approximate neighbor graph can be well embedded in CSPG. In addition, due to the lightness and flexibility of random partitioning, the addition and deletion of vectors is also simple and efficient:
[0050] Vector insertion: 1) The algorithm randomly inserts the vector to be inserted into m sparse approximate neighbor graphs. 2) According to the probability determined by the sampling ratio λ, it is determined whether the vector to be inserted is selected as a routing vector. If it is selected as a routing vector, the vector is inserted into the routing vector set Otherwise, no update is required
[0051] Vector deletion: 1) The algorithm performs a vector deletion operation on each sparse approximate neighbor graph containing the vector to be deleted. 2) If the deleted vector is a routing vector, remove it from the routing vector set. Delete it.
[0052] The CSPG framework is such a framework, which consists of different sparse approximate neighbor graphs and a set of routing vectors. Since each partitioned sub-dataset is obtained using a random partitioning strategy, they can maintain the same data distribution as the original data set, which also means that the sparse approximate neighbor graphs on different partitions have the same vector distribution, ensuring that their construction complexity and search complexity expectations are the same. In addition, based on this feature, an efficient two-stage vector search strategy is further proposed.
[0053] 2. Two-stage efficient search strategy
[0054] The search process of the CSPG framework is generally divided into two stages: the fast approximation stage and the precise search stage. Specifically: the goal of the first stage is to quickly approximate the query, so only a sparse approximate neighbor graph is used; the second stage considers the surrounding vectors of each partition near the query as finely as possible. The traditional greedy beam search method maintains a fixed-length candidate set queue on a single approximate neighbor graph. CSPG is modified on this basis, using different candidate set queue lengths ef1 and ef2 for the two stages of the search, where ef1 <ef2。
[0055] Table 1 Algorithm process
[0056]
[0057] The two-stage search strategy is generally shown in Table 1.
[0058] Phase 1: Fast approximation process for a single partition
[0059] In the first phase, the CSPG search algorithm is able to quickly approximate the area near the query with fewer expansions and larger backoffs. Figure 3 Take the case in for example: given a data set and single query The CSPG is constructed, and the first stage of fast approximation is performed using ef1=1 to approximate the area near the query (the dashed parallelogram area in the figure). Since the sparse approximate neighbor graph is even sparser, the algorithm only needs one hop to reach the query area. Figure 2 The ordinary greedy beam search method is used in the process, which requires 3 hops to reach the destination.
[0060] Each sparse approximate neighbor graph in CSPG is smaller and sparser than the single approximate neighbor graph on the entire dataset. This sparsity enables the first-stage search of CSPG to use a larger search step and fewer moves to approximate the query. On the other hand, CSPG uses a smaller candidate set length ef1, which further reduces some graph expansion calculations that are meaningless to the final result. In this case, the first-stage search is just an ordinary greedy search process with a smaller candidate set.
[0061] Phase 2: Refined and extended search across partitions
[0062] After the rapid approximation in the first stage, the point closest to the query can be obtained from the candidate set queue. Obviously, the algorithm has already executed to the area close to the query, which corresponds to the parallelogram area in the above example. At this time, the candidate set is reset and the candidate set length is set to ef2. The search in the second stage is different from the first stage, which is shown in line 13 of Algorithm 1. Specifically, when expanding to a certain vector u, if u is a routing vector, then its copies in all partitions will be added to the candidate set queue. This operation allows the search process to dynamically travel back and forth between different partitions, maximizing the search space of the vector, and at the same time making the results obtained in the search process taken from different sparse approximate neighbor graphs, thereby maximizing the search refinement near the query.
[0063] exist Figure 2 In , the naive greedy beam search strategy uses a candidate set queue length of ef = 3, and finally performs a 6-hop graph search process to reach the final destination v9. Consider again Figure 3 The search strategy of CSPG in the algorithm adopts the first-stage search with ef1=1 and the second-stage search with ef2=3. In the second stage, it enters the parallelogram area closer to the query, and then performs a more refined search process. This process generates a 3-hop graph expansion calculation. In the end, CSPG generates 1 expansion in the first stage and 3 expansions in the second stage, a total of 4 expansions. Using a simple greedy beam search generates 6 graph expansions. In actual calculations, if the appropriate number of partitions is set, this advantage will be further amplified compared to the traditional greedy beam search strategy.
[0064] Technical advantages:
[0065] In order to speed up the search process, the current ANNS graph index algorithm usually needs to maintain a series of long edges and short edges at the same time under limited degree constraints. Since the two types of edges are antagonistic and coupled with each other, it is difficult to achieve the optimal situation and is not conducive to actual parameter adjustment. Aiming at the commonality of efficient graph search, this solution proposes a cross-partition graph search strategy based on random partitioning, which provides a new idea for solving query speed and accuracy and maintaining long edges and short edges at the same time.
[0066] The current ANNS graph indexing algorithm usually performs differently on different data sets, which means it is difficult for an indexing algorithm to be optimal in all cases. This solution can optimize the query speed by 1.5 to 2 times for almost all graph algorithms with good performance in almost all accuracy ranges, which is a huge improvement for graph search algorithms.
[0067] Application scenarios:
[0068] The multi-zone traversal vector approximate retrieval method based on partitioning provided in this embodiment provides an optimization framework that aims to improve the performance of existing mainstream graph indexing algorithms and, as a unified search framework in the field of vector retrieval, has broad application potential. Specifically, in the field of information retrieval, the system can retrieve information similar to the query vector based on the vectorization results by vectorizing metadata and query data. In the field of pattern recognition, the characteristics or patterns of each entity can be represented as a series of numbers, and matching patterns can be modeled as similar vectors. In the recommendation system, the preferences and characteristics of creators and consumers can also be vectorized, and then accurate recommendations can be made based on the similarity between vectors. In addition, with the rise of large language models represented by (Chat) GPT, vector approximate search has also been widely used in large model retrieval enhancement generation technology. In the face of possible amnesia, fabrication, and unreliable reasoning problems in large models, with the help of vectorized knowledge and corpus, the vector approximate search algorithm can provide large models with corpus information that is highly relevant to user questions, thereby supporting large models to generate more reliable answers.
[0069] Commercial value:
[0070] As more and more vector query-driven AI application scenarios are implemented, such as Retrieval Augmented Generation (RAG), CSPG can further reduce the end-to-end latency of vector queries and improve query accuracy, which helps reduce the cost of deploying AI applications for enterprises / individuals.
[0071] Vector approximate search algorithms and their widely used graph algorithms are not only widely used in AI application scenarios, but also widely used in recommendation systems, e-commerce, computer vision, information retrieval and other fields. They are also widely used in biological information mining such as protein retrieval and financial risk management such as anomaly detection.
[0072] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A multi-region crossing vector approximate retrieval method based on partitioning, characterized in that: The following steps are involved: Obtain the data set to be searched, sample a part of the vectors of the data set with a sampling ratio λ as the routing vector, randomly divide the unsampled vectors in the data set into m partitions, and use the routing vector as the shared vector of each partition; construct a sparse approximate neighbor graph for each of the m partitions; Receive query information, perform a first-stage search with the first candidate set queue length, and approach an area near the query information; A second-stage search is performed in the approaching area with the second candidate set queue length. If the search is extended to the routing vector, the copies of the routing vector in all partitions are added to the candidate set queue, and a dynamic search across partitions is performed to finally obtain the search result; the second candidate set queue length is greater than the first candidate set queue length.
2. The method for approximate retrieval of multi-region crossing vectors based on division according to claim 1 is characterized in that: The first stage search only searches a sparse approximate neighbor graph.
3. The method for approximate retrieval of multi-region crossing vectors based on division according to claim 1 is characterized in that: The updating operation of the sparse approximate neighbor graph includes vector insertion and vector deletion; The vector insertion specifically includes: randomly inserting the vector to be inserted into m sparse approximate neighbor graphs, using the probability determined by the sampling ratio λ to determine whether the vector to be inserted is selected as a routing vector, and if it is selected as a routing vector, inserting the vector into each sparse approximate neighbor graph.
4. The method for approximate retrieval of multi-region crossing vectors based on division according to claim 3 is characterized in that: The vector deletion specifically includes: performing a vector deletion operation on each sparse approximate neighbor graph containing the vector to be deleted; if the vector to be deleted is a routing vector, deleting the routing vector from each sparse approximate neighbor graph.
5. The method for approximate retrieval of multi-region crossing vectors based on division according to claim 1 is characterized in that: The method is used in vector query-driven AI application scenarios, recommendation systems, e-commerce, computer vision, information retrieval or financial risk management fields.
6. A multi-region crossing vector approximate retrieval system based on partitioning, characterized in that: include: A random partitioning framework construction module is used to obtain a data set to be searched, sample a portion of vectors in the data set with a sampling ratio λ as routing vectors, randomly divide the unsampled vectors in the data set into m partitions, and use the routing vectors as shared vectors of each partition; and construct sparse approximate neighbor graphs for each of the m partitions. A first stage search module, used for receiving query information, performing a first stage search with a first candidate set queue length, and approaching an area near the query information; The second stage search module is used to perform a second stage search in the approaching area with a second candidate set queue length. If the search process is extended to the routing vector, the copies of the routing vector in all partitions are added to the candidate set queue, and a dynamic search across partitions is performed to finally obtain the search result; the second candidate set queue length is greater than the first candidate set queue length.
7. The multi-region crossing vector approximate retrieval system based on partitioning according to claim 6 is characterized in that: The first-stage search module searches only a sparse approximate neighbor graph.
8. The multi-region crossing vector approximate retrieval system based on partitioning according to claim 6 is characterized in that: In the random partitioning framework building module, the updating operation of the sparse approximate neighbor graph includes vector insertion and vector deletion; The vector insertion specifically includes: randomly inserting the vector to be inserted into m sparse approximate neighbor graphs, using the probability determined by the sampling ratio λ to determine whether the vector to be inserted is selected as a routing vector, and if it is selected as a routing vector, inserting the vector into each sparse approximate neighbor graph.
9. The multi-region crossing vector approximate retrieval system based on partitioning according to claim 8 is characterized in that: The vector deletion specifically includes: performing a vector deletion operation on each sparse approximate neighbor graph containing the vector to be deleted; if the vector to be deleted is a routing vector, deleting the routing vector from each sparse approximate neighbor graph.
10. The multi-region crossing vector approximate retrieval system based on partitioning according to claim 6, characterized in that: The system is used in vector query-driven AI application scenarios, recommendation systems, e-commerce, computer vision, information retrieval or financial risk management fields.