A Vector Retrieval Method and Device Based on a Nearest Neighbor Graph
By optimizing neighbor selection by using non-central chi-square distribution integral function and progressive addition method in vector retrieval, a more representative nearest neighbor graph is solved, and the problem of neighbor set selection effectiveness in high-dimensional vector space is improved, and the retrieval accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510397237.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing vector retrieval technology based on nearest neighbor graphs has a data aggregation effect in high-dimensional vector space, affecting the effectiveness of neighbor set selection and resulting in a decrease in navigation accuracy.
The integral function of the non-central chi-square distribution is used to determine the centrality of the vector, build a nearest neighbor graph, and optimize the edge selection through incremental addition of vectors and tuning edge cutting operations to form a more representative nearest neighbor graph.
It improves the accuracy and efficiency of vector retrieval, adapts to different vector data sets and query scenarios, and improves the robustness and reliability of search results.
Smart Images

Figure CN119917705B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vector retrieval, and particularly to a vector retrieval method and device based on a proximity graph. Background Art
[0002] With the development of artificial intelligence technology, vector retrieval technology has been applied to fields such as information retrieval and retrieval-augmented generation (RAG) to handle search tasks for massive amounts of unstructured data such as images and texts. As the best solution to the vector data search task at present, the vector retrieval technology based on a proximity graph has become a research hotspot in the academic and industrial circles.
[0003] However, there is a data aggregation effect in the vector space itself, that is, the higher the dimension, the more concentrated the distance distribution between vectors will be. This phenomenon will significantly affect the effectiveness of the step of selecting the neighbor set in the current mainstream proximity graph construction algorithms, and thus will affect the navigation accuracy of the resulting proximity graph.
[0004] Therefore, in view of the above problems, there is an urgent need to provide a new vector retrieval method based on a proximity graph. Summary of the Invention
[0005] The purpose of this application is to provide a vector retrieval method and device based on a proximity graph, which can improve the accuracy and efficiency of vector retrieval in specific scenarios.
[0006] To achieve the above purpose, this application provides the following solutions:
[0007] In a first aspect, this application provides a vector retrieval method based on a proximity graph, where the vector retrieval method based on a proximity graph includes:
[0008] According to the vector data set, use the integral function of the non-central chi-square distribution to determine the centrality degree of each vector in the entire vector data set; and obtain the centrality degree set according to the centrality degrees of all vectors ; the vector data set is , is the number of vectors, is the vector dimension;
[0009] According to the similarity distances between the vectors in the vector data set and the centrality degree set , use the method of gradually adding vectors to construct a proximity graph ;
[0010] Perform a traversal calculation of the centrality degrees of vectors on the proximity graph to obtain the updated centrality degree set of all vectors ;
[0011] Set of centrality degrees based on all updated vectors , perform a tuning and edge trimming operation on the nearest neighbor graph to obtain the final result nearest neighbor graph ;
[0012] Using the feature vector of the query object as the query vector, search on the final result nearest neighbor graph to obtain the vector retrieval result.
[0013] In a second aspect, the present application provides a vector retrieval device based on a nearest neighbor graph. The vector retrieval device based on a nearest neighbor graph includes:
[0014] A centrality degree determination module, configured to determine the centrality degree of each vector in the entire vector data set according to the vector data set by using the integral function of the non-central chi-square distribution; and obtain a set of centrality degrees according to the centrality degrees of all vectors ; the vector data set is , is the number of vectors, is the vector dimension;
[0015] A nearest neighbor graph construction module, configured to construct a nearest neighbor graph by using the progressive addition of vectors according to the similarity distances between the vectors in the vector data set and the set of centrality degrees ;
[0016] A centrality degree set determination module, configured to perform a traversal calculation of the vector centrality degree on the nearest neighbor graph to obtain a set of centrality degrees of all updated vectors , used to represent the centrality nearest neighbor relationship of the nearest neighbor graph ;
[0017] A final result nearest neighbor graph determination module, configured to perform a tuning and edge trimming operation on the nearest neighbor graph based on the set of centrality degrees of all updated vectors to obtain the final result nearest neighbor graph ;
[0018] A vector retrieval result determination module, configured to use the feature vector of the query object as the query vector and search on the final result nearest neighbor graph to obtain the vector retrieval result.
[0019] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the above-mentioned vector retrieval method based on a nearest neighbor graph.
[0020] According to the specific embodiments provided in this application, the following technical effects are achieved:
[0021] This application provides a vector retrieval method and device based on a proximity graph. By using the integral function of the non-central chi-square distribution according to the vector data set, the centrality degree of each vector in the entire vector data set is determined, and then a quantization index is provided for each vector to characterize its centrality feature. When constructing the proximity graph, the selected neighbors not only depend on the Euclidean distance, but also comprehensively consider the distribution characteristics of the data. Based on the set of centrality degrees of all updated vectors , the proximity graph is optimized and trimmed to obtain the final proximity graph . By dynamically adjusting the neighbor selection rule, it helps to find more representative neighbors, thus forming a better proximity graph in the vector space; effectively improving the global navigation performance, making the retrieval results more robust and reliable; and then enabling this application to adapt to different vector data sets and query scenarios. Furthermore, the high self-adaptability enables this application to maintain a high performance in various application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0023] Figure 1 It is a schematic flowchart of a vector retrieval method based on a proximity graph in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some, rather than all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.
[0025] To make the above objects, features, and advantages of this application more obvious and understandable, the following will further describe this application in detail with reference to the drawings and specific embodiments.
[0026] In an exemplary embodiment, as Figure 1 shown, a vector retrieval method based on a proximity graph is provided. The method includes the following S101 to S105. Among them:
[0027] S101. Determine the centrality degree of each vector in the entire vector dataset by using the integral function of the non-central chi-square distribution according to the vector dataset ; and obtain the centrality degree set according to the centrality degrees of all vectors ; the vector dataset is , is the number of vectors, is the vector dimension
[0028] S101 specifically includes:
[0029] S11. Obtain the data distribution of the vector dataset ; the mean center in the vector dataset is .
[0030] S12. Determine the distance from the th vector to the mean center . .
[0031] S13. Determine the sorting of the distances from all vectors to the mean center . .
[0032] S14. Map the sorting to the meaningful value range of the integral function of the non-central chi-square distribution with the degree of freedom and the non-centrality parameter to obtain the mapping value . .
[0033] S15. Determine the centrality degree of each vector in the entire vector dataset by using the integral function of the non-central chi-square distribution according to the mapping value , .
[0034] S102. Construct a nearest neighbor graph by using the similarity distances between the vectors in the vector dataset and the centrality degree set by means of the method of gradually adding vectors
[0035] S102 specifically includes:
[0036] Step 2.1. Take the vector corresponding to the maximum centrality degree in the centrality degree set as the entry point of the nearest neighbor graph .
[0037] Step 2.2: Obtain the currently to-be-inserted vector , and use the vector as a query vector to search on the constructed nearest neighbor graph . Maintain an ordered candidate set queue with a capacity of , and insert the entry point into the candidate set queue .
[0038] Step 2.3: Select the vector in the candidate set queue that is the closest to the vector and has not been visited yet. Obtain the neighbor vector set of the vector , mark the vector as a visited vector; and add each vector in the neighbor vector set of the vector to the candidate set queue , while retaining the top vectors in the candidate set queue that have the smallest distance to the query vector .
[0039] Step 2.4: Repeat Step 2.3 until all vectors in the candidate set queue have been visited; and use the current candidate set queue as the candidate neighbor set of the currently to-be-inserted vector .
[0040] Step 2.5: According to the candidate neighbor set, use a judgment condition to determine the neighbor set of the currently to-be-inserted vector ; the upper limit of the capacity of the neighbor set is , is a hyperparameter configured for constructing the nearest neighbor graph; the judgment condition is that for any vector in the currently traversed candidate neighbor vectors and any vector already inserted into the neighbor set, the following is satisfied:
[0041] .
[0042] Among them, is a relaxation parameter determined according to the centrality degree of the candidate neighbor vector , , is a hyperparameter configured for constructing the nearest neighbor graph, and it is recommended to be set to 0.2, is the similarity distance function.
[0043] Step 2.6, traverse the vector to be inserted currently 's neighbor set , and add the vector to be inserted currently in reverse to the neighbor sets of each vector in the neighbor set ; after adding neighbors in reverse, expand the capacity upper limit of the neighbor set to ; when, after the reverse addition operation, the number of vectors in the neighbor set is greater than the capacity , perform the operation of reselecting neighbors. At this time, the candidate neighbor set is the neighbor set with capacity out-of-bounds , is the hyperparameter for constructing the nearest neighbor graph configuration.
[0044] Step 2.7, repeat Step 2.2 to Step 2.6 until all vectors are inserted into the current nearest neighbor graph to obtain the nearest neighbor graph .
[0045] S103, perform a traversal calculation on the centrality degree of vectors in the nearest neighbor graph to obtain the updated set of centrality degrees of all vectors .
[0046] S103 specifically includes:
[0047] Step 3.1, traverse the nearest neighbor graph , count the number of times each vector is selected as a neighbor by other vectors in the nearest neighbor graph , and use it as the in-degree of the vector in the nearest neighbor graph , denoted as .
[0048] Step 3.2, according to the minimum in-degree and the maximum in-degree , evenly divide the value range of the in-degrees of all vectors into intervals;
[0049] Step 3.3, let the number of vectors assigned to the th interval be , and use the empirical cumulative distribution function to determine the data centrality degree of the vector in the nearest neighbor graph :
[0050] .
[0051] Among them, the data centrality degree of all vector data in the th interval of the interval is .
[0052] Step 3.4, traverse all vectors to obtain the updated centrality degree set , which is used to represent the proximity graph 's centrality proximity relationship.
[0053] S104, based on the centrality degree set of all vectors of the current proximity relationship , perform a trimming operation on the proximity graph to obtain the final result proximity graph .
[0054] S104 specifically includes:
[0055] Adopt the current proximity relationship to perform a trimming operation on the proximity graph to obtain the final result proximity graph . Traverse all vectors in the proximity graph , and reduce the capacity of the neighbor set of each vector from to .
[0056] S105, use the feature vector of the query object as the query vector, and search on the final result proximity graph to obtain the vector retrieval result.
[0057] S105 specifically includes:
[0058] Step 5.1, determine the vectors with the highest in-degree according to the proximity graph , and use the vectors as the candidate set; use the entry point as the search entry point, and select neighbors from the candidate set as the initial start queue for the search, being the relaxation parameter initialized for the search.
[0059] Step 5.2, let the query vector be , maintain an ordered candidate set queue with a capacity of , and fill the ordered candidate set queue using the initial start queue.
[0060] Step 5.3, obtain the unvisited vector with the minimum distance from the query vector from the ordered candidate set queue , obtain the neighbor vector set of the unvisited vector , and mark the unvisited vector as visited.
[0061] Step 5.4, add each neighbor vector in the neighbor vector set to the ordered candidate set queue , and only retain the first vectors with the smallest distance from the query vector in the ordered candidate set queue .
[0062] Step 5.5, repeat Step 5.3 and Step 5.4 until the ordered candidate set queue is fully accessed.
[0063] Step 5.6, use the first vectors with the smallest distance from the query vector in the ordered candidate set queue as the vector retrieval result where .
[0064] The vector retrieval method based on the proximity graph provided by the present application is described below through specific embodiments. Among them, the obtained patent abstract dataset is used to generate corresponding vector data; the generation of the feature vector set of the patent abstract dataset is performed using a pre-trained model; for the text data of the patent abstract, an existing BERT model can be used for training to extract features and transform them into a vector dataset corresponding to the patent abstract features , the vector dimension is , and the number of vectors is .
[0065] Calculate the degree of data centrality of each vector data in the vector dataset corresponding to the patent abstract features within the entire vector dataset. Calculate the distance from all vector data to the dataset mean center and then sort. Map the sorting order of the distance between the vector data and the mean center to the meaningful value range of the integral function of the non-central chi-square distribution with degrees of freedom and non-centrality parameter to obtain the mapping value . Determine the centrality degree of each vector in the entire vector dataset according to the mapping value using the integral function of the non-central chi-square distribution .
[0066] .
[0067] Among them, is the central chi-square distribution integral function, ,! is the factorial, and the set of centrality degrees is:
[0068] .
[0069] Construct a neighbor graph according to the similarity degree of each vector in the patent abstract feature vector set . The calculation method of the similarity distance can be the Euclidean distance between vectors:
[0070] .
[0071] Among them, is any -dimensional vector, is the -th component of the vector , is the -th component of the vector , The function returns the similarity distance of the vector .
[0072] Furthermore, the specific process of constructing the patent neighbor graph is as follows:
[0073] Step 1, select the vector data with the largest data central value as the entry point of the neighbor graph . Take the entry point as the starting point for constructing the neighbor graph, and use the method of gradually adding vector data to construct the neighbor graph . .
[0074] Step 2, select the vector to be inserted from the patent feature vector set. Search on the constructed neighbor graph with as the query vector, and maintain an ordered candidate set queue with a capacity of. Insert the entry point into the candidate set queue . .
[0075] Step 3, select the point in the candidate set queue that is the closest to the vector to be inserted and has not been visited yet, and obtain the neighbor vector set of the .
[0076] Step 4, add each vector in to , and only retain the first vectors in that have the smallest distance from the vector to be inserted .
[0077] Step 5, repeat Step 3 and Step 4 until there are no unvisited vectors. The candidate set queue obtained at this time is used as the candidate neighbor set of the vector to be inserted .
[0078] Step 6, let the neighbor set of the vector data to be inserted be , with the upper capacity limit being , being a hyperparameter configured for constructing the nearest neighbor graph; traverse the neighbor candidate set queue in sequence, and only add candidate neighbors that meet the following conditions to the neighbor set: let the currently traversed candidate neighbor vector be , and any vector already inserted into the neighbor set both satisfy:
[0079] .
[0080] Where is a relaxation parameter determined according to the data centrality of the vector data , , is a hyperparameter configured for constructing the nearest neighbor graph, recommended to be set to 0.2, is the similarity distance function.
[0081] Step 7, traverse the neighbor set of the vector to be inserted, and add in reverse to the neighbor set of each vector in the neighbor set. Enlarge the upper capacity limit of the neighbor set after adding neighbors in reverse to . When the number of vectors in the neighbor set is greater than the capacity after the reverse addition operation, perform the operation of reselecting neighbors. At this time, the candidate neighbor set is the neighbor set with capacity out of bounds, being a hyperparameter configured for constructing the nearest neighbor graph.
[0082] Step 8. Repeat Step 2 to Step 7 to insert all vectors into the neighbor graph, and finally obtain the neighbor graph based on the patent vector dataset. .
[0083] Step 9. Traverse the neighbor graph. , and count the number of times each vector in the dataset is selected as a neighbor by other vectors in the whole graph, that is, the in-degree of the vector in the neighbor graph, denoted as .
[0084] Step 10. According to the minimum in-degree and the maximum , evenly divide the value range of the in-degree of all vectors in the dataset into intervals. Let the number of vectors assigned to the th interval be , and use the empirical cumulative distribution function to determine the data centrality degree of the vector in the neighbor graph ; traverse all vectors to obtain the updated set of centrality degrees , which is used to characterize the central neighbor relationship of the neighbor graph .
[0085] Step 11. Let the number of vectors assigned to the th interval be , and use the empirical cumulative distribution function to determine the data centrality degree of the vector in the neighbor graph :
[0086] .
[0087] Among them, the data centrality of all vector data in the th interval is . Traverse all vectors in the dataset to obtain their corresponding data centrality , and obtain the updated set of centrality degrees , which is used to characterize the central neighbor relationship of the neighbor graph .
[0088] Step 12. Use the obtained numerical set of the vector data to perform edge trimming and optimization operations on the neighbor graph to obtain the final result neighbor graph . Traverse all vectors in the neighbor graph , and reduce the capacity of the neighbor set of each vector from to , which is a process of reselecting neighbors. At this time, the candidate neighbor set is the neighbor set with capacity out of bounds .
[0089] Step 13, the neighbor graph generation phase is now complete, and the neighbor graph can be used To query. First, generate a query object from the patent abstract to be queried. , which is the feature vector of the patent abstract text data to be queried generated by the pre-trained model.
[0090] Step 14: Once the feature vector of the patent abstract text to be queried is generated, use the query vector In the neighbor graph based on patent abstract feature vector Query feature vectors, and finally returns the patent abstract text that is most similar to the query patents.
[0091] By statistically analyzing the degree of vector centrality, we can select neighbors that better reflect the data distribution characteristics during the construction of the neighbor graph, thereby reducing missed detections during searches. Higher-quality neighbor selection is based on the distribution characteristics of the vector dataset, allowing the resulting neighbor graph to better reflect the distribution characteristics of the vector data, thereby improving retrieval accuracy. Using the degree of centrality to appropriately relax or contract neighbor selection during the construction of the neighbor graph can reduce unnecessary computation, optimize the search process, and effectively mitigate the "curse of dimensionality" problem common in vector spaces, thereby increasing retrieval speed.
[0092] In an exemplary embodiment, a vector retrieval device based on a neighbor graph is provided, comprising:
[0093] The centrality degree determination module is used to determine the centrality degree of each vector in the entire vector data set based on the vector data set and the integral function of the non-central chi-square distribution; and obtain the centrality degree set based on the centrality degree of all vectors ; The vector data set is , is the number of vectors, is the vector dimension, which is used to represent the degree of freedom.
[0094] The neighbor graph construction module is used to collect data based on the similarity distance and centrality degree between each vector in the vector data set. , using the method of progressively adding vectors to construct a neighbor graph .
[0095] The centrality degree set determination module is used to determine the neighbor graph Perform vector centrality degree traversal calculation to obtain the updated centrality degree set of all vectors , used to represent the neighbor graph The central neighbor relationship.
[0096] The final result neighbor graph determination module is used to determine the centrality degree set of all vectors based on the current neighbor relationship , and perform a trimming operation on the neighbor graph to obtain the final result neighbor graph .
[0097] The vector retrieval result determination module is used to use the feature vector of the query object as the query vector and search on the final result neighbor graph to obtain the vector retrieval result.
[0098] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface.
[0099] In this application, all actions of obtaining signals, information, or data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining the authorization given by the owner of the corresponding device.
[0100] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0101] In this article, specific examples are used to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A vector retrieval method based on a nearest neighbor graph, characterized in that The vector retrieval method based on the neighbor graph includes: Obtain the vector dataset 's data distribution; the vector dataset has a mean center of ; the vector dataset is the vector dataset corresponding to the patent abstract features , is the number of vectors, is the vector dimension; Determine the th vector to the mean center distance ; Determine the distances of all vectors to the mean center for sorting ; Sort is mapped to a meaningful value range of the integral function of the non-central chi-square distribution with degrees of freedom and non-centrality parameter to obtain a mapped value ; According to the mapping value , the integral function of the non-central chi-square distribution is used to determine the centrality degree of each vector in the entire vector dataset ; Obtain the centrality degree set according to the centrality degrees of all vectors ; According to the similarity distances and centrality degree sets among the vectors in the vector dataset , a method of progressively adding vectors is adopted to construct a nearest neighbor graph ; Traverse the nearest neighbor graph , count the number of times each vector is selected as a neighbor by other vectors in the nearest neighbor graph , and use it as the in-degree of the vector in the nearest neighbor graph , denoted as ; According to the minimum in-degree and the maximum in-degree , evenly divide the value range of the in-degree of all vectors into intervals; Let the number of vectors assigned to the interval be , and use the empirical cumulative distribution function to determine the centrality degree of the vectors in the neighborhood graph; Traverse all vectors to obtain the updated set of centrality degrees for characterizing the centrality neighbor relationship of the proximity graph The set of centrality degrees based on all updated vectors , perform edge pruning and optimization operations on the nearest neighbor graph to obtain the final nearest neighbor graph ; Using the feature vector of the query object as the query vector, search on the nearest neighbor graph of the final result to obtain the vector retrieval result.
2. The vector retrieval method based on the nearest neighbor graph according to claim 1, wherein According to the similarity distances and centrality degree sets among the vectors in the vector dataset , a method of progressively adding vectors is adopted to construct a nearest neighbor graph , which specifically includes: Step 2.1, take the vector corresponding to the maximum centrality degree in the centrality degree set as the entry point of the proximity graph ; Step 2.2, obtain the currently to-be-inserted vector , using the vector as the query vector to search on the constructed nearest neighbor graph, maintaining an ordered candidate set queue with a capacity of , and inserting the entry point into the candidate set queue ; Step 2.3, select the candidate set queue mid-distance vector the vector with the closest distance and not yet visited , obtain the vector 's neighbor vector set , mark the vector as the visited vector; and add each vector in the neighbor vector set of the vector to the candidate set queue , while retaining the first vectors with the smallest distance to the query vector in the candidate set queue ; where is the first neighbor vector of , and is the number of neighbor vectors of . Step 2.
4. Repeat Step 2.3 until all vectors in the candidate set queue are visited; and use the current candidate set queue as the candidate neighbor set of the vector to be inserted currently ; ; Step 2.5, according to the candidate neighbor set , determine the neighbor set of the currently to-be-inserted vector ; the upper limit of the capacity of the neighbor set is , , is a hyperparameter configured for constructing the nearest neighbor graph; the judgment condition is that the currently traversed candidate neighbor vector and any vector already inserted into the neighbor set both satisfy: ; Among them, is the relaxation parameter determined according to the centrality degree of the candidate neighbor vector ; is the hyperparameter for constructing the nearest neighbor graph configuration, , and is the similarity distance function; Step 2.6, traverse the vector to be inserted currently 's neighbor set , and add the vector to be inserted currently in reverse to the neighbor set of each vector in the neighbor set ; After adding neighbors in reverse, expand the capacity upper limit of the neighbor set to ; When the number of vectors in the neighbor set is greater than the capacity after the reverse addition operation , perform the operation of reselecting neighbors. At this time, the candidate neighbor set is the neighbor set with capacity out of bounds , is the hyperparameter configured for constructing the nearest neighbor graph; Step 2.7, repeat Steps 2.2 to 2.6 to insert all the vectors into the current nearest neighbor graph, and obtain the nearest neighbor graph .
3. The vector retrieval method based on a nearest neighbor graph according to claim 1, wherein Using the feature vector of the query object as the query vector, search on the nearest neighbor graph of the final result to obtain the vector retrieval result, which specifically includes: Step 5.1, according to the nearest neighbor graph Determine the vectors with the highest in-degree and use these vectors as the candidate set; use the entry point as the entry point, and select neighbors from the candidate set as the initial starting queue for the search; Initialize the relaxation parameter for the search Step 5.2, let the query vector be , maintain an ordered candidate set queue with a capacity of , and , and fill it using the initial startup queue ; Step 5.3, obtain from the ordered candidate set queue the unvisited vector with the smallest distance to the query vector , obtain the neighbor vector set of the unvisited vector , and mark the unvisited vector as the visited vector; Step 5.4, add each neighbor vector in the neighbor vector set to the ordered candidate set queue , and only retain the first vectors in the ordered candidate set queue that have the smallest distance to the query vector ; Step 5.5, repeat Step 5.3 and Step 5.4 until all vectors in the ordered candidate set queue are accessed; Step 5.6, take the first vectors in the ordered candidate set queue with the smallest distances to the query vector as the vector retrieval results, where .
4. A vector retrieval device based on a nearest neighbor graph, which is used to implement the nearest neighbor graph-based vector retrieval method according to any one of claims 1-3, characterized in that The vector retrieval device based on the neighbor graph includes: A centrality degree determination module, which is used to determine the centrality degree of each vector in the entire vector dataset according to the vector dataset by using the integral function of the non-central chi-square distribution; and obtain a centrality degree set according to the centrality degrees of all vectors ; the vector dataset is the vector dataset corresponding to the patent abstract features , is the number of vectors, is the vector dimension; A nearest neighbor graph construction module, which is used to construct a nearest neighbor graph by adopting a method of progressively adding vectors according to the similarity distances and the centrality degree set among the vectors in the vector dataset , and adopting a method of progressively adding vectors ; Centrality degree set determination module, which is used to perform vector centrality degree traversal calculation on the neighbor graph to obtain the centrality degree set of all updated vectors , which is used to represent the centrality neighbor relationship of the neighbor graph ; Final result neighbor graph determination module, which is used to determine the neighbor graph based on the updated set of centrality degrees of all vectors , optimize and trim the edges of the neighbor graph to obtain the final result neighbor graph ; A vector retrieval result determination module, which uses the feature vector of the query object as a query vector to perform a search on the final result nearest neighbor graph to obtain a vector retrieval result.
5. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the vector retrieval method based on the neighbor graph according to any one of claims 1-3.
Citation Information
Patent Citations
Approximate nearest neighbor search for single instruction, multiple thread (SIMT) or single instruction, multiple data (SIMD) type processors
US20210157606A1
Construction of nearest neighbor structures for graph machine learning technologies
US20240379227A1