Distribution outward vector retrieval method and device

By constructing a fusion graph and optimizing the nearest neighbor relationship using anchor vector sets, the performance problems of the existing technology in out-of-distribution query scenarios are solved, and more efficient and accurate vector retrieval is achieved.

CN120030192AActive Publication Date: 2025-05-23HANGZHOU DIANZI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510505842.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing vector search algorithm has poor performance in out-of-distribution query scenarios, resulting in slow search speed and low accuracy of search engines in text searches.

Method used

By obtaining the base vector set and the query vector set outside the distribution, a fusion graph is constructed, and the nearest neighbor relationship of the query outside the distribution is optimized by using the anchor vector set and the fusion graph to improve the search accuracy and efficiency.

Benefits of technology

It significantly improves the search accuracy and efficiency of distributed out-of-vector retrieval, can obtain more accurate results under the same search delay, and improves the throughput of search engines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030192A_ABST
    Figure CN120030192A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed outward vector retrieval method and device. The method comprises the following steps: firstly, constructing a corresponding anchor point vector for each base vector in a base vector set based on a query vector set, and constructing a fusion graph based on the base vectors and the anchor point vectors; and searching the to-be-queried vector through the fusion graph so as to obtain a result similar to the to-be-queried vector. According to the method, the neighborhood relation of the base vector set and the query vector set is comprehensively considered, self-adaptive balance is achieved by fusing the weights, the distributed outward vector retrieval task can be effectively dealt with, and the search precision and efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of vector retrieval, and in particular relates to a distributed external vector retrieval method and device. Background Art

[0002] Vector retrieval is a basic module for many modern applications (such as search engines, large language models, and recommendation systems). The vector retrieval algorithm provides a good trade-off between search latency and accuracy for the search stage by constructing an index that can describe the similarity relationship between basis vectors in a basis vector set.

[0003] However, the theory on which existing vector retrieval relies has an implicit premise, that is, the query vector set and the base vector set have the same distribution (that is, the query vector set is in-distribution data). When the query vector set and the base vector set have different distributions (that is, the query vector set is out-of-distribution data), the out-of-distribution query performance is an order of magnitude or more worse than the in-distribution query. Taking the image library of a search engine as an example, when a user enters an image as a query, it is an in-distribution query, and when a user enters text as a query, it is an out-of-distribution query. The out-of-distribution query problem of vector retrieval causes users to take longer to obtain results when searching for images with text, and the search results are inaccurate. This is mainly because the neighbors of out-of-distribution queries lack clustering, that is, the neighbors of the neighbors of out-of-distribution queries are likely not neighbors. In order to solve the problem of slow speed and poor accuracy of text-based graph search by search engines, existing studies have proposed out-of-distribution vector retrieval algorithms, such as RoarGraph and RobustVamana, which add edges between the nearest neighbors of the sampled query vector set (usually 10% of the base vector set size) to enhance the clustering of the graph index. However, this method can only optimize some nodes of the graph index, and most of the nodes are still not optimized for out-of-distribution queries. Therefore, the search engine based on this algorithm can only obtain limited performance improvement by text-based graph search. In response to the above problems, the present invention comprehensively considers the neighbor relationship of each node in the graph index under the distribution of the base vector set and the distribution of the query vector set, and proposes a new vector retrieval method for out-of-distribution queries. Summary of the invention

[0004] In view of the above problems, the present invention proposes a distributed out-of-vector retrieval method and device, which constructs a fusion graph based on the neighbor relationship between the base vector set and the query vector set that obey different distributions to achieve the retrieval task. Compared with the existing algorithm, it can better adapt to the distributed out-of-vector retrieval task and improve the search accuracy and efficiency.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0006] A distributed external vector retrieval method comprises the following steps:

[0007] Step (1) obtaining a base vector set and a query vector set, wherein the query vector set is out-of-distribution data relative to the base vector set;

[0008] Step (2) for each basis vector in the basis vector set, find the nearest query vectors in the query vector set, and generate an anchor vector for each basis vector based on the found query vectors; all anchor vectors constitute an anchor vector set;

[0009] Step (3) constructing a fusion graph based on the basis vector set and the anchor point vector set;

[0010] Step (4) obtains and converts the query text input by the search engine user into a vector to be queried, performs a search on the fusion graph to obtain a basis vector similar to the vector to be queried, and returns the obtained basis vector as the query result.

[0011] Preferably, in step (1), the query vector set is out-of-distribution data relative to the base vector set and is defined as follows:

[0012] Step (1-1) is to obtain the The basis vector set of independent and identically distributed basis vectors , and contains The query vector set of independent and identically distributed query vectors ; Query vector The query vector set Any element in the basis vector set In The neighbor set is ,in yes No. A neighbor, yes The nearest neighbor of . The query vector Nearest neighbor In the basis vector set In The neighbor set is ,in yes No. Nearest neighbors. Query vector For the basis vector set Neighborhood alignment quality Defined as:

[0013]

[0014] Query Vector Set Relative to the basis vector set Neighborhood alignment quality Defined as:

[0015]

[0016] Step (1-2) If the neighborhood alignment quality Less than the preset constant , then the query vector set Relative to the basis vector set The data is outside the distribution, otherwise it is inside the distribution.

[0017] Preferably, in step (2), the process of constructing the anchor vector set includes the following steps:

[0018] Step (2-1) for the basis vector set Any basis vector in , get In the query vector set In The query neighbor set is ,in The number of query neighbors required to construct the anchor vector set;

[0019] Step (2-2) converts each basis vector The corresponding query neighbor set , input aggregation function , and get an anchor vector , all anchor vectors constitute the anchor vector set .

[0020] Preferably, in step (3), the process of constructing the fusion graph includes the following steps:

[0021] Step (3-1) is based on the basis vector set Construct a basic neighbor graph with an out-degree upper limit of L , Each node v in represents a basis vector , Represents the neighbor relationship between nodes. For the basic neighbor graph Any two nodes and , use distance Evaluate and The similarity between For Node The corresponding basis vectors are, For Node The corresponding basis vectors;

[0022] Step (3-2) for the basic neighbor graph For each node v in The basic neighbors in as candidates and re-ranked;

[0023] The basic neighbor for The set of endpoints of directed edges starting from node v in ;

[0024] The reordering includes: for node v and basic neighbors Any node , using the fusion distance in the rearrangement process Evaluate and The similarity between is the fusion weight parameter, and are the basis vectors and anchor vectors corresponding to node v, and For Node The corresponding basis vectors and anchor vectors;

[0025] Step (3-3) Set the outbound limit , for each node v, from the basic neighbors In the example, similarity is sorted according to the fusion distance, and nodes are selected from large to small according to similarity. Try to connect the edges, when the node If the occlusion rule is met, the edges are connected, and the number of connected edges is equal to Stop selecting edges and get fused neighbors ;

[0026] when After all nodes in the graph are connected according to the occlusion rules, the fusion graph is obtained. ;

[0027] Fusion The set of endpoints of the directed edges starting at node v in .

[0028] Preferably, in step (4), performing a search on the fusion graph to obtain a basis vector similar to the query vector specifically includes the following steps:

[0029] Step (4-1) From the fusion graph Select nodes as the entry point, Nodes are placed into the candidate set middle;

[0030] Step (4-2) From the candidate set Select the unvisited distance query vector The nearest node As the current access node, and Distance between Evaluate and merge the neighbors of the currently visited node Insert candidate set ,in is the basis vector corresponding to the currently visited node v;

[0031] The vector to be queried and query vector set Independent and identically distributed;

[0032] Step (4-3): The candidate set For any node v in the Evaluate v and The similarity of the candidate set All nodes in the , sort the similarities from large to small, and delete the candidate set H with a ranking greater than Nodes, where is the basis vector corresponding to node v;

[0033] Step (4-4): Repeat steps (4-2) to (4-3) until the candidate set There are no unvisited nodes in the candidate set. forward The result is basis vectors.

[0034] Preferably, the base vector set in step (1) is composed of vectors converted from pictures or videos in the search engine image library, and the query vector set is composed of vectors converted from query texts input by search engine users;

[0035] Step (4) also includes the following steps:

[0036] Based on the basis vector obtained in step (4), the picture or video in the search engine image library corresponding to the basis vector is output.

[0037] Preferably, the distance and fusion distance middle, is any of the following three distance metrics: Euclidean distance, cosine similarity, and inner product.

[0038] The present invention also provides a distributed external vector retrieval device, which is used to implement the distributed external vector retrieval method, comprising:

[0039] An input module, used for obtaining a base vector set and a query vector set as input, wherein the query vector set is out-of-distribution data relative to the base vector set;

[0040] An anchor vector set building module, used for finding several nearest query vectors in the query vector set for each basis vector in the basis vector set, and generating an anchor vector for each basis vector based on the found query vectors; all anchor vectors constitute an anchor vector set;

[0041] A fusion graph construction module, used to construct a fusion graph based on a basis vector set and an anchor vector set;

[0042] The search result output module is used to perform a search on the fusion graph to obtain a basis vector close to the query vector, and return the obtained basis vector as the query result.

[0043] Preferably, the input module includes: a basis vector set input module for obtaining a basis vector set; a query vector set input module for obtaining a query vector set; and a query distribution difference determination module for determining whether the query vector set is out-of-distribution or in-distribution data.

[0044] A distributed out-of-vector retrieval method provided by the present invention constructs a fusion graph based on the neighbor relationship between a basis vector set and a query vector set that obey different distributions, and comprehensively considers the similarity relationship between the basis vector set and the query vector set, so as to better adapt to the distributed out-of-vector retrieval task and improve the search accuracy and search efficiency.

[0045] After applying the present invention, taking the text-to-image search application of a search engine as an example, after the user enters a text description, an image or video in the image library that is more appropriate to the text description can be obtained under the same search delay. For example, the recall rate of the k results obtained by the user's search is higher. In addition, the search engine applying the present invention can effectively improve the throughput. For example, the search engine system has a higher query rate per second. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic diagram of the process of the present invention.

[0047] Figure 2 It is a diagram of the benchmark test results of the present invention. DETAILED DESCRIPTION

[0048] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0049] (1) First, an original database and an original query set are obtained. In this embodiment, the original database is video data, and the original query set is text data;

[0050] (2) Generate a base vector set and a query vector set, where the base vector set is generated from the original database through the UniterVideo network model, and the query vector set is generated from the original query set through the Transformer network model. Include independent and identically distributed basis vectors, query vector set Include Independent and identically distributed query vectors. Query vector In the basis vector set In The neighbor set is ,in yes No. A neighbor, yes The nearest neighbor of . The query vector Nearest neighbor In the basis vector set In The neighbor set is ,in yes No. Nearest neighbors, query vector set For the basis vector set The following conditions must be met:

[0051]

[0052] in, Defined as:

[0053]

[0054] As a preference, The setting method is as follows: Get a data set with the same distribution as the basis vector set , calculate its basis vector set The NAQ of .

[0055] (3) For the basis vector set Any basis vector in , get exist In The query neighbor set is ,in is the parameter that needs to be determined based on the given basis vector set and query vector set. The corresponding query neighbor set , input aggregation function , output an anchor vector , the set of anchor vectors constitutes the anchor vector set ;

[0056] (4) As a preference, the aggregation function

[0057] (5) Based on basis vector set Constructing the basic neighbor graph , Each node v in represents a basis vector , Represents the neighbor relationship between nodes. For the basic neighbor graph Any two nodes and , use distance As a similarity measure, the upper limit of the out-degree of the basic neighbor graph is L. For Node The corresponding basis vectors are, For Node The corresponding basis vectors;

[0058] (6) For the basic neighbor graph For each node v in The basic neighbors in As candidates and re-ranked, for The set of endpoints of directed edges starting from node v in . Specifically, for node v and the basic neighbor Any node , using the fusion distance As a new similarity measure Rearrange the order, is the fusion weight parameter, and are the basis vectors and anchor vectors corresponding to node v, and For Node The corresponding basis vectors and anchor vectors. The fusion distance enables each node in the graph index to comprehensively represent the neighbor relationship under the basis vector set distribution and the query vector set distribution through distance. For example, for two nodes that are close under the basis vector distribution, , at this time, the fusion distance supports further consideration of the neighbor relationship under the query vector set distribution, enhancing the aggregation of query neighbors. On the contrary, for two nodes that are close to each other under the query vector set distribution, , thus ensuring that the nodes maintain basic neighbor relationships;

[0059] (7) As a preferred option, the parameter and fusion weight parameters Obtained by the following method: For a fusion vector set , where the fusion vector is , .parameter and fusion weight parameters Satisfy query vector set For the fused vector set Neighborhood alignment quality maximum;

[0060] (8) As a preference, set an upper limit on the outgoing degree , for each node v, from the basic neighbors Select the fusion distance The smallest node Try to connect the edges. When the nodes Satisfy the occlusion rules of the neighbor graph (such as the navigation expansion graph NSG) algorithm: , right Connect edges and update fusion neighbors { , and update the candidate ,in Fusion The set of endpoints of directed edges starting from node v in . The edges are reconnected by blocking rules so that the neighbors of each node are scattered in different directions, avoiding the situation where the neighbors are too concentrated and the search process cannot converge quickly.

[0061] (9) Repeat step (8) until the number of connected edges equals Stop selecting edges and get new neighbors , thus generating a fusion map ;

[0062] (10) From the fusion graph Select Nodes are used as entry points and placed in the candidate set In the candidate set Size ;

[0063] (11) From the candidate set Select the unvisited distance query vector The nearest node As the currently visited node, the distance metric is , and merge the neighbors of the currently visited node Place candidates, where is the basis vector corresponding to the currently visited node v;

[0064] (12) The candidate set Sort from small to large and delete the candidates with a ranking greater than Nodes of the candidate set For any node v in As a distance metric, is the basis vector corresponding to node v;

[0065] (13) Repeat (11) to (12) until the candidate set There are no unvisited nodes in the candidate set. forward The result is basis vectors.

[0066] The present invention also provides a distributed external vector retrieval device, which is used to implement the distributed external vector retrieval method, comprising:

[0067] An input module, used for obtaining a base vector set and a query vector set as input, wherein the query vector set is out-of-distribution data relative to the base vector set;

[0068] An anchor vector set construction module, used for finding the nearest query vectors in the query vector set for each basis vector in the basis vector set, and generating an anchor vector for each basis vector based on the found query vectors; all anchor vectors constitute an anchor vector set; a fusion graph construction module, used for constructing a fusion graph based on the basis vector set and the anchor vector set;

[0069] The search result output module is used to perform a search on the fusion graph to obtain a basis vector close to the query vector, and return the obtained basis vector as the query result.

[0070] Preferably, the input module includes: a basis vector set input module for obtaining a basis vector set; a query vector set input module for obtaining a query vector set; and a query distribution difference determination module for determining whether the query vector set is out-of-distribution or in-distribution data.

[0071] (14) To illustrate the effectiveness and efficiency of the present invention, the benchmark test results are as follows: Figure 2 The experimental data, baseline algorithm, experimental environment and experimental results involved in the experiment are described as follows:

[0072] Experimental data: This experiment uses five commonly used out-of-distribution datasets (Text-to-Image, LAION, WIT, WebVid, and CC3M) for benchmarking. The five out-of-distribution datasets are specifically introduced as follows: Text-to-Image is an out-of-distribution dataset from a visual search engine. The base vector set consists of image vectors generated by the Se-ResNext-101 model, and the query vector set consists of text vectors extracted from user-specified text queries using a variant of the DSSM model, with the distance metric being the inner product; LAION is a widely used text-image dataset. Both the base vector set and the query vector set are embedded using CLIP-ViT-B / 3, and cosine similarity is used as the metric; WIT is a dataset derived from a text-image application. The text-image pairs are extracted from Wikipedia pages. The base vector set is encoded using CLIP-ViT-B / 32, and the query vector set is generated using CLIP-ViT-B32-multilingual-v1, and cosine similarity is used as the metric; WebVid is a subtitle-video dataset from a material website. The data is encoded using CLIP-ViT-B / 32, and the distance metric used is cosine similarity; CC3M is a text-image pair dataset designed for caption-image systems. The basis vector is encoded using ViT-B / 16, and the additional text is converted to a query vector using BERT, using cosine similarity as the metric.

[0073] Baseline Algorithms: The benchmark experiments compare the proposed method with four advanced graph indexing algorithms, including HNSW and -MNG is a distributed inner graph index algorithm, while RoarGraph and RobustVamana are distributed outer graph index algorithms.

[0074] Experimental environment: All experiments were conducted on an Ubuntu 20.04 server equipped with an Intel(R) Xeon(R) Gold 5218 CPU (2.30GHz) and 125GB of memory. Matrix operations were enhanced using the Eigen library, and indexes were constructed in parallel using OpenMP with 20 threads.

[0075] Experimental results: Figure 2As shown in the figure, the experiment uses recall @100 (rows 1 and 2) and recall @10 (rows 3 and 4) to evaluate search accuracy, and uses query rate per second and number of distance calculations to evaluate search efficiency. In the graphs illustrating recall @k and query rate per second (rows 1 and 3), the curve located at the upper right indicates better performance. On the contrary, in the graphs describing recall @k and number of distance calculations (rows 2 and 4), the curve located at the lower right indicates better performance. The experimental results show that the performance of the present invention is better than all baseline algorithms on all data sets, and it has a better trade-off between recall @k and query rate per second / number of distance calculations, that is, under the same accuracy requirement, it has a shorter search delay and fewer calculations, and under the same search delay or number of calculations, the search results are more accurate.

Claims

1. A distributed external vector retrieval method, characterized in that: The following steps are involved: Step (1) obtaining a base vector set and a query vector set, wherein the query vector set is out-of-distribution data relative to the base vector set; Step (2) for each basis vector in the basis vector set, find the nearest query vectors in the query vector set, and generate an anchor vector for each basis vector based on the found query vectors; all anchor vectors constitute an anchor vector set; Step (3) constructing a fusion graph based on the basis vector set and the anchor point vector set; Step (4) obtains and converts the query text input by the search engine user into a vector to be queried, performs a search on the fusion graph to obtain a basis vector similar to the vector to be queried, and returns the obtained basis vector as the query result.

2. A distributed external vector retrieval method according to claim 1, characterized in that: In step (1), the query vector set is defined as out-of-distribution data relative to the base vector set as follows: Step (1-1) is for The basis vector set of independent and identically distributed basis vectors , and contains The query vector set of independent and identically distributed query vectors ; Query vector The query vector set Any element in the basis vector set In The neighbor set is ,in yes No. A neighbor, yes The nearest neighbor of Nearest neighbor In the basis vector set In The neighbor set is ,in yes No. a neighbor; The query vector is calculated as follows: For the basis vector set Neighborhood alignment quality : ; Query Vector Set Relative to the basis vector set Neighborhood alignment quality Defined as: ; Step (1-2) If the neighborhood alignment quality Less than the preset constant , then the query vector set Relative to the basis vector set The data is outside the distribution, otherwise it is inside the distribution.

3. A distributed external vector retrieval method according to claim 2, characterized in that: In step (2), the process of constructing the anchor vector set includes the following steps: Step (2-1) for the basis vector set Any basis vector in , get In the query vector set In The query neighbor set is ,in The number of query neighbors required to construct the anchor vector set; Step (2-2) converts each basis vector The corresponding query neighbor set , input aggregation function , and get an anchor vector , all anchor vectors constitute the anchor vector set .

4. A distributed external vector retrieval method according to claim 3, characterized in that: In step (3), the process of constructing the fusion graph includes the following steps: Step (3-1) is based on the basis vector set Construct a basic neighbor graph with an out-degree upper limit of L , Each node v in represents a basis vector , Represents the neighbor relationship between nodes. For the basic neighbor graph Any two nodes and , use distance Evaluate and The similarity between For Node The corresponding basis vectors are, For Node The corresponding basis vectors; Step (3-2) for the basic neighbor graph For each node v in The basic neighbors in as candidates and re-ranked; The basic neighbor for The set of endpoints of directed edges starting from node v in ; The reordering includes: for node v and basic neighbors Any node , using the fusion distance in the rearrangement process Evaluate and The similarity between is the fusion weight parameter, and For Node The corresponding basis vectors and anchor vectors, and are the basis vector and anchor point vector corresponding to node v; Step (3-3) Set the outbound limit , for each node v, from the basic neighbors In the example, similarity is sorted according to the fusion distance, and nodes are selected from large to small according to similarity. Try to connect the edges, when the node If the occlusion rule is met, the edges are connected, and the number of connected edges is equal to Stop selecting edges when , and get the fused neighbors ; when After all nodes in the graph are connected according to the occlusion rules, the fusion graph is obtained. ; Fusion graph The set of endpoints of the directed edges starting at node v in .

5. A distributed external vector retrieval method according to claim 4, characterized in that: In step (4), searching on the fusion graph to obtain basis vectors similar to the query vector specifically includes the following steps: Step (4-1) From the fusion graph Select nodes as the entry point, Nodes are placed into the candidate set middle; Step (4-2) From the candidate set Select the unvisited distance query vector The nearest node As the current access node, and The similarity between Evaluate and merge the neighbors of the currently visited node Insert candidate set ,in is the basis vector corresponding to the currently visited node v; The vector to be queried and query vector set Independent and identically distributed; Step (4-3) converts the candidate set For any node v in the Evaluate v and The similarity of the candidate set All nodes in the , sort the similarities from large to small, and delete the candidate set H with a ranking greater than Nodes, where is the basis vector corresponding to node v; Step (4-4) Repeat steps (4-2) to (4-3) until the candidate set There are no unvisited nodes in the candidate set. forward The result is basis vectors.

6. A distributed external vector retrieval method as claimed in claim 1, characterized in that: The base vector set in step (1) is composed of vectors converted from images or videos in the search engine image library, and the query vector set is composed of vectors converted from query texts input by search engine users; Step (4) also includes the following steps: Based on the basis vector obtained in step (4), the picture or video in the search engine image library corresponding to the basis vector is output.

7. A distributed external vector retrieval method as claimed in claim 4, characterized in that: The distance and fusion distance middle, is any of the following three distance metrics: Euclidean distance, cosine similarity, and inner product.

8. A distributed external vector retrieval device, used to implement the distributed external vector retrieval method as claimed in any one of claims 1 to 7, characterized in that: include: An input module, used for obtaining a base vector set and a query vector set as input, wherein the query vector set is out-of-distribution data relative to the base vector set; An anchor vector set building module, used for finding several nearest query vectors in the query vector set for each basis vector in the basis vector set, and generating an anchor vector for each basis vector based on the found query vectors; all anchor vectors constitute an anchor vector set; A fusion graph construction module, used to construct a fusion graph based on a basis vector set and an anchor vector set; The search result output module is used to perform a search on the fusion graph to obtain a basis vector similar to the query vector, and return the obtained basis vector as the query result.

9. A distributed external vector search device as claimed in claim 8, characterized in that: The input module includes: a base vector set input module for obtaining a base vector set; a query vector set input module for obtaining a query vector set; and a query distribution difference determination module for determining whether the query vector set is out-of-distribution or in-distribution data.

Citation Information

Patent Citations

  • Cross-modal data retrieval method, system and equipment based on Hash coding and medium

    CN112925962A

  • Multi-modal search method based on neighbor graph

    CN113656678A

  • K-nearest neighbor graph-based out-of-distribution detection technology

    CN116467650A

  • Partial multi-modal Hash method based on fine-grained feature fusion

    CN118981507A

  • Method and system of retrieving multimodal assets

    US20230306087A1