Learning-based query perception vector partition retrieval method

Through learning partition detection model and redundant partition construction strategy, the detection range is dynamically adjusted, and the problems of redundant calculation and long-tail distribution in traditional partition search methods are solved, achieving efficient and accurate approximate nearest neighbor search.

CN120256689APending Publication Date: 2025-07-04YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510277194.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In traditional partition search methods, there are negative effects caused by partition detection redundancy and long-tail distribution, resulting in large calculation overhead and low retrieval efficiency, making it difficult to meet real-time response requirements.

Method used

Using a learning-based query-aware vector partition retrieval method, we train a learning-based partition detection model and a redundant partition construction strategy, dynamically adjust the detection range, identify long-tail data points, and reduce redundant calculations. Combined with cross-partition and intrapartition retrieval mechanisms, the number of probe partitions is adaptively adjusted.

Benefits of technology

It significantly reduces the computational overhead and query delay, improves the retrieval efficiency and accuracy, and is suitable for approximate nearest neighbor searches of large-scale high-dimensional vector data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256689A_ABST
    Figure CN120256689A_ABST
Patent Text Reader

Abstract

The invention discloses a query perception vector partition retrieval method based on learning, which relates to the technical field of information retrieval, and comprises the following steps: training a learning type partition detection model by using a historical query vector and distance information between the historical query vector and the center of each partition, and constructing a learning type redundant partition construction strategy in the training process, comprising the steps that for a certain data point, if the number of partitions including the data point is larger than a partition number threshold value, the data point is marked as a long mantissa data point; and for each long mantissa data point, if the partition where the long mantissa data point is located is the partition with the highest detection probability, the long mantissa data point is copied to the partition with the second highest detection probability, and otherwise, the long mantissa data point is copied to the partition with the highest probability. According to the method, the problems of'partition detection redundancy 'and'negative influence caused by long-tail distribution' in a traditional partition retrieval method are solved, the calculation overhead is remarkably reduced, and the retrieval efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information retrieval technology, and in particular to a learning-based query-aware vector partition retrieval method. Background Art

[0002] With the continuous development of information retrieval technology, vectorized representation methods have been widely used in the retrieval of unstructured data such as text and images. By converting data into vectors, the semantic similarity between data can be reflected, providing a new way for large-scale data retrieval. However, with the rapid increase in data size and data dimension, the traditional exact nearest neighbor search method is difficult to meet the requirements of real-time response due to the huge amount of calculation. Therefore, the approximate nearest neighbor search method came into being. It can significantly reduce query latency and improve retrieval efficiency at the expense of some retrieval accuracy, and has gained widespread attention in many practical application scenarios.

[0003] At present, most approximate nearest neighbor search methods adopt a partition-based strategy, that is, a clustering algorithm (such as a mean-based clustering algorithm) is used to divide the data set into several partitions. In the query phase, the indexing method sorts the query vector according to the distance between it and the center of each partition, and selects a fixed number of partitions for retrieval. Although the traditional partition-based search method has improved the retrieval efficiency to a certain extent, it still has obvious limitations, which are specifically reflected in the following two aspects:

[0004] First, in traditional partitioning methods, the detection process usually determines the partition to be detected by calculating the distance between the query vector and the center of each partition, that is, detecting multiple partitions closest to the query. However, this method ignores the actual distribution of data within the partition, and this method often produces "partition detection redundancy". Therefore, the traditional partitioning method may lead to the detection of partitions with low correlation with the query vector during the retrieval process, thereby introducing a large amount of unnecessary computational burden and reducing retrieval efficiency.

[0005] Secondly, traditional partitioning strategies usually use hard partitioning, that is, each data point can only belong to one fixed partition. However, in practical applications, the nearest neighbors of the query vector are often distributed in multiple partitions, showing an obvious long-tail distribution phenomenon, and some partitions containing nearest neighbors have only one nearest neighbor. In order to ensure that all relevant data can be covered, traditional methods usually use a fixed number of detection partitions, but this "one-size-fits-all" strategy cannot adaptively adjust the detection range for different queries, which often leads to additional computational overhead while improving the retrieval recall rate.

[0006] Therefore, how to ensure a high retrieval accuracy while adaptively adjusting the detection range and reducing the negative impact of long-tail distribution has become an important issue that needs to be urgently solved in the current field of approximate nearest neighbor search. Summary of the Invention

[0007] The present invention aims to provide a learning-based query-aware vector partition retrieval method for optimizing the approximate nearest neighbor search process of large-scale high-dimensional vector data. By combining a learning model and a redundant partition strategy, this method solves the problems of "partition probing redundancy" and "long-tailed distribution of nearest neighbor data across partitions" existing in traditional partition methods, and significantly reduces query latency and computational overhead while ensuring search accuracy.

[0008] The technical solution adopted by the present invention is as follows:

[0009] A learning-based query-aware vector partition retrieval method, comprising:

[0010] Using historical query vectors and the distance information between historical query vectors and the centers of each partition, training a learning-based partition probing model to predict the partition most relevant to the query vector and output the probing probability of each partition;

[0011] During the training process of the learning-based partition probing model, constructing a learning-based redundant partition construction strategy, including: for a certain data point, if the number of partitions including this data point is greater than the partition number threshold, then mark this data point as a long-tailed data point; for each long-tailed data point, if the partition where it is located is the partition with the highest probing probability, then copy it to the partition with the second highest probing probability, otherwise, copy it to the partition with the highest probability;

[0012] Inputting the user query vector into the trained learning-based partition probing model, obtaining the user query probing probability of each partition, and performing partition retrieval according to the user query probing probability to obtain the partition retrieval result.

[0013] As a further description of the above technical solution, the partition retrieval process includes a cross-partition retrieval stage and an intra-partition retrieval stage;

[0014] The query-aware retrieval mechanism in the cross-partition retrieval stage is: regarding the partitions corresponding to the user query probing probabilities greater than the set probability threshold as the partitions to be probed;

[0015] The query-aware retrieval mechanism in the intra-partition retrieval stage is: for each partition to be probed, using an internal index to perform a nearest neighbor search on the data within this partition to obtain the candidate results of this partition; merging the candidate results of all partitions to be probed, reordering according to the distance between the candidate results and the user query vector, the smaller the distance, the higher the ranking, and selecting the top k candidate results as the final partition retrieval result.

[0016] As a further description of the above technical solution, the output of the learning-based partition detection model is a probability vector with the number of partitions as its dimension. Each element in the probability vector represents the detection probability of the corresponding partition, reflecting the relevance of the partition to the query in the model prediction.

[0017] As a further description of the above technical solution, the learning-based partition detection model includes neural networks for the query branch, centroid distance branch, and fusion branch; the neural network of the query branch is used to extract the features of the query vector; the centroid distance branch is used to convert the distance information into corresponding feature vectors; the fusion branch is used to concatenate the features of the query vector and the features of the query vector to generate the final detection probability.

[0018] As a further description of the above technical solution, during the process of training the learning-based partition detection model, the partition where the true nearest neighbor of each query is located is represented by a binary vector, where the partition containing the nearest neighbor is marked as 1 and the remaining partitions are marked as 0.

[0019] As a further description of the above technical solution, the cross-entropy loss function is used to optimize the model parameters when training the learning-based partition detection model. Specifically:

[0020]

[0021] where B is the total number of partitions, represents the detection probability of the b-th partition predicted by the model for a training query vector q, is the true partition detection label, is the total loss of the detection models for each partition on the query vector q.

[0022] Compared with the prior art, the beneficial effects of the present invention are:

[0023] (1) By training the learning-based partition detection model, it is possible to directly predict the partition most relevant to the query vector based on the features of the query vector and the distance information from the partition center, and output the detection probability of each partition. By setting a probability threshold, only the partitions highly relevant to the query are detected, avoiding the redundant calculations of detecting all partitions in the traditional method and significantly reducing the computational overhead.

[0024] (2) Through the learning-based partition detection model, it is possible to dynamically adjust the number of detected partitions, flexibly select the partitions to be detected according to the characteristics of the query vector, ensuring both a high retrieval recall rate and avoiding unnecessary computational overhead, and significantly improving the retrieval efficiency.

[0025] (3) By means of the learning-based redundant partition construction strategy, long-tail data points are identified and copied into redundant partitions related to their nearest neighbor distributions. In this way, the number of partitions to be probed during query is reduced, effectively alleviating the negative impact brought by the long-tail distribution and further reducing the computational overhead.

[0026] (4) By combining the learning-based partition probing model and the redundant partition construction strategy, it is possible to more accurately predict the partitions related to the query vector and ensure coverage of all possible nearest neighbor data points; in the query-aware retrieval mechanism, through two stages of cross-partition retrieval and in-partition retrieval, the probing range can be adaptively adjusted, and efficient nearest neighbor search can be performed within each probing partition. Finally, the candidate results of all partitions are merged and re-ranked to ensure that the returned retrieval results have high accuracy and recall rates.

[0027] (5) By means of the approximate nearest neighbor search method, the query latency can be significantly reduced at the expense of some retrieval accuracy, which is applicable to the retrieval scenario of large-scale high-dimensional vector data; by dynamically adjusting the number of probing partitions and reducing the negative impact brought by the long-tail distribution, efficient query response can be achieved in large-scale data retrieval, providing strong technical support for practical applications.

[0028] To make the above objects, features, and advantages of the present invention more obvious and understandable, specific embodiments of the present invention are hereinafter given, in conjunction with the accompanying drawings, and are described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 It is the overall architecture diagram of the learning-based query-aware vector partition retrieval method of the present invention;

[0031] Figure 2 It is the structural diagram of the learning-based redundant partition construction strategy of the present invention;

[0032] Figure 3 It is the structural diagram of the query-aware retrieval mechanism of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention.

[0034] Please refer to Figure 1 , the embodiments of the present invention provide a learning-based query-aware vector partition retrieval method. As shown in the attached Figure 1 figure, the following details three core modules of the method framework and their technical methods: the learning-based partition detection model, the learning-based redundant partition construction strategy, and the query-aware retrieval mechanism.

[0035] 1. Learning-based partition detection model

[0036] The learning-based partition detection model is the core of the framework. Its goal is to, based on the initial partitioning of the clustered data partitions, determine which partitions need to be detected according to the characteristics of the query vector, thereby reducing the detection of irrelevant partitions and reducing the computational overhead. Specifically as follows:

[0037] (1) Input and output parts of the learning-based partition detection model

[0038] The learning-based partition detection model accepts two inputs:

[0039] 1) Query vector q, that is, the query data to be searched;

[0040] 2) Distance information I between the query vector q and the centers of each partition, which is used to assist the learning-based partition detection model in judging the relevance between the query and each partition.

[0041] The goal of the learning-based partition detection model is to determine whether each partition needs to be detected, and its output is a probability vector with the number of partitions as the dimension. Each element in the probability vector represents the detection probability of the corresponding partition reflecting the relevance of this partition to the query in the model prediction.

[0042] (2) Structural part of the learning-based partition detection model

[0043] This learning-based partition detection model respectively converts the two inputs (query vector q and the distance information I between the query vector q and the centers of each partition) into feature vectors x q and x I , and then connects the two feature vectors to generate the predicted detection probability as follows:

[0044]

[0045] Among them, φ q 、φI and φ p are the neural networks of the query branch, centroid distance branch, and fusion branch respectively. denotes vector concatenation. Specifically, the query branch is used to extract the feature x of the query vector q q ; the centroid distance branch is used to convert the distance information I into the corresponding feature vector x I ; the fusion branch concatenates the outputs x q 、x I of the query branch and centroid distance branch to generate the final detection probability Through this multi-branch structure, the learning-based partition detection model can effectively combine the query vector q and the distance information I for efficient detection prediction.

[0046] (3) Training of the learning-based partition detection model

[0047] During the model training process, in order to enable the learning-based partition detection model to accurately predict the detection partition of each query, the training process represents the partition where the true neighbor of each query is located with a binary vector p q , that is, the partition containing the neighbor is marked as 1, and the remaining partitions are marked as 0. Therefore, the training problem can be transformed into a multi-variable binary classification problem.

[0048] During the training process, the cross-entropy loss function is used to optimize the model parameters. The model continuously adjusts the network parameters to make the predicted detection probability gradually approach the true neighbor distribution of the query. This process realizes the precise clipping of the partition, thereby reducing unnecessary calculations. The loss function is:

[0049]

[0050] where B refers to the total number of partitions, represents the detection probability of the b-th partition predicted by the model for a training query vector q, and its value is a probability between 0 and 1; is the true partition detection label, and its value is 0 or 1. Therefore, is the total loss of the detection model for each partition on the query vector q.

[0051] 2. Learning-based redundant partition construction strategy

[0052] To mitigate the negative impact of the long-tail distribution on retrieval efficiency, the present invention also proposes a learning-based redundant partition construction strategy. In a dataset, there are query vectors and data points to be queried. That is to say, the data points to be queried are the data points that need to be in the index. In vector approximate nearest neighbor search, the query vector needs to find its similar data points in the index. Generally speaking, in partition-based approximate nearest neighbor search, the nearest neighbors of a data point of a query vector are usually distributed in multiple partitions. Long-tail data points usually refer to the data points where the nearest neighbors of the query vector are scattered in partitions. Based on the learning-based redundant partition construction strategy constructed by the present invention, during the training process of the learning-based partition detection model, it is possible to judge which data points belong to long-tail data points and perform redundant storage between partitions on them, thereby reducing the number of partitions to be detected during query.

[0053] The learning-based redundant partition construction strategy of the present invention is divided into two steps: selecting long-tail data points and copying data points to replicated partitions.

[0054] (1) Selecting long-tail data points

[0055] For each data point in the dataset, the learning-based partition detection model can output the partition detection probability vector corresponding to the data point v Therefore, the present invention can regard the partitions with values greater than 0.5 in the probability vector as the partitions to be detected, and regard the total number of partitions to be detected as the detection partition number of the data point v. When selecting long-tail data points, the data points with a higher detection partition number are more likely to belong to long-tail vector data. Specifically, the present invention obtains the detection partition number for all data points respectively, sorts the data points according to the detection partition number, and takes the data points with the largest 3% detection partition number as long-tail data points. Therefore, in the present invention, based on the inference of vector data by the learning-based partition detection model, the data points with a higher detection partition number in the model prediction are marked as long-tail data points.

[0056] (2) Copying data points to replicated partitions

[0057] When copying data points, in addition to the partition where the data point itself is located, the long-tail data points need to be copied one by one to a replicated partition corresponding to the data point.

[0058] We found that the replicated partitions of a data point usually have a high degree of overlap with the partition distributions of its nearest neighbors, that is, most replicated partitions have a high probability in the prediction results of the learning-based partition detection model. Therefore, for each data point labeled as a long tail, first obtain the partition where the data point itself is located, as well as the partitions with the highest and the second highest prediction probabilities of the learning-based partition detection model. If the partition where the long-tail data point itself is located is the partition with the highest prediction probability of the partition detection model, then copy the data point to the partition with the second highest probability; otherwise, copy the data point to the partition with the highest probability. This learning-based redundant partition construction strategy can effectively reduce the number of partitions that actually need to be probed during query, and further improve the query efficiency.

[0059] Take Figure 2 as an example, where v 1 , v 2 , v 3 these three data points are used to describe the algorithm process. Through prediction, the number of probed partitions for these three data points is 1, 4, and 3 respectively. In this example, it is considered that the number of probed partitions of v 1 is very small, so it is not a long-tail data point and does not need to be redundant, while v 2 and v 3 are long-tail data points. Next is the part of copying data points. The traditional approach is to copy data points by ranking the distances to the partition centers. This method tends to copy data points to the partitions with relatively close distances to the partition centers, and due to considerations such as data locality, it does not completely copy according to the single metric of distance. For example, v 2 is copied to the three partitions ranked 1, 3, and 4 in terms of the distance to the partition center. The approach of the present invention is to copy the data point to another partition with the highest prediction probability of the model except the partition where the data point itself is located. For the data point v 2 , the prediction probability of its partition (i.e., the located partition) ranks first, so it is copied to the partition with the second highest prediction probability. For the data point v 3 , it is not in the partition with the highest prediction probability, so it is copied to the partition with the highest prediction probability.

[0060] 3. Query-Aware Retrieval Mechanism

[0061] The query-aware retrieval mechanism processes a given query according to the trained partition detection model, and obtains a retrieval result better than the traditional clustering partition method, as shown in Figure 3 . The specific process is divided into two stages: the cross-partition retrieval stage and the intra-partition retrieval stage.

[0062] (1) Cross-Partition Retrieval Stage

[0063] In the cross-partition retrieval stage, for a given user query vector q, the trained learning-based partition detection model is used for prediction to obtain the detection probabilities of the query on each partition. Then, according to the set probability threshold σ (e.g., σ = 0.5), the partitions to be detected are screened out (e.g., select the partitions whose detection probabilities for the b-th partition predicted by all models for a training query vector q are greater than the threshold σ as the partitions to be detected). This can ensure that only those partitions relevant to the query are detected, avoiding the situation where all partitions need to be detected in the traditional method, thus significantly reducing the computational overhead.

[0064] (2) Intra-partition retrieval stage

[0065] In the intra-partition retrieval stage, for each selected detection partition, an internal index is used to perform an efficient nearest neighbor search on the data within the partition to obtain the candidate results of the partition. Finally, after merging the candidate results of all detection partitions, they are re-sorted according to the distance between the data points (candidate results) and the user query vector q, and the top k data points are selected as the final Top-k retrieval results.

[0066] The query-aware retrieval process provided by the embodiments of the present invention enables the framework to adaptively adjust the number of detected partitions according to the characteristics of the query. While ensuring a high recall rate, it significantly reduces the query latency and computational amount. Compared with the traditional fixed partition number detection strategy, the framework of the present invention can flexibly adjust the detected partition range according to the specific requirements of the query, thereby achieving a more efficient retrieval process.

[0067] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A learning-based query-aware vector partition retrieval method, characterized in that Including: Utilize the historical query vector and the distance information between the historical query vector and each partition center to train a learning-based partition detection model, predict the partition most relevant to the query vector, and output the detection probability of each partition; During the training process of the learning-based partition detection model, construct a learning-based redundant partition construction strategy, including: for a certain data point, if the number of partitions containing this data point is greater than the partition number threshold, mark this data point as a long-tail data point; for each long-tail data point, if the partition it belongs to is the partition with the highest detection probability, copy it to the partition with the second-highest detection probability, otherwise, copy it to the partition with the highest probability; Input the user query vector into the trained learning-based partition detection model, output the user query detection probability of each partition, and perform partition retrieval based on the user query detection probability to obtain the partition retrieval result.

2. The learning-based query-aware vector partition retrieval method according to claim 1, wherein The partition retrieval process includes a cross-partition retrieval stage and an intra-partition retrieval stage; The query-aware retrieval mechanism in the cross-partition retrieval stage is: use the partitions corresponding to the user query detection probability greater than the set probability threshold as the partitions to be detected; The query-aware retrieval mechanism in the intra-partition retrieval stage is: for each partition to be detected, use the internal index to perform a nearest neighbor search on the data within this partition to obtain the candidate results of this partition; merge the candidate results of all partitions to be detected, re-rank according to the distance between the candidate results and the user query vector, the smaller the distance, the higher the ranking, and select the top k candidate results as the final partition retrieval result.

3. The learning-based query-aware vector partition retrieval method according to claim 1, wherein The output of the learning-based partition detection model is a probability vector with the dimension of the number of partitions. Each element in the probability vector represents the detection probability of the corresponding partition, reflecting the relevance of this partition to the query in the model prediction.

4. The learning-based query-aware vector partition retrieval method according to claim 1, wherein The learning-based partition detection model includes a neural network with a query branch, a centroid distance branch, and a fusion branch; the neural network of the query branch is used to extract the features of the query vector; The centroid distance branch is used to convert the distance information into the corresponding feature vector; the fusion branch is used to splice the features of the query vector and the features of the query vector to generate the final detection probability.

5. The learning-based query-aware vector partition retrieval method according to claim 1, wherein During the training process of the learning-based partition detection model, the partition where the true nearest neighbor of each query is located is represented by a binary vector, the partitions containing the nearest neighbor are marked as 1, and the remaining partitions are marked as 0.

6. The learning-based query-aware vector partition retrieval method according to claim 1, wherein When training the learning-based partition detection model, use the cross-entropy loss function to optimize the model parameters, specifically: where B is the total number of partitions, represents the detection probability of the b-th partition predicted by the model for a training query vector q, is the true partition detection label, is the total loss of the detection models for each partition on the query vector q.