Adaptive data-driven high-dimensional vector distance estimation method
Through the adaptive data-driven high-dimensional vector distance estimation method, high-dimensional data is projected into low-dimensional space and the orthogonal transformation matrix is optimized, and the dimensions are dynamically adjusted, which solves the problem of high distance calculation complexity in high-dimensional data processing, and achieves efficient and accurate query performance.
Patent Information
- Application Number
- CN202510277190.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-04
AI Technical Summary
In high-dimensional data processing, the distance comparative calculation complexity of the prior art has high results, resulting in performance bottlenecks in high-dimensional approximate nearest neighbor search algorithms, which is difficult to meet practical application requirements, and the existing methods are difficult to balance between accuracy and efficiency.
Adaptive data-driven high-dimensional vector distance estimation method is adopted to project data into low-dimensional space through orthogonal transformation based on data distribution, and optimize the orthogonal transformation matrix in combination with principal component analysis, dynamically adjust the dimensions to reduce the calculation amount while maintaining estimation accuracy.
It significantly improves the query efficiency and accuracy of high-dimensional data processing, can improve search efficiency by more than 40% on multiple data sets, and achieves a leading level of accuracy. It is suitable for seamless integration of existing approximate nearest neighbor search algorithms.
Smart Images

Figure CN120256910A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high-dimensional data processing, and in particular to an adaptive data-driven high-dimensional vector distance estimation method. Background Art
[0002] In the field of high-dimensional data processing, nearest neighbor search is one of the core tasks, which is widely used in scenarios such as large-scale language models, recommendation systems, information retrieval, and data mining. However, as the data dimension increases, traditional indexing methods (such as R-Tree, KD-Tree) perform severely degraded in high-dimensional scenarios, and the computational overhead increases significantly, resulting in low efficiency and inability to meet the actual application requirements. This phenomenon is called the "curse of dimensionality". To solve this problem, researchers have proposed approximate nearest neighbor search methods to improve query efficiency by reducing computational overhead while maintaining high query quality. Existing approximate nearest neighbor search methods mainly include techniques based on graph indexing, vector quantization, tree structures, and hash mapping. The core goal of these methods is to quickly generate a candidate set and screen the candidate points to identify the final nearest neighbor. However, during the screening process, almost all approximate nearest neighbor search algorithms involve distance comparison calculations, that is, a heap is allocated for each target point to store its current nearest neighbor; the distance between the target point and the candidate point is calculated and compared with the candidate point with the farthest distance in the heap; if the distance of the new candidate point is smaller, the heap is updated, otherwise the candidate point is discarded, as shown in the appendix. Figure 1 Since the complexity of distance calculation is O(D) (D is the dimension of the data in the Euclidean space), in high-dimensional cases, the distance calculation comparison operation has become the main performance bottleneck of this type of experiment. Some studies have shown that on the 256-dimensional DEEP dataset, 77% of the total running time is spent on distance comparison calculations, which greatly limits the performance improvement of approximate nearest neighbor search algorithms.
[0003] To reduce the computational overhead of distance comparison, some researchers have tried to approximate the distance calculation to avoid exact calculation in all dimensions. For example, the product quantization method encodes the data vectors so that the distance calculation can be performed in a compressed space, thereby reducing the computational complexity. However, the main problem of this method is that the accuracy drops significantly, so it is more suitable for the preliminary screening of candidate points rather than the final exact nearest neighbor judgment.
[0004] To further optimize the distance comparison calculation, researchers have proposed an optimization method for distance comparison operations based on random orthogonal transformation. The core idea of this method includes:
[0005] First, through random orthogonal transformation (orthogonally projecting the original data to reduce the number of dimensions required for calculation), then performing adaptive sampling (dynamically selecting some dimensions for calculation during the query phase instead of using all dimensions), and it is proved that its distance estimation is unbiased under random projection and the error probability is controlled. Although the distance comparison operation optimization method based on random orthogonal transformation improves the efficiency of distance comparison operations to a certain extent, its greatest limitation lies in the lack of data awareness. That is, the orthogonal transformation of this method is randomly selected and cannot be optimized according to the specific distribution of the data, resulting in difficulty in achieving optimal query performance on some data sets.
[0006] Although some studies have tried to reduce the computational overhead by approximately calculating distances, these methods often lead to a decrease in accuracy and are difficult to meet the requirements for accuracy in practical applications. Summary of the Invention
[0007] The object of the present invention is to provide an adaptive data-driven high-dimensional vector distance estimation method, aiming to optimize the distance comparison operation in high-dimensional approximate nearest neighbor search. When calculating the exact distance using traditional distance comparison operations, the computational amount is huge, seriously affecting the query efficiency. This method is based on orthogonal transformation and hypothesis testing, significantly reducing the computational overhead while maintaining the accuracy of distance estimation. This method can also be used as a pluggable component and seamlessly integrated with existing approximate nearest neighbor search algorithms (such as IVF and HNSW) to improve the search speed while maintaining a high recall rate.
[0008] To achieve the above object, the present invention is implemented according to the following technical solutions:
[0009] The adaptive data-driven high-dimensional vector distance estimation method of the present invention includes the following steps:
[0010] Perform orthogonal transformation based on the data distribution on the original data, projecting the data from the high-dimensional space to the low-dimensional space;
[0011] Calculate the distance estimation value between two vectors in the low-dimensional space such that the expected value of the distance estimation value is equal to the true distance;
[0012] Optimize the distance estimation value to minimize the error between the estimated distance and the true distance;
[0013] Adopt an adaptive dimension expansion strategy based on the data distribution to dynamically adjust the dimensions of distance calculation to reduce the computational amount while ensuring accuracy.
[0014] The beneficial effects of the present invention are:
[0015] The present invention is an adaptive data-driven high-dimensional vector distance estimation method. Compared with the prior art, the present invention significantly improves the query efficiency and accuracy in high-dimensional data processing through the following technical means:
[0016] Orthogonal transformation based on data distribution: By projecting high-dimensional data into a low-dimensional space, the complexity of distance calculation is reduced, and at the same time, the unbiasedness of distance estimation is maintained, ensuring that the estimated value is equal to the expected value of the true distance.
[0017] Unbiased estimation and optimization: By optimizing the orthogonal transformation matrix through the principal component analysis method, while maximizing the variance in the low-dimensional space, minimizing the error between the estimated distance and the true distance, so as to provide a more accurate distance estimation in the low-dimensional space.
[0018] Adaptive dimension expansion strategy: Adopt the hypothesis testing method to dynamically adjust the dimensions required for distance calculation, and adaptively determine the required number of dimensions according to the specific distribution of the data and the query requirements, further reducing the calculation amount.
[0019] Seamless integration with existing algorithms: This method can be used as a pluggable component and seamlessly integrated with existing approximate nearest neighbor search algorithms (such as HNSW and IVF), significantly improving the query speed while maintaining a high recall rate.
[0020] Experimental results show that the present invention can improve the search efficiency by more than 40% on multiple data sets, and at the same time achieve the leading accuracy level in the field, with significant technical advantages and practical application value. Description of the Drawings
[0021] Figure 1 is the flowchart of the distance comparison operation based on the large root heap;
[0022] Figure 2 is the overall architecture diagram of the distance estimation method based on adaptive data drive;
[0023] Figure 3 is the comparison chart of the number of queries per second of several distance estimation methods based on the recall rate;
[0024] Figure 4 is the comparison chart of the number of queries per second of several distance estimation methods based on the average comparison distance. Detailed Embodiment
[0025] The present invention will be further described below in conjunction with the drawings and specific embodiments. The illustrative embodiments and descriptions of this invention are used to explain the present invention, but not to limit the present invention.
[0026] The present invention performs an orthogonal transformation based on data distribution to reduce the complexity of distance calculation while ensuring that the distance calculation result in the low-dimensional space is unbiased, and optimizes the unbiased estimation to improve the accuracy of distance calculation. This method mainly consists of the following steps: unbiased estimation (performing an orthogonal transformation based on data distribution, projecting the data from the high-dimensional space to the low-dimensional space, and approximately calculating the Euclidean distance in the low-dimensional space to reduce the computational complexity of distance calculation operations in the high-dimensional space), optimization of the estimation (further reducing the error between the estimated distance and the true distance by optimizing the unbiased estimation algorithm), and dynamic adjustment of the dimension (adopting an adaptive dimension expansion strategy based on data distribution to dynamically adjust the dimension of distance calculation to minimize the amount of calculation while ensuring accuracy), as shown in the appendix Figure 2 as follows.
[0027] The adaptive data-driven high-dimensional vector distance estimation method of the present invention includes the following steps:
[0028] Performing an orthogonal transformation based on data distribution on the original data to project the data from the high-dimensional space to the low-dimensional space;
[0029] Calculating the distance estimation value between two vectors in the low-dimensional space such that the expected value of the distance estimation value is equal to the true distance;
[0030] Optimizing the distance estimation value to minimize the error between the estimated distance and the true distance;
[0031] Adopting an adaptive dimension expansion strategy based on data distribution to dynamically adjust the dimension of distance calculation to reduce the amount of calculation while ensuring accuracy.
[0032] Unbiased estimation algorithm:
[0033] To estimate the distance in a lower dimension, the present invention first performs an orthogonal transformation on the original data and projects the data into a new coordinate system. Different from the random orthogonal transformation, the present invention adopts an orthogonal transformation based on data distribution, and the estimated distance obtained after the orthogonal transformation is equal to the expected value of the true distance, that is, while reducing the complexity of distance calculation operations, the unbiasedness of the distance calculation result is ensured. Specifically, for any dimension d, given an orthogonal basis W d belonging to the space R d = [w1, w2,..., w d ∈ R D×d , and for any i ≠ j, there is w i T w j = 0 and w i T w i= 1. Assume two independent and identically distributed D-dimensional vectors X1 and X2. Based on the adaptive data-driven distance estimation method, the distance between them is estimated by the following formula:
[0034]
[0035] where Var(w k T X) represents the variance of w k T X.
[0036] Optimization of the estimation algorithm:
[0037] To further optimize the accuracy of distance estimation, that is, to minimize the error between the estimated distance and the true distance, the following optimization objective is set:
[0038]
[0039] where △X = X1 - X 2, W D T W D = I.
[0040] Further derivation reveals that the above optimization objective can be transformed into:
[0041]
[0042] where W d T W d = I, E[XX T represents the matrix composed of all data objects,
[0043] Meanwhile, the optimal solution of this optimization objective can be obtained through the principal component analysis method, that is, to obtain the orthogonal transformation matrix W d , such that the variance of the first d dimensions is maximized while minimizing the error between the estimated distance and the true distance, enabling this method to provide more accurate distance estimation in the low-dimensional space.
[0044] Let λ k represent the k-th largest eigenvalue of the matrix E[XX T , that is:
[0045] E[XX T w k = λ k w k , k = 1, 2,..., d
[0046] In addition, the following equation holds:
[0047] Var(w k T X) = E[w k T XX T w k = w k T E[XX T w k = λ k
[0048] In summary, the final distance calculation formula is as follows:
[0049]
[0050] Adaptive Dimension Expansion Strategy Based on Data Distribution
[0051] In practical applications, different data points require different dimensions to ensure the estimation accuracy. The present invention uses a hypothesis testing method to dynamically determine the dimensions required for distance calculation.
[0052] Hypothesis Testing Strategy:
[0053] Null hypothesis H0: dis′ < r (estimated distance is less than the threshold)
[0054] Alternative hypothesis H1: dis′ ≥ r (estimated distance is greater than or equal to the threshold)
[0055] In the initial stage, the present invention uses a smaller dimension d0 for calculation and calculates the estimation error:
[0056]
[0057] where ∈ e is the error threshold, and P s is the significance level.
[0058] If H0 is rejected, the dimension d = d + Δd is increased until the required confidence level is reached.
[0059] Embodiment:
[0060] The algorithm flow of the distance estimation method based on adaptive data-driven is as follows:
[0061] Input: The target vector q after linear transformation based on an orthogonal matrix, the alternative vector o after linear transformation based on an orthogonal matrix, the distance threshold r, and the dimension expansion step Δd.
[0062] Initialization: Set the initial dimension number d = 0.
[0063] Iteration:
[0064] 1. Increase the dimension number d = d + Δd
[0065] 2. Calculate the estimated distance dis'.
[0066] 3. Conduct a hypothesis test to determine whether the null hypothesis holds (i.e., dis' ≤ r).
[0067] 4. If the hypothesis test passes, return dis'; otherwise, continue to iterate until d > D.
[0068] Output: Return the estimated distance dis'.
[0069] Experimental results:
[0070] By integrating the adaptive data-driven high-dimensional vector distance estimation method proposed in the present invention, the current traditional distance estimation method, and the distance comparison operation optimization method based on random orthogonal transformation, which is currently at the leading level in the industry, into the approximate nearest neighbor search algorithms HNSW and IVF respectively, and on three datasets, DEEP, GIST, and Tiny5M, by setting two sets of hyperparameters K (K = 20 / 100) to conduct experimental comparisons of the number of queries per second of several algorithms under different recall rates and average comparison distances (when the recall rate or the average number of comparisons is the same, the more queries per second, the shorter the single query time and the better the performance). Among them, HNSW and IVF represent two approximate nearest neighbor search algorithms based on the traditional distance estimation method; HNSW + and IVF + represent two approximate nearest neighbor search algorithms based on the distance estimation method of random orthogonal transformation; HNSW ++ and IVF ++ represent two approximate nearest neighbor search algorithms that are optimized by decoupling the heap of each point, so that the heap is not only used to provide the current nearest k and the candidate point set (i.e., the threshold), but also can be used to save and search the compared candidate points, and are based on the distance estimation method of random orthogonal transformation; HNSW * and IVF * represent two approximate nearest neighbor search algorithms based on the adaptive data-driven high-dimensional vector distance estimation method; HNSW ** and IVF ** represent two approximate nearest neighbor search algorithms that are optimized by decoupling the heap of each point, so that the heap is not only used to provide the current nearest k and the candidate point set (i.e., the threshold), but also can be used to save and search the compared candidate points, and are based on the adaptive data-driven high-dimensional vector distance estimation method. Finally, the following conclusions are obtained through experiments:
[0071] (1) At the same recall rate, the number of queries per second of the high-dimensional vector distance estimation method based on adaptive data-driven proposed by the present invention is significantly better than that of the traditional distance estimation method and the distance estimation method based on random orthogonal transformation, with higher query speed and execution efficiency. At the same time, when the number of queries per second is the same, the recall rate of this method is higher and the accuracy is better, as shown in the appendix Figure 3 as follows.
[0072] (2) At the same average comparison distance, the number of queries per second of the high-dimensional vector distance estimation method based on adaptive data-driven proposed by the present invention is better than that of the traditional distance estimation method and the distance estimation method based on random orthogonal transformation, with higher query speed and execution efficiency, as shown in the appendix Figure 4 as follows.
[0073] (3) Combining appendix Figure 3 and Figure 4, it can be seen that in most scenarios (especially when dealing with high-dimensional data), the algorithm performance of the HNSW approximate nearest neighbor search algorithm is better than that of the IVF approximate nearest neighbor search algorithm. And the performance improvement of the optimized algorithm HNSW ++ of the HNSW approximate nearest neighbor search algorithm is relatively obvious.
[0074] The technical solution of the present invention is not limited to the limitations of the above specific embodiments. Any technical deformation made according to the technical solution of the present invention falls within the protection scope of the present invention.
Claims
1. An adaptive data-driven high-dimensional vector distance estimation method, characterized in that, Including the following steps: Perform an orthogonal transformation based on the data distribution on the original data, projecting the data from a high-dimensional space to a low-dimensional space; Calculate the distance estimate between two vectors in the low-dimensional space such that the expected value of the distance estimate is equal to the true distance; Optimize the distance estimate to minimize the error between the estimated distance and the true distance; Adopt an adaptive dimension expansion strategy based on the data distribution to dynamically adjust the dimension of distance calculation to reduce the computational amount while ensuring accuracy.
2. The adaptive data-driven high-dimensional vector distance estimation method according to claim 1, wherein: The orthogonal transformation of the original data based on the data distribution includes, for any dimension d, given an orthogonal basis W d belonging to the space R d = [w1, w2, …, w d ∈ R D×d , and for any i ≠ j, there is w i T w j = 0 and w i T w i = 1; Estimate the distance between two independently and identically distributed D-dimensional vectors X1 and X2 through the following formula: Among them Var(w k T X) represents the variance of w k T X.
3. The adaptive data-driven high-dimensional vector distance estimation method according to claim 1, wherein: The optimization of the distance estimate includes setting the optimization objective as: where △X = X1 - X 2, W D T W D = I; Transform the optimization objective into: Among which W d T W d = I, E[XX T represents the matrix composed of all data objects; Obtain the optimal solution through the principal component analysis method, that is, obtain the orthogonal transformation matrix W through the principal component analysis method d , maximizing the variance of the first d dimensions while minimizing the error between the estimated distance and the true distance.
4. The adaptive data-driven high-dimensional vector distance estimation method according to claim 1, wherein: The adoption of the adaptive dimension expansion strategy based on the data distribution includes: Set the null hypothesis H0: dis′ < r and the alternative hypothesis H1: dis′ ≥ r; where the null hypothesis estimates that the distance is less than the threshold, and the alternative hypothesis estimates that the distance is greater than the threshold; In the initial stage, use a smaller dimension d0 for calculation and calculate the estimation error: where ε e is the error threshold, and P s is the significance level If H0 is rejected, increase the dimension d = d + Δd until the required confidence level is reached.