High-dimensional k-nearest neighbor micro-cluster search method based on self dimension

Through the high-dimensional k-nearest neighbor microcluster search method based on its own dimensions, the problem of inefficient similarity retrieval in high-dimensional space is solved, fast and efficient query is achieved, and the requirements of result density attributes are met.

CN120045742APending Publication Date: 2025-05-27NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510108948.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing similarity search methods are inefficient in high-dimensional space and ignore the characteristics of the data points themselves, resulting in poor query performance, especially in application scenarios where the result density attribute needs to be considered.

Method used

The high-dimensional k-nearest neighbor microcluster search method based on its own dimensions is adopted, and the projection vector is decentralized through principal component analysis method, the own dimension of each object is calculated, and the data points are divided and indexed according to their own dimensions to meet the density requirements.

Benefits of technology

Shorten query time, improve search efficiency, save storage space, and provide accurate results in k-nearest neighbor problems under density constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045742A_ABST
    Figure CN120045742A_ABST
Patent Text Reader

Abstract

The invention provides a high-dimensional k-nearest neighbor micro-cluster search method based on self dimensions, and relates to the technical field of similarity retrieval. The method comprises the steps that firstly, objects in a vector database are projected, the dimension of each object is calculated, all the objects in the vector database are partitioned, and an index is constructed for each data block; projecting the query vector into the same space of a vector database, and calculating a reference dimension of the query vector; setting an upper limit of a distance between a cluster center of a k neighbor micro-cluster query result and a query vector, and setting left and right pointers from a reference dimension of the query vector as a query starting point to perform range query in data blocks corresponding to the pointers; checking whether the k-nearest neighbor micro-cluster query result meets the k-nearest neighbor micro-cluster query distance requirement and density requirement or not, updating the distance upper limit, continuing to query until the left pointer and the right pointer reach the left boundary and the right boundary of the data block where the query result is located, and returning the k-nearest neighbor micro-cluster result set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of similarity retrieval, and in particular to a high-dimensional k-nearest neighbor microcluster search method based on its own dimension. Background Art

[0002] Similarity retrieval is a basic problem in the field of computer science. Similarity search can generally be divided into two categories: k-nearest neighbor retrieval and range retrieval. With the development of AI technology, people can abstract unstructured data such as pictures and texts into high-dimensional vectors. The advent of the big data era and the explosion of massive data have made similarity retrieval applications increasingly widespread. For example, recommendation systems, picture retrieval, text retrieval, etc. The common feature of these applications is that the data volume is large and the dimension is high.

[0003] Existing similarity retrieval methods can be divided into four types: tree-based, graph-based, hash-based, and quantization-based.

[0004] (1) Tree-based similarity retrieval method

[0005] By constructing a tree-like data structure, vectors are divided and organized according to certain rules. However, as the dimension increases, the efficiency will be significantly reduced due to the "curse of dimensionality".

[0006] (2) Graph-based method

[0007] Construct a neighbor graph for the data and search on the graph through a greedy idea. However, the cost of constructing the graph index is huge, and it is difficult to maintain the graph index when batch inserting and deleting data.

[0008] (3) Hash-based method

[0009] Project vectors into specific hash buckets through a family of hash functions and search in the hash buckets corresponding to the query vectors. Although this method is easy to construct, the query performance depends to a large extent on the design of the family of hash functions. In addition, there are problems such as hash boundaries and hash collisions.

[0010] (4) Quantization-based method

[0011] Divide the vectors in the high-dimensional space into several sub-vectors (sub-spaces), and then independently quantize each sub-vector. It can not only significantly reduce the storage space but also quickly calculate the similarity between vectors. However, the quantization granularity of the product quantization will significantly affect the query accuracy and performance.

[0012] Existing tree-based methods perform poorly in high-dimensional spaces. For graph-based methods, graph indexing is difficult to construct and maintain. Hash-based and quantization-based methods reduce the dataset to the same dimension through hash function families and vector partitioning methods, making construction simple. However, this approach of reducing all data to a unified dimension ignores the characteristics of the data points themselves, causing query performance to largely depend on the choice of dimension.

[0013] Currently, there are also methods to reduce the objects in the database to different dimensions by setting a fixed loss. However, it is difficult to set a reasonable loss in this way. If the loss is set too large, most data points will be concentrated in the lower few dimensions, while if the loss is set too small, the dimensionality reduction effect will not be obvious. Therefore, different loss thresholds can be set for each object according to its distance from the origin (i.e., the norm of the vector).

[0014] In addition, although existing k-nearest neighbor search and range search algorithms can solve most problems. However, these two search algorithms only focus on the distance between the final result and the query point, ignoring the density attribute of the returned results themselves, which makes k-nearest neighbor search and range search not perfectly applicable in some application scenarios. For example, regarding the research technical direction of each user as a vector, when a user wants to select a certain technical direction for learning, not only the similarity between the technical direction and the user's query needs should be considered, but also the activity of this technical direction. This activity can be regarded as the density with the number of objects covered by a hypersphere centered on the result vector itself as a reference. Results with a density reaching the set value ρ are considered to meet the density requirements. The results that meet the requirements are regarded as a hypersphere with a density reaching the set value ρ, that is, a microcluster. k-nearest neighbor microcluster query aims to return k nearest microclusters in such application scenarios. If traditional similarity retrieval is used, the returned results will not meet the density requirements. Summary of the Invention

[0015] The technical problem to be solved by the present invention is to provide a high-dimensional k-nearest neighbor microcluster search method based on its own dimension in view of the above-mentioned deficiencies of the prior art. By reducing data points to different dimensions according to their own characteristics, the query time is shortened, the index construction overhead is reduced, and it is verified whether the density within the microcluster centered on the returned results themselves meets the requirements to solve the k-nearest neighbor microcluster problem.

[0016] To solve the above technical problems, the technical solution adopted by the present invention is: A high-dimensional k-nearest neighbor microcluster search method based on its own dimension, comprising the following steps:

[0017] Step 1: Obtain the original vectors of all objects from the vector database, project the vectors in the original vector space to the projection space by principal component analysis after de-centering, and retain all information;

[0018] The object vector in the original space is de-centered and projected onto the projection space through the principal component analysis method, and its form is shown in formula (1):

[0019]

[0020] where o is the object vector after projection, C is the transformation matrix obtained by the principal component analysis method, o′ is the representation of the object vector in the original space, and is the mean value of the object vector in the original space;

[0021] Step 2: Calculate the self-dimension of each object in the projected vector database;

[0022] Define the loss of a d-dimensional object vector o after truncating the first d ′ dimensions to obtain the vector as shown in formula (2):

[0023]

[0024] where, represents the vector obtained after truncating the first d ′ dimensions of the vector o, represents the loss after truncating the first d ′ dimensions of the vector o, and o i represents the value of the i-th dimension of the vector o;

[0025] According to the predefined loss parameter ε, define the self-dimension sd(o) of the object vector o as its shortest truncation length, as shown in formula (3) specifically:

[0026]

[0027] where sd(o) is the self-dimension of the data point o;

[0028] Step 3: Divide and block all objects in the vector database according to the self-dimension of the object vector;

[0029] Step 3-1: Logically divide all objects in the vector database into d categories according to their self-dimensions;

[0030] Step 3-2: Block all objects in the vector database by category, physically divide all objects in the database into m data blocks, and set the threshold τ for the number of objects contained in each data block;

[0031] First, divide the objects into d data blocks according to the logical d categories, and then, according to the threshold τ, merge the blocks with the number of objects less than the threshold τ in the data block with its adjacent block, and split the blocks with the number of objects exceeding τ in the data block;

[0032] Step 3-3: Determine the dimension of the data block and truncate the objects within each data block;

[0033] Determine the dimension for each data block such that the dimension of each object is not lower than that of its corresponding data block, and truncate the objects within each data block. The truncation dimension is the dimension of the corresponding data block;

[0034] Step 4: Use a high-dimensional index capable of performing range queries to construct an index for each data block respectively. The dimension of the index is the same as that of the data block;

[0035] Step 5: Project the query vector into the same space of the vector database and calculate the benchmark dimension of the query vector;

[0036] The benchmark dimension bd(q) of the query vector q is shown in the following formula (4):

[0037]

[0038] where δ max represents the maximum loss after truncating the objects within each data block;

[0039] Step 6: Set an upper limit R for the distance between the cluster center of the k-nearest neighbor micro-cluster query result and the query vector, determine the left and right boundaries of the data block where the k-nearest neighbor micro-cluster query result is located, and set left and right pointers starting from the benchmark dimension of the query vector to perform a range query within the data block corresponding to the pointer. The radius of the range query is R;

[0040] When the self-dimension of the k-nearest neighbor micro-cluster result p is less than the benchmark dimension of the query vector q, formulas (6) and (7) hold:

[0041] δ(q[1,……,bd(q)-1])>δ max (6)

[0042] (‖p[sd(p),……d]‖ - ‖q[sd(p),……d]‖) 2 ≤‖p,q‖≤R (7)

[0043] From formulas (6) and (7), we can obtain:

[0044] δ(q[1,……sd(p)]) = ‖q[sd(p),……d]‖ ≤ R + ‖p[sd(p),……d]‖ ≤ R + δ max (8)

[0045] δ(q[1,……,sd(p)]) - δ(q[1,……,bd(q)-1]) ≤ R (9)

[0046] Under the limit that the upper bound of the distance between the farthest cluster center in the k-nearest neighbor microcluster and the query vector is R, the minimum self-dimension of the k-nearest neighbor microcluster query result is shown in formula (10):

[0047] sd(ans) L =min{i|1≤i≤sd(q)&δ(q[1,……,i])-δ(q[1,……,bd(q)-1])≤R}(10)

[0048] The maximum self-dimension of the k-nearest neighbor microcluster query result is shown in formula (11):

[0049] sd(ans) R =max{i|i≥sd(q)&δ(p[1,……,bd(q)])-δ(p[1,……,i])≤R&sd(p)=i}(11)

[0050] Among them, sd(ans) L and sd(ans) R represent the maximum and minimum values of the self-dimension of the k-nearest neighbor microcluster query result respectively, and δ(q[1,……,i]) represents the loss of the query vector q intercepted in the first i dimensions;

[0051] Set the left pointer to point to the data block with dimension sd(q)-1, and the right pointer to point to the data block with dimension sd(q), and the step size of each pointer movement is 1;

[0052] Step 7. Check whether the k-nearest neighbor microcluster query result meets the k-nearest neighbor microcluster query distance requirement and density requirement, put the qualified query results into the k-nearest neighbor microcluster result set, and update the upper bound R of the distance between the k-nearest neighbor microcluster result and the query vector;

[0053] Step 7-1. Check whether the distance between the k-nearest neighbor microcluster query result and the query vector exceeds the upper bound R of the distance in the original vector space;

[0054] Step 7-2. Check whether the query result meets the density requirement of the k-nearest neighbor microcluster;

[0055] Taking the query result itself as the center and the radius r of the microcluster as the distance for range query, check whether the number of query return results reaches the microcluster density requirement. During the range search process, determine the left and right boundaries of the data block according to Step 5, perform range query on the data blocks within the boundaries, and finally verify the results of the range query in the original space;

[0056] Step 7-3: Put the qualified query results into the k-nearest neighbor micro-cluster result set. When the number of query results in the k-nearest neighbor micro-cluster result set reaches k, select the query result with the largest distance from the query vector as the distance upper limit, and update the distance upper limit R between the k-nearest neighbor micro-cluster result and the query vector.

[0057] Step 8: Move the left and right pointers to both sides of the query starting point, and return to Step 5 for the next iteration until both the left and right pointers have reached the left and right boundaries of the data block where the query results are located, and then return the k-nearest neighbor micro-cluster result set.

[0058] After studying the existing similarity search methods, the method of the present invention discovers the limitations of existing high-dimensional space similarity queries in certain application backgrounds. Ordinary k-nearest neighbor queries only focus on the distance relationship, but ignore the density attribute of the returned results themselves. In some practical applications, the density attribute is required to represent the popularity or diversity of the returned results, and there are certain requirements for the density attribute of the results. Therefore, the method of the present invention converts the query problem into a k-nearest neighbor problem under density constraint conditions, that is, the k-nearest neighbor micro-cluster problem. A micro-cluster refers to a hypersphere with a certain data point in the database as the center and a radius of r. The data points contained within this hypersphere represent the density of this micro-cluster. The self-dimension refers to the dimension that each object's corresponding vector can intercept under a certain loss ratio according to its own modulus length. According to the characteristics of the data points themselves, the objects in the database are reduced to different self-dimensions, and indexes are constructed separately according to different dimensions, reducing the time and space overhead of index construction and accelerating the query speed.

[0059] The beneficial effects of adopting the above technical solutions are as follows: A high-dimensional k-nearest neighbor micro-cluster search method based on self-dimension provided by the present invention: (1) Firstly, it proposes and solves the k-nearest neighbor micro-cluster search problem, and returns a set of solutions to the k-nearest neighbor problem in the form of micro-clusters, with restrictions on both the distance between the query point and the results and the density of the micro-clusters centered on the results themselves; (2) It allows users to select appropriate micro-cluster radii and micro-cluster densities as restrictive conditions according to different application requirements to solve practical problems such as searching for popular directions; (3) According to the different characteristics of the data points themselves, the data points are reduced to different dimensions to construct indexes, shortening the search time, improving the search efficiency, and saving storage space; Description of the Drawings

[0060] Figure 1 It is a flowchart of a high-dimensional k-nearest neighbor micro-cluster search method based on self-dimension provided by an embodiment of the present invention;

[0061] Figure 2 It is a flowchart for verifying the density of points in the k-nearest neighbor micro-cluster result set provided by an embodiment of the present invention;

[0062] Figure 3 Schematic diagram of high-dimensional k-nearest neighbor microcluster search based on its own dimension provided by an embodiment of the present invention. Detailed implementation manners

[0063] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0064] In this embodiment, taking a certain query object as an example, the high-dimensional space k-nearest neighbor microcluster search method based on its own dimension of the present invention is used to query the object. In the embodiment of the present invention, the database object and the query object are pictures, and the original vectors of the database object and the query object are provided by the ANN_SIFT10K dataset and the ANN_SIFT1M dataset. Both of these datasets are composed of vectors extracted from natural images, and contain 10,000 128-dimensional vectors and 1,000,000 128-dimensional vectors respectively. And each dataset contains 1,000 128-dimensional query vectors.

[0065] In this embodiment, a high-dimensional k-nearest neighbor microcluster search method based on its own dimension, as Figure 1 shown, includes the following steps:

[0066] Step 1: Obtain the original vectors of all objects from the vector database, project the vectors in the original vector space to the projection space by principal component analysis after de-centralizing, and retain all information;

[0067] The form of projecting the object vectors in the original space to the projection space by principal component analysis is shown in formula (1):

[0068]

[0069] where o is the object vector after projection, C is the transformation matrix obtained by principal component analysis, o′ is the representation of the object vector in the original space, is the mean value of the object vector in the original space;

[0070] All dimensional information is retained during the principal component analysis process, so the dimension of the projection space is the same as that of the original space;

[0071] Step 2: Calculate the own dimension of each object in the projected vector database;

[0072] Define the loss of a d-dimensional object vector o after truncating the first d ′ dimensions to obtain the vector as shown in formula (2):

[0073]

[0074] Among them, represents the vector obtained after the vector o is truncated by the first d ′ dimensions, represents the loss after the vector o is truncated by the first d ′ dimensions, and o i represents the value of the i-th dimension of the vector o;

[0075] According to the predefined loss parameter ε, the self-dimension sd(o) of the object vector o is defined as its shortest truncation length, as shown in formula (3) specifically:

[0076]

[0077] where sd(o) is the self-dimension of the data point o;

[0078] In this embodiment, the loss parameter ε is set to 10%;

[0079] The physical meaning of the self-dimension is that after arranging the dimensions of the entire space in descending order of variance, how many dimensions each object can use to express its own characteristics. The farther the vector corresponding to the object is from the origin (i.e., the larger the modulus of the vector), the more information needs to be represented, and vice versa;

[0080] Taking the vector [5, 7, 4, 3, 10, 1] as an example, when the loss parameter is 10%, the modulus of the vector is When the first 5 dimensions are truncated, the loss is meeting the requirements of formula (3). When the first 4 dimensions are truncated, the loss is not meeting the requirements of formula (3). Therefore, the self-dimension of this vector is 5;

[0081] Step 3: Divide and block all objects in the vector database according to the self-dimension of the object vector. The specific steps are as follows:

[0082] Step 3-1: Logically divide all objects in the vector database into d categories according to their self-dimensions;

[0083] In this embodiment, the vectors in the ANN_SIFT10K and ANN_SIFT1M data sets are all 128-dimensional and are divided into 128 categories;

[0084] Step 3-2: Block all objects in the vector database by category, physically divide all objects in the database into m data blocks, and set the threshold τ for the objects included in each data block;

[0085] First, objects are divided into d data blocks according to the logical d categories. Then, according to the threshold τ, the blocks whose number of objects does not reach the threshold τ are merged with their adjacent blocks, and the blocks whose number of objects exceeds τ are split.

[0086] In this embodiment, the number of objects in each data block is limited to 2000. This avoids the situation where the amount of data in the index corresponding to a data block is small, saving space overhead when building the index; it also avoids the situation where the amount of data in the index corresponding to a data block is too large, so that the query has good query performance.

[0087] Step 3-3, determine the dimension of the data block, and truncate the objects in each data block; determine the dimension of each data block so that the dimension of each object is not less than the dimension of the data block to which it belongs, and truncate the objects in each data block, and the truncated dimension is the dimension of the data block to which it belongs;

[0088] Step 4: Use a high-dimensional index that can perform range queries to build an index for each data block. The dimension of the index is consistent with the dimension of the data block.

[0089] In this embodiment, a pivot-based high-dimensional index structure EPT* is used to construct an index for each data block;

[0090] Step 5: Project the query vector into the same space of the vector database and calculate the base dimension of the query vector;

[0091] The base dimension bd(q) of the query vector q is given by the following formula (4):

[0092]

[0093] Among them, δ max Indicates the maximum value of the loss after object truncation in each data block;

[0094] Step 6: Set the upper limit R of the distance between the cluster center of the k-nearest-neighbor micro-cluster query result and the query vector, determine the left and right boundaries of the data block where the k-nearest-neighbor micro-cluster query result is located, set the left and right pointers from the base dimension of the query vector as the query starting point, and perform a range query in the data block corresponding to the pointers. The radius of the range query is R.

[0095] In this embodiment, because the distance between the farthest cluster center in the k nearest neighbor micro-cluster and the query vector cannot be determined at the beginning, the distance upper limit R is set to ∞ when the distance upper limit is initially set, and the distance upper limit is shrunk in the subsequent search process. When the dimension of the k nearest neighbor micro-cluster result p is smaller than the reference dimension of the query vector q, formulas (6) and (7) are established:

[0096] δ(q[1,……,bd(q)-1])>δ max (6)

[0097] (‖p[sd(p),……d]‖ - ‖q[sd(p),……d]‖) 2 ≤‖p,q‖≤R (7)

[0098] From formulas (6) and (7), it can be obtained that:

[0099] δ(q[1,……sd(p)]) = ‖q[sd(p),……d]‖ ≤ R + ‖p[sd(p),……d]‖ ≤ R + δ max (8)

[0100] δ(q[1,……,sd(p)]) - δ(q[1,……,bd(q)-1]) ≤ R (9)

[0101] Therefore, under the limitation that the upper limit of the distance between the farthest cluster center in the k-nearest neighbor micro-cluster and the query vector is R, the minimum self-dimension of the k-nearest neighbor micro-cluster query result is as shown in formula (10):

[0102] sd(ans) L = min{i|1 ≤ i ≤ sd(q) & δ(q[1,……,i]) - δ(q[1,……,bd(q)-1]) ≤ R} (10)

[0103] Similarly, the maximum self-dimension of the k-nearest neighbor micro-cluster query result is as shown in formula (11):

[0104] sd(ans) R = max{i|i ≥ sd(q) & δ(p[1,……,bd(q)]) - δ(p[1,……,i]) ≤ R & sd(p) = i} (11)

[0105] Among them, sd(ans) L and sd(ans) R respectively represent the maximum and minimum values of the self-dimension of the k-nearest neighbor micro-cluster query result, and δ(q[1,……,i]) represents the loss of the query vector q after intercepting the first i dimensions;

[0106] In this embodiment, the set left pointer points to the data block with dimension sd(q)-1, and the right pointer points to the data block with dimension sd(q), and the step size of each pointer movement is 1;

[0107] Step 7, check whether the k-nearest neighbor micro-cluster query result meets the k-nearest neighbor micro-cluster query distance requirement and density requirement, put the qualified query results into the k-nearest neighbor micro-cluster result set, and update the upper limit R of the distance between the k-nearest neighbor micro-cluster result and the query vector;

[0108] Step 7-1: Check whether the distance between the k-nearest neighbor microcluster query result and the query vector in the original vector space exceeds the distance upper limit R;

[0109] In this embodiment, the range query within the data block only ensures that the distance between the returned cluster center and the query vector does not exceed the distance upper limit R within the dimension corresponding to the data block. Verify whether the distance requirement is met in the original space, that is, for the query result in the low-dimensional space corresponding to the data block, calculate its distance from the query point in the original space and determine whether this distance exceeds R.

[0110] Step 7-2: Check whether the query result meets the density requirement of the k-nearest neighbor microcluster;

[0111] Perform a range query with the query result itself as the center and the radius r of the microcluster as the distance to check whether the number of query return results reaches the microcluster density requirement. During the range search process, the left and right boundaries of the data block will be determined according to Step 5, and a range query will be performed on the data blocks within the boundaries. Finally, verify the results of the range query in the original space. The process in this embodiment is as Figure 2 shown;

[0112] Since the dimension of the data block is different from that of the original space, it cannot be guaranteed that the distance between the query result within the data block range and the query point in the original space still does not exceed R. Therefore, it is necessary to recalculate the distance between the two points in the original space to determine whether it exceeds R.

[0113] Step 7-3: Put the query results that meet the requirements into the k-nearest neighbor microcluster result set. When the number of query results in the k-nearest neighbor microcluster result set reaches k, select the query result with the largest distance from the query vector as the distance upper limit and update the distance upper limit R between the k-nearest neighbor microcluster result and the query vector;

[0114] In this embodiment, the k-nearest neighbor microcluster result set represents the set of microclusters that are currently closest to the query vector. When the number of query results in the result set reaches k, it means that the distance between the cluster center that is farthest from the query vector in the k-nearest neighbor microcluster result and the query vector will not exceed the distance between the cluster center and the query vector in the current result set. Therefore, it is necessary to update the distance upper limit;

[0115] Step 8: Move the left and right pointers to both sides of the query starting point, return to Step 5 for the next iteration until both the left and right pointers have reached the left and right boundaries of the data block where the query results are located, and return the k-nearest neighbor microcluster result set.

[0116] In this embodiment, the k-nearest neighbor microcluster search can be represented by Figure 3 and the black solid dots p 1 ~p 9 represent the objects in the database, and the white hollow dots represent the query object q. MC1 , MC 2 and MC 3 are three micro-clusters with a radius of r centered at points p 1 , p 6 and p 8 respectively. When performing a k-nearest neighbor micro-cluster search on a query point, where k = 2 and the micro-cluster density requirement is 3, MC 1 and MC 3 are the two micro-clusters closest to the query point and with densities meeting the requirements, while the density of MC 2 is only 2 and does not meet the micro-cluster density requirement, so it cannot be used as a return result.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A high-dimensional k-nearest-neighbor micro-cluster search method based on its own dimension, characterized by: The following steps are involved: Step 1: Get the original vectors of all objects from the vector database, and use the principal component analysis method to decentralize and project the vectors in the original vector space into the projection space, while retaining all information; Step 2: Calculate the own dimension of each object in the projected vector database; Step 3: Divide and block all objects in the vector database according to the dimension of the object vector itself; Step 4: Use a high-dimensional index that can perform range queries to build an index for each data block. The dimension of the index is consistent with the dimension of the data block. Step 5: Project the query vector into the same space of the vector database and calculate the base dimension of the query vector; Step 6: Set the upper limit R of the distance between the cluster center of the k-nearest-neighbor micro-cluster query result and the query vector, determine the left and right boundaries of the data block where the k-nearest-neighbor micro-cluster query result is located, set the left and right pointers from the base dimension of the query vector as the query starting point, and perform a range query in the data block corresponding to the pointers. The radius of the range query is R. Step 7: Check whether the k-nearest-neighbor micro-cluster query result meets the k-nearest-neighbor micro-cluster query distance requirement and density requirement, put the query result that meets the requirement into the k-nearest-neighbor micro-cluster result set, and update the upper limit R of the distance between the k-nearest-neighbor micro-cluster result and the query vector; Step 8: Move the left and right pointers to both sides of the query starting point, and return to step 5 for the next iteration until the left and right pointers have reached the left and right boundaries of the data block where the query result is located, and return the k-nearest neighbor micro-cluster result set.

2. The high-dimensional k-nearest-neighbor micro-cluster search method based on its own dimension according to claim 1, characterized in that: The form of the decentralized projection of the object vector in the original space to the projection space by the principal component analysis method in step 1 is shown in formula (1): Among them, o is the projected object vector, C is the transformation matrix obtained by principal component analysis, and o′ is the representation of the object vector in the original space. is the mean of the object vector in the original space.

3. The high-dimensional k-nearest-neighbor micro-cluster search method based on its own dimension according to claim 2, characterized in that: The step 2 comprises: Define a d-dimensional object vector o before truncation d ′ After the dimensions are The loss is shown in formula (2): in, Denotes the vector o before interception d ′ The vector obtained after the dimension is Represents the vector o before interception d ′ The loss after the dimension is o i Represents the value of the i-th dimension of vector o; According to the predefined loss parameter ε, the dimension sd(o) of the object vector o is defined as its shortest truncation length, as shown in formula (3): Among them, sd(o) is the dimension of the data point o itself.

4. The high-dimensional k-nearest-neighbor micro-cluster search method based on its own dimension according to claim 3, characterized in that: The step 3 comprises: Step 3-1, logically divide all objects in the vector database into d categories according to their own dimensions; Step 3-2, divide all objects in the vector database into blocks according to categories, physically divide all objects in the database into m data blocks, and set a threshold τ for the objects contained in each data block; First, objects are divided into d data blocks according to the logical d categories. Then, according to the threshold τ, the blocks whose number of objects does not reach the threshold τ are merged with their adjacent blocks, and the blocks whose number of objects exceeds τ are split. Step 3-3, determine the dimension of the data block, and truncate the objects in each data block; The dimension of each data block is determined so that the dimension of each object is not less than the dimension of the data block to which it belongs, and the objects in each data block are truncated, and the truncated dimension is the dimension of the data block to which it belongs.

5. The high-dimensional k-nearest-neighbor micro-cluster search method based on its own dimension according to claim 4, characterized in that: The base dimension of the query vector in step 5 is as shown in the following formula (4): Where bd(q) is the base dimension of query q; δ max Indicates the maximum value of the loss after object truncation in each data block.

6. The high-dimensional k-nearest-neighbor micro-cluster search method based on its own dimension according to claim 5, characterized in that: The step 6 comprises: When the dimension of the k-nearest-neighbor micro-cluster result p is smaller than the baseline dimension of the query vector q, formulas (6) and (7) hold: δ(q[1,……,bd(q)-1])>δ max (6)(‖p[sd(p),……d]‖-‖q[sd(p),……d]‖) 2 ≤‖p,q‖≤R (7) From formula (6) and formula (7), we can get: δ(q[1,……sd(p)])=‖q[sd(p),……d]‖≤R+‖p[sd(p),……d]‖≤R+δ max (8) δ(q[1,…,sd(p)])-δ(q[1,…,bd(q)-1])≤R (9) Under the restriction that the upper limit of the distance between the farthest cluster center in the k-nearest-neighbor micro-cluster and the query vector is R, the minimum self-dimension of the k-nearest-neighbor micro-cluster query result is shown in formula (10): sd(ans) L =min{i|1≤i≤sd(q)&δ(q[1,……,i])-δ(q[1,……,bd(q)-1])≤R} (10) The maximum self-dimension of the k-nearest-neighbor micro-cluster query result is shown in formula (11): sd(ans) R =max{i|i≥sd(q)&δ(p[1,……,bd(q)])-δ(p[1,……,i])≤R&sd(p)=i}(11) Among them, sd(ans) L and sd(ans) R They represent the maximum and minimum values ​​of the k-nearest-neighbor micro-cluster query result dimensions, respectively. δ(q[1,…,i]) represents the loss of the query vector q after capturing the first i dimensions. The left pointer is set to point to a data block with a dimension of sd(q)-1, and the right pointer is set to point to a data block with a dimension of sd(q), and the step size of each pointer movement is 1.

7. The high-dimensional k-nearest-neighbor micro-cluster search method based on its own dimension according to claim 6, characterized in that: The step 7 comprises: Step 7-1, check in the original vector space whether the distance between the k nearest neighbor micro-cluster query result and the query vector exceeds the upper limit R of the distance; Step 7-2: Check whether the query result meets the density requirement of k-nearest neighbor micro-clusters; Step 7-3: put the query results that meet the requirements into the k-nearest neighbor micro-cluster result set. When the number of query results in the k-nearest neighbor micro-cluster result set reaches k, select the query result with the largest distance to the query vector as the upper limit of the distance, and update the upper limit R of the distance between the k-nearest neighbor micro-cluster result and the query vector.

8. The high-dimensional k-nearest-neighbor micro-cluster search method based on its own dimension according to claim 7, characterized in that: The step 7-2 comprises: Perform a range query with the query result itself as the center and the radius r of the micro-cluster as the distance to check whether the number of query results meets the micro-cluster density requirement. During the range search, determine the left and right boundaries of the data block according to step 5, perform a range query on the data blocks within the boundaries, and finally verify the results of the range query in the original space.