A k-nearest neighbor search method based on KD-tree and octree

By using a combined multi-layer index structure of KD tree and octree for data storage and k-nearest neighbor search in simulated flight, the problem of low data search efficiency and accuracy in real-time monitoring of fuel capacity status in the prior art is solved, and efficient and accurate data search results are achieved.

CN116342793BActive Publication Date: 2025-06-27XIAN AVIATION COMPUTING TECH RES INST OF AVIATION IND CORP OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211612374.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-15
Publication Date
2025-06-27
Estimated Expiration
2042-12-15

AI Technical Summary

Technical Problem

When the prior art conducts real-time monitoring of fuel capacity status during simulated flights, it is difficult to efficiently search k nearest neighbors of specified data, resulting in low search accuracy and low efficiency.

Method used

Using a combined multi-layer index structure based on KD tree and octree, data is stored and k-nearest neighbor searched. Through the binary strategy of KD tree and the spatial division of octree, subspace containing points to be searched can be quickly positioned to improve search efficiency and accuracy.

Benefits of technology

It achieves higher search efficiency and accuracy, and can quickly and accurately search k neighbor data at specified points, significantly saving search time and improving the accuracy of data interpolation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116342793B_ABST
    Figure CN116342793B_ABST
Patent Text Reader

Abstract

The present invention provides a k-nearest neighbor search method based on a KD tree and an octree, including: S1. Import data; S2. Construct a KD tree: Analyze the spatial distribution characteristics of the data, and establish an index for the data according to the partitioning criterion by establishing the KD tree and octree structures; S3. Construct an octree: Calculate the median of the specified point in three dimensions, and use the median as the partitioning plane to further divide the data space into 8 subspaces, and repeat the partitioning until the amount of data in the subspace is less than 8; S4. k-nearest neighbor search: Establish a priority queue with a length of k, quickly locate the subspace where the data to be searched is located according to the spatial partitioning information stored in the root node, use the specified point of the data to be searched as the center, and use the distance to the k-th nearest point in the subspace as the radius, compare the distance between the data to be searched and the partitioning plane to search for the intersecting subspace, and complete the search for the k-nearest neighbor data of the specified point through backtracking search. The present invention can quickly and accurately search for the k-nearest neighbor data of the specified point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of airborne software data processing, and particularly relates to a k-nearest neighbor search method based on KD-tree and octree. Background Art

[0002] The fuel capacity in the aircraft fuel tank is an important information to ensure flight safety. With the development of hardware technology, the fuel capacity status during the simulated flight process can be monitored in real time. A large amount of data will be generated in this process. In this context, there is a problem of k-nearest neighbor data search for specified data, and the accuracy of the search directly affects the later data interpolation results.

[0003] The existing method is to simply split the data table to narrow the data range, and then perform data screening and distance calculation. This method is time-consuming and laborious, does not consider the spatial distribution relationship of multi-dimensional data, the search result does not have global nature, and the accuracy is low. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a k-nearest neighbor search method based on KD-tree and octree. This method adopts a multi-layer index structure combining KD-tree and octree to store data, which can make full use of the advantages of the search tree and has higher search efficiency and accuracy.

[0005] The embodiments of the present application provide the following technical solutions: A k-nearest neighbor search method based on KD-tree and octree, including the following steps:

[0006] S1. Import data: Organize and import the data set, and store the data in the data set into an array;

[0007] S2. Construct KD-tree: Analyze the spatial distribution characteristics of the data, and establish an index for the data according to the partitioning criterion to establish the KD-tree and octree structures. The specific partitioning criterion is: for the specified point (h, α, β), to ensure that the established KD-tree is a balanced binary tree, calculate the variance of each dimension to judge the partitioning dimension. When the variance D[h] of the three dimensions h, α, β is greater than the variance D[α] or the variance D[h] is greater than the variance D[β], perform multiple partitions on the data in the h dimension to establish the KD-tree structure;

[0008] S3. Construct octree: Continue to calculate the median of the three dimensions in each subspace obtained in step S2, use the median of the three dimensions as the octree partitioning plane to further divide the data space into 8 subspaces, and repeat the partitioning until the data volume of the subspace is less than 8, and store the space partitioning information in the root node;

[0009] S4. Perform k-nearest neighbor search using a tree structure: Establish a priority queue of length k. Based on the space partitioning information stored in the root node, quickly locate the subspace where the data to be searched is located. With the specified point (h, α, β) of the data to be searched as the center and the distance to the k-th nearest point in the subspace as the radius, search for the intersecting subspace according to the distance between the data to be searched and the partitioning plane. Calculate the distances between the data in the intersecting subspace and the data to be searched. Through backtracking search, continuously update the search queue until there are no new intersecting subspaces, thus completing the k-nearest neighbor data search for the specified point (h, α, β).

[0010] According to an embodiment of the present application, in step S1, before importing the data, first establish an array arrTree[m][n] for storing the entire data set; where m represents the number of layers and n represents the layer index, and then store the data in the data set into this array.

[0011] According to an embodiment of the present application, in step S2, during the process of constructing the KD tree, calculate the median of the data in the h dimension As the partitioning root node, perform multiple partitions on the data space in the h dimension; where, compare the size of h i , α i , β i in the data (h i and the partitioning plane . If partition the data (h i , α i , β i ) to the left subtree of the current root node, otherwise partition the data to the right subtree. Then calculate the median partitioning planes of the left and right subtrees respectively, compare the size of the data and the partitioning plane, and partition the left and right subtrees, and so on to construct the KD tree.

[0012] According to an embodiment of the present application, the calculation criterion for the median is: First, sort the data in ascending order in the h dimension. If the number of data is even, the median is the n / 2-th number. If it is odd, the median is the (n + 1) / 2-th number.

[0013] According to an embodiment of the present application, the calculation formula for the variance D[h] is:[[]]

[0014]

[0015] where n refers to the number of data in the current h dimension, (h1, h2, h3…h n ) is the data in the h dimension, is the average value in the h dimension, and the average value calculation formula is:[[]]

[0016]

[0017] 6. The k-nearest neighbor search method based on KD-tree and octree according to claim 1, wherein in step S4, with the specified point (h, α, β) of the data to be searched as the center, and the distance d to the k-th nearest point to the subspace k as the radius, search for the intersecting subspace according to the distance d' between the data to be searched and the partitioning plane. If d k > d', it indicates that the data to be searched intersects with the current adjacent subspace.

[0018] According to an embodiment of the present application, in step S4, the distance between the point (h i , α i , β i ) in the subspace and the specified point (h, α, β) adopts the Euclidean distance, and the calculation formula is:

[0019]

[0020] The present invention fully considers the spatial distribution characteristics of the existing data set, stores data and performs k-nearest neighbor search by constructing a combined multi-layer index structure of KD-tree and octree. By establishing an index structure between data, it can quickly locate the subspace containing the point to be searched. The KD-tree adopts a binary strategy, and the time complexity is about log(nlogn), which greatly improves the search efficiency. At the same time, through the spatial partitioning information, recursively search for the spatial nearest neighbor points to avoid falling into local optimality and improve the search accuracy. The present invention has higher search efficiency and accuracy and is an efficient spatial index method. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a flow chart of the construction of KD-tree and octree in the embodiment of the present invention;

[0023] Figure 2 It is a schematic diagram of k-nearest neighbor search for three-dimensional data in the embodiment of the present invention;

[0024] Figure 3 It is a schematic diagram of the combined structure of KD-tree and octree in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The following will describe the embodiments of the present application in detail with reference to the drawings.

[0026] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will describe the present invention in detail with reference to the accompanying drawings and in combination with the embodiments, and clearly and completely describe the technical solutions of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0027] The present invention aims at the data search problem existing in airborne electromechanical software, specifically:

[0028] The process of interpolation calculation output for the known database input conditions (h, α, β). For a given set of (h, α, β) input data, 8 sets of adjacent data can be searched in the database, which are respectively (h i , α i , β i ), (h i+1 , α i , β i ), (h j , α i+1 , β i ), (h j+1 , α i+1 , β i ), (h n , α i , β i+1 ), (h n+1 , α i , β i+1 ), (h k , α i+1 , β i+1 ), (h k , α i+1 , β i+1 ), which satisfy the conditions h i ≤ h ≤ h i+1 , h j ≤ h ≤ h j+1 , h k ≤ h ≤ h k+1 , h n ≤ h ≤ h n+1 , α i ≤ α ≤ α i+1 , β i ≤ β ≤ β i+1 . The data distribution is as Figure 2 shown.

[0029] In the database, perform a k-nearest neighbor (k = 8) data search for the specified point (h, α, β). By using the method of the present invention to construct a tree structure for storage, fully considering the spatial distribution characteristics of the data, a combined structure of KD-tree and octree is used to store the data, and an index structure is established. The k-nearest neighbor data of the specified point can be quickly and accurately searched, effectively saving the search time, improving the search accuracy, and significantly enhancing the search efficiency. This method can also be used for similar platforms. The technical solution of the present invention is as follows.

[0030] Combined with the data distribution characteristics, the present invention uses a multi-layer index structure combining KD-tree and octree for data storage, and the structure is as Figure 3 shown.

[0031] S1. Import data: Organize and import the data set, and store the data in the data set into an array;

[0032] S2. Construct a KD-tree: Analyze the spatial distribution characteristics of the data, and establish a KD-tree and an octree structure to index the data according to the division criterion. The specific division criterion is: for the specified point (h, α, β), to ensure that the constructed KD-tree is a balanced binary tree, calculate the variance of each dimension to determine the division dimension. When the variance D[h] of the three dimensions h, α, β is greater than the variance D[α] or the variance D[h] is greater than the variance D[β], perform multiple divisions on the data in the h dimension to establish a KD-tree structure.

[0033] A KD-tree is a structure that generalizes a binary search tree to multi-dimensional space data, realizing the organization and storage of multi-dimensional space data. The KD-tree uses hyperplanes to divide a space into multiple non-overlapping subspaces. Each layer divides the contained space into two subspaces. The top-level node is divided according to one dimension, and the next layer of nodes is divided according to another dimension. The attributes of all dimensions of the KD-tree cycle between layers.

[0034] By analyzing the data characteristics, h ∈ (1500, 2500), α ∈ (-10, 15), β ∈ (-10, 10), the value intervals of α and β are 1 or 2, the value interval of h is between 8 and 30, and for each (α, β), a series of h ∈ (1500, 2500) corresponds on the h dimension. To ensure the balance of the KD-tree, compare the variances of the data in each dimension, and select the dimension with the largest variance as the division dimension of the data space. According to the spatial distribution characteristics of the data, it is necessary to perform multiple spatial divisions on the data in the h dimension in the early stage. After multiple divisions, the variances of each dimension of the subspace gradually approach until the variance of the h dimension is less than or equal to the variance of the α or β dimension. At this time, construct an octree for each subspace, and set a pointer value at the leaf node of the KD-tree to point to the associated octree.

[0035] S3. Construct KD - tree: Analyze the spatial distribution characteristics of the data, and establish KD - tree and octree structures to index the data according to the partitioning criteria. The specific partitioning criteria are as follows: For a specified point (h, α, β), to ensure that the constructed KD - tree is a balanced binary tree, calculate the variance of each dimension to determine the partitioning dimension. When the variance D[h] of the h - dimension is greater than the variance D[α] or the variance D[h] is greater than the variance D[β], perform multiple partitions on the data in the h - dimension to establish the KD - tree structure.

[0036] An octree is a tree - shaped data structure used to describe three - dimensional space. Use the tree - shaped structure to recursively divide the model, take the median of three different dimensions (h, α, β) as the partitioning hyperplane, and divide the three - dimensional space into 8 sub - spaces. Then, according to the number of targets contained in each sub - space, determine whether to continue to divide the sub - space into 8 equal parts until the set partitioning parameter is reached. By setting the partitioning parameter, the depth of the octree can be controlled, reducing the storage space overhead. Thus, the tree - shaped structure is established.

[0037] S4. Use the tree - shaped structure for k - nearest neighbor search: Establish a priority queue of length k. According to the spatial partitioning information stored in the root node, quickly locate the sub - space where the data to be searched is located. Taking the specified point (h, α, β) of the data to be searched as the center and the distance to the k - th nearest point in the sub - space as the radius, search for the intersecting sub - spaces according to the distance between the data to be searched and the partitioning plane, calculate the distance between the data in the intersecting sub - spaces and the data to be searched, and continuously update the search queue through backtracking search until there are no new intersecting sub - spaces, completing the k - nearest neighbor data search for the specified point (h, α, β).

[0038] Finally, use the tree - shaped structure for nearest neighbor search. To search for the first 8 points closest to (h, α, β), a priority queue of length 8 needs to be established. Taking (h, α, β) as the center, record the distance of the 8 - th closest point to (h, α, β) in the queue. By comparing the distance between (h, α, β) and the current root - node partitioning plane, determine whether there are intersecting sub - spaces, and continuously update the distance through backtracking search. Continuously update the priority queue during the backtracking process until the 8 closest points are found.

[0039] In a specific embodiment, as Figure 1 shown, this embodiment provides a k - nearest neighbor search method based on KD - tree and octree, including the following steps:

[0040] 1. Initialization: Establish an array arrTree[m][n] to store the entire database, where m represents the number of layers and n represents the layer index. All elements are cleared and used to store index information;

[0041] 2. Import database data: Import the data in the entire database into memory;

[0042] 3. Construct KD tree: According to the spatial distribution characteristics of data, in order to ensure the balance of binary tree, calculate the median of h-dimensional data As the partition root node, the data space is partitioned multiple times in the h dimension, and the space partition information is stored in the root node. i ,α i ,β i ) i and partition plane The size of The data (h i ,α i ,β i ) is divided into the left subtree of the current root node, otherwise the data is divided into the right subtree, and then the median partition planes of the left and right subtrees are calculated respectively, and the size of the data and the partition plane are compared, and the left and right subtrees are divided, and the KD tree is constructed in this way. The median calculation rule is: first sort the data from small to large in the h dimension, if the number of data is even, the median is the n / 2th number, if it is odd, the median is the (n+1) / 2th number;

[0043] 4. Calculate the variance of each dimension to determine the division dimension: After each division, calculate the variance D[h], D[α], D[β] of the three dimensions h, α, and β. The calculation formula for the variance D[h] is:

[0044]

[0045] Among them, n refers to the number of data in the current h dimension, (h1,h2,h3…h n ) is the data of h dimension, is the average of the h dimension, and the average calculation formula is:

[0046]

[0047] The calculation of D[α] and D[β] is similar. When D[h]>D[α] or D[h]>D[β], continue to divide the space in the h dimension and repeat step 3 multiple times. Otherwise, set the pointer value in the leaf node of the KD tree and proceed to step 5 to point to the associated octree;

[0048] 5. Construct octree: Continue to calculate the median of each dimension in the subspace obtained in step 4, taking the median of three dimensions As the octree divides the plane, the points in each subspace satisfy the relationship: Construct the octree in this way;

[0049] 6. Calculate the data volume of the subspace: It is stipulated that the division stops when the data volume in each octree leaf node is less than 8, and the space division information is stored in the root node. Otherwise, repeat steps 4 and 5 multiple times;

[0050] 7. k-nearest neighbor search: Establish a priority queue with a length of k. According to the space division information of the root node, quickly locate the subspace where the data to be searched is located. Taking the data to be searched as the center, the distance d to the k-th nearest point in the subspace k is used as the radius. Trace back to the root node of the current subspace to obtain the space division information, and compare the distance d′ between the data to be searched and the division plane to search for the intersecting subspace. If d k > d′, it means that it intersects with the current adjacent subspace. Calculate the distance between the data in the intersecting subspace and the data to be searched, and continuously update the priority queue until there are no new intersecting subspaces. The distance between a point (h i , α i , β i ) and the specified point (h, α, β) adopts the Euclidean distance, and the calculation formula is:

[0051]

[0052] As mentioned above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A k-nearest neighbor search method based on KD-tree and octree, characterized in that, It includes the following steps: S1. Import data: Organize and import the data set, and store the data in the data set into an array; S2. Construct a KD tree: Analyze the spatial distribution characteristics of the data, and establish a KD tree and an octree structure to index the data according to the division criterion. The specific division criterion is: for the specified point (h, α, β), to ensure that the constructed KD tree is a balanced binary tree, calculate the variance of each dimension to determine the division dimension. When the variance D[h] of the three dimensions h, α, β is greater than the variance D[α] or the variance D[h] is greater than the variance D[β], divide the data multiple times in the h dimension to establish a KD tree structure; S3. Construct an octree: Continue to calculate the median of the three dimensions of each subspace in the subspace obtained in step S2, and use the median of the three dimensions as the octree division plane to further divide the data space into 8 subspaces. Repeat the division until the number of data in the subspace is less than 8, and store the space division information in the root node; S4. Use the tree structure for k-nearest neighbor search: Establish a priority queue with a length of k. According to the space division information stored in the root node, quickly locate the subspace where the data to be searched is located. With the specified point (h, α, β) of the data to be searched as the center and the distance to the kth nearest point in the subspace as the radius, search for the intersecting subspaces according to the distance between the data to be searched and the division plane, calculate the distance between the data in the intersecting subspaces and the data to be searched, and continuously update the search queue through backtracking search until there are no new intersecting subspaces, and complete the k-nearest neighbor data search for the specified point (h, α, β).

2. The k-nearest neighbor search method based on KD-tree and octree according to claim 1, wherein In step S1, an array arrTree[m][n] for storing the entire data set is established before importing the data; where m represents the number of layers and n represents the layer index, and then the data in the data set is stored in this array.

3. The k-nearest neighbor search method based on KD-tree and octree according to claim 2, wherein In step S2, during the process of constructing the KD tree, calculate the median of the data in the h dimension As the dividing root node, the data space is divided multiple times in the h dimension; among them, compare the size of h i , α i , β i ) in h i and the median . If , divide the data (h i , α i , β i ) into the left subtree of the current root node, otherwise divide the data into the right subtree. Then calculate the median dividing planes of the left and right subtrees respectively, compare the size of the data and the dividing planes, and divide the left and right subtrees. By analogy, construct the KD tree.

4. The k-nearest neighbor search method based on KD-tree and octree according to claim 3, characterized in that, The calculation criterion for the median is: First, sort the data in the h dimension from small to large. If the number of data is even, the median is the n / 2th number. If it is odd, the median is the (n + 1) / 2th number.

5. The k-nearest neighbor search method based on KD-tree and octree according to claim 3, characterized in that, The calculation formula for the variance D[h] is: Among them, n refers to the number of data in the current h dimension, and (h1, h2, h3…h n ) is the data in the h dimension, is the average value in the h dimension, and the average value calculation formula is:

6. The k-nearest neighbor search method based on KD-tree and octree according to claim 1, characterized in that, In step S4, with the specified point (h, α, β) of the data to be searched as the center and the distance d to the k-th closest point to the subspace as the radius, search for the intersecting subspace according to the distance d between the data to be searched and the partitioning plane. If d k > d ′ , it indicates that the data to be searched intersects with the current adjacent subspace. k > d ′ ​ 7. The k-nearest neighbor search method based on KD-tree and octree according to claim 6, wherein, In step S4, the distance between the point (h i , α i , β i ) in the subspace and the specified point (h, α, β) is the Euclidean distance, and the calculation formula is as follows:

Citation Information

Patent Citations

  • Hybrid tree parallel construction method based on GPU

    CN104463940A

  • A massive point cloud spatial management method based on octree-like encoding

    CN109345619A