Arbitrary dimension vector set multilevel load balance space convergence structure construction method

By building a multi-level load balancing spatial convergence structure full of binary tree index trees, the storage and management problems of large-scale vector data in any dimension are solved, and efficient data management and rapid retrieval effects are achieved.

CN120216729APending Publication Date: 2025-06-27KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510313473.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively manage and store large-scale vector data in any dimension, especially in terms of spatial complexity and data fragmented storage.

Method used

The construction method of a multi-level load-balancing spatial convergence structure is adopted. Through the construction of a full binary tree index tree, the depth and fan-out degree of the tree structure are optimized to realize data balance management and multi-resolution hierarchy construction.

Benefits of technology

It realizes efficient storage and management of large-scale vector data in any dimension, eliminates the cost of addressing, optimizes space utilization, and supports fast retrieval and multi-resolution rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216729A_ABST
    Figure CN120216729A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method for a multi-level load balancing space convergence structure of an arbitrary-dimension vector set. The construction method comprises the steps of initialization of control parameters, parameter optimization, application of a storage space, calculation of node positions, storage of original data, recursive division of a data set, construction of a full binary index tree and optimization of an index tree structure. According to the method, balance management and multi-resolution hierarchical construction are carried out on a comparable vector data set of any dimension, the depth of a tree structure can be optimized under the condition that the fan-out degree of the tree structure is reasonably controlled, and a full binary tree index tree is constructed to weaken or eliminate the addressing cost; for a data set, partition retrieval can be completed by adopting a single complete data file in an internal sorting mode or fragmentation organization management of data is realized in a file directory hierarchy mode, so that compact utilization of a storage space and rapid retrieval of data with different granularities are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and particularly relates to a method for constructing a multi-level load balancing space convergence structure for a vector set of any dimension. Background Art

[0002] With the rapid development of software and hardware technologies, the amount of data obtained shows an explosive growth trend, and point cloud data is a typical representative of such data. For the storage, retrieval, and analysis of large-scale data, high requirements are imposed on computing resources. In order to fully and efficiently use data structures such as large-scale point cloud data that require high efficiency for management. Point cloud data usually contains attribute information such as coordinates and colors, and various more general attributes are usually attached in practical applications. Therefore, storing and managing large-scale vector data of any dimension is a very important task.

[0003] In classical spatial data structures, quadtrees are more suitable for two-dimensional cases, and octrees are suitable for three-dimensional cases. Using such structures will result in different data densities in the divided subspaces due to the uneven distribution of data in space. KD-trees have a high retrieval efficiency in the management of point data, especially multi-dimensional point data. Overmars and van Leeuwen suggested that data be retained in leaf nodes (Overmars, M. H.,&van Leeuwen, J. (1982). Dynamicmulti-dimensional data structures based on quad- and k-d trees. Acta Informatica, 17 (3), 267–285.). Existing KD-tree structures mostly focus on the time complexity during construction and retrieval, and less discussion is made on the spatial complexity level. Moreover, retaining data in leaf nodes will lead to fragmented storage of data, and it is difficult to make good use of the buffer characteristics of modern hardware during regional searches. For large-scale data of any dimension, if divided into K-ary trees according to dimensions, it will lead to a rapid increase in the fan-out degree. If the fan-out degree is controlled, it will lead to an increase in depth. How to make the structure space for storing data converge quickly is a problem.

[0004] Therefore, it is very necessary to invent a structure that can weaken or eliminate the addressing cost for a large-scale vector data set of any dimension, while also taking into account the uniform distribution of data and rapid convergence. Summary of the Invention

[0005] To solve the above technical problems, the purpose of the present invention is to provide a construction method for a multi-level load balancing space convergence structure of an arbitrary-dimensional vector set, which can perform balanced management and multi-resolution level construction on a comparable vector data set of any dimension, optimize the depth of the tree structure while reasonably controlling the fan-out degree of the tree structure, construct a full binary tree index tree to weaken or eliminate the addressing cost, and the data set can adopt a single complete data file to complete partition retrieval through internal sorting or achieve fragmented organization and management of data through the method of file directory levels, so as to achieve compact utilization of storage space and fast retrieval of different data granularities.

[0006] The purpose of the present invention is achieved as follows, including the following steps: S100. Initialize control parameters: Set a pre-control parameter K, which is used to control the maximum number of points or elements accommodated in the leaf nodes of the finally constructed full binary index tree; among them, the value of K is input by the user, and this value is a theoretical value, and the actual value can be actually solved according to the balance of the data set; Due to the unstructured characteristics of large-scale arbitrary-dimensional vector data sets, and they exist in a discrete form without corresponding topological connection characteristics, it is necessary to specify parameters to control the number of points or elements accommodated in the leaf nodes of the index tree. After this parameter is determined, the number of intermediate nodes and leaf nodes required to construct the index tree can be determined synchronously. Regardless of the dimension size of the points or elements in the data set, the storage space required to construct the index tree structure, the number of nodes (intermediate nodes and leaf nodes) of the tree structure, and the depth are only related to the number of points or elements, and are independent of the dimension of the data set; S200. Parameter optimization: Let the number of points or elements in the data set to be processed be T, calculate T / K to obtain the estimated number of leaf nodes of the full binary index tree. In order to achieve balanced configuration of the full binary tree, it is necessary to round up T / K to the power of 2 N , that is, let T / K = 2 N , then , take the integer part of N to the left to get , for example, if T / K = 5, then the obtained after rounding is 2, and recalculate , and take the integer part of it to the right to get the upper limit of the number of points / elements accommodated in each leaf node; S300. Apply for storage space: Apply for memory space or a memory-mapped file, where the number of node parts is , multiply this number by the size of a single node structure to obtain the actual required storage space. Here, the space does not include file header information; the storage space of the file header can be determined according to the actual situation; S400. Calculate the node position: According to the formula Calculate the starting position of the nodes at the specified layer, where i is the layer number, i≥1, and the number of nodes between the starting positions of adjacent layers is the total number of nodes in the lower layer; for example, the starting position of the nodes in the first layer is 1, the starting position of the nodes in the second layer is 2, and the starting position of the nodes in the third layer is 4; S500. Store the original data: Construct a memory-mapped file for the original data set and copy the original data to this memory-mapped file; S600. Recursive partitioning of the data set and construction of the full binary index tree: Select the direction with the largest dimension in the multi-dimensional vector as the benchmark for sorting, divide the data set into two equal parts in this dimension, and use the corresponding range information to fill the relevant information of the corresponding layer of the index tree; if using compact encoding, keep the original range, otherwise perform spatial compression; repeat the above operations on the two subsets obtained by partitioning in their largest dimensions, continue sorting, evenly cutting, and filling the information of the next layer of the index tree; perform iterative calculations until the number of vectors in the subspace is less than or equal to the value (the upper limit of the number of points / elements that each leaf node can accommodate calculated in S200), thereby completing the construction of the entire full binary index structure; S700. Index tree structure optimization: If the original data set is stored as a single file, when constructing the index tree, sort the subspaces according to the recursive partitioning, and a continuous storage structure of the data is automatically formed inside the file; if the data set of each node is stored as an independent file, organize it into a file directory structure according to the levels of the index tree.

[0007] The corresponding LOD information and the addressing method of the corresponding data page are encoded in the final index tree; in the data file, the sorting method of the data is recorded in the form of a file header, which can quickly obtain the data in the corresponding area and achieve fast addressing; for the areas corresponding to the space partitioning, the data in the data pages is continuously stored in units of areas, and this continuous storage is independent of the resolution.

[0008] Preferably, in step S600 when constructing the full binary index tree, the full binary tree represents the space where the data set is located. Starting from the top, sequentially select the direction with the largest dimension in the data set, sort the data in this direction, and evenly divide it into two parts from the middle as the two child nodes of the full binary tree structure; for the two child nodes, respectively select the direction with the largest dimension in their sub-data sets for even division to obtain 4 child nodes, and so on, recursively divide until the number of points or elements in the leaf nodes is less than or equal to the upper limit of the number of points or elements that each leaf node can accommodate calculated in S200; enabling the structure to converge quickly.

[0009] Preferably, in step S600, the space scale of the sub-dataset included in the full binary index tree node is stored, and the space scale is compact or non-compact; among them, the formation process of the compact space is: when recursively partitioning the dataset, after evenly cutting the data into two parts, the coordinate range of the midpoint is used as the boundary in the two sub-spaces and saved in the node to form a compact space; the formation process of the non-compact space is: if for the two sub-spaces cut, one retains the space range of the boundary points, and the other is collapsed, that is, it is encoded with its actual coordinate range to form a non-compact space; these two strategies can be controlled according to specific applications.

[0010] Preferably, step S700 further includes calculating statistical parameters. Specifically, the statistical parameters of the corresponding sub-space are stored in the full binary index tree node. There are many types of parameter calculations. Calculating the simple average of the sub-space dataset is used for LOD calculation; calculating the expectation and variance is used to detect outlier data; calculating the distribution density of the point set or elements in the sub-space is used for calculating the uniformity of data acquisition and distribution uniformity; randomly and discretely sampling the dataset in the sub-space at a lower ratio to establish an effective LOD expression mechanism.

[0011] Preferably, step S700 further includes optimizing the addressing method. When the index tree is constructed, one recursive cut will cause the space to be evenly cut into two parts along the long dimension direction until the leaf nodes, and the space is evenly cut. According to the characteristics of uniform cutting, the corresponding memory version applies for the corresponding space by calculating the number of nodes, and the position of the nodes at the specified layer can be obtained through simple calculation, thus avoiding defining corresponding pointers to point to the corresponding nodes. This operation can save a large amount of addressing space and eliminate the addressing cost; the corresponding file version is processed in the same way; optimizing the addressing converts the full binary index into a linear array method, and the addressing cost is eliminated through direct calculation.

[0012] Compared with the prior art, the present invention has the following technical effects: 1. The present invention makes full use of the characteristics of the full binary tree, eliminates the fields required for addressing, and greatly reduces the consumption of storage resources by the addressing part; 2. The space division converges quickly. Taking the dataset as the division object and partitioning the data, the space where it is located converges quickly; 3. It can be used to uniformly process comparable vector datasets of any dimension. The structure of the index tree only relates to the number of points / elements and is independent of the dimension of the dataset; 4. Use a unified method to achieve spatial division and processing of data of any dimension; 5. The space after the index tree is constructed can be compact or non-compact. Compact partitioning is a continuous modeling expression for the dataset; 6. The dataset has a built-in sorting mode, ensuring data continuity in the unified file mode. When retrieving data in a specified subspace, buffer switching can be reduced, accelerating data retrieval. 7. In the file storage mode, data can be discretely stored in a tree structure, making the LOD fast scheduling and rendering more efficient. 8. The data corresponding to different resolution levels is continuously stored. For any region at any resolution level in the index tree, the continuously stored dataset can be obtained. This mechanism can effectively utilize the buffer mechanism of modern computing systems to achieve fast retrieval and loading processing of the dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a flowchart of the method of the present invention; Figure 2 It is a schematic structural diagram of a four-layer full binary index tree; Figure 3 It is a schematic diagram of the recursive partitioning and index tree construction of a two-dimensional point cloud dataset; Figure 4 It is a schematic diagram of the sequential storage of the index tree; Figure 5 It is a schematic diagram of constructing a file directory according to the node level; Figure 6 It is a schematic diagram of the recursive partitioning and index tree construction of a three-dimensional point cloud dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The present invention will be further described below in conjunction with the embodiments and the drawings, but the present invention is not limited in any way. Any transformation or replacement based on the teachings of the present invention falls within the protection scope of the present invention.

[0015] Embodiment 1 This embodiment is based on a two-dimensional large-scale vector dataset to illustrate how to achieve efficient management and retrieval of the dataset through a multi-level load balancing space convergence structure; As shown in the attached Figures 1 to 5 figure, the construction method of the multi-level load balancing space convergence structure for the two-dimensional large-scale vector dataset in this embodiment includes the following steps: S100. Initialize the control parameters: Set the pre-control parameter K = 2, which is used to control the maximum number of data points accommodated in the leaf nodes of the finally constructed full binary index tree; S200. Parameter optimization: Assume that the number of points in the dataset to be processed is T = 20. Calculate the estimated number of leaf nodes L = T / K according to the total number of points T in the dataset to obtain the estimated number of leaf nodes L of the full binary index tree. In order to achieve the balanced configuration of the full binary tree, let T / K = 2 N , then , and take the integer part of N to the left to get is 3, calculate , and take rounding up to the right to obtain the upper limit of the number of points or elements that each leaf node can accommodate, which is 3; as Figure 2 shown, where each leaf node contains at most 3 elements; S300, Apply for storage space: According to the number of nodes and the structure size of the full binary index tree, allocate the corresponding memory or memory-mapped file space, and store the node data, including: subspace range information, statistical parameters (such as mean, variance), and data page pointers; construct a memory-mapped file for the original data set, and copy the original data into this memory-mapped file; S400, Calculate node positions: Calculate the starting position of the nodes at the specified layer according to the formula , where i is the layer number, i ≥ 1. For example, the starting position of the nodes in the first layer is 1, the starting position of the nodes in the second layer is 2, and the starting position of the nodes in the third layer is 4. The number of nodes in the middle of the starting positions of adjacent layers is the total number of nodes in the lower layer; S500, Store the original data: Construct a memory-mapped file for the original data set, and copy the original data into this memory-mapped file; S600, Recursive partitioning of the data set and construction of the full binary index tree: As Figure 3 shown, starting from the root node, select the direction with the largest length in the two-dimensional data set (such as the X or Y axis), sort according to this dimension and evenly cut it into two parts, which are used as the data sets of the child nodes respectively, and record the range of the subspace corresponding to the nodes; the left data set is {(X 11 , Y 11 ), (X 12 , Y 12 ), …(X 1n , Y 1n ), and the right data set is {(X 21 , Y 21 ), (X 22 , Y 22 ), …(X 2n , Y 2n )}; for compact encoding, retain the true range; for non-compact encoding, appropriately contract the range of the secondary dimension; calculate the mathematical mean, variance, etc. of each node to support subsequent outlier detection and data screening; repeat the above operations on the two subsets divided on their maximum dimension, continue to sort, evenly cut, and fill in the information of the next layer of the index tree; perform iterative calculations until the number of vectors in the subspace is less than or equal to 3, thus completing the construction of the entire full binary index structure; S700, Index tree structure optimization: Achieve address optimization for nodes through linear encoding, eliminate the addressing fields of the traditional index tree, and reduce space consumption; asFigure 4 As shown, for leaf nodes, the original data pages are stored in the order after recursive cutting for direct retrieval; the original data set is stored as a single file, and data partitioning and fast retrieval are achieved through internal sorting; as Figure 5 shown, a file directory is constructed according to the node levels, and the data of each node is stored as an independent file to support dynamic loading and efficient scheduling; In this embodiment, the retrieval efficiency of the two-dimensional data set and the utilization rate of storage resources are significantly improved, and the efficient management of large-scale vector data sets is realized.

[0016] Embodiment 2 This embodiment is based on a three-dimensional large-scale vector data set and illustrates how to achieve the efficient management and retrieval of the data set through a multi-level load balancing space convergence structure; As shown in the appendix Figure 1 、 Figure 2 and Figure 6 shown, this embodiment is a method for constructing a multi-level load balancing space convergence structure for a three-dimensional large-scale vector data set, including the following steps: S100. Initialize control parameters: Set the pre-control parameter K = 2, which is used to control the maximum number of points or elements accommodated in the leaf nodes of the finally constructed full binary index tree; S200. Parameter optimization: Let the number of points or elements in the data set to be processed be T = 20, calculate L = T / K, and obtain the estimated number of leaf nodes L of the full binary index tree as 10. Let T / K = 2 N , then , round N down to get , calculate , and round it up to get the upper limit of the number of points or elements accommodated in each leaf node = 3; as Figure 2 shown, where each leaf node contains at most 3 elements; S300. Apply for storage space: Apply for the memory space or memory-mapped file of the full binary index tree nodes. The node information includes: subspace range (ranges in the three dimensions of X, Y, and Z), statistical parameters (such as mathematical mean, variance, point density), data page pointer or file path, where the number of node parts is , multiply this number by the size of a single node structure to obtain the actual required storage space, and this space does not include the file header information; S400. Calculate node positions: Calculate the starting position of the nodes at the specified layer according to the formula , where i is the layer number, i ≥ 1. For example, the starting position of the nodes in the first layer is 1, the starting position of the nodes in the second layer is 2, and the starting position of the nodes in the third layer is 4. The number of nodes in the middle of the starting positions of adjacent layers is the total number of nodes in the lower layer; S500, Store the original data: Construct a memory-mapped file for the original data set and copy the original data into this memory-mapped file; S600, Recursive partitioning of the data set and construction of a full binary index tree: As Figure 6 shown, starting from the root node, select the dimension (X, Y, or Z) with the largest length in the three-dimensional data set, sort the data according to this dimension and evenly cut it into two parts, which are used as the data sets of the child nodes respectively, to achieve recursive cutting of the three-dimensional data; for the compact encoding mode, retain the actual range information after cutting; for the non-compact encoding mode, collapse the range of the non-dominant dimension; calculate the mean and variance of the data in the subspace corresponding to each node for outlier detection and optimized screening; record the subspace point density information to support LOD (Level of Detail) optimization; repeat the above operations on the two subsets divided in their largest dimension, continue to sort, evenly cut, and fill the information of the next layer of the index tree; perform iterative calculations until the number of vectors in the subspace is less than or equal to 3, thus completing the construction of the entire full binary index structure; S700, Optimization of the index tree structure: Achieve address optimization for nodes through linear encoding, eliminate the addressing fields of the traditional index tree, and reduce space consumption; for leaf nodes, store the original data pages in the order after recursive cutting for direct retrieval; if the three-dimensional data set is stored as a complete file, when constructing the index tree, sort the subspaces according to the recursive partitioning, and a continuous storage structure of the data is automatically formed inside the file; if the data set of each node is stored as an independent file, it should be organized as a file directory structure according to the levels of the index tree; intermediate nodes store the sampled and simplified data for fast LOD rendering; Through this embodiment, large-scale three-dimensional vector data sets are efficiently organized and managed, the data storage is compact, the depth of the index tree is optimized, the fast retrieval ability of three-dimensional data is significantly improved, and at the same time, multi-resolution loading and dynamic rendering are supported.

[0017] Through the above two embodiments, the present invention can also expand the construction method of a multi-level load-balanced space convergence structure for vector sets of any dimension.

Claims

1. A method for constructing a multi-level load balancing spatial convergence structure for arbitrary dimensional vector sets, characterized in that The following steps are involved: S100, initializing control parameters: setting a pre-control parameter K, which is used to control the maximum number of points or elements contained in a leaf node in the final constructed full binary index tree; S200, parameter optimization: Let the number of points or elements in the data set to be processed be T, calculate T / K, and get the estimated number of leaf nodes of the full binary index tree, let T / K=2 N ,but , round N to the left to get ,calculate ,Will Round right to get the upper limit of the number of points or elements that each leaf node can accommodate; S300, apply for storage space: apply for memory space or memory mapping file, where the number of nodes is , multiply this number by the size of a single node structure to get the actual storage space required, where the space does not include file header information; S400, calculate node position: according to the formula Calculate the starting position of the nodes in the specified layer, where i is the number of layers, i ≥ 1, and the number of nodes between the starting positions of two adjacent layers is the total number of nodes in the lower layer; S500, storing original data: constructing a memory mapping file for the original data set, and copying the original data to the memory mapping file; S600, recursive partitioning of data set and construction of full binary index tree: select the direction with the largest dimension in the multidimensional vector as the basis for sorting, divide the data set into two on this dimension, and use the corresponding range information to fill in the relevant information of the corresponding layer of the index tree; if compact coding is used, keep the original range, otherwise perform space compaction; repeat the above operations on the largest dimension of the two divided subsets, continue sorting, balanced cutting, and fill in the next layer of information of the index tree; iterate the calculation until the number of vectors in the subspace is less than or equal to value, thus completing the construction of the entire full binary index structure; S700, index tree structure optimization: If the original data set is stored as a single file, when building the index tree, the subspaces are sorted according to the recursive partitioning, and a continuous storage structure of the data is automatically formed inside the file; if the data set of each node is stored as an independent file, it is organized into a file directory structure according to the hierarchy of the index tree.

2. The method for constructing a multi-level load balancing spatial convergence structure of an arbitrary dimensional vector set according to claim 1 is characterized in that When constructing a full binary index tree in step S600, the full binary tree represents the space where the data set is located. The direction with the largest dimension in the data set is selected from top to bottom, the data is sorted in this direction, and it is evenly divided into two parts from the middle as two child nodes of the full binary tree structure; the two child nodes are evenly divided in the direction with the largest dimension in their child data sets to obtain 4 child nodes, and so on, recursively dividing until the number of points or elements in the leaf node is less than or equal to the upper limit of the number of points or elements that each leaf node can accommodate calculated in S200.

3. The method for constructing a multi-level load balancing spatial convergence structure of an arbitrary dimensional vector set according to claim 2 is characterized in that In step S600, the full binary index tree node stores the spatial scale of the sub-dataset contained in the node, and the spatial scale is compact or non-compact; wherein, the formation process of the compact space is: when the data set is recursively divided, the data is evenly cut into two parts, and then the coordinate range of the middle point in the two subspaces is used as the boundary and saved in the node to form a compact space; the formation process of the non-compact space is: if for the two subspaces cut, one retains the spatial range of the boundary point, and the other is collapsed, that is, encoded with its actual coordinate range, to form a non-compact space.

4. The method for constructing a multi-level load balancing spatial convergence structure of an arbitrary dimensional vector set according to claim 1 is characterized in that Step S700 also includes calculating statistical parameters, specifically, storing the statistical parameters of the corresponding subspace in the full binary index tree node, calculating the simple average of the subspace data set for LOD calculation; calculating the expectation and variance for detecting outlier data; calculating the distribution density of the point set or element in the subspace for data collection uniformity and distribution uniformity calculation; performing random discrete sampling on the data set in the subspace at a lower ratio to establish an LOD expression mechanism.

5. The method for constructing a multi-level load balancing spatial convergence structure of an arbitrary dimensional vector set according to claim 1 is characterized in that Step S700 also includes optimizing the addressing method, specifically: according to the uniform cutting characteristics, the corresponding memory version applies for the corresponding space through the calculated number of nodes, calculates the position of the node of the specified layer, and eliminates the addressing cost; the corresponding file version is processed in the same way.