High-performance distributed spatial algorithm parallelization processing method for forest and grass spatial data

Through adaptive QuadTree division, distributed R-Tree index and Z-Order encoding, the problem of insufficient processing capabilities of forest and grass space data in centralized computing mode is solved, and efficient, stable and consistent distributed data processing is achieved, which is suitable for rapid query and analysis of forest and grass space data.

CN120256111APending Publication Date: 2025-07-04国家林业和草原局中南调查规划院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335331.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When facing big data in forest and grass space, the traditional centralized computing model has limited processing capabilities and storage capabilities, which cannot meet data processing needs, and there is a single point of failure risk, which affects the continuity and stability of the project, making it difficult to maintain data consistency and topology.

Method used

Adaptive QuadTree dynamic division algorithm and distributed R-Tree index are adopted, combined with Z-Order space fill curve coding, adaptive division and load balancing of data blocks, dynamic task mapping and load balancing, efficient storage and querying are used for distributed databases, and data fusion is used for distributed aggregation algorithm.

Benefits of technology

Improves computing throughput, reduces index construction time, ensures data consistency and integrity, realizes high availability computing, and supports rapid processing and query of large-scale forest and grass space data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256111A_ABST
    Figure CN120256111A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of forest and grass space, and particularly discloses a forest and grass space data-oriented high-performance distributed spatial algorithm parallelization processing method, which comprises the steps of data preprocessing, data division and index construction, task scheduling and parallel computing, and result integration and storage and data exchange. According to the method, an adaptive QuadTree dynamic division algorithm is adopted, the scale of data blocks is adjusted based on data density, load balance of calculation tasks is ensured, and waste of calculation node resources is avoided; in combination with a dynamic task mapping algorithm based on a load factor, task adaptive scheduling is realized among computational nodes, the utilization rate of computational resources is maximized, and the computational throughput is improved; a distributed R-Tree index is adopted, index construction is optimized through Bulk-Loading, and the index construction time is greatly shortened; in combination with Z-Order space filling curve coding, space data is efficiently stored and retrieved in a distributed database (such as HBase and Cassandra), and rapid space query is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of forest and grass space, and particularly relates to a high-performance distributed spatial algorithm parallel processing method for forest and grass space data. Background Art

[0002] With the rapid development of information technology, the forest and grass field has gradually entered the big data era. In many projects such as forest and grass resource monitoring, ecological environment assessment, forest fire early warning and prevention, the amount of data involved has shown an explosive growth.

[0003] When facing such a large amount of forest and grass space data, the traditional centralized computing mode seems powerless. On the one hand, the processing capacity of centralized computing is limited, and it is difficult to efficiently analyze and process large-scale data in a short time. On the other hand, the storage capacity of centralized computing also faces challenges and cannot meet the growing data storage requirements. Moreover, once the central computing node fails, the entire data processing process will be paralyzed, seriously affecting the continuity and stability of the project.

[0004] In addition, forest and grass space data has significant complexity and geographical relevance. The forest and grass ecosystems in different regions interact with each other, and there are complex spatial relationships and logical connections between the data. This requires that in the data processing process, not only the storage and calculation efficiency of the data should be considered, but also the consistency and topology of the data should be maintained.

[0005] The emergence of the distributed spatial computing technology for forest and grass space data provides new ideas and methods to solve these problems. It can disperse and store large-scale forest and grass space data on multiple computing nodes, and through parallel computing, make full use of the computing resources of each node, greatly improving the efficiency and speed of data processing. At the same time, the distributed computing system has strong fault tolerance and scalability. Even if some nodes fail, it will not affect the normal operation of the entire system, and new nodes can be conveniently added to meet the growing data processing requirements. Therefore, in forest and grass big data projects, adopting the distributed spatial computing technology for forest and grass space data has become an inevitable trend, which is of great significance for promoting scientific research, resource management and ecological protection in the forest and grass field.

[0006] In view of this, the inventor proposes a high-performance distributed spatial algorithm parallel processing method for forest and grass space data to solve the above problems. Summary of the Invention

[0007] The purpose of the present invention is to provide a high-performance distributed spatial algorithm parallel processing method for forest and grass space data to solve the problems raised in the above background art.

[0008] To achieve the above purpose, the present invention provides the following technical solutions:

[0009] A high-performance distributed spatial algorithm parallel processing method for forestry and grassland spatial data, including:

[0010] Adopt a parallel data reading module to read the original forestry and grassland spatial data from multiple data sources to obtain preliminary data slices; use a data cleaning algorithm to standardize the preliminary data slices to obtain standardized data;

[0011] Adopt an adaptive spatial partitioning algorithm to partition the standardized data to obtain partitioned data blocks; adopt a distributed spatial indexing algorithm to locate and locally query the spatial data in the partitioned data blocks to obtain an index structure;

[0012] According to the partitioned data blocks and the index structure, use task subdivision, data redundancy, and local combination strategies to assemble local computing units; use a dynamic task mapping and load balancing method to allocate the local computing units to multiple computing nodes to perform spatial calculations in parallel to obtain preliminary calculation results;

[0013] Perform coding identification processing on the preliminary calculation results using a unified spatial coding method to obtain encoded data; use a distributed aggregation algorithm to perform boundary correction and data fusion processing on the encoded data to obtain globally consistent and standardized final calculation results.

[0014] Preferably, the adaptive spatial partitioning algorithm is used to adaptively partition the standardized data into multiple data blocks according to the density and distribution characteristics of the forestry and grassland spatial data. The expression of the adaptive spatial partitioning algorithm is:

[0015]

[0016] Among them, C is a regional constant used to reflect the uniformity of data distribution;

[0017] M is the available memory size of a single node;

[0018] K is an empirical constant used to balance the partitioning granularity and system load.

[0019] Preferably, the calculation steps of the distributed spatial indexing algorithm are as follows:

[0020] First, for the data points in each of the partitioned data blocks, calculate the minimum bounding rectangle. The formula is:

[0021]

[0022] The maximum node capacity M_max within each node;

[0023] The number of sorting and grouping S in the STR algorithm depends on the size of the data block;

[0024] Then, using the Bulk-Loading method, the data is sorted according to the spatial position and a hierarchical tree structure is constructed;

[0025] Finally, in a distributed environment, a global index is formed among nodes through a local index merging mechanism or a multi-level index structure is used to achieve fast querying and updating of the index structure.

[0026] Preferably, the dynamic task mapping and load balancing method is used to dynamically allocate the local computing units to each computing node to achieve load balancing and reduce communication latency. The specific steps of the dynamic task mapping and load balancing method are as follows:

[0027] Set the computing power parameter Pi of each node, and at the same time calculate the task load Li of the current node;

[0028] Define the load balancing objective function

[0029] Δ = max(Li / Pi) - min(Li / Pi)

[0030] where Pi is the computing power of the node;

[0031] wj is the task weight, which can be estimated according to the task data volume and computational complexity;

[0032] ∈ is the load balancing tolerance threshold, which is used to determine whether task reallocation is required;

[0033] Minimize Δ through the scheduling algorithm;

[0034] For each task in the local computing unit, it is allocated to a node according to the task weight wj, such that:

[0035] f: E → {nodes}, such that min Δ

[0036] where E is the minimum operation unit of dynamic task allocation, representing the subset of tasks to be scheduled;

[0037] f is the core logic of load balancing, which realizes efficient resource utilization through mathematical optimization and real-time feedback.

[0038] Preferably, the unified spatial coding method uses the Z-order curve to encode the spatial coordinates (x, y) into a one-dimensional value, which is convenient for unified identification and fast indexing of cross-node data, and the encoded data is obtained

[0039] Discretize each spatial point coordinate (x, y) and convert it into an integer representation with a fixed precision, where the precision is p bits;

[0040] Using the bit-interleaving algorithm, the binary bits of x and y are interleaved and combined to obtain the Z-order encoding:

[0041] Z = Interleave(bin(x), bin(y))

[0042] Among them, the Z value is the generated Z-Order encoding value, which is used to uniquely identify the spatial position and achieve fast sorting and regional positioning of data;

[0043] bin(x) and bin(y) respectively represent the binary strings of the discretized x coordinate and y coordinate;

[0044] Interleave() represents the bit-interleaving function, which alternately combines the binary bits of bin(x) and bin(y);

[0045] The precision p determines the resolution of the spatial encoding. The larger the p value, the higher the encoding precision, but the computational complexity also increases accordingly.

[0046] Preferably, the distributed aggregation algorithm adopts a spatial data fusion algorithm based on weighted average of overlapping regions, which is used to perform boundary correction and data fusion processing on the preliminary results calculated by each node to obtain a globally consistent final result. The specific steps are as follows:

[0047] For the overlapping regions existing between adjacent data blocks, the result values A and B of the two data blocks in the overlapping part are calculated respectively;

[0048] Set the overlapping weight factors α and β, where the weights can be determined according to their respective regional coverage rates or data credibility;

[0049] The fusion formula is:

[0050]

[0051] Repeat the above process for all overlapping regions, and finally fuse the results of each node through the distributed aggregation algorithm to obtain a globally consistent and standardized final calculation result;

[0052] Among them, Rfused represents

[0053] The overlapping weight factors α and β can be set according to the overlapping area ratio, data quality or sensor accuracy;

[0054] The weight standard adopted during fusion can be optimized according to the specific application scenario.

[0055] Preferably, the parallelized data reading module adopts a distributed file system and a multi-thread mechanism to simultaneously read data from multiple data sources and obtain preliminary data slices;

[0056] The standardization process includes outlier detection, noise removal, and data format standardization for the preliminary data slices to obtain the standardized data.

[0057] Preferably, the method further includes storing the final calculation results in a unified spatial encoding using a distributed database system, and realizing data exchange with other platforms through a standard data interface to obtain data interoperability and application expansion capabilities.

[0058] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0059] (1) The present invention adopts an adaptive QuadTree dynamic partitioning algorithm to adjust the data block size based on data density, ensuring balanced computing task load and avoiding waste of computing node resources; combined with a dynamic task mapping algorithm based on load factor, it realizes adaptive task scheduling among computing nodes, maximizes the utilization rate of computing resources, and improves computing throughput; adopts a distributed R-Tree index, uses Bulk-Loading to optimize index construction in batches, and significantly reduces the index construction time; combined with Z-Order space filling curve encoding, it efficiently stores and retrieves spatial data in distributed databases such as HBase and Cassandra to achieve fast spatial queries.

[0060] (2) The present invention adopts a fusion algorithm based on weighted average of overlapping regions to optimize and fuse the boundaries of adjacent data blocks, reduce boundary data discontinuity, and improve data integrity; combined with multi-scale hierarchical computing, it can automatically match the best fusion strategy based on different spatial scales to ensure data consistency; adopts a distributed data redundancy storage strategy to ensure that computing tasks can continue to run when some nodes fail through master-slave data backup; combined with a task dynamic migration mechanism, when a computing node fails, the task can be automatically migrated to other idle nodes to achieve highly available computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a flowchart of a high-performance distributed spatial algorithm parallelization processing method for forest and grassland spatial data of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0063] Embodiment 1:

[0064] Please refer toFigure 1 As shown in Figure 1 , a high-performance distributed spatial algorithm parallel processing method for forest and grassland spatial data includes:

[0065] Data preprocessing: Using a parallel data reading module (such as using the distributed file system HDFS and multi-threading mechanism) to read the original forest and grassland spatial data from multiple data sources to obtain preliminary data slices; using a data cleaning algorithm to perform standardization processing on the preliminary data slices, and the standardization processing includes outlier detection, noise removal, and format conversion processing to obtain standardized data;

[0066] The parallel data reading module uses a distributed file system (such as HDFS) and multi-threading mechanism to simultaneously read data from multiple data sources in a short time to obtain preliminary data slices;

[0067] The standardization processing steps include performing outlier detection, noise removal, and data format standardization processing on the preliminary data slices to obtain the standardized data;

[0068] Data partitioning and index construction: Using an adaptive spatial partitioning algorithm (such as based on a grid or spatial tree structure, such as QuadTree or R-Tree) to perform spatial partitioning on the standardized data to obtain partitioned data blocks; using a distributed spatial index algorithm (such as based on the R-Tree index structure) to locate and locally query the spatial data in the partitioned data blocks to obtain an index structure;

[0069] The adaptive spatial partitioning algorithm is used to adaptively partition the standardized data into multiple data blocks according to the density and distribution characteristics of the forest and grassland spatial data. The expression of the adaptive spatial partitioning algorithm is:

[0070]

[0071] where C is a regional constant used to reflect the uniformity of data distribution;

[0072] M is the available memory size of a single node (or a computing resource metric);

[0073] K is an empirical constant used to balance the partitioning granularity and system load.

[0074] If the number of data points N in a certain node > T, then divide this area into four sub-areas according to the center point, and repeat the steps until the number of data points in each sub-area is less than or equal to T, so as to obtain fine-grained data blocks;

[0075] The value of C can be determined according to historical data analysis, and generally ranges from 1 to 3;

[0076] The value of M is determined according to the current distributed node configuration;

[0077] K is calibrated through experiments so that the amount of data in each sub-region is neither too large nor too small.

[0078] This algorithm fully considers the differences in the density of forest and grass data in different geographical regions, realizes adaptive and dynamic regional division, ensures data locality, and avoids overloading the node computing load, thus providing an efficient basis for subsequent parallel computing;

[0079] The calculation steps of the distributed spatial index algorithm are as follows:

[0080] First, for the data points in each of the divided data blocks, calculate the minimum bounding rectangle (MBR), and the formula is:

[0081]

[0082] The maximum node capacity M_max in each node (affecting the branching factor of the R-Tree node);

[0083] The number of sorting groups S in the STR algorithm, and S depends on the size of the data block;

[0084] Then, use the Bulk-Loading method (such as the Sort-Tile-Recursive algorithm STR) to sort the data according to the spatial position and construct a hierarchical tree structure;

[0085] Finally, in a distributed environment, global indexes are formed among nodes through a local index merging mechanism or a multi-level index structure is used to achieve fast query and update of the index structure;

[0086] This method enables each node to efficiently construct an index locally and achieve a global index through distributed merging, greatly reducing the cross-node query latency and ensuring the consistency of the data topological relationship during the parallel computing process;

[0087] For task scheduling and parallel computing, based on the divided data blocks and the index structure, use task subdivision, data redundancy, and local combination strategies (i.e., division, communication, combination, and mapping strategies) to assemble local computing units; use dynamic task mapping and load balancing methods (for example, by means of the dynamic scheduling mechanism of the Spark computing framework) to allocate the local computing units to multiple computing nodes for parallel execution of spatial computing to obtain preliminary computing results;

[0088] The dynamic task mapping and load balancing method is used to dynamically allocate the local computing units to each computing node to achieve load balancing and reduce communication latency. The specific steps of the dynamic task mapping and load balancing method are as follows:

[0089] Set the computing power parameter Pi for each node (such as the number of CPU cores, memory size, etc.), and at the same time calculate the current node task load Li (which can be represented by the number of tasks or the total task weight);

[0090] Define the load balancing objective function:

[0091] Δ=max(Li / Pi)-min(Li / Pi)

[0092] where Pi is the computing power of the node;

[0093] wj is the task weight, which can be estimated according to the task data volume and computational complexity;

[0094] ∈ is the load balancing tolerance threshold, which is used to determine whether task reallocation is required;

[0095] Minimize Δ through the scheduling algorithm;

[0096] For each task in the local computing unit, allocate it to a node according to the task weight wj, so that:

[0097] f:E→{node}, such that minΔ

[0098] where E is the minimum operation unit of dynamic task allocation, representing the subset of tasks to be scheduled;

[0099] f is the core logic of load balancing, which realizes efficient resource utilization through mathematical optimization and real-time feedback;

[0100] During the task execution process, monitor the load of each node in real time. If it is found that the local load is too high, dynamically migrate some tasks to idle nodes to finally obtain the preliminary calculation result;

[0101] Dynamic task mapping enables the distributed system to automatically achieve balanced allocation in the face of node performance differences and dynamic task changes, ensuring that each node fully utilizes resources, improving the overall parallel computing efficiency and reducing the communication bottleneck caused by uneven load;

[0102] Result integration, storage and data exchange. Use a unified spatial coding method to encode and identify the preliminary calculation results to obtain the encoded data; use the distributed aggregation algorithm to perform boundary correction and data fusion processing on the encoded data to obtain the globally consistent and standardized final calculation result;

[0103] The unified spatial coding method uses the Z-order curve (Z-Order) to encode the spatial coordinates (x, y) into a one-dimensional value, which is convenient for unified identification and fast indexing of cross-node data to obtain the encoded data

[0104] Discretize each spatial point coordinate (x, y) and convert it into an integer representation with a fixed precision, where the precision is p bits, representing the number of binary digits after the coordinate conversion;

[0105] Using the bit interleaving algorithm, cross - merge the binary bits of x and y to obtain the Z - order encoding:

[0106] Z = Interleave(bin(x), bin(y))

[0107] Where the Z value is the generated Z - Order encoding value, which is used to uniquely identify the spatial position and achieve fast sorting and regional positioning of data;

[0108] Where bin(x) and bin(y) respectively represent the binary strings of the discretized x - coordinate and y - coordinate;

[0109] Interleave() represents the bit interleaving function, which alternately merges the binary bits of bin(x) and bin(y);

[0110] The precision p determines the resolution of the spatial encoding. The larger the p value, the higher the encoding precision, but the computational complexity also increases accordingly;

[0111] The discretization strategy adopted should consider the actual spatial range and data distribution;

[0112] Using a space - filling curve to map two - dimensional spatial information to one - dimensional helps to maintain spatial proximity during distributed storage and query, and greatly improves the efficiency of data exchange, fusion, and cross - platform interoperability;

[0113] The described distributed aggregation algorithm adopts a spatial data fusion algorithm based on weighted average of overlapping regions, which is used to perform boundary correction and data fusion processing on the preliminary results calculated by each node to obtain a globally consistent final result. The specific steps are as follows:

[0114] For the overlapping regions existing between adjacent data blocks, calculate the result values A and B of the two data blocks in the overlapping part respectively;

[0115] Set the overlapping weight factors α and β, where the weights can be determined according to their respective regional coverage rates or data credibility;

[0116] The fusion formula is:

[0117]

[0118] Repeat the above process for all overlapping regions. Finally, fuse the results of each node through the distributed aggregation algorithm to obtain a globally consistent and standardized final calculation result;

[0119] Among them, the overlapping weight factors α and β can be set according to the overlapping area ratio, data quality, or sensor accuracy;

[0120] The weight criteria adopted during fusion can be optimized according to specific application scenarios (such as fire warning, ecological assessment);

[0121] This method can effectively eliminate data redundancy and inconsistency problems caused by partition boundaries, achieve smooth transition and precise fusion of multi-node calculation results, and ensure the continuity and accuracy of the final data.

[0122] Specifically, the method further includes storing the final calculation results in a distributed database system (such as HBase or Cassandra) according to a unified spatial encoding, and realizing data exchange with other platforms through a standard data interface to obtain data interconnection and application extension capabilities.

[0123] As can be seen from the above, by using the adaptive QuadTree dynamic partitioning algorithm, adjusting the data block size based on data density to ensure balanced computing task load and avoid waste of computing node resources; combined with the dynamic task mapping algorithm based on the load factor, realizing adaptive task scheduling among computing nodes, maximizing the utilization rate of computing resources, and improving computing throughput.

[0124] Using a distributed R-Tree index, optimizing index construction with Bulk-Loading to significantly reduce index construction time; combined with Z-Order space filling curve encoding, efficiently storing and retrieving spatial data in a distributed database (such as HBase, Cassandra) to achieve fast spatial queries;

[0125] Using a fusion algorithm based on weighted average of overlapping regions to optimize the fusion of boundaries of adjacent data blocks, reduce boundary data discontinuity, and improve data integrity; combined with multi-scale hierarchical calculation, the best fusion strategy can be automatically matched based on different spatial scales to ensure data consistency;

[0126] Adopting a distributed data redundancy storage strategy to ensure that computing tasks can continue to run even when some nodes fail through master-slave data backup; combined with a task dynamic migration mechanism, when a certain computing node fails, the task can be automatically migrated to other idle nodes to achieve highly available computing;

[0127] Through the HDFS+Spark parallel computing architecture, it supports the processing of forest and grass data at the petabyte level and is applicable to the analysis of forest and grass ecological data across the country; the distributed index supports multi-level spatial queries and can be applied to analysis tasks with different precision requirements (such as global forest change monitoring VS local forest area health assessment).

[0128] This method can run on cloud computing platforms such as Hadoop / Spark, local server clusters, and edge computing devices, and is applicable to different computing environments; through an extensible task scheduling framework, it supports GPU-accelerated computing and improves the application ability of deep learning models in ecological environment prediction.

[0129] Example 2:

[0130] Forest fire early warning system based on distributed spatial computing:

[0131] 1. Data acquisition and preprocessing:

[0132] Data source:

[0133] Multi-source remote sensing data is adopted, including Landsat satellite images with a resolution of 30 meters, UAV aerial photography data, and ground meteorological monitoring data.

[0134] The data is stored in a distributed file system (such as HDFS), and the total amount of data is about 1TB.

[0135] Preprocessing steps:

[0136] Parallel data reading: Using HDFS and multi-threaded mechanisms, raw data is read from multiple data sources simultaneously to obtain preliminary data slices.

[0137] Data cleaning and format conversion: Cloud detection, noise removal, and filtering of water bodies and shadow areas are performed to ensure data consistency and generate standardized data.

[0138] 2. Adaptive spatial partitioning:

[0139] Algorithm description:

[0140] The adaptive QuadTree partitioning algorithm is adopted to dynamically partition regions according to the spatial density of forest area data points.

[0141] The initial region is set as the root node, and the number of data points N in this region is counted;

[0142] A threshold T is set, where:

[0143]

[0144] The parameters are as follows:

[0145] C = 2 (region constant, determined according to historical statistics);

[0146] M = 4096MB (available memory per node);

[0147] K = 100 (empirical constant);

[0148] Then T≈2×4096 / 100≈82 data points.

[0149] For each region, when N>T, divide the region into four sub-regions according to the center point until the number of data points in all sub-regions does not exceed 82, thereby obtaining the partitioned data block C.

[0150] 3. Distributed Spatial Index Construction

[0151] Index construction process:

[0152] Adopt the distributed R-Tree construction algorithm to construct an index for each data block C:

[0153] Calculate the minimum bounding rectangle (MBR) for each data point in C.

[0154]

[0155] Adopt the Bulk-Loading method and the STR algorithm to sort the data according to the spatial position to construct a hierarchical tree structure, and set the maximum capacity Mmax of each node to 50 data points.

[0156] Merge the local indexes between each node to form a global index to achieve fast positioning and local query.

[0157] 4. Dynamic Task Scheduling and Parallel Computing

[0158] Task allocation and load balancing:

[0159] Adopt the dynamic task mapping algorithm based on the load factor to allocate the local computing units (consisting of data blocks and indexes) to each computing node.

[0160] The computing power parameter Pi of each node (for example, node 1: 8 cores, node 2: 6 cores, etc.);

[0161] The task weight wj is estimated based on the data volume and computational complexity;

[0162] The objective function is

[0163]

[0164] Dynamically monitor the load of each node, migrate tasks in real time, and finally perform parallel computing on the fire risk indicators on each node to obtain preliminary calculation results.

[0165] Spatial computing content:

[0166] Conduct comprehensive analysis of multiple indicators such as high temperature, dryness, and vegetation coverage for each data block, and use spatial relationships to judge the fire risk.

[0167] 5. Result Fusion and Encoding Storage

[0168] Overlapping region data fusion:

[0169] For example, use a fusion algorithm based on weighted average of overlapping regions to process the results of adjacent data blocks.

[0170] For the fire risk values A and B obtained separately in the overlapping area of adjacent blocks, set the weight factors

[0171] α = 0.6 and β = 0.4

[0172]

[0173] Repeat the processing of all overlapping regions and integrate to generate a global fire risk map.

[0174] Spatial encoding and storage:

[0175] Use an encoding algorithm based on the Z-Order curve to encode each spatial point.

[0176] Set the discretization precision to p = 16 bits.

[0177] Z = Interleave(bin(x), bin(y))

[0178] The encoded data H is stored in a distributed database (such as HBase) and interconnected with the fire control system through a standard data interface to achieve real-time early warning data exchange.

[0179] As can be seen from the above,

[0180] Distributed parallel processing reduces the total processing time by about 50% compared with centralized processing.

[0181] The accuracy rate of fire risk area identification reaches 95%, greatly shortening the fire early warning time (alarm more than 5 minutes in advance).

[0182] When some nodes fail, the system can still maintain about 80% of the computing load to ensure the stable operation of the early warning system.

[0183] Example 3:

[0184] Ecological environment assessment and forest resource management system based on distributed spatial computing

[0185] 1. Data acquisition and preprocessing

[0186] Data source:

[0187] Integrate satellite remote sensing data, UAV high-resolution images, ground sensor data and historical forest and grass survey data;

[0188] The data scale is about 500GB and is stored in HDFS.

[0189] Preprocessing steps:

[0190] Parallel data reading: Use HDFS and multi-threading technology to synchronously read data and obtain the preliminary data slice A;

[0191] Data cleaning and format standardization: Remove noise, perform image registration and spectral correction on A to obtain the standardized data B.

[0192] Example of parameter setting:

[0193] In the adaptive partitioning algorithm, the parameter C is set to 1.5, the single-node memory M is 2048 MB, and the empirical constant K is 80, then T = 1.5×2048 / 80 ≈ 38 data points

[0194] 2. Adaptive spatial partitioning:

[0195] Partitioning process:

[0196] Use the adaptive QuadTree algorithm to partition the standardized data B according to the threshold T≈38 to ensure that the number of data points in each data block C does not exceed 38, taking into account data locality and balanced load.

[0197] 3. Distributed spatial index construction:

[0198] Index construction steps:

[0199] For the data in each data block C, calculate the MBR and use the STR algorithm to construct an R-Tree index,

[0200] Set the maximum capacity of each node Mmax = 30 data points, the number of sorting and grouping S = 10 (adjusted according to the data block size), and construct the index structure D to support fast spatial queries.

[0201] 4. Distributed task scheduling and parallel computing:

[0202] Task decomposition and allocation:

[0203] Use the dynamic task mapping algorithm to allocate the data block C to each computing node, and the task content is mainly to calculate vegetation indices (such as NDVI), water body coverage rate, and other ecological environment indicators.

[0204] Each node dynamically schedules according to the computing power parameter Pi (such as some nodes are 4-core and some are 6-core) and the task weight wj to achieve load balancing and obtain the preliminary calculation result G.

[0205] NDVI calculation formula within the node:

[0206] NDVI = (NIR - Red) / (NIR + Red)

[0207] Perform parallel processing to obtain the vegetation indices of each region.

[0208] 5. Result Fusion and Unified Coding

[0209] Result Fusion:

[0210] For the overlapping regions of different data block boundaries, use the weighted average method for overlapping regions for fusion.

[0211] Set the weight factors of the overlapping region as α = 0.5 and β = 0.5, then:

[0212] Rfused = (0.5·A + 0.5·B) / 1.0

[0213] Obtain the globally unified vegetation index and ecological environment index map.

[0214] Unified Spatial Coding and Storage:

[0215] Adopt the Z - Order curve coding algorithm to encode the two - dimensional spatial coordinates into a one - dimensional identifier, and set the discretization accuracy as p = 14 bits for fast sorting and spatial retrieval;

[0216] The encoded data is stored in a distributed database (such as Cassandra), and data sharing is realized with the regional ecological environment monitoring platform through a standard interface.

[0217] As can be seen from the above, distributed processing significantly improves the data processing speed. Compared with the traditional centralized method, the overall processing time is reduced by about 60%.

[0218] The generated high - resolution vegetation index map and ecological index map can accurately reflect the regional ecological status, and the accuracy of determining the healthy and degraded regions reaches more than 90%, providing a scientific basis for resource management and environmental protection decision - making.

[0219] Data standardization and unified coding achieve cross - platform data interoperability, facilitating long - term monitoring and dynamic resource management, and improving management efficiency and emergency response capabilities.

[0220] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A high-performance distributed spatial algorithm parallel processing method for forestry and grassland spatial data, characterized in that Including: Read the original forest and grassland spatial data from multiple data sources to obtain preliminary data slices; Use a data cleaning algorithm to perform standardization processing on the preliminary data slices to obtain standardized data; Adopt an adaptive space partitioning algorithm to partition the standardized data to obtain partitioned data blocks; adopt a distributed spatial indexing algorithm to locate and locally query the spatial data in the partitioned data blocks to obtain an index structure; According to the partitioned data blocks and the index structure, use task subdivision, data redundancy, and local combination strategies to assemble local computing units; use a dynamic task mapping and load balancing method to allocate the local computing units to multiple computing nodes to perform spatial computing in parallel to obtain preliminary computing results; Perform coding identification processing on the preliminary computing results using a unified spatial coding method to obtain encoded data; use a distributed aggregation algorithm to perform boundary correction and data fusion processing on the encoded data to obtain a globally consistent and standardized final computing result.

2. The high-performance distributed spatial algorithm parallelization processing method for forestry and grassland spatial data according to claim 1, wherein The adaptive space partitioning algorithm is used to adaptively partition the standardized data into multiple data blocks according to the density and distribution characteristics of the forest and grassland spatial data. The expression of the adaptive space partitioning algorithm is: where C is a regional constant used to reflect the uniformity of data distribution; M is the available memory size of a single node; K is an empirical constant used to balance the partitioning granularity and system load.

3. The high-performance distributed spatial algorithm parallelization processing method for forestry and grassland spatial data according to claim 1, wherein The calculation steps of the distributed spatial indexing algorithm are as follows: First, for the data points in each of the partitioned data blocks, calculate the minimum bounding rectangle. The formula is: The maximum node capacity M_max within each node; The number of sorting groups S in the STR algorithm, and S depends on the size of the data block; Then use the Bulk-Loading method to sort the data according to the spatial position and construct a hierarchical tree structure; Finally, in a distributed environment, each node forms a global index through a local index merging mechanism or uses a multi-level index structure to achieve fast query and update of the index structure.

4. The high-performance distributed spatial algorithm parallelization processing method for forestry and grassland spatial data according to claim 1, wherein The dynamic task mapping and load balancing method is used to dynamically allocate the local computing units to each computing node to achieve load balancing and reduce communication latency. The specific steps of the dynamic task mapping and load balancing method are: Set the computing power parameter Pi of each node, and at the same time calculate the current node task load Li; Define the load balancing objective function Δ = max(Li / Pi) - min(Li / Pi) where Pi is: the node computing power; wj is: the task weight; ∈ is: the load balancing tolerance threshold; Minimize Δ through a scheduling algorithm; For each task in the local computing unit, allocate it to a node according to the task weight wj, so that: f: E → {node}, so that minΔ where E is the minimum operation unit of dynamic task allocation, representing the subset of tasks to be scheduled; f is the core logic of load balancing, and realizes efficient use of resources through mathematical optimization and real-time feedback.

5. The high-performance distributed spatial algorithm parallelization processing method for forestry and grassland spatial data according to claim 1, wherein The unified spatial coding method uses the Z-order curve to encode the spatial coordinates (x, y) into a one-dimensional value, which is convenient for the unified identification and fast indexing of cross-node data, and obtains the encoded data Discretize each spatial point coordinate (x, y) and convert it into an integer representation with a fixed precision, where the precision is p bits; Using the bit interleaving algorithm, cross-combine the binary bits of x and y to obtain the Z-order coding: Z = Interleave(bin(x), bin(y)) Among them, the Z value is the generated Z-Order coding value, which is used to uniquely identify the spatial position and realize the fast sorting and regional positioning of data; bin(x) and bin(y) respectively represent the binary strings of the discretized x coordinate and y coordinate; Interleave() represents the bit interleaving function, which alternately combines the binary bits of bin(x) and bin(y); The precision p determines the resolution of the spatial coding. The larger the p value, the higher the coding precision, but the computational complexity also increases accordingly.

6. The high-performance distributed spatial algorithm parallelization processing method for forestry and grassland spatial data according to claim 1, characterized in that, The distributed aggregation algorithm uses a spatial data fusion algorithm based on weighted average of overlapping regions to perform boundary correction and data fusion processing on the preliminary results calculated by each node, and obtains a globally consistent final result. The specific steps are as follows: For the overlapping regions existing between adjacent data blocks, calculate the result values A and B of the two data blocks in the overlapping part respectively; Set the overlapping weight factors α and β, where the weights are determined according to their respective regional coverage rates or data credibility; The fusion formula Rfused is: Repeat the above process for all overlapping regions, and finally fuse the results of each node through the distributed aggregation algorithm to obtain a globally consistent and standardized final calculation result; Among them, the overlapping weight factors α and β are set according to the overlapping area ratio, data quality or sensor accuracy.

7. The high-performance distributed spatial algorithm parallel processing method for forestry and grassland spatial data according to claim 1, characterized in that The parallelized data reading module uses a distributed file system and a multi-thread mechanism to simultaneously read data from multiple data sources and obtain preliminary data slices; The standardization processing step includes performing outlier detection, noise removal and data format standardization processing on the preliminary data slices to obtain the standardized data.

8. The high-performance distributed spatial algorithm parallelization processing method for forestry and grassland spatial data according to claim 1, wherein The method further includes using a distributed database system to store the final calculation result according to the unified spatial coding, and realizing data exchange with other platforms through a standard data interface, so as to obtain data interoperability and application extension capabilities.