Forestry point cloud data processing method and device based on lake and warehouse integration

CN120523883BActive Publication Date: 2026-09-08JIULING (SHANGHAI) INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510533067.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2026-09-08
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

[0004]本申请的主要目的在于提供一种基于湖仓一体的林业点云数据处理方法和装置、存储介质及电子设备,以解决相关技术中通过人工的方式对海量林业点云数据进行处理和存储,导致数据处理效率比较低的问题

Benefits of technology

[0025]在本申请实施例中,采用以下步骤:获取目标区域的待处理的原始点云数据,并将原始点云数据存储至数据湖中;通过分布式计算引擎依据预设的多个数据处理管道对原始点云数据进行处理,得到目标区域对应的林分因子数据;依据业务需求对林分因子数据进行数据建模处理,得到目标数据模型,并依据目标数据模型,构建林业数仓,其中,林业数仓用于为目标对象提供所需数据,解决了相关技术中通过人工的方式对海量林业点云数据进行处理和存储,导致数据处理效率比较低的技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523883B_ABST
    Figure CN120523883B_ABST
Patent Text Reader

Abstract

The application discloses a kind of forestry point cloud data processing method and device based on lake warehouse integration. It is related to data processing technical field, the method comprises: obtaining the original point cloud data to be processed of target area, and the original point cloud data is stored to data lake;By the distributed computing engine, the original point cloud data is processed according to the preset multiple data processing pipelines, and the corresponding stand factor data of the target area is obtained;According to the business requirement, the stand factor data is processed by data modeling, and the target data model is obtained, and according to the target data model, forestry data warehouse is constructed, wherein the forestry data warehouse is used to provide the required data for target object. Through the present application, the problem that the massive forestry point cloud data is processed and stored by manual method in the related art, resulting in relatively low data processing efficiency, is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a method and apparatus for processing forestry point cloud data based on lake-warehouse integration. Background Technology

[0002] In the field of forestry remote sensing information processing technology, with the widespread application of LiDAR technology, the scale of forestry point cloud data acquisition has exploded, with TB to PB levels of data becoming the norm. The storage and management of massive point cloud data still rely on traditional file systems, and existing forestry data warehouses lack the ability to perform multi-dimensional modeling and analysis of point cloud data, resulting in low data analysis and reuse rates and an inability to meet ever-changing business needs and in-depth analysis.

[0003] There is currently no effective solution to the problem of low data processing efficiency caused by manually processing and storing massive amounts of forestry point cloud data in related technologies. Summary of the Invention

[0004] The main objective of this application is to provide a forestry point cloud data processing method, device, storage medium, and electronic device based on lake-warehouse integration, in order to solve the problem of low data processing efficiency caused by manually processing and storing massive amounts of forestry point cloud data in related technologies.

[0005] To achieve the above objectives, according to one aspect of this application, a forestry point cloud data processing method based on a data lake-database integration is provided. The method includes: acquiring raw point cloud data of a target area to be processed and storing the raw point cloud data in a data lake; processing the raw point cloud data using a distributed computing engine based on multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area; performing data modeling processing on the forest stand factor data according to business requirements to obtain a target data model; and constructing a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data to the target object.

[0006] Furthermore, before processing the raw point cloud data through a distributed computing engine based on multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area, the method further includes: determining multiple data processing components based on the operation behavior of the target object; constructing data processing steps for processing the point cloud data based on the multiple data processing components; and configuring the multiple data processing pipelines based on the data processing steps and preset triggering methods.

[0007] Furthermore, before processing the original point cloud data through a distributed computing engine using multiple preset data processing pipelines, the method further includes: performing data optimization processing on the original point cloud data to enable the processed original point cloud data to support distributed computing; processing the original point cloud data through a distributed computing engine using multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area includes: processing the processed original point cloud data through the distributed computing engine using the multiple data processing pipelines to obtain the forest stand factor data corresponding to the target area.

[0008] Furthermore, the forest stand factor data corresponding to the target area is obtained by processing the original point cloud data through a distributed computing engine based on multiple preset data processing pipelines. This includes: cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size; determining a second preset size for each data processing pipeline based on the resource allocation information of each data processing pipeline; merging the multiple first point cloud data blocks based on the second preset size of each data processing pipeline to obtain multiple second point cloud data blocks of the second preset size for each data processing pipeline; and processing the multiple second point cloud data blocks through each data processing pipeline to obtain the forest stand factor data.

[0009] Furthermore, after acquiring the raw point cloud data to be processed in the target area, the method further includes: converting the raw point cloud data according to columnar storage to obtain converted raw point cloud data, and storing the converted raw point cloud data in the raw data layer of the target system; after performing data modeling processing on the forest stand factor data according to business requirements to obtain the target data model, the method further includes: storing the forest stand factor data and the target data model in the service data layer of the target system, wherein the target object accesses the required data through the service data layer.

[0010] Furthermore, after cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, the method further includes: converting the multiple first point cloud data blocks into a format to obtain multiple first point cloud data blocks in binary format; bucketing the multiple first point cloud data blocks in binary format according to the storage space of preset buckets to obtain processed point cloud data blocks; and storing the processed point cloud data blocks in the processing data layer of the target system.

[0011] Furthermore, after processing the plurality of second point cloud data blocks through each data processing pipeline to obtain the forest stand factor data, the method further includes: obtaining intermediate data results output when processing the plurality of second point cloud data blocks; constructing a data processing lineage based on the intermediate data results, the original point cloud data, and the plurality of first point cloud data blocks, wherein data tracking is achieved through the data processing lineage.

[0012] Further, constructing the data processing lineage based on the intermediate data results, the original point cloud data, and the plurality of first point cloud data blocks includes: obtaining the first metadata corresponding to the original point cloud data, the second metadata corresponding to the plurality of first point cloud data blocks, and the third metadata corresponding to the intermediate data results; determining the data processing relationship between the first metadata, the second metadata, and the third metadata; and obtaining the data processing lineage based on the first metadata, the second metadata, the third metadata, and the data processing relationship.

[0013] Furthermore, the target data model is obtained by performing data modeling processing on the forest stand factor data according to business needs, including: determining a first analysis dimension for the forest stand factor data based on the business needs; determining a second analysis dimension for the forest stand factor data based on the type of data accessed by the target object; constructing a fact table based on the forest stand factor data; constructing a first dimension table corresponding to the first analysis dimension based on the first analysis dimension and the forest stand factor data; and constructing a second dimension table corresponding to the second analysis dimension based on the second analysis dimension and the forest stand factor data; and obtaining the target data model based on the fact table, the first dimension table, and the second dimension table.

[0014] To achieve the above objectives, according to another aspect of this application, a forestry point cloud data processing device based on a data lake is provided. The device includes: a first acquisition unit, configured to acquire raw point cloud data of a target area to be processed and store the raw point cloud data in a data lake; a first processing unit, configured to process the raw point cloud data using a distributed computing engine according to multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area; and a second processing unit, configured to perform data modeling processing on the forest stand factor data according to business requirements to obtain a target data model, and construct a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data to the target object.

[0015] Furthermore, the device further includes: a determining unit, configured to determine multiple data processing components based on the operational behavior of the target object before processing the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area; a constructing unit, configured to construct data processing steps for processing the point cloud data based on the multiple data processing components; and a configuring unit, configured according to the data processing steps and a preset triggering method to obtain the multiple data processing pipelines.

[0016] Furthermore, the device further includes: an optimization unit, configured to perform data optimization processing on the original point cloud data before processing the original point cloud data through a distributed computing engine according to a plurality of preset data processing pipelines, so that the processed original point cloud data supports distributed computing; the first processing unit is further configured to process the processed original point cloud data through the distributed computing engine according to the plurality of data processing pipelines to obtain forest stand factor data corresponding to the target area.

[0017] Further, the first processing unit includes: a cutting module, used to cut the original point cloud data to obtain a plurality of first point cloud data blocks of a first preset size; a first determining module, used to determine a second preset size corresponding to each data processing pipeline based on the resource allocation information of each data processing pipeline; a merging module, used to merge the plurality of first point cloud data blocks based on the second preset size corresponding to each data processing pipeline to obtain a plurality of second point cloud data blocks of a second preset size corresponding to each data processing pipeline; and a processing module, used to process the plurality of second point cloud data blocks through each data processing pipeline to obtain the forest stand factor data.

[0018] Furthermore, the device further includes: a first conversion unit, configured to, after acquiring the raw point cloud data to be processed in the target area, perform format conversion on the raw point cloud data according to columnar storage to obtain converted raw point cloud data, and store the converted raw point cloud data in the raw data layer of the target system; and a first storage unit, configured to, after performing data modeling processing on the forest stand factor data according to business requirements to obtain a target data model, store the forest stand factor data and the target data model in the service data layer of the target system, wherein the target object accesses the required data through the service data layer.

[0019] Furthermore, the device further includes: a second conversion unit, configured to, after cutting the original point cloud data to obtain a plurality of first point cloud data blocks of a first preset size, convert the plurality of first point cloud data blocks into a binary format; a third processing unit, configured to, according to a preset storage space of buckets, perform bucketing on the plurality of first point cloud data blocks in binary format to obtain processed point cloud data blocks; and a second storage unit, configured to store the processed point cloud data blocks in the processing data layer of the target system.

[0020] Furthermore, the device further includes: a second acquisition unit, configured to acquire intermediate data results output during the processing of the plurality of second point cloud data blocks after processing them through each data processing pipeline to obtain the forest stand factor data; and a construction unit, configured to construct a data processing lineage based on the intermediate data results, the original point cloud data, and the plurality of first point cloud data blocks, wherein data tracking is achieved through the data processing lineage.

[0021] Furthermore, the construction unit includes: an acquisition module, used to acquire first metadata corresponding to the original point cloud data, second metadata corresponding to the plurality of first point cloud data blocks, and third metadata corresponding to the intermediate data result; a second determination module, used to determine the data processing relationship between the first metadata, the second metadata, and the third metadata; and a third determination module, used to obtain the data processing lineage based on the first metadata, the second metadata, the third metadata, and the data processing relationship.

[0022] Further, the second processing unit includes: a fourth determining module, used to determine a first analytical dimension of the stand factor data based on the business requirements; a fifth determining module, used to determine a second analytical dimension of the stand factor data based on the type of data accessed by the target object; a construction module, used to construct a fact table based on the stand factor data, construct a first dimension table corresponding to the first analytical dimension based on the first analytical dimension and the stand factor data, and construct a second dimension table corresponding to the second analytical dimension based on the second analytical dimension and the stand factor data; and a sixth determining module, used to obtain the target data model based on the fact table, the first dimension table, and the second dimension table.

[0023] According to another aspect of the present invention, an electronic device is also provided, comprising: a memory storing an executable program; and a processor for running the program, wherein the program executes the forestry point cloud data processing method based on lake-warehouse integration as described above when it runs.

[0024] According to another aspect of the present invention, a computer-readable storage medium is also provided, the storage medium storing a program, wherein, when the program is running, the device where the storage medium is located controls the execution of any of the above-mentioned forestry point cloud data processing methods based on lake-warehouse integration.

[0025] In this embodiment, the following steps are adopted: acquiring the raw point cloud data of the target area to be processed and storing the raw point cloud data in a data lake; processing the raw point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area; performing data modeling processing on the forest stand factor data according to business needs to obtain the target data model, and constructing a forestry data warehouse based on the target data model. The forestry data warehouse is used to provide the required data for the target object, which solves the technical problem in related technologies where the processing and storage of massive forestry point cloud data is done manually, resulting in low data processing efficiency.

[0026] This solution introduces distributed computing and data processing pipelines to achieve efficient processing and analysis of massive forestry point cloud data, significantly improving data processing efficiency and quality. Leveraging the powerful parallel processing capabilities of the distributed computing engine, raw point cloud data can be quickly processed to obtain corresponding stand factor data, drastically shortening data processing time. By constructing target data models tailored to specific business needs, stand factor data can be rapidly transformed into analytical results from various business perspectives, avoiding repetitive data processing work, reducing data management costs, and ultimately achieving the technical effect of improving overall data utilization. Attached Figure Description

[0027] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0028] Figure 1 This is a flowchart of a forestry point cloud data processing method based on lake-warehouse integration provided in the embodiments of this application. Figure 1 ;

[0029] Figure 2 This is a flowchart of a forestry point cloud data processing method based on lake-warehouse integration provided in the embodiments of this application. Figure 2 ;

[0030] Figure 3 This is a schematic diagram of a forestry point cloud data processing method based on lake-warehouse integration provided in the embodiments of this application;

[0031] Figure 4 This is a schematic diagram of a forestry point cloud data processing device based on a lake-warehouse integration, provided according to an embodiment of this application.

[0032] Figure 5 This is a schematic diagram of a forestry point cloud data processing system based on lake-warehouse integration, provided according to an embodiment of this application.

[0033] Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0035] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0036] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding access points are provided for users to choose to authorize or refuse. For example, interfaces are set up between this system and relevant users or organizations, providing users with corresponding access points to choose to agree to or refuse automated decision-making results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.

[0037] According to an embodiment of this application, an embodiment of a forestry point cloud data processing method based on lake-warehouse integration is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0038] This application provides, as follows: Figure 1 The method for processing forestry point cloud data based on the integration of lake and warehouse is shown. Figure 1 This is the flowchart of the forestry point cloud data processing method based on lake-warehouse integration according to Embodiment 1 of this application. Figure 1 The processing method includes:

[0039] Step S101: Obtain the raw point cloud data to be processed in the target area and store the raw point cloud data in the data lake.

[0040] Optionally, by specifying the specific geographical area to be analyzed (i.e., the target area mentioned above), raw point cloud data of the target area can be collected using airborne LiDAR. Airborne LiDAR creates a highly detailed spatial dataset by emitting laser pulses and measuring the reflection time, including information such as tree location, height, and crown width, forming point cloud data. Raw point cloud data is typically stored in .LAS or .LAZ file format. The collected data can be directly stored in a data lake so that raw point cloud data can be retrieved from the data lake later.

[0041] Step S102: The raw point cloud data is processed by the distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area.

[0042] Optionally, after obtaining the raw point cloud data, multiple data processing pipelines can be configured to enhance the flexibility of data processing configuration and achieve full automation of the data computation process. The raw point cloud data is distributed to multiple pre-defined data processing pipelines via a distributed computing engine for data processing, such as point cloud thinning, point cloud denoising, point cloud normalization, vegetation point separation, and feature calculation, thereby obtaining the forest stand factor data corresponding to the target area. Forest stand factor data includes, but is not limited to, tree height, canopy area, trunk diameter, and forest density. These forest stand factors are key indicators describing forest structure and health status, and are crucial for the quantitative analysis of forest resources.

[0043] In an alternative embodiment, the distributed computing engine may be a high-performance distributed computing engine to achieve parallel processing of large-scale data.

[0044] Step S103: Based on business needs, perform data modeling processing on the forest stand factor data to obtain the target data model, and construct a forestry data warehouse based on the target data model. The forestry data warehouse is used to provide the required data for the target object.

[0045] Optionally, current business needs are determined, such as estimating forest carbon storage, monitoring tree growth status, or assessing forest health. Different business needs determine the direction of data model construction and key indicators. Based on business needs, dimensional modeling is performed on stand factor data, combining stand factor data with dimensions such as time, space, and tree species to construct star-shaped or snowflake-shaped models (i.e., the target data model mentioned above). For example, to assess the carbon sink potential of forests, a multidimensional data model can be constructed with "carbon storage" as the fact table and time, space, and tree species as dimension tables. Based on the constructed target data model, a forestry data warehouse is built so that users (i.e., the target objects mentioned above) can develop and implement specific business applications based on the forestry data warehouse, such as forest resource management platforms and ecological monitoring systems. These applications can directly extract the required data from the data model for real-time or historical analysis, providing decision-makers with intuitive and accurate data support.

[0046] In summary, by introducing distributed computing and data processing pipelines, efficient processing and analysis of massive forestry point cloud data were achieved, significantly improving data processing efficiency and quality. Leveraging the powerful parallel processing capabilities of the distributed computing engine, raw point cloud data was rapidly processed to obtain corresponding stand factor data, significantly shortening data processing time. By constructing target data models tailored to specific business needs, stand factor data could be quickly transformed into analytical results from various business perspectives, avoiding repetitive data processing work, reducing data management costs, and ultimately achieving the technical effect of improving overall data utilization.

[0047] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in this application embodiment, before processing the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area, the method further includes: determining multiple data processing components based on the operation behavior of the target object; constructing data processing steps for processing point cloud data based on the multiple data processing components; and configuring multiple data processing pipelines according to the data processing steps and preset triggering methods.

[0048] In an optional embodiment, the forestry point cloud data processing method based on lake-granary integration provided in this application can be implemented using a forestry point cloud data processing system based on lake-granary integration. This system can pre-configure eight key processing components, including thinning, segmentation, and feature extraction. These components cover most of the core functions from point cloud data preprocessing to stand factor extraction. The advantage of these pre-configured components is that they are carefully designed and optimized, enabling rapid response to common data processing needs and reducing the time and resources users invest in component development. Furthermore, the system supports users in self-writing and uploading their own designed data processing components according to specific business needs or preferences.

[0049] This processing system provides an intuitive front-end interface where users (the target object) can construct data processing steps by dragging and dropping preset or custom processing components. These steps are displayed in the front-end interface as a directed acyclic graph. Therefore, multiple data processing components can be determined based on the target object's actions (e.g., drag-and-drop operations), and data processing steps for processing point cloud data can be constructed based on these components. Finally, multiple data processing pipelines are obtained by configuring the data processing steps and preset triggering methods.

[0050] The system supports two main triggering mechanisms: event-driven and scheduled tasks. The event-driven mechanism allows the data processing flow to begin immediately upon being triggered by an operation in the management interface, such as automatically starting data processing when a new batch of point cloud data is uploaded. Scheduled tasks, on the other hand, provide scheduling configuration based on Cron expressions, allowing users to preset the data processing cycle, such as executing specific data processing flows at fixed times daily or weekly, ensuring timely data updates and analysis.

[0051] It should be noted that when a data processing component or stage malfunctions, the system automatically executes a retry mechanism. The preset maximum number of retries is 3, but users can also customize the number of retries according to specific circumstances to ensure the continuity and reliability of the data processing flow. Furthermore, regardless of whether any step in the data processing is successful or not, the system will promptly notify the user of the data processing results via email or message, facilitating real-time monitoring of the data processing status and enabling the identification and resolution of problems. After the data processing flow is completed, the system automatically performs a correctness check on the results to ensure that the generated stand factor data meets the expected quality standards, avoiding analytical biases caused by data errors.

[0052] Through a component-based architecture and custom component uploads, users can flexibly configure data processing workflows according to different business scenarios and needs. The combination of a visual DAG editor and triggering mechanisms makes data processing configuration both intuitive and efficient. Users can complete the design and scheduling of data processing workflows without writing complex script code, thereby achieving the technical effect of improving data processing efficiency.

[0053] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in this application embodiment, before processing the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines, the method further includes: performing data optimization processing on the original point cloud data so that the processed original point cloud data supports distributed computing; processing the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area includes: processing the processed original point cloud data through a distributed computing engine according to multiple data processing pipelines to obtain forest stand factor data corresponding to the target area.

[0054] In an optional embodiment, before the raw point cloud data enters the distributed computing engine for processing, to ensure efficient and stable processing in the distributed environment, the raw point cloud data can be optimized. For example, the raw point cloud data can be converted from LAS or LAZ format to a format more suitable for distributed computing, such as Parquet, or the raw point cloud data can be compressed. Furthermore, the Python Pdal point cloud processing package and the Python Open3D point cloud processing package can be optimized to run on a distributed computing engine. After obtaining the processed raw point cloud data, the distributed computing engine processes the raw point cloud data through multiple data processing pipelines to obtain the forest stand factor data corresponding to the target area.

[0055] By pre-optimizing the raw point cloud data, not only is basic support provided for distributed computing, but the efficiency and stability of the entire data processing flow are also significantly improved.

[0056] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in this application embodiment, the process of processing the original point cloud data according to multiple preset data processing pipelines by a distributed computing engine to obtain forest stand factor data corresponding to the target area includes: cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size; determining a second preset size corresponding to each data processing pipeline based on the resource allocation information of each data processing pipeline; merging the multiple first point cloud data blocks according to the second preset size corresponding to each data processing pipeline to obtain multiple second point cloud data blocks of the second preset size corresponding to each data processing pipeline; and processing the multiple second point cloud data blocks through each data processing pipeline to obtain forest stand factor data.

[0057] In an optional embodiment, the raw point cloud data is first cut into multiple first point cloud data blocks according to a preset first size (e.g., 25m*25m). This process helps to decompose large files into smaller blocks, facilitating subsequent parallel processing. It should be noted that the first preset size can be set according to actual needs.

[0058] Then, based on the actual performance of each computing node in the distributed cluster (such as hardware indicators like CPU and memory), the resource allocation information for each pipeline (i.e., the data processing pipeline mentioned above) is determined. Subsequently, the most suitable data block size for each pipeline to process, i.e., the second preset size, can be determined based on the resource allocation information to ensure the reasonable allocation and maximum utilization of resources.

[0059] After obtaining the second preset size, multiple first point cloud data blocks are merged to form second point cloud data blocks suitable for each data processing pipeline. It should be noted that the merging strategy must consider spatial continuity and the performance of the computing nodes, ensuring that the merged data block size meets computational requirements without being too large (leading to memory overflow) or too small (leading to low computational efficiency). Finally, each data processing pipeline processes the multiple second point cloud data blocks, performing actions such as thinning, denoising, vegetation point separation, and feature extraction, ultimately obtaining the forest stand factor data.

[0060] The above steps effectively avoid the waste of computing resources caused by mismatched data block sizes, and significantly improve the efficiency and response speed of distributed computing.

[0061] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in the embodiments of this application, after obtaining the raw point cloud data to be processed in the target area, the method further includes: converting the format of the raw point cloud data according to columnar storage to obtain the converted raw point cloud data, and storing the converted raw point cloud data in the raw data layer of the target system; after performing data modeling processing on the forest stand factor data according to business needs to obtain the target data model, the method further includes: storing the forest stand factor data and the target data model in the service data layer of the target system, wherein the target object accesses the required data through the service data layer.

[0062] In an optional embodiment, after acquiring the raw point cloud data to be processed for the target area, the raw point cloud data is first converted to a columnar storage format according to the requirements of columnar storage. Common raw point cloud data is stored in LAS / LAZ format, but to leverage the advantages of columnar storage for efficient querying and computation, it needs to be converted to a columnar storage format, such as Parquet. Then, the converted raw point cloud data is stored in the raw data layer of the target system. In forestry analysis, statistical analysis is often required for a specific column of data (e.g., tree height, crown width). Columnar storage can directly read the data in the required column, reducing I / O operations and significantly improving query efficiency. Furthermore, columnar storage encodes and compresses data in the same column, achieving a higher compression ratio and reducing storage costs, especially for TB-PB level data. The columnar storage structure is also easier to implement in a distributed computing environment for parallel read / write operations, contributing to the efficient execution of parallel computing tasks.

[0063] In an optional embodiment, after the distributed computing engine processes the point cloud data to generate the target data model, the forest stand factor data and the target data model are stored in the service data layer of the integrated lake-warehouse forestry point cloud data processing system (i.e., the aforementioned target system). This layer primarily provides fast access and multidimensional analysis capabilities. The service data layer typically employs an OLAP (Online Analytical Processing) storage structure to facilitate rapid querying and data analysis by users. It should be noted that the aforementioned forestry data warehouse can be deployed within the service data layer.

[0064] The introduction of columnar storage significantly improves the storage efficiency of point cloud data, reduces storage space requirements, and enhances data reading and query speed. By serving the forest stand factor data and data models stored in the data layer, users can flexibly perform multidimensional analysis, improve data reuse, reduce the cost of repetitive analysis, and meet the needs of different business scenarios and analysts.

[0065] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in the embodiments of this application, after cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, the method further includes: converting the multiple first point cloud data blocks into a format to obtain multiple first point cloud data blocks in binary format; dividing the multiple first point cloud data blocks in binary format into buckets according to the storage space of preset buckets to obtain processed point cloud data blocks; and storing the processed point cloud data blocks in the processing data layer of the target system.

[0066] In an optional embodiment, during point cloud data processing, the segmentation process generates a large number of small files, which severely impact IO efficiency and hinder parallel computing efficiency. Therefore, after segmenting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, these point cloud data blocks are further converted into binary format. Typically, binary format refers to a more efficient data storage format, such as a sequence file, which offers significant advantages over the original LAS / LAZ or text formats in terms of read / write speed, memory usage, and network transmission efficiency.

[0067] Based on the preset bucket size, the first point cloud data block in binary format is bucketed. It should be noted that the preset bucket size can be designed based on cluster storage capacity and computational efficiency requirements. Multiple first point cloud data blocks in binary format within the same bucket are merged into a larger file, thereby reducing the number of small files. Finally, the point cloud data blocks, after being optimized by SequenceFile merging, are stored in the processing data layer of the target system. The processing data layer is an intermediate layer used to store processed data in the integrated lake-warehouse forestry point cloud data processing system (i.e., the aforementioned target system).

[0068] By using and optimizing SequenceFile, storage space requirements were significantly reduced, storage efficiency was improved, and data processing efficiency and data analysis response speed were also enhanced.

[0069] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in the embodiments of this application, after processing multiple second point cloud data blocks through each data processing pipeline to obtain forest stand factor data, the method further includes: obtaining intermediate data results output when processing multiple second point cloud data blocks; constructing a data processing lineage based on the intermediate data results, the original point cloud data, and multiple first point cloud data blocks, wherein data tracking is achieved through the data processing lineage.

[0070] In an optional embodiment, a series of intermediate data results are generated when each data processing pipeline processes multiple second point cloud data blocks. These include intermediate results from various stages such as thinning, noise reduction, vegetation point separation, and feature extraction. Based on these intermediate data results, the original point cloud data, and the multiple first point cloud data blocks, a data processing lineage is constructed, enabling the tracking of the complete processing lineage from the original point cloud data to the final stand factor data. It should be noted that the data processing lineage can consist of multiple nodes and edges.

[0071] For example, nodes in the data processing lineage represent various entities in the data processing process, such as the original point cloud data, the thinned point cloud data, and the data after feature extraction. Edges represent the transformation relationships between data entities, recording each step of the operation from the original data to the intermediate results and then to the final stand factor data. For example, the edge from the original point cloud data to the thinned data can record the parameters of the thinning algorithm, such as the grid size.

[0072] By constructing a data processing lineage, each step of data processing becomes visible and traceable, and the data lineage tracking mechanism ensures the accessibility of intermediate data results, providing a foundation for data reuse and multidimensional analysis. For example, if it is necessary to recalculate the stand factors of a certain region, the required intermediate data results can be directly located from the data processing lineage, avoiding redundant steps of processing from scratch and saving time and computing resources.

[0073] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in the embodiments of this application, constructing the data processing lineage based on intermediate data results, original point cloud data, and multiple first point cloud data blocks includes: obtaining first metadata corresponding to the original point cloud data, second metadata corresponding to the multiple first point cloud data blocks, and third metadata corresponding to the intermediate data results; determining the data processing relationship between the first metadata, second metadata, and third metadata; and obtaining the data processing lineage based on the first metadata, second metadata, third metadata, and data processing relationship.

[0074] In an optional embodiment, intermediate data results, raw point cloud data, and first metadata, second metadata, and third metadata corresponding to multiple first point cloud data blocks are obtained. For unstructured point cloud data, the corresponding metadata can be obtained according to the storage path and file name of the point cloud processing process. Then, the data processing relationship between the first metadata, second metadata, and third metadata is determined. For example, the data processing relationship between the first metadata and the second metadata is to thin out the first metadata to obtain the second cloud data.

[0075] In an optional embodiment, nodes in the data processing lineage are constructed based on first metadata, second metadata, and third metadata. Each node represents a state or entity in the data processing flow. For example, raw point cloud data, thinned point cloud data, and data after vegetation point separation can all serve as nodes in the data processing lineage. Edges of the data processing lineage are constructed based on the data processing type corresponding to the metadata, representing the transformation relationships and dependencies between data entities. Edge attributes can include specific data processing operations (such as algorithm parameters for thinning and noise reduction) and contextual information about the transformation process (such as the size and format changes of the data before and after processing). Through the definition of nodes and edges, the data processing lineage for the entire data processing flow is constructed. The data processing lineage ensures the traceability and manageability of the data processing process, allowing each step of processing from raw data to intermediate results and finally to forest stand factor data to be intuitively displayed and queried.

[0076] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in this application embodiment, the data modeling processing of forest stand factor data according to business needs to obtain the target data model includes: determining a first analysis dimension of the forest stand factor data according to business needs; determining a second analysis dimension of the forest stand factor data according to the type of data accessed by the target object; constructing a fact table based on the forest stand factor data; constructing a first dimension table corresponding to the first analysis dimension based on the first analysis dimension and the forest stand factor data; and constructing a second dimension table corresponding to the second analysis dimension based on the second analysis dimension and the forest stand factor data; and obtaining the target data model based on the fact table, the first dimension table, and the second dimension table.

[0077] In an optional embodiment, a first analytical dimension for analyzing stand factor data is determined based on specific business needs. Examples include time dimensions (year, season), geographic dimensions (administrative region, sub-compartment division), and tree species dimensions. A second analytical dimension for the stand factor data is determined based on the type of data accessed by the user (i.e., the target object mentioned above). The second analytical dimension reflects common patterns of data access and use, and can be spatial extent, sampling frequency, or data type (e.g., tree height, crown width, diameter at breast height). By analyzing access patterns, it is possible to predict which dimensions are most frequently queried, thereby optimizing the data model to support rapid responses to such queries.

[0078] Then, a fact table is constructed based on the stand factor data. A first-dimensional table corresponding to the first analytical dimension is constructed based on the first analytical dimension and the stand factor data. Similarly, a second-dimensional table corresponding to the second analytical dimension is constructed based on the second analytical dimension and the stand factor data. The fact table is the core table in the data warehouse, containing key indicators and measures of the stand factor data, such as stand volume, tree height, and crown width. The first-dimensional table is constructed around the first analytical dimension (e.g., time or location), providing detailed information for that dimension, such as specific dates, seasons, geographical coordinates, or administrative divisions, used to refine the data in the fact table to the specific dimension level. The second-dimensional table is constructed based on the second analytical dimension (e.g., access type or data type), covering the information required under different access modes, such as the spatial resolution of the data and the identification of specific tree species, enabling data to be filtered and analyzed from multiple perspectives. Finally, by integrating the fact table and all dimension tables, the target data model described above is obtained.

[0079] By integrating fact tables and dimension tables, the data model supports multidimensional analysis, allowing users to easily gain in-depth insights into forestry point cloud data from multiple perspectives such as time, space, and tree species, thereby improving the flexibility and adaptability of data use.

[0080] In an alternative embodiment, it can be achieved through, as follows: Figure 2 The flowchart shown illustrates the processing of forestry point cloud data based on a lake-warehouse integrated system: Step 1: Receive forestry spatiotemporal data, primarily point cloud data (.las format) and geographic information data (.jason format). Step 2: The distributed computing engine processes the collected data, including the following steps: segmentation, thinning, denoising, normalization, vegetation point separation, feature extraction, and stand factor extraction, ultimately obtaining stand factor data. During segmentation, due to the generation of numerous small files, a bucketing mechanism is required. Simultaneously, based on machine resource allocation, the appropriate size of the point cloud file blocks (the optimal size ratio to fully utilize machine performance) is calculated and dynamically merged. In each data processing step, metadata information and lineage relationships are extracted from the segmented files to achieve metadata management.

[0081] Step 3: Store the extracted stand factor data in a data warehouse, along with basic forestry data. Based on the analysis dimensions, such as time and region, aggregate the stand factor data (crown width, tree height, diameter at breast height, volume, etc.) to create a model. Use a data visualization platform to analyze the stand factors, enabling multidimensional analysis of the stand data across various dimensions.

[0082] In an alternative embodiment, it can be achieved through, as follows: Figure 3The diagram shown illustrates the processing of forestry point cloud data based on the lake-warehouse integration: First, the data is uploaded, storing the received raw point cloud data in the raw data layer.

[0083] Then, the raw point cloud data is processed through a data processing pipeline, including thinning, vegetation point separation, feature extraction, and stand factor extraction to obtain stand factor data. Simultaneously, during point cloud data processing, small files that impact I / O are bucketed (SequenceFile), dynamically segmented, and dynamically merged (using a spatial index-based merging mechanism). The size of the point cloud data is also adjusted based on machine performance to ensure stable computation. Unified modeling of the point cloud data metadata is implemented, with real-time metadata acquisition and data lineage tracking. The thinned and segmented point clouds are stored in the processing data layer using a bucketing mechanism (SequenceFile).

[0084] Finally, a forestry multidimensional analysis warehouse (based on data warehouse multidimensional analysis) is established. Based on stand factor data (region, tree height, crown width, etc.) and combined with analysis scenarios, dimensional modeling is performed to construct a forestry data warehouse, supporting rapid querying and flexible analysis of stand factor data. Stand factors and forestry structured data are stored in the service data layer, enabling multidimensional analysis of stand data based on various dimensions.

[0085] The forestry point cloud data processing method based on lake-database integration provided in this application embodiment acquires the raw point cloud data to be processed in the target area and stores the raw point cloud data in a data lake; processes the raw point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area; performs data modeling processing on the forest stand factor data according to business needs to obtain the target data model, and constructs a forestry data warehouse based on the target data model. The forestry data warehouse is used to provide the required data for the target object, which solves the technical problem of low data processing efficiency caused by manually processing and storing massive amounts of forestry point cloud data in related technologies.

[0086] This solution introduces distributed computing and data processing pipelines to achieve efficient processing and analysis of massive forestry point cloud data, significantly improving data processing efficiency and quality. Leveraging the powerful parallel processing capabilities of the distributed computing engine, raw point cloud data can be quickly processed to obtain corresponding stand factor data, drastically shortening data processing time. By constructing target data models tailored to specific business needs, stand factor data can be rapidly transformed into analytical results from various business perspectives, avoiding repetitive data processing work, reducing data management costs, and ultimately achieving the technical effect of improving overall data utilization.

[0087] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0088] This application also provides a forestry point cloud data processing device based on lake-warehouse integration. It should be noted that this lake-warehouse integrated forestry point cloud data processing device can be used to execute the lake-warehouse integrated forestry point cloud data processing method provided in this application. The following describes the lake-warehouse integrated forestry point cloud data processing device provided in this application.

[0089] According to an embodiment of this application, an apparatus for implementing the above-described forestry point cloud data processing method based on lake-warehouse integration is also provided, such as... Figure 4 As shown, the device includes: a first acquisition unit 401, a first processing unit 402, and a second processing unit 403.

[0090] The first acquisition unit 401 is used to acquire the raw point cloud data to be processed in the target area and store the raw point cloud data in the data lake;

[0091] The first processing unit 402 is used to process the original point cloud data through a distributed computing engine based on multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area.

[0092] The second processing unit 403 is used to perform data modeling processing on forest stand factor data according to business needs, obtain target data model, and construct forestry data warehouse based on target data model. The forestry data warehouse is used to provide the required data for the target object.

[0093] The forestry point cloud data processing device based on the integration of a forestry data warehouse and a data lake provided in this application embodiment acquires the raw point cloud data to be processed in the target area through a first acquisition unit 401 and stores the raw point cloud data in a data lake; a first processing unit 402 processes the raw point cloud data through a distributed computing engine based on multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area; a second processing unit 403 performs data modeling processing on the forest stand factor data according to business needs to obtain a target data model, and constructs a forestry data warehouse based on the target data model. The forestry data warehouse is used to provide the required data for the target object, which solves the technical problem in related technologies where the processing and storage of massive forestry point cloud data is done manually, resulting in low data processing efficiency.

[0094] This solution introduces distributed computing and data processing pipelines to achieve efficient processing and analysis of massive forestry point cloud data, significantly improving data processing efficiency and quality. Leveraging the powerful parallel processing capabilities of the distributed computing engine, raw point cloud data can be quickly processed to obtain corresponding stand factor data, drastically shortening data processing time. By constructing target data models tailored to specific business needs, stand factor data can be rapidly transformed into analytical results from various business perspectives, avoiding repetitive data processing work, reducing data management costs, and ultimately achieving the technical effect of improving overall data utilization.

[0095] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in this application embodiment, the device further includes: a determination unit, used to determine multiple data processing components based on the operation behavior of the target object before processing the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area; a construction unit, used to construct data processing steps for processing point cloud data based on the multiple data processing components; and a configuration unit, used to configure multiple data processing pipelines according to the data processing steps and preset triggering methods.

[0096] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in the embodiments of this application, the device further includes: an optimization unit, used to optimize the original point cloud data before processing it through a distributed computing engine according to multiple preset data processing pipelines, so that the processed original point cloud data supports distributed computing; the first processing unit is also used to process the processed original point cloud data through a distributed computing engine according to multiple data processing pipelines to obtain forest stand factor data corresponding to the target area.

[0097] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in this application embodiment, the first processing unit includes: a cutting module, used to cut the original point cloud data to obtain multiple first point cloud data blocks of a first preset size; a first determining module, used to determine a second preset size corresponding to each data processing pipeline based on the resource allocation information of each data processing pipeline; a merging module, used to merge the multiple first point cloud data blocks according to the second preset size corresponding to each data processing pipeline to obtain multiple second point cloud data blocks of a second preset size corresponding to each data processing pipeline; and a processing module, used to process the multiple second point cloud data blocks through each data processing pipeline to obtain stand factor data.

[0098] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in this application embodiment, the device further includes: a first conversion unit, used to convert the format of the original point cloud data to be processed according to columnar storage after acquiring the original point cloud data to be processed in the target area, to obtain the converted original point cloud data, and to store the converted original point cloud data in the original data layer of the target system; and a first storage unit, used to store the forest stand factor data and the target data model in the service data layer of the target system after performing data modeling processing on the forest stand factor data according to business needs to obtain the target data model, wherein the target object accesses the required data through the service data layer.

[0099] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in this application embodiment, the device further includes: a second conversion unit, used to perform format conversion on the multiple first point cloud data blocks after cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, to obtain multiple first point cloud data blocks in binary format; a third processing unit, used to perform bucketing processing on the multiple first point cloud data blocks in binary format according to the storage space of preset buckets, to obtain processed point cloud data blocks; and a second storage unit, used to store the processed point cloud data blocks to the processing data layer of the target system.

[0100] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in the embodiments of this application, the device further includes: a second acquisition unit, used to acquire intermediate data results output when processing multiple second point cloud data blocks through each data processing pipeline to obtain forest stand factor data; and a construction unit, used to construct a data processing lineage based on the intermediate data results, the original point cloud data, and multiple first point cloud data blocks, wherein data tracking is achieved through the data processing lineage.

[0101] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in this application embodiment, the construction unit includes: an acquisition module, used to acquire the first metadata corresponding to the original point cloud data, the second metadata corresponding to multiple first point cloud data blocks, and the third metadata corresponding to the intermediate data results; a second determination module, used to determine the data processing relationship between the first metadata, the second metadata, and the third metadata; and a third determination module, used to obtain the data processing lineage based on the first metadata, the second metadata, the third metadata, and the data processing relationship.

[0102] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in this application embodiment, the second processing unit includes: a fourth determining module, used to determine a first analysis dimension of the forest stand factor data according to business needs; a fifth determining module, used to determine a second analysis dimension of the forest stand factor data according to the type of data accessed by the target object; a construction module, used to construct a fact table based on the forest stand factor data, construct a first dimension table corresponding to the first analysis dimension based on the first analysis dimension and the forest stand factor data, and construct a second dimension table corresponding to the second analysis dimension based on the second analysis dimension and the forest stand factor data; and a sixth determining module, used to obtain a target data model based on the fact table, the first dimension table, and the second dimension table.

[0103] It should be noted that the first acquisition unit 401, the first processing unit 402, and the second processing unit 403 mentioned above correspond to steps S101 to S103. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules or units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n).

[0104] This application also provides a forestry point cloud data processing system based on lake-warehouse integration. It should be noted that this system can be used to execute the forestry point cloud data processing method based on lake-warehouse integration provided in this application. The following describes the forestry point cloud data processing system based on lake-warehouse integration provided in this application.

[0105] According to embodiments of this application, a system for implementing the above-described forestry point cloud data processing method based on lake-warehouse integration is also provided, such as... Figure 5As shown, the data structure includes: Raw Zone: preserving unprocessed LAS / LAZ format point clouds, optimized for columnar storage (Parquet format conversion), achieving a compression ratio of 2.1:1; Processed Zone: storing point clouds after thinning and segmentation, employing a bucketing mechanism (SequenceFile), improving IO read / write efficiency by 3 times; Serving Zone: containing stand factors and forestry structured data, establishing an OLAP-optimized storage structure. Distributed end-to-end computation is used in the processed data layer (raw point cloud → segmentation → data thinning → noise reduction → normalization → vegetation point separation → feature extraction → stand factor extraction). Stand factors and forestry structured data include a star schema core table and pre-computed high-frequency queries (results such as average tree height in sub-compartments / segmented areas): Fact table: stand factor facts (tree height, diameter at breast height, crown width, etc.); Dimension table: time dimension (year / quarter), spatial dimension (administrative division), tree species dimension (pine / fir / broadleaf, etc.). The processing system supports incremental update mechanisms as well as rapid front-end visualization and multidimensional analysis.

[0106] Embodiments of this application may provide an electronic device. Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 (Only one is shown) processor 602, memory 604, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0107] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the above-described methods. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0108] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: acquire raw point cloud data of the target area to be processed and store the raw point cloud data in a data lake; process the raw point cloud data through a distributed computing engine based on multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area; perform data modeling processing on the forest stand factor data according to business needs to obtain a target data model, and construct a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object.

[0109] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: Before processing the raw point cloud data through the distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area, the method also includes: determining multiple data processing components based on the operation behavior of the target object; constructing data processing steps for processing the point cloud data based on the multiple data processing components; and configuring multiple data processing pipelines according to the data processing steps and preset triggering methods.

[0110] The processor can invoke information and applications stored in the memory via a transmission device to execute the following steps: Before processing the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines, the method further includes: performing data optimization processing on the original point cloud data so that the processed original point cloud data supports distributed computing; processing the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area includes: processing the processed original point cloud data through a distributed computing engine according to multiple data processing pipelines to obtain forest stand factor data corresponding to the target area.

[0111] The processor can access information and applications stored in the memory via a transmission device to execute the following steps: Processing the raw point cloud data through a distributed computing engine based on multiple preset data processing pipelines to obtain stand factor data corresponding to the target area includes: cutting the raw point cloud data to obtain multiple first point cloud data blocks of a first preset size; determining a second preset size corresponding to each data processing pipeline based on the resource allocation information of each data processing pipeline; merging the multiple first point cloud data blocks according to the second preset size corresponding to each data processing pipeline to obtain multiple second point cloud data blocks of the second preset size corresponding to each data processing pipeline; and processing the multiple second point cloud data blocks through each data processing pipeline to obtain stand factor data.

[0112] The processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: After acquiring the raw point cloud data to be processed in the target area, the method further includes: converting the raw point cloud data according to columnar storage to obtain converted raw point cloud data, and storing the converted raw point cloud data in the raw data layer of the target system; After performing data modeling processing on the stand factor data according to business requirements to obtain the target data model, the method further includes: storing the stand factor data and the target data model in the service data layer of the target system, wherein the target object accesses the required data through the service data layer.

[0113] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: after cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, the method further includes: converting the multiple first point cloud data blocks into a format to obtain multiple first point cloud data blocks in binary format; according to the storage space of preset buckets, performing bucketing on the multiple first point cloud data blocks in binary format to obtain processed point cloud data blocks; and storing the processed point cloud data blocks in the processing data layer of the target system.

[0114] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: after processing multiple second point cloud data blocks through each data processing pipeline to obtain forest stand factor data, the method further includes: obtaining intermediate data results output when processing multiple second point cloud data blocks; constructing a data processing lineage based on the intermediate data results, the original point cloud data and multiple first point cloud data blocks, wherein data tracking is achieved through the data processing lineage.

[0115] The processor can invoke information and applications stored in the memory through the transmission device to perform the following steps: constructing a data processing lineage based on intermediate data results, raw point cloud data, and multiple first point cloud data blocks, including: obtaining first metadata corresponding to the raw point cloud data, second metadata corresponding to the multiple first point cloud data blocks, and third metadata corresponding to the intermediate data results; determining the data processing relationship between the first metadata, second metadata, and third metadata; and obtaining the data processing lineage based on the first metadata, second metadata, third metadata, and data processing relationship.

[0116] The processor can access information and applications stored in the memory via a transmission device to execute the following steps: Data modeling of stand factor data based on business requirements to obtain a target data model, including: determining a first analytical dimension of the stand factor data based on business requirements; determining a second analytical dimension of the stand factor data based on the type of data accessed by the target object; constructing a fact table based on the stand factor data; constructing a first dimension table corresponding to the first analytical dimension based on the first analytical dimension and the stand factor data; and constructing a second dimension table corresponding to the second analytical dimension based on the second analytical dimension and the stand factor data; and obtaining the target data model based on the fact table, the first dimension table, and the second dimension table.

[0117] Those skilled in the art will understand that Figure 6 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones (such as Android phones, iOS phones, etc.), tablets, PDAs, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 6 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 6 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 6 The different configurations shown.

[0118] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0119] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the forestry point cloud data processing method based on lake-warehouse integration provided in Embodiment 1.

[0120] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0121] This application also provides a computer program product that, when executed on a data processing device, is suitable for executing the steps of a forestry point cloud data processing method based on lake-warehouse integration.

[0122] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0123] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0124] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0125] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0126] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0127] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0128] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A forestry point cloud data processing method based on lake-warehouse integration, characterized in that, include: Acquire the raw point cloud data to be processed for the target area and store the raw point cloud data in the data lake; The raw point cloud data is processed by a distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area. Based on business needs, the forest stand factor data is processed by data modeling to obtain a target data model. Based on the target data model, a forestry data warehouse is constructed, wherein the forestry data warehouse is used to provide the required data for the target object. Specifically, the raw point cloud data is processed by a distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area, including: The original point cloud data is segmented to obtain multiple first point cloud data blocks of a first preset size; Based on the resource allocation information of each data processing pipeline, determine the second preset size corresponding to each data processing pipeline; Based on the second preset size corresponding to each data processing pipeline, the plurality of first point cloud data blocks are merged to obtain a plurality of second point cloud data blocks of the second preset size corresponding to each data processing pipeline; The forest stand factor data is obtained by processing the plurality of second point cloud data blocks through each data processing pipeline; After segmenting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, the method further includes: The multiple first point cloud data blocks are converted to obtain multiple first point cloud data blocks in binary format; Based on the preset storage space of the buckets, the multiple first point cloud data blocks in the binary format are bucketed to obtain the processed point cloud data blocks. The processed point cloud data blocks are stored in the processing data layer of the target system.

2. The method according to claim 1, characterized in that, Before processing the raw point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area, the method further includes: Based on the operational behavior of the target object, multiple data processing components are identified; Based on the aforementioned multiple data processing components, a data processing step for processing point cloud data is constructed; The multiple data processing pipelines are obtained by configuring the data processing steps and preset triggering methods.

3. The method according to claim 1, characterized in that, Before processing the raw point cloud data through a distributed computing engine based on multiple preset data processing pipelines, the method further includes: The raw point cloud data is optimized to enable distributed computing. The raw point cloud data is processed by a distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area, including: The distributed computing engine processes the raw point cloud data according to the multiple data processing pipelines to obtain the forest stand factor data corresponding to the target area.

4. The method according to claim 1, characterized in that, After acquiring the raw point cloud data to be processed for the target area, the method further includes: The original point cloud data is converted according to columnar storage to obtain converted original point cloud data, and the converted original point cloud data is stored in the original data layer of the target system. After performing data modeling processing on the forest stand factor data according to business needs to obtain the target data model, the method further includes: The forest stand factor data and the target data model are stored in the service data layer of the target system, wherein the target object accesses the required data through the service data layer.

5. The method according to claim 1, characterized in that, After processing the plurality of second point cloud data blocks through each data processing pipeline to obtain the forest stand factor data, the method further includes: Obtain intermediate data results output when processing the plurality of second point cloud data blocks; Based on the intermediate data results, the original point cloud data, and the multiple first point cloud data blocks, a data processing lineage is constructed, wherein the data is tracked through the data processing lineage.

6. The method according to claim 5, characterized in that, Based on the intermediate data results, the original point cloud data, and the plurality of first point cloud data blocks, the data processing lineage is constructed as follows: Obtain the first metadata corresponding to the original point cloud data, the second metadata corresponding to the plurality of first point cloud data blocks, and the third metadata corresponding to the intermediate data results; Determine the data processing relationship between the first metadata, the second metadata, and the third metadata; The data processing lineage is obtained based on the first metadata, the second metadata, the third metadata, and the data processing relationship.

7. The method according to claim 1, characterized in that, Based on business requirements, the forest stand factor data is processed through data modeling to obtain the target data model, which includes: Based on the aforementioned business requirements, the first analytical dimension for the forest stand factor data is determined; Based on the type of data accessed by the target object, a second analytical dimension for the forest stand factor data is determined; A fact table is constructed based on the forest stand factor data; a first dimension table corresponding to the first analysis dimension is constructed based on the first analysis dimension and the forest stand factor data; and a second dimension table corresponding to the second analysis dimension is constructed based on the second analysis dimension and the forest stand factor data. Based on the fact table, the first dimension table, and the second dimension table, the target data model is obtained.

8. A forestry point cloud data processing device based on lake-warehouse integration, characterized in that, include: The first acquisition unit is used to acquire the raw point cloud data to be processed in the target area and store the raw point cloud data in the data lake; The first processing unit is used to process the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area. The second processing unit is used to perform data modeling processing on the forest stand factor data according to business needs to obtain a target data model, and to construct a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object; The first processing unit includes: a cutting module for cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size; a first determining module for determining a second preset size corresponding to each data processing pipeline based on the resource allocation information of each data processing pipeline; a merging module for merging the multiple first point cloud data blocks according to the second preset size corresponding to each data processing pipeline to obtain multiple second point cloud data blocks of a second preset size corresponding to each data processing pipeline; and a processing module for processing the multiple second point cloud data blocks through each data processing pipeline to obtain the forest stand factor data. The device further includes: a second conversion unit, configured to, after cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, convert the multiple first point cloud data blocks into a binary format; a third processing unit, configured to, according to the storage space of preset buckets, perform bucketing on the multiple first point cloud data blocks in binary format to obtain processed point cloud data blocks; and a second storage unit, configured to store the processed point cloud data blocks in the processing data layer of the target system.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device where the computer-readable storage medium is located to perform the forestry point cloud data processing method based on lake-warehouse integration as described in any one of claims 1 to 7.

10. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program executes the forestry point cloud data processing method based on lake-warehouse integration as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Point cloud compression storage method and device based on object storage

    CN116095181A

  • Method and apparatus for recovering point cloud data

    US20190206071A1