Forestry point cloud data processing method and device based on lake and warehouse integration

Through a integrated method based on the lake and warehouse, a distributed computing engine and data processing pipeline are used to process forestry point cloud data, generate forest stand factor data and build digital warehouses, solving the problem of low processing efficiency of massive forestry point cloud data, and achieving efficient analysis and data utilization improvement.

CN120523883AActive Publication Date: 2025-08-22JIULING (SHANGHAI) INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510533067.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-22
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

In the prior art, the efficiency of processing and storing massive forestry point cloud data through manual methods is low, resulting in low data analysis and reuse rates, which cannot meet the changing business needs and in-depth analysis.

Method used

The integrated method based on the lake and warehouse is adopted to process the original point cloud data through a distributed computing engine and multiple data processing pipelines, generate the forest stand factor data, and build a forestry data warehouse to support specific business needs.

Benefits of technology

It significantly improves data processing efficiency and quality, shortens processing time, improves the overall utilization rate of data, reduces management costs, and meets variable business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523883A_ABST
    Figure CN120523883A_ABST
Patent Text Reader

Abstract

The invention discloses a lake and warehouse integration-based forestry point cloud data processing method and device. Relates to the technical field of data processing, and the method comprises the steps: obtaining to-be-processed original point cloud data of a target area, and storing the original point cloud data in a data lake; processing the original point cloud data through a distributed computing engine according to a plurality of preset data processing pipelines to obtain forest stand factor data corresponding to the target area; data modeling processing is conducted on the forest stand factor data according to business requirements, a target data model is obtained, a forestry data warehouse is constructed according to the target data model, and the forestry data warehouse is used for providing needed data for a target object. Through the forestry point cloud data processing method and device, the problem that the data processing efficiency is low due to the fact that massive forestry point cloud data are processed and stored in a manual mode in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method and device for processing forestry point cloud data based on lake-warehouse integration. Background Art

[0002] In the field of forestry remote sensing information processing, with the widespread application of LiDAR technology, the scale of forestry point cloud data collection has exploded, with data volumes ranging from terabytes to petabytes becoming the norm. The storage and management of massive point cloud data still relies on traditional file systems. Existing forestry data warehouses lack multidimensional modeling and analysis capabilities for point cloud data, resulting in low data analysis and reuse rates, making them unable to meet changing business needs and in-depth analysis.

[0003] Currently, no effective solution has been proposed to the problem that related technologies use manual methods to process and store massive forestry point cloud data, resulting in relatively low data processing efficiency. Summary of the Invention

[0004] The main purpose of this application is to provide a forestry point cloud data processing method and device based on lake-warehouse integration, storage medium and electronic equipment, so as to solve the problem in related technologies of manually processing and storing massive forestry point cloud data, resulting in relatively low data processing efficiency.

[0005] To achieve the above objectives, according to one aspect of the present application, a forestry point cloud data processing method based on a lake-warehouse integration is provided. The method comprises: obtaining raw point cloud data to be processed in a target area and storing the raw point cloud data in a data lake; processing the raw point cloud data using a distributed computing engine according to multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area; performing data modeling on the forest stand factor data based on business needs to obtain a target data model, and constructing a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object.

[0006] Furthermore, before the original point cloud data is processed by a distributed computing engine according to a plurality of preset data processing pipelines to obtain the forest stand factor data corresponding to the target area, the method also includes: determining a plurality of data processing components according to the operation behavior of the target object; constructing data processing steps for processing the point cloud data according to the plurality of data processing components; and configuring according to the data processing steps and a preset triggering method to obtain the plurality of data processing pipelines.

[0007] Furthermore, before processing the original point cloud data through the distributed computing engine according to multiple preset data processing pipelines, the method also includes: performing data optimization processing on the original point cloud data so that the processed original point cloud data supports distributed computing; processing the original point cloud data through the distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area includes: processing the processed original point cloud data through the distributed computing engine according to the multiple data processing pipelines to obtain the forest stand factor data corresponding to the target area.

[0008] Furthermore, the original point cloud data is processed by a distributed computing engine according to a plurality of preset data processing pipelines to obtain the forest stand factor data corresponding to the target area, including: cutting the original point cloud data to obtain a plurality of first point cloud data blocks of a first preset size; determining a second preset size corresponding to each data processing pipeline according to resource allocation information of each data processing pipeline; merging the plurality of first point cloud data blocks according to the second preset size corresponding to each data processing pipeline to obtain a plurality of second point cloud data blocks of a second preset size corresponding to each data processing pipeline; and processing the plurality of second point cloud data blocks through each data processing pipeline to obtain the forest stand factor data.

[0009] Furthermore, after obtaining the original point cloud data to be processed in the target area, the method also includes: performing format conversion on the original point cloud data based on column storage to obtain the converted original point cloud data, and storing the converted original point cloud data in the original data layer of the target system; after performing data modeling processing on the forest stand factor data according to business needs to obtain the target data model, the method also includes: storing the forest stand factor data and the target data model in the service data layer of the target system, wherein the target object accesses the required data through the service data layer.

[0010] Furthermore, after the original point cloud data is cut and processed to obtain multiple first point cloud data blocks of a first preset size, the method also includes: performing format conversion on the multiple first point cloud data blocks to obtain multiple first point cloud data blocks in binary format; performing bucket processing on the multiple first point cloud data blocks in binary format according to the storage space of the preset bucket to obtain processed point cloud data blocks; and storing the processed point cloud data blocks in the processing data layer of the target system.

[0011] Furthermore, after processing the multiple second point cloud data blocks through each data processing pipeline to obtain the forest stand factor data, the method also includes: obtaining the intermediate data results output when processing the multiple second point cloud data blocks; constructing data processing lineage based on the intermediate data results, the original point cloud data and the multiple first point cloud data blocks, wherein data tracking is achieved through the data processing lineage.

[0012] Furthermore, constructing data processing lineage based on the intermediate data results, the original point cloud data and the multiple first point cloud data blocks includes: obtaining first metadata corresponding to the original point cloud data, second metadata corresponding to the multiple first point cloud data blocks and third metadata corresponding to the intermediate data results; determining a data processing relationship between the first metadata, the second metadata and the third metadata; and obtaining the data processing lineage based on the first metadata, the second metadata, the third metadata and the data processing relationship.

[0013] Furthermore, data modeling processing is performed on the stand factor data according to business needs to obtain a target data model, including: determining a first analysis dimension of the stand factor data according to the business needs; determining a second analysis dimension of the stand factor data according to the type of data accessed by the target object; constructing a fact table according to the stand factor data, constructing a first dimension table corresponding to the first analysis dimension according to the first analysis dimension and the stand factor data, and constructing a second dimension table corresponding to the second analysis dimension according to the second analysis dimension and the stand factor data; and obtaining the target data model based on the fact table, the first dimension table and the second dimension table.

[0014] To achieve the above-mentioned purpose, according to another aspect of the present application, a forestry point cloud data processing device based on a lake-warehouse integration is provided. The device comprises: a first acquisition unit, configured to acquire the raw point cloud data to be processed in the target area and store the raw point cloud data in a data lake; a first processing unit, configured to process the raw point cloud data using a distributed computing engine according to a plurality of preset data processing pipelines to obtain forest stand factor data corresponding to the target area; a second processing unit, configured to perform data modeling on the forest stand factor data according to business requirements to obtain a target data model, and to construct a forestry data warehouse based on the target data model, wherein the forestry data warehouse is configured to provide the required data for the target object.

[0015] Furthermore, the device also includes: a determination unit, which is used to determine multiple data processing components based on the operation behavior of the target object before processing the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area; a construction unit, which is used to construct data processing steps for processing the point cloud data based on the multiple data processing components; and a configuration unit, which is used to configure according to the data processing steps and a preset trigger method to obtain the multiple data processing pipelines.

[0016] Furthermore, the device also includes: an optimization unit, which is used to perform data optimization processing on the original point cloud data before processing the original point cloud data through the distributed computing engine according to multiple preset data processing pipelines, so that the processed original point cloud data supports distributed computing; the first processing unit is also used to process the processed original point cloud data through the distributed computing engine according to the multiple data processing pipelines to obtain the forest stand factor data corresponding to the target area.

[0017] Furthermore, the first processing unit includes: a cutting module, used to cut the original point cloud data to obtain multiple first point cloud data blocks of a first preset size; a first determination module, used to determine the second preset size corresponding to each data processing pipeline based on the resource allocation information of each data processing pipeline; a merging module, used to merge the multiple first point cloud data blocks according to the second preset size corresponding to each data processing pipeline to obtain multiple second point cloud data blocks of the second preset size corresponding to each data processing pipeline; a processing module, used to process the multiple second point cloud data blocks through each data processing pipeline to obtain the forest stand factor data.

[0018] Furthermore, the device also includes: a first conversion unit, which is used to, after obtaining the original point cloud data to be processed in the target area, convert the format of the original point cloud data according to column storage to obtain the converted original point cloud data, and store the converted original point cloud data in the original data layer of the target system; a first storage unit, which is used to perform data modeling processing on the forest stand factor data according to business needs, and after obtaining the target data model, store the forest stand factor data and the target data model in the service data layer of the target system, wherein the target object accesses the required data through the service data layer.

[0019] Furthermore, the device also includes: a second conversion unit, which is used to perform format conversion on the multiple first point cloud data blocks after cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, so as to obtain multiple first point cloud data blocks in binary format; a third processing unit, which is used to perform bucket processing on the multiple first point cloud data blocks in binary format according to the storage space of the preset buckets, so as to obtain processed point cloud data blocks; and a second storage unit, which is used to store the processed point cloud data blocks in the processing data layer of the target system.

[0020] Furthermore, the device also includes: a second acquisition unit, used to obtain the intermediate data results output when processing the multiple second point cloud data blocks after processing the multiple second point cloud data blocks through each data processing pipeline to obtain the forest stand factor data; a construction unit, used to construct data processing lineage based on the intermediate data results, the original point cloud data and the multiple first point cloud data blocks, wherein data tracking is achieved through the data processing lineage.

[0021] Furthermore, the construction unit includes: an acquisition module, used to obtain first metadata corresponding to the original point cloud data, second metadata corresponding to the multiple first point cloud data blocks, and third metadata corresponding to the intermediate data results; a second determination module, used to determine the data processing relationship between the first metadata, the second metadata and the third metadata; a third determination module, used to obtain the data processing lineage based on the first metadata, the second metadata, the third metadata and the data processing relationship.

[0022] Furthermore, the second processing unit includes: a fourth determination module, used to determine the first analysis dimension of the stand factor data based on the business needs; a fifth determination module, used to determine the second analysis dimension of the stand factor data based on the type of data accessed by the target object; a construction module, used to construct a fact table based on the stand factor data, construct a first dimension table corresponding to the first analysis dimension based on the first analysis dimension and the stand factor data, and construct a second dimension table corresponding to the second analysis dimension based on the second analysis dimension and the stand factor data; a sixth determination module, used to obtain the target data model based on the fact table, the first dimension table and the second dimension table.

[0023] According to another aspect of an embodiment of the present invention, an electronic device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein when the program is running, any one of the above-mentioned forestry point cloud data processing methods based on lake-warehouse integration is executed.

[0024] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is also provided, which stores a program, wherein when the program is running, the device where the storage medium is located is controlled to execute any one of the above-mentioned forestry point cloud data processing methods based on lake-warehouse integration.

[0025] In an embodiment of the present application, the following steps are adopted: obtaining the original point cloud data to be processed in the target area, and storing the original point cloud data in a data lake; processing the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area; performing data modeling processing on the forest stand factor data according to business needs to obtain a target data model, and constructing a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object, solving the technical problem in the related technology of manually processing and storing massive forestry point cloud data, resulting in relatively low data processing efficiency.

[0026] This solution, through the introduction of distributed computing and data processing pipelines, enables efficient processing and analysis of massive amounts of forestry point cloud data, significantly improving data processing efficiency and quality. Leveraging the powerful parallel processing capabilities of the distributed computing engine, raw point cloud data can be rapidly processed to obtain corresponding stand factor data, significantly reducing data processing time. By building target data models tailored to specific business needs, stand factor data can be quickly converted into analytical results from various business perspectives, avoiding repetitive data processing and reducing data management costs, thereby achieving the technical effect of improving overall data utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0028] Figure 1 This is the process of the forestry point cloud data processing method based on lake warehouse integration provided in the embodiment of this application Figure 1 ;

[0029] Figure 2 This is the process of the forestry point cloud data processing method based on lake warehouse integration provided in the embodiment of this application Figure 2 ;

[0030] Figure 3 Schematic diagram of a forestry point cloud data processing method based on lake-warehouse integration provided in an embodiment of the present application;

[0031] Figure 4 2 is a schematic diagram of a forestry point cloud data processing device based on lake-warehouse integration according to an embodiment of the present application;

[0032] Figure 5 Schematic diagram of a forestry point cloud data processing system based on lake and warehouse integration according to an embodiment of the present application;

[0033] Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0034] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0036] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation portals for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation portals for users to choose to agree or refuse the automated decision-making results; if the user chooses to refuse, the expert decision-making process will be entered.

[0037] According to an embodiment of the present application, an embodiment of a forestry point cloud data processing method based on lake-warehouse integration is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0038] This application provides Figure 1 The forestry point cloud data processing method based on lake-warehouse integration is shown. Figure 1 This is the process of the forestry point cloud data processing method based on lake warehouse integration according to Example 1 of this application Figure 1 The treatment method includes:

[0039] Step S101: Obtain the original point cloud data to be processed in the target area and store the original point cloud data in the data lake.

[0040] Alternatively, to define the specific geographic scope of the analysis (i.e., the target area mentioned above), airborne laser ranging (LiDAR) can be used to collect raw point cloud data of the target area to be processed. Airborne LiDAR creates a highly detailed spatial dataset by emitting laser pulses and measuring the reflection time, including information such as tree location, height, and crown width, forming point cloud data. Raw point cloud data is typically stored in .LAS or .LAZ file formats. The collected data can be stored directly in a data lake so that the raw point cloud data can be retrieved from the data lake later.

[0041] Step S102 : Processing the original point cloud data by a distributed computing engine according to a plurality of preset data processing pipelines to obtain forest stand factor data corresponding to the target area.

[0042] Optionally, after obtaining the above-mentioned raw point cloud data, multiple data processing pipelines can be configured to improve the flexibility of data processing configuration and realize full-process automation of data calculation. The raw point cloud data is distributed to multiple preset data processing pipelines through a distributed computing engine for data processing, such as point cloud thinning, point cloud denoising, point cloud normalization, vegetation point separation, feature calculation, etc., to obtain the forest stand factor data corresponding to the target area. Forest stand factor data includes but is not limited to tree height, crown area, trunk diameter, forest density, etc. These forest stand factors are key indicators that describe forest structure and health status, and are crucial for the quantitative analysis of forest resources.

[0043] In an optional embodiment, the distributed computing engine may select a high-performance distributed computing engine to implement parallel processing of large-scale data.

[0044] Step S103: perform data modeling on the stand factor data according to business needs to obtain a target data model, and build a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object.

[0045] Optionally, determine the current business needs, for example, whether it is necessary to estimate forest carbon reserves, monitor forest growth status, or assess forest health status. Different business needs determine the construction direction and key indicators of the data model. Based on business needs, perform dimensional modeling on the stand factor data, combine the stand factor data with dimensions such as time, space, and tree species, and construct a star model or snowflake model (i.e., the target data model mentioned above). For example, to evaluate the carbon sequestration potential of forests, a multidimensional data model can be constructed with "carbon reserves" as the fact table and time, space, tree species, etc. as dimension tables. Based on the constructed target data model, a forestry data warehouse is constructed so that users (i.e., the target objects mentioned above) can develop and implement specific business applications based on the forestry data warehouse, such as forest resource management platforms and ecological monitoring systems. These applications can directly extract the required data from the data model and perform real-time or historical analysis to provide intuitive and accurate data support for decision makers.

[0046] In summary, the introduction of distributed computing and data processing pipelines enables efficient processing and analysis of massive amounts of forestry point cloud data, significantly improving data processing efficiency and quality. Leveraging the powerful parallel processing capabilities of the distributed computing engine, raw point cloud data can be rapidly processed to obtain corresponding stand factor data, significantly reducing data processing time. By building target data models tailored to specific business needs, stand factor data can be quickly converted into analytical results from various business perspectives, avoiding repetitive data processing and reducing data management costs, thereby achieving the technical effect of improving overall data utilization.

[0047] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in an embodiment of the present application, before the original point cloud data is processed by a distributed computing engine according to multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area, the method also includes: determining multiple data processing components based on the operation behavior of the target object; constructing data processing steps for processing point cloud data based on the multiple data processing components; and configuring according to the data processing steps and preset triggering methods to obtain multiple data processing pipelines.

[0048] In an optional embodiment, the forestry point cloud data processing method based on the lake-warehouse integration provided in the embodiment of the present application can be used in a forestry point cloud data processing system based on the lake-warehouse integration, in which 8 key processing components can be pre-set, including thinning, cutting, feature extraction, etc. These components cover most of the core functions from point cloud data preprocessing to forest stand factor extraction. The advantage of preset components is that they are carefully designed and optimized, and can quickly respond to common data processing needs, reducing the user's time and resource investment in component research and development. In addition, the processing system supports users to self-write and upload data processing components of their own design according to specific business needs or preferences.

[0049] This processing system provides an intuitive front-end interface where users (i.e., the aforementioned target objects) can construct data processing steps by dragging and dropping pre-set or custom processing components. The front-end interface displays the data processing steps as a directed acyclic graph. Consequently, multiple data processing components can be determined based on the target object's operational behavior (e.g., a drag operation), and data processing steps for processing point cloud data can be constructed based on these multiple data processing components. Finally, multiple data processing pipelines are configured based on the data processing steps and preset triggering methods.

[0050] The processing system supports two primary triggering mechanisms: event-driven and scheduled tasks. The event-driven mechanism allows data processing to begin immediately upon a certain action in the management interface, such as automatically starting data processing upon the upload of a new batch of point cloud data. Scheduled tasks, on the other hand, provide scheduling configurations based on Cron expressions, allowing users to pre-set data processing cycles, such as executing specific data processing processes at fixed times daily or weekly, ensuring timely data updates and analysis.

[0051] It should be noted that when a data processing component or stage fails, the processing system automatically executes a retry mechanism. The preset maximum number of retries is three, but users can customize the number of retries based on their specific circumstances to ensure the continuity and reliability of the data processing process. Furthermore, regardless of the success or failure of any data processing step, the processing system promptly notifies users of the results via email or message notifications, facilitating real-time monitoring of data processing status and identifying and resolving issues. After the data processing process completes, the system automatically verifies the correctness of the results to ensure that the generated stand factor data meets the expected quality standards, thus avoiding analytical bias caused by data errors.

[0052] Through component-based architecture and custom component upload, users can flexibly configure data processing flows according to different business scenarios and requirements. The combination of a visual DAG editor and a trigger mechanism makes data processing configuration both intuitive and efficient. Users can complete the design and scheduling of data processing flows without writing complex script code, thereby achieving the technical effect of improving data processing efficiency.

[0053] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in an embodiment of the present application, before the original point cloud data is processed by a distributed computing engine according to a preset multiple data processing pipelines, the method also includes: performing data optimization processing on the original point cloud data so that the processed original point cloud data supports distributed computing; processing the original point cloud data by a distributed computing engine according to a preset multiple data processing pipelines to obtain forest stand factor data corresponding to the target area includes: processing the processed original point cloud data by a distributed computing engine according to multiple data processing pipelines to obtain forest stand factor data corresponding to the target area.

[0054] In an optional embodiment, before the original point cloud data enters the distributed computing engine for processing, in order to ensure that the data can be processed efficiently and stably in a distributed environment, the original point cloud data can be optimized. For example, the original point cloud data can be converted from LAS or LAZ format to a format more suitable for distributed computing, such as Parquet. The original point cloud data can also be subjected to code compression and other processing. For another example, the Python Pdal point cloud processing package is optimized so that it can run on distributed computing based on the distributed computing engine, and the Python open3D point cloud processing package is optimized so that it can run on distributed computing based on the distributed computing engine. After obtaining the processed original point cloud data, the processed original point cloud data is processed by the distributed computing engine based on multiple data processing pipelines to obtain the forest stand factor data corresponding to the target area.

[0055] By pre-optimizing the original point cloud data, it not only provides basic support for distributed computing, but also significantly improves the efficiency and stability of the entire data processing process.

[0056] Optionally, in the lake-warehouse integrated forestry point cloud data processing method provided in an embodiment of the present application, the original point cloud data is processed by a distributed computing engine according to multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area, including: cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size; determining the second preset size corresponding to each data processing pipeline based on the resource allocation information of each data processing pipeline; merging the multiple first point cloud data blocks based on the second preset size corresponding to each data processing pipeline to obtain multiple second point cloud data blocks of the second preset size corresponding to each data processing pipeline; processing the multiple second point cloud data blocks through each data processing pipeline to obtain forest stand factor data.

[0057] In an optional embodiment, the original point cloud data is first segmented into a predetermined first size (e.g., 25m*25m) to obtain a plurality of first point cloud data blocks. This process helps break down large files into smaller blocks, facilitating subsequent parallel processing. It should be noted that the first predetermined size can be set according to actual needs.

[0058] Then, based on the actual performance of each computing node in the distributed cluster (such as hardware indicators such as CPU and memory), the resource allocation information of each pipeline (i.e., the above-mentioned data processing pipeline) is determined. Then, based on the resource allocation information, the data block size that is most suitable for processing in each pipeline can be determined, i.e., the second preset size, to ensure the reasonable allocation and maximum utilization of resources.

[0059] After obtaining the second preset size, the multiple first point cloud data blocks are merged to form a second point cloud data block suitable for each data processing pipeline. It should be noted that the merging strategy must consider spatial continuity and the performance of the computing nodes to ensure that the size of the merged data block meets the computing requirements while not being too large to cause memory overflow or too small to cause low computing efficiency. Finally, the multiple second point cloud data blocks are processed through each data processing pipeline, for example, thinning, denoising, vegetation point separation, feature extraction, etc., to ultimately obtain the forest stand factor data.

[0060] The above steps effectively avoid the waste of computing resources caused by data block size mismatch, and greatly improve the efficiency and response speed of distributed computing.

[0061] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in an embodiment of the present application, after obtaining the original point cloud data to be processed in the target area, the method further includes: converting the format of the original point cloud data based on columnar storage to obtain the converted original point cloud data, and storing the converted original point cloud data in the original data layer of the target system; after performing data modeling processing on the forest stand factor data according to business needs to obtain the target data model, the method further includes: storing the forest stand factor data and the target data model in the service data layer of the target system, wherein the target object accesses the required data through the service data layer.

[0062] In an optional embodiment, after obtaining the raw point cloud data to be processed in the target area, the format of the raw point cloud data is first converted according to the requirements of column storage. Common raw point cloud data is stored in LAS / LAZ format, but in order to take advantage of the advantages of column storage for efficient query and calculation, it needs to be converted to a column storage format, such as Parquet. Then, the converted raw point cloud data is stored in the raw data layer of the target system. In forestry analysis, it is often necessary to perform statistical analysis on a certain column of data (such as tree height, crown width). Column storage can directly read the data of the required column, reduce I / O operations, and significantly improve query efficiency. In addition, column storage encodes and compresses data in the same column, which can achieve a higher compression ratio and reduce storage costs. This is especially obvious for TB-PB level data. The column storage structure is easier to implement parallel reading and writing in a distributed computing environment, which helps to execute parallel computing tasks efficiently.

[0063] In an optional embodiment, after the distributed computing engine completes the processing of the point cloud data and generates the target data model, the forest stand factor data and the target data model are stored in the service data layer of the forestry point cloud data processing system based on the lake-warehouse integration (i.e., the above-mentioned target system). This layer is mainly used to provide fast access and multi-dimensional analysis capabilities. The service data layer usually adopts an OLAP (online analytical processing) storage structure to facilitate users to perform fast queries and data analysis. It should be noted that the above-mentioned forestry data warehouse can be directly deployed in the service data layer.

[0064] The introduction of columnar storage significantly improves the storage efficiency of point cloud data, reduces storage space requirements, and increases data reading and query speed. Through the forest stand factor data and data models stored in the service data layer, users can flexibly perform multidimensional analysis, improve data reuse, reduce the cost of repeated analysis, and meet the needs of different business scenarios and analysts.

[0065] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in an embodiment of the present application, after cutting and processing the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, the method also includes: format conversion of the multiple first point cloud data blocks to obtain multiple first point cloud data blocks in binary format; bucketing the multiple first point cloud data blocks in binary format according to the storage space of the preset bucket to obtain processed point cloud data blocks; and storing the processed point cloud data blocks in the processing data layer of the target system.

[0066] In an optional embodiment, during the point cloud data processing process, the cutting process generates a large number of small files. These small files can seriously affect the efficiency of IO operations and hinder the efficiency of parallel computing. Therefore, after the original point cloud data is cut and processed to obtain multiple first point cloud data blocks of a first preset size, these point cloud data blocks are further converted into a binary format. Generally, binary formats refer to more efficient data storage formats, such as Sequence Files, which have significant advantages over the original LAS / LAZ format or text format in terms of read and write speed, memory usage, and network transmission efficiency.

[0067] According to the preset storage bucket size, the first point cloud data block in binary format is bucketed. It should be noted that the preset bucket size can be designed based on the cluster storage capacity and computing efficiency requirements. Multiple first point cloud data blocks in binary format in the same bucket will be merged into a larger file, thereby achieving the purpose of reducing the number of small files. Finally, the processed point cloud data blocks that have been merged and optimized by SequenceFile are stored in the processing data layer of the target system. The processing data layer is an intermediate layer for storing processed data in the forestry point cloud data processing system based on the lake and warehouse integration (i.e., the target system mentioned above).

[0068] By using and optimizing SequenceFile, storage space requirements are significantly reduced, storage efficiency is improved, and data processing efficiency and data analysis response speed are increased.

[0069] Optionally, in the lake-warehouse integrated forestry point cloud data processing method provided in an embodiment of the present application, after processing multiple second point cloud data blocks through each data processing pipeline to obtain forest stand factor data, the method also includes: obtaining the intermediate data results output when processing multiple second point cloud data blocks; constructing data processing lineage based on the intermediate data results, the original point cloud data and multiple first point cloud data blocks, wherein data tracking is achieved through data processing lineage.

[0070] In an optional embodiment, when each data processing pipeline processes multiple second point cloud data blocks, a series of intermediate data results are generated. For example, intermediate results output from various stages such as thinning, noise reduction, vegetation point separation, and feature extraction are generated. Based on the intermediate data results, the original point cloud data, and the multiple first point cloud data blocks, a data processing lineage is constructed, thereby tracing the complete data processing lineage from the original point cloud data to the final forest stand factor data. It should be noted that the data processing lineage can be composed of multiple nodes and edges.

[0071] For example, nodes in a data processing lineage represent entities in the data processing process, such as the original point cloud data, the thinned point cloud data, and the data after feature extraction. Edges represent the transformation relationships between data entities, recording each step from the original data to the intermediate results and finally to the final stand factor data. For example, the edge from the original point cloud data to the thinned data can record the parameters of the thinning algorithm, such as the grid size.

[0072] By building a data processing lineage, each step of data processing becomes visible and traceable. The data lineage tracking mechanism ensures the accessibility of intermediate data results, providing a foundation for data reuse and multidimensional analysis. For example, if a forest stand factor for a specific area needs to be recalculated, the required intermediate data results can be directly located from the data processing lineage, avoiding the redundant steps of starting from scratch and saving time and computing resources.

[0073] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in an embodiment of the present application, constructing data processing lineage based on intermediate data results, original point cloud data and multiple first point cloud data blocks includes: obtaining first metadata corresponding to the original point cloud data, second metadata corresponding to multiple first point cloud data blocks and third metadata corresponding to the intermediate data results; determining the data processing relationship between the first metadata, the second metadata and the third metadata; and obtaining data processing lineage based on the first metadata, the second metadata, the third metadata and the data processing relationship.

[0074] In an optional embodiment, first metadata, second metadata, and third metadata corresponding to the intermediate data results, the original point cloud data, and multiple first point cloud data blocks are obtained. For unstructured point cloud data, the corresponding metadata can be obtained based on the storage path and file name of the point cloud processing process. Then, a data processing relationship between the first metadata, the second metadata, and the third metadata is determined. For example, the data processing relationship between the first metadata and the second metadata is to thin out the first metadata to obtain the second cloud data.

[0075] In an optional embodiment, the nodes in the data processing lineage are constructed based on the first metadata, the second metadata and the third metadata, and the nodes represent a state or entity in the data processing process. For example, the original point cloud data, the thinned point cloud data, the data after vegetation point separation, etc. can all be used as nodes in the data processing lineage. The edges of the data processing lineage are constructed based on the data processing type corresponding to the metadata, which represents the conversion relationship and dependency between data entities. The attributes of the edges can include specific data processing operations (such as thinning and noise reduction algorithm parameters) and contextual information of the conversion process (such as the size of the data before and after processing, format changes, etc.). Through the definition of nodes and edges, the data processing lineage of the entire data processing flow is constructed. The data processing lineage ensures the traceability and manageability of the data processing process, so that each step of processing from the original data to the intermediate results to the final stand factor data can be intuitively displayed and queried.

[0076] Optionally, in the forestry point cloud data processing method based on lake-warehouse integration provided in an embodiment of the present application, data modeling processing is performed on the stand factor data according to business needs to obtain a target data model including: determining the first analysis dimension of the stand factor data according to business needs; determining the second analysis dimension of the stand factor data according to the type of data accessed by the target object; constructing a fact table based on the stand factor data, constructing a first dimension table corresponding to the first analysis dimension based on the first analysis dimension and the stand factor data, and constructing a second dimension table corresponding to the second analysis dimension based on the second analysis dimension and the stand factor data; obtaining the target data model based on the fact table, the first dimension table and the second dimension table.

[0077] In an optional embodiment, the first analysis dimension for analyzing the stand factor data is determined based on specific business needs. For example, the time dimension (year, season), the geographical location dimension (administrative district, small class division), the tree species dimension, etc. The second analysis dimension of the stand factor data is determined based on the type of data accessed by the user (i.e., the target object mentioned above). The second analysis dimension reflects the common patterns of data access and use, which can be spatial range, sampling frequency, data type (such as tree height, crown width, diameter at breast height), etc. By analyzing the access pattern, it is possible to predict which dimensions are queried most frequently, thereby optimizing the data model to support rapid responses to such queries.

[0078] Next, a fact table is constructed based on the stand factor data. A first dimension table corresponding to the first analysis dimension is constructed based on the first analysis dimension and the stand factor data. Furthermore, a second dimension table corresponding to the second analysis dimension is constructed based on the second analysis dimension and the stand factor data. The fact table is the core table in the data warehouse, containing key indicators and metrics for the stand factor data, such as stand volume, tree height, and crown width. The first dimension table is constructed around the first analysis dimension (such as time or location). It provides detailed information on that dimension, such as specific dates, seasons, geographic coordinates, or administrative divisions, and is used to refine the data in the fact table to the level of specific dimensions. The second dimension table is constructed based on the second analysis dimension (such as access type or data type). It covers the information required under different access modes, such as the spatial resolution of the data and the identification of specific tree species, allowing data to be filtered and analyzed from multiple perspectives. Finally, by integrating the fact table and all dimension tables, the target data model described above is obtained.

[0079] By integrating fact tables and dimension tables, the data model supports multidimensional analysis. Users can easily gain in-depth insights into forestry point cloud data from multiple perspectives such as time, space, and tree species, thereby improving the flexibility and adaptability of data use.

[0080] In an alternative embodiment, the Figure 2 The flowchart shown realizes the processing of forestry point cloud data based on the lake-warehouse integration: Step 1: Receive forestry spatiotemporal data, mainly point cloud data (.las format) and geographic information data (jason format). Step 2: The distributed computing engine processes the collected data, including the following steps: segmentation, thinning, denoising, normalization, vegetation point separation, feature extraction, forest stand factor extraction and other processing operations, and finally obtains the forest stand factor data. Among them, after segmentation, due to the generation of a large number of small files, a bucketing mechanism is required for the data. At the same time, according to the machine resource ratio, the size of the appropriate point cloud file block is calculated (the optimal size ratio that can give full play to the performance of the machine) and dynamically merged; in each step of the data processing, the metadata information of the cut files will be extracted, as well as their blood relationship extraction, to achieve metadata management.

[0081] Step 3: Store the extracted stand factor data in a data warehouse, along with the forestry basic data. Based on the analysis dimensions, such as time and region, aggregate stand factor data (crown width, tree height, DBH, standing volume, etc.) to create a model. Use a data visualization platform to analyze stand factors, enabling multidimensional analysis of stand data based on various dimensions.

[0082] In an alternative embodiment, the Figure 3The schematic diagram shown realizes the processing of forestry point cloud data based on lake-warehouse integration: first, the data is uploaded and the received original point cloud data is stored in the original data layer.

[0083] Then, through the data processing pipeline, the processing of the original point cloud data is realized, for example, thinning, vegetation point separation, feature extraction, and forest stand factor extraction to obtain forest stand factor data. At the same time, in the process of point cloud data processing, small files that affect IO are bucketed (SequenceFile), dynamically segmented, and dynamically merged (merging mechanism based on spatial index). At the same time, the point cloud data size is cut and merged based on the performance of the machine to enable stable calculation. The metadata of the point cloud data is uniformly modeled, metadata is collected in real time, and data lineage is tracked. The point cloud after thinning and segmentation is stored in the processing data layer using a bucketing mechanism (SequenceFile).

[0084] Finally, a forestry multidimensional analysis warehouse (data warehouse-based multidimensional analysis) was established. Based on stand factor data (region, tree height, crown width, etc.), combined with analysis scenarios, dimensional modeling was performed to construct a forestry data warehouse, supporting rapid query and flexible analysis of stand factor data. Stand factors and forestry structured data were stored in the service data layer, enabling multidimensional analysis of stand data based on various dimensions.

[0085] The lake-warehouse integrated forestry point cloud data processing method provided in the embodiment of the present application obtains the original point cloud data to be processed in the target area and stores the original point cloud data in the data lake; processes the original point cloud data according to multiple preset data processing pipelines through a distributed computing engine to obtain the forest stand factor data corresponding to the target area; performs data modeling on the forest stand factor data according to business needs to obtain a target data model, and constructs a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object, solving the technical problem in the related technology of manually processing and storing massive forestry point cloud data, resulting in relatively low data processing efficiency.

[0086] This solution, through the introduction of distributed computing and data processing pipelines, enables efficient processing and analysis of massive amounts of forestry point cloud data, significantly improving data processing efficiency and quality. Leveraging the powerful parallel processing capabilities of the distributed computing engine, raw point cloud data can be rapidly processed to obtain corresponding stand factor data, significantly reducing data processing time. By building target data models tailored to specific business needs, stand factor data can be quickly converted into analytical results from various business perspectives, avoiding repetitive data processing and reducing data management costs, thereby achieving the technical effect of improving overall data utilization.

[0087] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0088] The present application also provides a lake-warehouse integrated forestry point cloud data processing device. It should be noted that the lake-warehouse integrated forestry point cloud data processing device of the present application can be used to execute the lake-warehouse integrated forestry point cloud data processing method provided in the present application. The following describes the lake-warehouse integrated forestry point cloud data processing device provided in the present application.

[0089] According to an embodiment of the present application, a device for implementing the above-mentioned forestry point cloud data processing method based on lake-warehouse integration is also provided, such as Figure 4 As shown, the device includes: a first acquiring unit 401, a first processing unit 402 and a second processing unit 403.

[0090] The first acquisition unit 401 is used to acquire the raw point cloud data to be processed in the target area and store the raw point cloud data in the data lake;

[0091] The first processing unit 402 is configured to process the original point cloud data using a distributed computing engine according to a plurality of preset data processing pipelines to obtain forest stand factor data corresponding to the target area;

[0092] The second processing unit 403 is used to perform data modeling processing on the forest stand factor data according to business needs, obtain a target data model, and build a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object.

[0093] The lake-warehouse integrated forestry point cloud data processing device provided in the embodiment of the present application obtains the original point cloud data to be processed in the target area through the first acquisition unit 401, and stores the original point cloud data in the data lake; the first processing unit 402 processes the original point cloud data according to the preset multiple data processing pipelines through the distributed computing engine to obtain the forest stand factor data corresponding to the target area; the second processing unit 403 performs data modeling processing on the forest stand factor data according to business needs to obtain the target data model, and constructs a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object, which solves the technical problem in the related technology of manually processing and storing massive forestry point cloud data, resulting in relatively low data processing efficiency.

[0094] This solution, through the introduction of distributed computing and data processing pipelines, enables efficient processing and analysis of massive amounts of forestry point cloud data, significantly improving data processing efficiency and quality. Leveraging the powerful parallel processing capabilities of the distributed computing engine, raw point cloud data can be rapidly processed to obtain corresponding stand factor data, significantly reducing data processing time. By building target data models tailored to specific business needs, stand factor data can be quickly converted into analytical results from various business perspectives, avoiding repetitive data processing and reducing data management costs, thereby achieving the technical effect of improving overall data utilization.

[0095] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in an embodiment of the present application, the device also includes: a determination unit for determining multiple data processing components based on the operation behavior of the target object before processing the original point cloud data through a distributed computing engine according to multiple preset data processing pipelines to obtain forest stand factor data corresponding to the target area; a construction unit for constructing data processing steps for processing point cloud data based on multiple data processing components; and a configuration unit for configuring according to the data processing steps and preset triggering methods to obtain multiple data processing pipelines.

[0096] Optionally, in the lake-warehouse integrated forestry point cloud data processing device provided in an embodiment of the present application, the device also includes: an optimization unit, which is used to perform data optimization processing on the original point cloud data before the original point cloud data is processed by a distributed computing engine according to multiple preset data processing pipelines, so that the processed original point cloud data supports distributed computing; the first processing unit is also used to process the processed original point cloud data according to multiple data processing pipelines through a distributed computing engine to obtain forest stand factor data corresponding to the target area.

[0097] Optionally, in the lake-warehouse integrated forestry point cloud data processing device provided in an embodiment of the present application, the first processing unit includes: a cutting module, used to cut and process the original point cloud data to obtain multiple first point cloud data blocks of a first preset size; a first determination module, used to determine the second preset size corresponding to each data processing pipeline based on the resource allocation information of each data processing pipeline; a merging module, used to merge and process the multiple first point cloud data blocks based on the second preset size corresponding to each data processing pipeline to obtain multiple second point cloud data blocks of the second preset size corresponding to each data processing pipeline; a processing module, used to process the multiple second point cloud data blocks through each data processing pipeline to obtain forest stand factor data.

[0098] Optionally, in the forestry point cloud data processing device based on the lake-warehouse integration provided in the embodiment of the present application, the device also includes: a first conversion unit, which is used to convert the format of the original point cloud data to be processed according to column storage after obtaining the original point cloud data to be processed in the target area, to obtain the converted original point cloud data, and store the converted original point cloud data in the original data layer of the target system; a first storage unit, which is used to perform data modeling processing on the forest stand factor data according to business needs, and after obtaining the target data model, store the forest stand factor data and the target data model in the service data layer of the target system, wherein the target object accesses the required data through the service data layer.

[0099] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in an embodiment of the present application, the device also includes: a second conversion unit, for performing format conversion on the multiple first point cloud data blocks after cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, to obtain multiple first point cloud data blocks in binary format; a third processing unit, for performing bucket processing on the multiple first point cloud data blocks in binary format according to the storage space of a preset bucket, to obtain processed point cloud data blocks; a second storage unit, for storing the processed point cloud data blocks to the processing data layer of the target system.

[0100] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in an embodiment of the present application, the device also includes: a second acquisition unit, used to obtain the intermediate data results output when processing multiple second point cloud data blocks after processing multiple second point cloud data blocks through each data processing pipeline to obtain forest stand factor data; a construction unit, used to construct data processing lineage based on the intermediate data results, original point cloud data and multiple first point cloud data blocks, wherein data tracking is achieved through data processing lineage.

[0101] Optionally, in the forestry point cloud data processing device based on lake-warehouse integration provided in an embodiment of the present application, the construction unit includes: an acquisition module for acquiring first metadata corresponding to the original point cloud data, second metadata corresponding to multiple first point cloud data blocks, and third metadata corresponding to the intermediate data results; a second determination module for determining the data processing relationship between the first metadata, the second metadata and the third metadata; a third determination module for obtaining data processing lineage based on the first metadata, the second metadata, the third metadata and the data processing relationship.

[0102] Optionally, in the lake-warehouse integrated forestry point cloud data processing device provided in an embodiment of the present application, the second processing unit includes: a fourth determination module, used to determine the first analysis dimension of the stand factor data based on business needs; a fifth determination module, used to determine the second analysis dimension of the stand factor data based on the type of data accessed by the target object; a construction module, used to construct a fact table based on the stand factor data, construct a first dimension table corresponding to the first analysis dimension based on the first analysis dimension and the stand factor data, and construct a second dimension table corresponding to the second analysis dimension based on the second analysis dimension and the stand factor data; a sixth determination module, used to obtain the target data model based on the fact table, the first dimension table and the second dimension table.

[0103] It should be noted that the first acquisition unit 401, the first processing unit 402, and the second processing unit 403 described above correspond to steps S101 to S103, and the examples and application scenarios implemented by the three units and the corresponding steps are the same, but are not limited to the contents disclosed in the above embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n).

[0104] The present application also provides a lake-warehouse integrated forestry point cloud data processing system. It should be noted that the lake-warehouse integrated forestry point cloud data processing system of the present application can be used to execute the lake-warehouse integrated forestry point cloud data processing method provided in the present application. The following describes the lake-warehouse integrated forestry point cloud data processing system provided in the present application.

[0105] According to an embodiment of the present application, a system for implementing the above-mentioned forestry point cloud data processing method based on lake-warehouse integration is also provided, such as Figure 5As shown, the system includes the raw data layer (Raw Zone), which retains unprocessed point clouds in LAS / LAZ format and uses columnar storage optimization (Parquet format conversion) to achieve a compression ratio of 2.1:1. The processed data layer (Processed Zone) stores thinned and segmented point clouds, using a bucketing mechanism (Sequence File) to improve I / O read and write efficiency by three times. The serving data layer (Serving Zone) stores stand factors and forestry structured data, using an OLAP-optimized storage structure. The processing data layer utilizes a distributed, full-link computing operation (raw point cloud → segmentation → thinned data → noise reduction → normalization → vegetation point separation → feature extraction → stand factor extraction). Stand factor and forestry structured data include star-structured core tables and pre-computed high-frequency query results (such as average tree height for small groups / segments). The fact table contains stand factor facts (such as tree height, diameter at breast height, and crown width); and the dimension tables include time dimensions (year / quarter), spatial dimensions (administrative divisions), and tree species dimensions (pine / fir / broadleaf, etc.). The processing system supports an incremental update mechanism as well as fast front-end visualization and multi-dimensional analysis.

[0106] An embodiment of the present application may provide an electronic device, Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 Only one is shown) processor 602, memory 604, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0107] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0108] The processor can call the information and applications stored in the memory through the transmission device to perform the following steps: obtain the original point cloud data to be processed in the target area, and store the original point cloud data in the data lake; process the original point cloud data through the distributed computing engine according to multiple preset data processing pipelines to obtain the forest stand factor data corresponding to the target area; perform data modeling on the forest stand factor data according to business needs to obtain the target data model, and build a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object.

[0109] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: before the original point cloud data is processed by the distributed computing engine according to the preset multiple data processing pipelines to obtain the forest stand factor data corresponding to the target area, the method also includes: determining multiple data processing components based on the operation behavior of the target object; constructing data processing steps for processing the point cloud data based on the multiple data processing components; configuring according to the data processing steps and the preset triggering method to obtain multiple data processing pipelines.

[0110] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: before the original point cloud data is processed by the distributed computing engine according to the preset multiple data processing pipelines, the method also includes: performing data optimization processing on the original point cloud data so that the processed original point cloud data supports distributed computing; processing the original point cloud data by the distributed computing engine according to the preset multiple data processing pipelines to obtain the forest stand factor data corresponding to the target area includes: processing the processed original point cloud data by the distributed computing engine according to the multiple data processing pipelines to obtain the forest stand factor data corresponding to the target area.

[0111] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: processing the original point cloud data through the distributed computing engine according to the preset multiple data processing pipelines to obtain the forest stand factor data corresponding to the target area, including: cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size; determining the second preset size corresponding to each data processing pipeline based on the resource allocation information of each data processing pipeline; merging the multiple first point cloud data blocks based on the second preset size corresponding to each data processing pipeline to obtain multiple second point cloud data blocks of the second preset size corresponding to each data processing pipeline; processing the multiple second point cloud data blocks through each data processing pipeline to obtain the forest stand factor data.

[0112] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: after obtaining the original point cloud data to be processed in the target area, the method also includes: converting the format of the original point cloud data based on column storage to obtain the converted original point cloud data, and storing the converted original point cloud data in the original data layer of the target system; after performing data modeling processing on the forest stand factor data according to business needs to obtain the target data model, the method also includes: storing the forest stand factor data and the target data model in the service data layer of the target system, wherein the target object accesses the required data through the service data layer.

[0113] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: after cutting the original point cloud data to obtain multiple first point cloud data blocks of a first preset size, the method also includes: converting the format of the multiple first point cloud data blocks to obtain multiple first point cloud data blocks in binary format; bucketing the multiple first point cloud data blocks in binary format according to the storage space of the preset bucket to obtain processed point cloud data blocks; and storing the processed point cloud data blocks in the processing data layer of the target system.

[0114] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: after processing multiple second point cloud data blocks through each data processing pipeline to obtain forest stand factor data, the method also includes: obtaining the intermediate data results output when processing multiple second point cloud data blocks; constructing data processing lineage based on the intermediate data results, original point cloud data and multiple first point cloud data blocks, wherein data tracking is achieved through data processing lineage.

[0115] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: constructing data processing lineage based on the intermediate data results, the original point cloud data and multiple first point cloud data blocks, including: obtaining first metadata corresponding to the original point cloud data, second metadata corresponding to the multiple first point cloud data blocks and third metadata corresponding to the intermediate data results; determining the data processing relationship between the first metadata, the second metadata and the third metadata; obtaining data processing lineage based on the first metadata, the second metadata, the third metadata and the data processing relationship.

[0116] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: perform data modeling processing on the stand factor data according to business needs to obtain the target data model including: determining the first analysis dimension of the stand factor data according to business needs; determining the second analysis dimension of the stand factor data according to the type of data accessed by the target object; constructing a fact table based on the stand factor data, constructing a first dimension table corresponding to the first analysis dimension based on the first analysis dimension and the stand factor data, and constructing a second dimension table corresponding to the second analysis dimension based on the second analysis dimension and the stand factor data; obtaining the target data model based on the fact table, the first dimension table and the second dimension table.

[0117] It can be understood by those skilled in the art that Figure 6 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 6 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 6 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 6 Different configurations shown.

[0118] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0119] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the forestry point cloud data processing method based on lake-warehouse integration provided in the first embodiment.

[0120] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0121] The present application also provides a computer program product which, when executed on a data processing device, is suitable for executing the steps of a forestry point cloud data processing method based on lake-warehouse integration.

[0122] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0123] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0124] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0125] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0126] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0127] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0128] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A forestry point cloud data processing method based on lake-warehouse integration, characterized in that: include: Obtaining the original point cloud data to be processed in the target area, and storing the original point cloud data in the data lake; Processing the original point cloud data through a distributed computing engine according to a plurality of preset data processing pipelines to obtain forest stand factor data corresponding to the target area; The stand factor data is subjected to data modeling processing according to business needs to obtain a target data model, and a forestry data warehouse is constructed based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object.

2. The method according to claim 1, characterized in that Before processing the original point cloud data by a distributed computing engine according to a plurality of preset data processing pipelines to obtain forest stand factor data corresponding to the target area, the method further includes: Determine multiple data processing components based on the target object's operational behavior; constructing data processing steps for processing point cloud data based on the plurality of data processing components; The multiple data processing pipelines are configured according to the data processing steps and the preset triggering method.

3. The method according to claim 1, characterized in that Before processing the raw point cloud data using the distributed computing engine according to the preset multiple data processing pipelines, the method further includes: performing data optimization processing on the original point cloud data so that the processed original point cloud data supports distributed computing; The raw point cloud data is processed by a distributed computing engine according to a plurality of preset data processing pipelines to obtain forest stand factor data corresponding to the target area, including: The processed original point cloud data is processed by the distributed computing engine according to the multiple data processing pipelines to obtain forest stand factor data corresponding to the target area.

4. The method according to claim 2, characterized in that The raw point cloud data is processed by a distributed computing engine according to a plurality of preset data processing pipelines to obtain forest stand factor data corresponding to the target area, including: Performing a cutting process on the original point cloud data to obtain a plurality of first point cloud data blocks of a first preset size; Determining a second preset size corresponding to each data processing pipeline based on resource allocation information of each data processing pipeline; Merging the plurality of first point cloud data blocks according to a second preset size corresponding to each data processing pipeline to obtain a plurality of second point cloud data blocks of the second preset size corresponding to each data processing pipeline; The plurality of second point cloud data blocks are processed through each data processing pipeline to obtain the forest stand factor data.

5. The method according to claim 1, wherein After obtaining the original point cloud data to be processed in the target area, the method further includes: Performing format conversion on the original point cloud data according to column storage to obtain converted original point cloud data, and storing the converted original point cloud data in an original data layer of a target system; After performing data modeling processing on the stand factor data according to business requirements to obtain a target data model, the method further includes: The forest stand factor data and the target data model are stored in the service data layer of the target system, wherein the target object accesses required data through the service data layer.

6. The method according to claim 4, characterized in that After the original point cloud data is segmented to obtain a plurality of first point cloud data blocks of a first preset size, the method further includes: Performing format conversion on the plurality of first point cloud data blocks to obtain a plurality of first point cloud data blocks in a binary format; performing bucket processing on the plurality of first point cloud data blocks in the binary format according to a preset storage space of the bucket to obtain processed point cloud data blocks; The processed point cloud data blocks are stored in the processing data layer of the target system.

7. The method according to claim 4, characterized in that After processing the plurality of second point cloud data blocks through each data processing pipeline to obtain the stand factor data, the method further includes: Acquire intermediate data results output when processing the plurality of second point cloud data blocks; A data processing lineage is constructed based on the intermediate data result, the original point cloud data, and the plurality of first point cloud data blocks, wherein data tracking is achieved through the data processing lineage.

8. The method according to claim 7, characterized in that Constructing a data processing lineage based on the intermediate data result, the original point cloud data, and the plurality of first point cloud data blocks includes: Acquire first metadata corresponding to the original point cloud data, second metadata corresponding to the plurality of first point cloud data blocks, and third metadata corresponding to the intermediate data result; determining a data processing relationship among the first metadata, the second metadata, and the third metadata; The data processing lineage is obtained based on the first metadata, the second metadata, the third metadata and the data processing relationship.

9. The method according to claim 1, characterized in that The stand factor data is processed for data modeling according to business requirements to obtain a target data model including: Determining a first analysis dimension for the stand factor data according to the business requirements; Determining a second analysis dimension for the stand factor data based on the type of the target object access data; Constructing a fact table based on the stand factor data, constructing a first dimension table corresponding to the first analysis dimension based on the first analysis dimension and the stand factor data, and constructing a second dimension table corresponding to the second analysis dimension based on the second analysis dimension and the stand factor data; The target data model is obtained based on the fact table, the first dimension table and the second dimension table.

10. A forestry point cloud data processing device based on lake and warehouse integration, characterized in that: include: A first acquisition unit is configured to acquire raw point cloud data to be processed in a target area and store the raw point cloud data in a data lake; A first processing unit is configured to process the original point cloud data using a distributed computing engine according to a plurality of preset data processing pipelines to obtain forest stand factor data corresponding to the target area; The second processing unit is used to perform data modeling processing on the forest stand factor data according to business needs to obtain a target data model, and to build a forestry data warehouse based on the target data model, wherein the forestry data warehouse is used to provide the required data for the target object.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the forestry point cloud data processing method based on lake-warehouse integration as described in any one of claims 1 to 9.

12. An electronic device, characterized in that: include: a memory storing an executable program; A processor is used to run the program, wherein when the program is running, it executes the forestry point cloud data processing method based on lake-warehouse integration as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Point cloud compression storage method and device based on object storage

    CN116095181A

  • Method and apparatus for recovering point cloud data

    US20190206071A1