Visual point cloud segmentation assembly line construction system and method based on deep learning

The visual point cloud segmentation pipeline constructed through modular design and deep learning technology solves the problems of low data labeling efficiency and poor system adaptability in the existing technology, and realizes an efficient and flexible point cloud segmentation pipeline, improving segmentation accuracy and user interaction experience.

CN120388268APending Publication Date: 2025-07-29BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510150357.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

When the prior art processes large-scale and complex point cloud data, there are problems such as low data labeling efficiency, lack of open source visual point cloud segmentation system and insufficient automation, resulting in high cost of training of segmentation models and poor adaptability, making it difficult to meet the preprocessing needs of different data sets.

Method used

Modular design and deep learning technology are adopted to build a visual point cloud segmentation pipeline based on deep learning, including the front-end display layer, core service layer, storage service layer, resource management layer and infrastructure layer. Through efficient collaboration of data analysis, processing, segmentation and storage modules, combined with dynamic visual modules, real-time feedback is provided to realize automated and flexible point cloud segmentation pipelines.

Benefits of technology

It improves the automation level and processing efficiency of point cloud segmentation, enhances user interaction experience, improves segmentation accuracy and system compatibility, adapts to different data formats and application scenarios, and reduces manual intervention and debugging costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388268A_ABST
    Figure CN120388268A_ABST
Patent Text Reader

Abstract

The invention discloses a visual point cloud segmentation assembly line construction system and method based on deep learning, and the system combines a deep learning model and a dynamic visualization technology, and achieves the automatic segmentation and real-time result correction in point cloud data processing. The traditional point cloud processing flow is broken through in design, the processing precision and efficiency are remarkably improved through the deep learning model, and meanwhile, the dependence on manual annotation is reduced. The system not only optimizes the user interaction experience, but also enhances the flexibility and adaptability, provides powerful support for the accurate application of the point cloud data, and has important technical innovation and wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional data processing, and more particularly to a system and method for constructing a visualization point cloud segmentation pipeline based on deep learning. Background Art

[0002] In recent years, as an important information carrier, three-dimensional point cloud data has shown extensive application value in many fields such as urban modeling, cultural heritage protection, and geographical mapping. However, despite the continuous progress of point cloud segmentation technology, it still faces multiple challenges in processing large-scale and complex-structured point cloud data, and more in-depth research is urgently needed to achieve more efficient segmentation and understanding.

[0003] The traditional point cloud segmentation process is costly and involves multiple aspects such as dataset collection and processing, model environment configuration, and training deployment. In this development mode, the efficiency of segmentation annotation is extremely low. At the same time, due to the differences in data collection devices and application scenarios, the scalability and maintainability of the code are also limited. This results in a long development cycle and high learning cost for three-dimensional point cloud segmentation systems, making it difficult to meet the preprocessing requirements of different datasets.

[0004] Although some current visualization technologies have alleviated the standardization problems of computing resources and heterogeneous environments to a certain extent, there are still a series of unsolved challenges, which are specifically manifested as follows:

[0005] (1) Low data annotation efficiency

[0006] The low data annotation efficiency makes it difficult to train the point cloud segmentation model quickly and accurately. The key to point cloud model training lies in data annotation. However, the annotation process of point cloud data in large-scale scenarios is both cumbersome and time-consuming. Currently, most projects rely on manual annotation, which not only consumes a large amount of manpower and time but also may lead to uneven data quality, affecting the performance of the segmentation model. The limitation of manual annotation is that it is difficult to ensure consistency and efficiency. Therefore, there is an urgent need for an efficient automated pipeline for point cloud segmentation algorithms to simplify the annotation process. Through the automated process, the efficiency and accuracy of data annotation can be significantly improved, human errors can be reduced, a more reliable data foundation can be provided for model training, and the adaptability and stability of the segmentation model on different datasets can be enhanced.

[0007] (2) Lack of open-source visualization point cloud segmentation systems

[0008] Existing visualization systems are difficult to meet the requirements of efficient and flexible point cloud segmentation, and there is a lack of open-source solutions suitable for complex scenarios. Although visualization technology plays an important role in point cloud segmentation, there is still a lack of efficient and open-source 3D point cloud semantic segmentation systems in the market. Existing systems such as CloudCompare, although providing basic point cloud visualization functions, are not designed specifically for segmentation tasks and are difficult to meet the high-precision and high-efficiency requirements when dealing with complex geometries. In addition, many open-source systems only support data files in a single format, limiting their applicability in practical applications. With the diversification of point cloud data sources, users need to process data in different formats, and existing systems are difficult to meet this requirement. Even if some systems provide diversified data processing functions, the components are fixed, lacking flexibility and reusability, and are difficult to adapt to the dynamic model production requirements. In addition, the maintenance and update of some systems are not timely enough, resulting in lagging functions and bringing troubles to users. Therefore, there is an urgent need to develop an interactive visualization 3D point cloud segmentation system to provide higher flexibility and operation control, and better support users' segmentation tasks in different scenarios and data requirements.

[0009] (3) Specific challenges

[0010] When constructing a visualization point cloud segmentation pipeline, the complexity of data processing, the generalization of the model, and the degree of automation are the three core challenges faced.

[0011] First of all, the complexity of data processing is the primary difficulty in pipeline design. As a high-dimensional and large-scale data carrier, point cloud data often contains rich spatial information and irregular structural characteristics. Facing large-scale point cloud data, the system must efficiently process multiple links from data upload, preprocessing, model segmentation to post-processing to ensure the smoothness and stability of the entire process. Due to the large volume and complex structure of the data, it not only increases the processing time but also poses higher requirements for system resource management.

[0012] Secondly, the generalization of the model is crucial. Since point cloud data comes from a wide range of sources, there are often significant differences in the structure, density, and distribution characteristics of different data sets. If the segmentation model cannot flexibly adapt to these characteristics, it will be difficult to maintain accuracy and stability in practical applications. Therefore, it is necessary to introduce a model with good generalization ability in the pipeline to meet the requirements of various data sets, thereby improving the segmentation effect and efficiency.

[0013] Finally, the degree of automation directly affects the usability and productivity of the pipeline. In the traditional segmentation process, data processing, model training, parameter tuning, and other steps rely on a large amount of manual intervention, resulting in low process efficiency. How to improve the degree of automation, reduce manual intervention, and achieve a more efficient point cloud segmentation pipeline has become a key issue. Through system optimization, the full automation of data flow processing can be achieved, which will significantly reduce human interference and improve the system processing efficiency and model consistency. Summary of the Invention

[0014] The present invention provides a method and system for constructing a visual point cloud segmentation pipeline, aiming to improve the automation degree, accuracy, and processing efficiency of point cloud segmentation through modular design and deep learning technology, and at the same time improve the user interaction experience through visualization means.

[0015] A visual point cloud segmentation pipeline construction system based on deep learning includes:

[0016] The front-end display layer provides users with diverse interaction methods, including browser-based WebUI, RESTful interfaces supporting API access, three-dimensional point cloud online rendering based on WebGL, and the Open3D tool on the desktop.

[0017] The core service layer is used for point cloud data processing, including a data parsing module, a data processing module, a deep learning point cloud segmentation module, a manual correction module, a data storage module, and a dynamic visualization module.

[0018] The storage service layer is used for storing and managing point cloud data, including a file system, a MongoDB database, and a cache middleware.

[0019] A point cloud file with a unique identifier corresponds to a set of JSON strings and is stored in the MongoDB database.

[0020] The data storage strategy adopted by this system stores the original data. For the processed data, the invalid points will be deleted to ensure that no invalid data is retained.

[0021] The resource management layer is used for dynamically allocating and managing computing and storage resources, and the computing and storage resources include CPU resources, memory resources, storage resources, network resources, and GPU computing resources.

[0022] The infrastructure layer is used to provide the operating environment for computing resources, including public clouds, virtual machines, and physical machines.

[0023] Through a hierarchical architecture and modular design, the system achieves efficient collaboration among modules while retaining a high degree of flexibility and scalability. Each module can not only operate independently but also be seamlessly integrated through standardized interfaces. The system forms a closed-loop processing flow in each stage from data parsing, enhancement, segmentation, correction to storage and visualization, thus meeting the comprehensive requirements for high performance, ease of use and flexibility in point cloud processing scenarios.

[0024] In terms of the system architecture, it is re-planned and designed rather than directly applying the existing system architecture paradigm.

[0025] In the past data processing process, the implementation of various functions was in a scattered state, with only manual annotation tools and data processing scripts for deep learning models. In this system, these originally independent functions have been organically integrated. For the modules at each level, although the functions of some of them have similar implementation forms in the prior art, targeted improvement and innovation measures have been implemented in this system, so that this system has significant uniqueness and superiority in function and performance, effectively overcoming many defects and deficiencies of the existing technology and achieving substantial progress and improvement in technology.

[0026] Furthermore, the data parsing module converts point cloud data in different formats into a unified standard data dictionary and sets up a file upload and parsing process;

[0027] The standard data dictionary at least includes the coordinates of the point cloud data, the features of the point cloud data, and the data source identifier as an extended field of the standard data dictionary;

[0028] The features of the point cloud data at least include intensity, color information, normal vector, curvature, depth information, number of echoes, classification label; among which the classification label includes class label and instance label;

[0029] The data source identifier of the point cloud data at least includes acquisition timestamp, data quality index, and spatial reference information;

[0030] The data processing module performs at least normalization, normal estimation, chunking, data enhancement, downsampling, and batch processing operations on the data dictionary output by the data parsing module, and transmits the enhanced data to the dynamic visualization module to display the data processing effect in real time through this module;

[0031] The deep learning point cloud segmentation module automatically performs semantic segmentation on the point cloud data processed by the data processing module based on a deep learning model to obtain a segmentation result;

[0032] The manual correction module is used for users to interactively adjust the segmentation result and is connected to the dynamic visualization module;

[0033] A data storage module, responsible for managing and storing all processing results;

[0034] The data parsing module, data processing module, deep learning point cloud segmentation module, manual correction module, and dynamic visualization module are interconnected through API interfaces;

[0035] The dynamic visualization module visually displays the data operations of the data parsing module, data processing module, deep learning point cloud segmentation module, and manual correction module through API interfaces;

[0036] The data storage module receives the data transmitted by the data parsing module, data processing module, deep learning point cloud segmentation module, and manual correction module through API interfaces and is responsible for storing this data;

[0037] The data corrected by the manual correction module is transmitted to the deep learning point cloud segmentation module for segmentation optimization and is also transmitted to the deep learning point cloud segmentation module for segmentation optimization;

[0038] The dynamic visualization module provides real-time visual feedback through a unified point cloud rendering process.

[0039] Compared with traditional methods, the data parsing module has significant differences. In the past, it could only process.pcd format data and its functions were relatively scattered. Now, to handle multi-source heterogeneous data, it has added the file upload and format conversion processes, and has newly added the parsing ability for various formats such as.h5 and.npy. By writing preprocessing scripts, key attributes can be accurately extracted. These attributes are organized into a unified data dictionary structure, effectively improving data compatibility.

[0040] At the same time, the original data processing steps are split and incorporated into the preprocessing function, effectively optimizing the data quality.

[0041] Especially in the operation process, visual interaction is innovatively added, that is, multiple modules are connected to the dynamic visualization module, enabling users to view and evaluate the three-dimensional distribution of the original data during parsing, providing a basis for subsequent enhancement operations, and greatly enhancing the controllability and intelligence of the processing.

[0042] The segmentation results obtained by the newly added point cloud segmentation module are converted into class labels and mapped to the label list of point cloud visualization. When users perform manual correction, they make corrections through the point cloud visualization labels, thus assisting the original annotation tool to achieve automatic annotation, significantly improving the annotation efficiency and accuracy.

[0043] In the past, data processing and data parsing were not clearly distinguished and there was no unified input format. The data processing of the technical solution of the present invention takes a unified data dictionary as the input, improves the data quality and diversity through a series of operations, and is independent of the data parsing module, facilitating maintenance and expansion.

[0044] In summary, the entire system is based on an in-depth analysis of the problems of the existing system. Through innovative architecture design and in-depth improvement and optimization of the functions of each module, an efficient, flexible and highly innovative point cloud data processing system has been successfully created, rather than a simple patchwork of existing modules.

[0045] The data processing module uses a data dictionary in a unified format as the input data, ensuring that the data processing module can efficiently and seamlessly process data from different acquisition devices or scenarios, eliminating compatibility issues caused by the diversity of data formats; the center alignment processing ensures the spatial consistency of subsequent processing, and the random discarding and denoising operations clean up invalid points and noise points, simulating data loss and interference scenarios in a real environment and improving the robustness of the model;

[0046] The data parsing and data processing are split into independent modules. The data parsing focuses on converting point cloud data in different formats into a unified standard data dictionary to ensure data compatibility and consistency, providing a basis for subsequent processing. The data processing is based on the parsed standard format data, and through operations such as center alignment, random discarding and denoising, scale transformation, cropping and sampling, etc., to improve data quality and diversity, providing better data support for subsequent tasks such as point cloud segmentation.

[0047] Through the independent design of the modules, when the data format or parsing requirements change, only the data parsing module needs to be adjusted, without directly affecting the running logic of the data processing module;

[0048] Similarly, if we want to improve or expand the strategies and methods of data processing, it will not interfere with the data parsing process, thus improving the maintainability and scalability of the entire system, making it more conducive to flexible configuration and optimization for different application scenarios and data characteristics.

[0049] Furthermore, the data processing module performs grid sampling and sphere cropping on the point cloud data dictionary;

[0050] The grid sampling divides the point cloud data into blocks according to the specified grid size based on the point cloud coordinates;

[0051] The sphere cropping extracts a point cloud subset through a spherical region centered on a certain point in the point cloud data, and crops the point cloud coordinates and features into multiple independent sub-regions.

[0052] Grid sampling reduces the point cloud density while retaining local structural features; sphere cropping facilitates subsequent segmentation, feature extraction and rendering tasks.

[0053] Further, the deep learning point cloud segmentation module is based on the Point Transformer V3 segmentation model and combines with the WHU-Urban3D dataset. It is trained by setting training configurations and loss functions, and the trained Point Transformer V3 segmentation model is used for point cloud segmentation.

[0054] The training configuration parameters include a batch size of 48, 16 worker threads, the AdamW optimizer is used, the OneCycleLR scheduler is selected to dynamically adjust the learning rate, the training period is 300 epochs, and validation is performed every 50 epochs; the loss functions selected are CrossEntropyLoss and LovaszLoss.

[0055] Further, during the training process of the deep learning point cloud segmentation module, an automatic annotation method based on constructing an annotation model with a small number of labeled samples is used to annotate the point cloud.

[0056] It includes the following steps:

[0057] Step A1: Overall annotation model training;

[0058] Collect a small number of overall point cloud samples, which are used as training samples after being accurately annotated manually. Use deep learning algorithms to train the overall annotation model so that the overall annotation model can initially identify the corresponding object regions in the large-scale point cloud and label category tags.

[0059] Step A2: Manual correction;

[0060] Manually verify the initial recognition results of the model and manually correct problems such as blurred object contours and misclassification.

[0061] Step A3: Local training;

[0062] From the annotation results of the overall annotation model after manual correction, select local regions with poor annotation effects for specified categories; and perform manual annotation again according to the inherent attributes of the specified categories as training samples, and use deep learning algorithms to train the overall annotation model for the second time.

[0063] Perform targeted secondary training according to the unique attributes of these categories respectively to improve the annotation accuracy in key areas of the model.

[0064] In the past, it was not closely combined with manual correction visualization and automatic annotation. In this solution, it is connected with dynamic visualization, which can assist automatic annotation and improve efficiency and accuracy.

[0065] Further, the dynamic visualization module adopts a phased dynamic visualization processing mechanism to provide real-time visualization feedback through a unified point cloud rendering process, and generates real-time rendering results in the data parsing, data processing, deep learning point cloud segmentation, and manual correction stages;

[0066] By extending the function of the THREE.PCDLoader class in Three.js, the PCD file is asynchronously loaded and parsed, and the point cloud data is converted into a THREE.BufferGeometry object for rendering; the point cloud data from different sources is converted into a standardized data dictionary format.

[0067] In the past, the visualization display was simple and without a phased mechanism. The technical solution of the present invention adopts phased dynamic visualization, realizes rendering through Three.js, etc., supports multiple solutions to ensure flexibility, and is convenient for users to evaluate the effects of each stage.

[0068] In a second aspect, a method for constructing a visualization point cloud segmentation pipeline based on deep learning, based on the above-mentioned system for constructing a visualization point cloud segmentation pipeline based on deep learning, through a hierarchical architecture and modular design, the visualization point cloud segmentation is set as follows:

[0069] Setting of the front-end display layer, which provides diverse interaction methods for users, including WebUI based on the browser, RESTful interfaces supporting API access, online rendering of three-dimensional point clouds based on WebGL, and setting of the Open3D tool on the desktop;

[0070] Setting of the core service layer, which is used for point cloud data processing, including setting of a data parsing module, a data processing module, a deep learning point cloud segmentation module, a manual correction module, a data storage module, and a dynamic visualization module;

[0071] Setting of the storage service layer, which is used for storing and managing point cloud data, including a file system, a MongoDB database, and a cache middleware;

[0072] Setting of the resource management layer, which is used for dynamically allocating and managing computing and storage resources, and the computing and storage resources include CPU resources, memory resources, storage resources, network resources, and GPU computing resources;

[0073] Setting of the infrastructure layer, which is used for providing an operating environment for computing resources, including public clouds, virtual machines, and physical machines.

[0074] In a third aspect, a computer storage medium stores a computer program, and the computer program is called by a processor to implement:

[0075] The steps of a method for constructing a visualization point cloud segmentation pipeline based on deep learning.

[0076] In a fourth aspect, an electronic terminal includes at least: one or more processors; and a memory storing one or more computer programs; wherein, the processor invokes the computer programs to execute:

[0077] Steps of a method for constructing a visualization point cloud segmentation pipeline based on deep learning.

[0078] Beneficial effects

[0079] Compared with existing methods, the advantages of the present invention are:

[0080] In view of the problem that when dealing with large-scale complex scene point cloud data, due to the lack of dynamic visualization feedback, it is impossible to timely judge whether the data preprocessing (such as denoising, sampling) reaches the expected effect, resulting in subsequent segmentation tasks may be based on inaccurate data, affecting the final result, the present invention proposes a phased dynamic visualization processing mechanism to display the processing effect in real time at each stage of data preprocessing, enhancement, segmentation, etc. Users can comprehensively observe the data features through interactive operations and adjust parameters in a timely manner. This innovative mechanism runs through the entire data processing process, forming an organic whole, effectively improving the processing efficiency and accuracy, reducing the debugging cost, and having significant innovation compared with traditional visualization methods, providing a new visualization solution for point cloud data processing.

[0081] Combined with the deep learning model Point Transformer V3, data processing strategies and multi-stage dynamic visualization technology, it can efficiently and accurately process large-scale point cloud data. Through modular design, the system integrates the uploading, preprocessing, segmentation, visualization and result correction links of point cloud data into an automated pipeline, and each stage is displayed through visualization, improving the segmentation accuracy and user interaction experience. During the segmentation process, users can reduce the workload of manual annotation and improve the accuracy of the segmentation results through the visualization correction function. Data preprocessing ensures the efficiency of the segmentation task by enhancing data quality and optimizing model adaptability.

[0082] Compared with the prior art, the present application innovatively constructs and integrates each module to create an efficient point cloud segmentation pipeline, thereby significantly improving the segmentation accuracy, processing efficiency and system compatibility, and being able to better adapt to different point cloud data formats and application scenarios. Brief description of the drawings

[0083] Figure 1 It is a schematic diagram of the point cloud segmentation pipeline of the technical solution of the present invention;

[0084] Figure 2 It is a schematic diagram of the system architecture of the technical solution of the present invention;

[0085] Figure 3 It is an internal schematic diagram of the core service layer in the technical solution of the present invention; Figure 4 It is a schematic diagram of the processing flow of the data dictionary output by the data parsing module by the data processing module; Figure 5 It is a schematic diagram of the file upload workflow; Figure 6 It is a schematic diagram of the module connection for point cloud data processing. Specific implementation manners

[0086] Next, the technical solution of the present invention will be further described in conjunction with the accompanying drawings and embodiments.

[0087] Embodiment 1

[0088] A visualization point cloud segmentation pipeline construction system based on deep learning, as Figure 2 shown, includes:

[0089] A front-end display layer, which provides diverse interaction methods for users, including WebUI based on browsers, RESTful interfaces supporting API access, online rendering of three-dimensional point clouds based on WebGL, and Open3D tools on the desktop;

[0090] A core service layer, as Figure 3 shown, for point cloud data processing, including a data parsing module, a data processing module, a deep learning point cloud segmentation module, an artificial correction module, a data storage module, and a dynamic visualization module;

[0091] A storage service layer, for storing and managing point cloud data, including a file system, a MongoDB database, and a cache middleware;

[0092] A point cloud file with a unique identifier corresponds to a set of JSON strings and is stored in the MongoDB database;

[0093] The data storage strategy adopted by this system stores the original data. For the processed data, the invalid points in it will be deleted to ensure that no invalid data is retained; a resource management layer, for dynamically allocating and managing computing and storage resources, and the computing and storage resources include CPU resources, memory resources, memory resources, network resources, and GPU computing resources;

[0094] An infrastructure layer, for providing an operating environment for computing resources, including public clouds, virtual machines, and physical machines.

[0095] Through a hierarchical architecture and modular design, the system achieves efficient collaboration among modules while retaining a high degree of flexibility and scalability. Each module can not only operate independently but also be seamlessly integrated through standardized interfaces. The system forms a closed-loop processing flow in each stage from data parsing, enhancement, segmentation, correction to storage and visualization, thus meeting the comprehensive requirements for high performance, usability and flexibility in point cloud processing scenarios.

[0096] In terms of the system architecture, it is redesigned and planned rather than directly applying the existing system architecture paradigm.

[0097] In the past data processing process, the implementation of various functions was in a scattered state, with only manual annotation tools and data processing scripts for deep learning models. In this system, these originally independent functions have been organically integrated. For the modules at each level, although the functions of some of them have similar implementation forms in the prior art, targeted improvements and innovation measures have been implemented in this system, so that this system has significant uniqueness and superiority in function and performance, effectively overcoming many defects and deficiencies of the prior art and achieving substantial progress and improvement in technology.

[0098] The data parsing module converts point cloud data in different formats into a unified standard data dictionary and sets up a file upload and parsing process;

[0099] The standard data dictionary at least includes the coordinates of the point cloud data, the features of the point cloud data, and the data source identifier as an extended field of the standard data dictionary;

[0100] The features of the point cloud data at least include intensity, color information, normal vector, curvature, depth information, number of echoes, classification label; where the classification label includes class label and instance label;

[0101] The data source identifier of the point cloud data at least includes acquisition timestamp, data quality index, and spatial reference information;

[0102] For example, a complete unified standard raw data dictionary may be as follows:

[0103] {

[0104] "coord": [coordinate information of the point cloud data],

[0105] "feat": [feature information of the point cloud data],

[0106] "source_device": "Device model X",

[0107] "acquisition_time":"2023-01-01 12:00:00",

[0108] "quality_metrics":{

[0109] "completeness":0.95,

[0110] "noise_level":0.05

[0111] },

[0112] "spatial_reference":"Geographic coordinate system WGS84"

[0113] }

[0114] The data processing module performs denoising, downsampling, and rotation operations on the data dictionary output by the data parsing module, and at the same time communicates with the dynamic visualization module to display the data processing effect in real time through this module;

[0115] The deep learning point cloud segmentation module automatically performs semantic segmentation on the point cloud data processed by the data processing module based on a deep learning model to obtain a segmentation result;

[0116] The manual correction module is used for users to interactively adjust the segmentation result and is connected to the dynamic visualization module;

[0117] The data storage module is responsible for managing and storing all processing results;

[0118] The data parsing module, data processing module, deep learning point cloud segmentation module, manual correction module, and dynamic visualization module are interconnected through API interfaces;

[0119] The dynamic visualization module visually displays the data operations of the data parsing module, data processing module, deep learning point cloud segmentation module, and manual correction module through API interfaces;

[0120] The data storage module receives the data transmitted by the data parsing module, data processing module, deep learning point cloud segmentation module, and manual correction module through API interfaces and is responsible for storing this data;

[0121] The data after being corrected by the manual correction module is transmitted to the deep learning point cloud segmentation module for segmentation optimization;

[0122] The dynamic visualization module provides real-time visual feedback through a unified point cloud rendering process;

[0123] Compared with traditional methods, the data parsing module has significant differences. In the past, it could only process.pcd format data and its functions were relatively scattered. Now, to handle multi-source heterogeneous data, it has added the ability to parse various formats such as.h5 and.npy by adding a file upload and format conversion process. By writing preprocessing scripts, key attributes can be accurately extracted. These attributes are organized into a unified data dictionary structure, effectively improving data compatibility.

[0124] At the same time, the original data processing steps are split and incorporated into the preprocessing function, effectively optimizing the data quality.

[0125] Especially in the operation process, visual interaction is innovatively added, that is, multiple modules are connected to the dynamic visualization module, enabling users to view and evaluate the three-dimensional distribution of the original data during parsing, providing a basis for subsequent enhanced operations, and greatly enhancing the controllability and intelligence of the processing.

[0126] The segmentation results obtained by the newly added point cloud segmentation module are converted into category labels and mapped to the label list of point cloud visualization. When users perform manual correction, they make corrections through the point cloud visualization labels, thereby assisting the original annotation tool to achieve automatic annotation, significantly improving the annotation efficiency and accuracy.

[0127] In the prior art, the preprocessing process is usually divided into two independent script modules for execution: Script 1 is mainly responsible for the preprocessing operations of large-scale scene point cloud data (usually tens of millions or even hundreds of millions of point clouds), including data parsing, coordinate translation, normal estimation, and point cloud chunking for large-scale scenes (such as chunking the original point cloud into point cloud data of hundreds of thousands), and storing the processing results in a specific format (such as npy format); Script 2 undertakes the preprocessing function of the network model, specifically including operations such as data parsing, denoising, and data augmentation. At the same time, to meet the input requirements of the network model, on the basis of the chunking of Script 1, secondary chunking sampling is required according to the limitations of the model input (such as the number of point clouds). In the application practice of traditional systems, to maintain system stability, the original architecture of Script 1 is usually kept unchanged.

[0128] Through in-depth research by the inventors of this application, it is found that there is a problem of too high coupling between data parsing and data preprocessing functions in traditional systems, resulting in a complex and redundant processing flow. In particular, the chunking process in Script 1 has a crucial impact on the entire preprocessing process, but its chunking strategy does not fully consider the requirements of subsequent model inputs. And the model input chunking sampling in Script 2 performs secondary processing on the basis of the chunking of Script 1, increasing additional computational overhead and process complexity. Based on the above technical problems, this technical solution proposes an improved preprocessing architecture, which significantly improves the data processing efficiency by optimizing the processing flow.

[0129] In summary, the entire system is based on an in-depth analysis of the problems of the existing system. Through innovative architecture design and in-depth improvement and optimization of the functions of each module, an efficient, flexible and highly innovative point cloud data processing system has been successfully created, rather than a simple patchwork of existing modules.

[0130] The data processing module uses a data dictionary in a unified format as input data, ensuring that the data processing module can efficiently and seamlessly process data from different acquisition devices or scenarios, eliminating compatibility issues caused by the diversity of data formats; the center-aligned processing ensures the spatial consistency of subsequent processing, and the random discarding and denoising operations clean up invalid points and noise points, simulating data loss and interference scenarios in a real environment, and improving the robustness of the model.

[0131] The data parsing and data preprocessing functions are decoupled and designed as independent modules respectively. The data parsing module is responsible for reading the original point cloud data and converting it into a unified standard data dictionary to ensure data compatibility and consistency, providing a basis for subsequent processing. The data preprocessing module, based on the parsed standard format data, enhances the data quality and diversity through data augmentation operations, providing better data support for subsequent tasks such as point cloud segmentation.

[0132] Through the independent design of the modules, when the data format or parsing requirements change, only the data parsing module needs to be adjusted, without directly affecting the running logic of the data processing module.

[0133] Similarly, if you want to improve or expand the strategies and methods of data processing, it will not interfere with the data parsing process, thus improving the maintainability and scalability of the entire system, making it more conducive to flexible configuration and optimization for different application scenarios and data characteristics.

[0134] The point cloud data processing operations involved in this data processing module cover multiple aspects. These operations at least include normalization, normal vector estimation, chunking, data augmentation, and downsampling and batch processing, etc. Each operation has unique functions and importance and belongs to an important part of point cloud data processing.

[0135] For the normalization operation, traverse the scene acquisition coordinate data, decompose it along the x, y, and z axes, and subtract the minimum value of the corresponding axis from each axis value to complete the normalization, normalizing the point cloud data coordinates to a unified coordinate system, which helps the model effectively learn the relative position relationship of the point cloud data and provides a consistent reference for subsequent processing.

[0136] Normal estimation operation: Construct an Open3D PointCloud object. Based on the point cloud coordinates, call the estimate_normals method, and use a KDTreeSearchParamHybrid instance to determine the search neighborhood to estimate point normals. Extract the normals and store them in the nx, ny, and nz arrays for later use in the data structure. This is crucial for understanding the geometric structure of the point cloud data and subsequent 3D processing tasks such as segmentation, classification, and reconstruction.

[0137] Chunking operation: Determine the range of the point cloud on the x and y axes. Calculate the number of grid cells based on the range and step size, and generate a list of coordinates. Traverse the cells to find the points belonging to them. Exclude some cells based on the number of points and labels. Sample the qualified cells to generate small chunks and add them to the chunk list. Stack the small chunks into an array for subsequent processing such as bounding box calculation, and store the result file. Splitting the point cloud data into small chunks facilitates subsequent processing and analysis, helps to focus on local features and details, and improves processing efficiency.

[0138] Data augmentation operation: Rotate the point cloud around the z-axis by (-1° to 1°) with a 50% probability; randomly scale the point cloud by a factor of 0.9 to 1.1; flip the point cloud in the xy plane with a 50% probability; jitter the point cloud with a standard deviation of 0.005 and a range of ±0.02. Increase the diversity of the point cloud data through rotation and other means, enhance the robustness of the model, reduce the risk of overfitting, and improve the generalization ability of the model.

[0139] Downsampling operation: Use grid sampling with a grid size of 0.05, FNV hash type, training mode, retain coordinates, features, and segmentation labels, and return the grid coordinates. Use a specific algorithm to reduce the number of points in the point cloud data, while reducing the computational complexity and retaining the basic structure of the data to the greatest extent, saving computational resources.

[0140] Batch processing operation: Specify the keys to retain and the feature keys, integrate multiple data samples into a batch, retain coordinates, features, and segmentation labels; organize the batch data and mix the 'offset' key values by probability; serialize the data and convert it into a tensor for use in deep learning models to improve the utilization rate of computational resources and ensure the coherence, stability, and effectiveness of training.

[0141] The data processing module performs grid sampling and sphere clipping on the point cloud data dictionary;

[0142] Grid sampling divides the point cloud data into chunks according to the specified grid size based on the point cloud coordinates;

[0143] Sphere clipping extracts subsets of the point cloud through a spherical region centered on a certain point in the point cloud data, and clips the point cloud coordinates and features into multiple independent sub-regions.

[0144] The deep learning point cloud segmentation module is based on the Point Transformer V3 segmentation model and combines with the WHU-Urban3D dataset. It is trained by setting training configurations and loss functions, and the trained Point Transformer V3 segmentation model is used for point cloud segmentation;

[0145] The training configuration parameters include a batch size of 48, 16 worker threads, the AdamW optimizer is adopted, the OneCycleLR scheduler is selected to dynamically adjust the learning rate, the training period is 300 epochs, and validation is performed every 50 epochs; The loss functions selected are CrossEntropyLoss and LovaszLoss.

[0146] The operating environment required for the visualization 3D point cloud segmentation system is shown in the following table, including operating system, browser, and development environment device requirements;

[0147] Table 1 Operating System Support Table

[0148] Operating System Development Environment Production Environment Ubuntu 22.04 and above 22.04 and above macOS 12.21 and above 12.21 and above

[0149] Table 2 Browser Support Table

[0150] Browser Development Environment Production Environment Google Chrome (recommended) 131.0.6778.205 and above 131.0.6778.205 and above Safari 15.3 and above 15.3 and above

[0151] Table 3 Development Environment Device Requirements Table

[0152]

[0153] When applying the technical solution of the present invention for model segmentation, the segmentation accuracy and IoU of each category are shown in Table 4 below;

[0154] The mean intersection over union (mIoU) is 0.5686, the mean accuracy of classes (mAcc) is 0.7105, and the overall point cloud accuracy (allAcc) is 0.8177;

[0155] Table 4 Segmentation Results of the Present Technical Solution

[0156]

[0157] The implementation plan of this module selects Point Transformer V3 as the segmentation model. By introducing serialization technology and space-filling curves, Point Transformer V3 converts unstructured point cloud data into a structured format. Combining with the self-attention mechanism, it can effectively capture the local and global features of point cloud data, and is particularly suitable for the segmentation task of large-scale point cloud data. Point Transformer V3 performs excellently on multiple public datasets, especially showing obvious advantages over traditional methods in terms of segmentation accuracy and computational efficiency, and is suitable for the automated point cloud segmentation requirements in this project.

[0158] During the training process of the deep learning point cloud segmentation module, an automatic annotation method based on constructing an annotation model with a small number of labeled samples is used to annotate the point cloud;

[0159] It includes the following steps:

[0160] Step A1: Overall annotation model training;

[0161] Collect a small number of overall point cloud samples, which are used as training samples after being accurately annotated manually. Use deep learning algorithms to train the overall annotation model, so that the overall annotation model can initially identify the corresponding object regions in the large-scale point cloud and annotate category labels;

[0162] Step A2: Manual correction;

[0163] Manually verify the initial recognition results of the model and manually correct the problems of blurred object contours and misclassification;

[0164] Step A3: Local training;

[0165] From the annotation results of the overall annotation model after manual correction, select the local regions with poor annotation effects for the specified categories; and re-annotate manually according to the inherent attributes of the specified categories as training samples, and use deep learning algorithms to retrain the overall annotation model.

[0166] According to the unique attributes of these categories, conduct targeted secondary training to improve the annotation accuracy in key areas of the model.

[0167] Compared with the existing methods, there are the following differences:

[0168] 1. Traditional few-shot annotation model training usually directly constructs a general annotation model based on a small number of sample labels. During training, the model only learns features from limited samples and is used for actual annotation, and the annotation accuracy is significantly insufficient.

[0169] Taking urban point cloud annotation as an example, in the past, a small number of labeled point cloud samples of buildings, streets, etc. were simply selected to train the model, expecting it to handle all relevant annotation tasks. However, due to the complexity of urban point clouds and the small number of samples, when the model faces fine local feature annotation, the annotation accuracy is often not high.

[0170] Although existing segmentation combined with annotation methods attempt to combine segmentation and annotation, they generally lack a coherent process of first overall control, manual fine correction, and then local in-depth optimization.

[0171] For example, some practices first perform automatic segmentation and then directly annotate the segmented regions, without paying attention to the key role of manual correction between the whole and the local, and also failing to make full use of a small number of labeled samples throughout, resulting in poor accuracy of the annotation results.

[0172] The dynamic visualization module adopts a phased dynamic visualization processing mechanism, provides real-time visualization feedback through a unified point cloud rendering process, and generates real-time rendering results in the stages of data parsing, data processing, deep learning point cloud segmentation, and manual correction;

[0173] Through the extended function of the THREE.PCDLoader class in Three.js, asynchronously load and parse PCD files, convert the point cloud data into a THREE.BufferGeometry object for rendering; convert point cloud data from different sources into a standardized data dictionary format.

[0174] Although this solution uses Three.js as the main visualization tool, it can also be implemented through alternative solutions such as Open3D and Potree.js, ensuring the flexibility and compatibility of the system.

[0175] Help users intuitively evaluate the effects of each stage. Users can comprehensively observe the characteristics of the point cloud data through interactive view rotation, scaling, and normalization.

[0176] The above division of functional modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. At the same time, the above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0177] The process of point cloud segmentation using the technical solution of the present invention is as Figure 1 shown below:

[0178] 1. File processing: Support the upload and parsing of multiple point cloud file formats. The system automatically processes standard formats (such as PCD files) through built-in parsing tools, and uses special scripts for conversion for non-standard formats (such as H5 files).

[0179] 2. Data preprocessing: Denoise, enhance, and standardize the uploaded point cloud data, and display the processing results in real time.

[0180] 3. Point cloud segmentation: Use the Point Transformer V3 model to automatically segment the preprocessed point cloud data and output semantic labels and category information.

[0181] 4. Visualization display: Through an interactive visualization interface, display the processed point cloud data and its segmentation results, and support dynamic adjustment by users.

[0182] 5. Result correction: Users can make detailed corrections to the segmentation results to ensure that the output labels are consistent with the actual requirements.

[0183] Example 2

[0184] A method for constructing a visualization point cloud segmentation pipeline based on deep learning. Through a hierarchical architecture and modular design, the visualization point cloud segmentation is set as follows:

[0185] Front-end display layer setting: Provide users with diverse interaction methods, including browser-based WebUI, RESTful interfaces supporting API access, WebGL-based online rendering of 3D point clouds, and setting of Open3D tools for the desktop.

[0186] Core service layer setting: Used for point cloud data processing, including data parsing module, data processing module, deep learning point cloud segmentation module, manual correction module, data storage module, and dynamic visualization module setting.

[0187] Storage service layer setting: Used for storing and managing point cloud data, including file system, MongoDB database, and cache middleware.

[0188] Resource management layer setting: Used for dynamically allocating and managing computing and storage resources, where the computing and storage resources include CPU resources, memory resources, storage resources, network resources, and GPU computing resources.

[0189] Infrastructure layer setting: Used to provide the operating environment for computing resources, including public cloud, virtual machines, and physical machines.

[0190] It should be understood that the implementation process of each setting can refer to the content description of the foregoing system.

[0191] Example 3

[0192] A computer-readable storage medium stores a computer program, and the computer program is called by a processor to implement:

[0193] Steps of the above method for constructing a visualization point cloud segmentation pipeline based on deep learning.

[0194] For the specific implementation process of each step, please refer to the description of the foregoing method.

[0195] It should be understood that in the embodiments of the present invention, the so-called processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0196] The readable storage medium is a computer-readable storage medium, which may be an internal storage unit of the software and hardware device described in any of the foregoing embodiments, such as the hard disk or memory of the controller. The readable storage medium may also be an external storage device of the controller, such as a plug-in hard disk equipped on the controller, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the readable storage medium may also include both the internal storage unit of the controller and the external storage device. The readable storage medium is used to store the computer program and other programs and data required by the controller. The readable storage medium may also be used to temporarily store data that has been output or will be output.

[0197] Embodiment 4

[0198] An electronic terminal, at least including: one or more processors; and a memory storing one or more computer programs; wherein, the processor calls the computer program to execute:

[0199] Steps of the method for constructing a visualization point cloud segmentation pipeline based on deep learning.

[0200] For the specific implementation process of each step, please refer to the description of the foregoing method.

[0201] Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned readable storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0202] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program codes. The present application is a device that generates, according to the instructions executed by a processor in the flowchart of the method, device (system), and computer program product of the present application, for implementing the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram. These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes of the flowchart and / or one or more boxes of the block diagram.

[0203] It should be emphasized that the examples described in the present invention are illustrative rather than restrictive. Therefore, the present invention is not limited to the examples described in the specific embodiments. Any other embodiments obtained by those skilled in the art based on the technical solution of the present invention, without departing from the spirit and scope of the present invention, whether modified or replaced, equally fall within the protection scope of the present invention.

Claims

1. A visualization point cloud segmentation pipeline construction system based on deep learning, characterized in that, Including: The front-end display layer provides users with diverse interaction methods, including browser-based WebUI, RESTful interfaces supporting API access, 3D point cloud online rendering based on WebGL, and the Open3D tool for the desktop. The core service layer is used for point cloud data processing, including a data parsing module, a data processing module, a deep learning point cloud segmentation module, a manual correction module, a data storage module, and a dynamic visualization module. The storage service layer is used for storing and managing point cloud data, including a file system, a MongoDB database, and a cache middleware. The resource management layer is used for dynamically allocating and managing computing and storage resources, and the computing and storage resources include CPU resources, memory resources, storage resources, network resources, and GPU computing resources. The infrastructure layer is used to provide a running environment for computing resources, including public clouds, virtual machines, and physical machines.

2. The system according to claim 1, wherein The data parsing module converts point cloud data in different formats into a unified standard data dictionary and sets up a file upload and parsing process. The standard data dictionary at least contains the coordinates of the point cloud data, and the features of the point cloud data and the data source identifier are extended fields of the standard data dictionary. The features of the point cloud data at least include intensity, color information, normal vector, curvature, depth information, number of echoes, classification label. Among them, the classification label includes a category label and an instance label. The data source identifier of the point cloud data at least includes an acquisition timestamp, a data quality index, and spatial reference information. The data processing module performs at least normalization, normal vector estimation, chunking, data augmentation, downsampling, and batch processing operations on the data dictionary output by the data parsing module, and transmits the enhanced data to the dynamic visualization module to display the data processing effect in real time through this module. The deep learning point cloud segmentation module automatically performs semantic segmentation on the point cloud data processed by the data processing module based on a deep learning model to obtain a segmentation result. The manual correction module is used for users to interactively adjust the segmentation result and is connected to the dynamic visualization module. The data storage module is responsible for managing and storing all processing results. The data parsing module, the data processing module, the deep learning point cloud segmentation module, the manual correction module, and the dynamic visualization module are interconnected through API interfaces. The dynamic visualization module visually displays the data operations of the data parsing module, the data processing module, the deep learning point cloud segmentation module, and the manual correction module through API interfaces. The data storage module receives the data transmitted by the data parsing module, the data processing module, the deep learning point cloud segmentation module, and the manual correction module through API interfaces, and is also responsible for storing the transmitted data. The data corrected by the manual correction module is transmitted to the deep learning point cloud segmentation module for segmentation optimization. The dynamic visualization module provides real-time visual feedback through a unified point cloud rendering process.

3. The system according to claim 2, wherein The data processing module performs grid sampling and sphere clipping on the point cloud data dictionary. The grid sampling performs chunking processing on the point cloud data according to the specified grid size based on the point cloud coordinates. The sphere clipping extracts a point cloud subset from a spherical region centered at a certain point in the point cloud data, and clips the point cloud coordinates and features into multiple independent sub-regions.

4. The system according to claim 2, wherein The deep learning point cloud segmentation module is based on the PointTransformer V3 segmentation model, and combines with the WHU-Urban3D dataset. It is trained by setting training configurations and loss functions, and uses the trained Point Transformer V3 segmentation model for point cloud segmentation. The training configuration parameters include a batch size of 48, 16 worker threads, the AdamW optimizer, the OneCycleLR scheduler to dynamically adjust the learning rate, 300 epochs for training, and validation every 50 epochs. The loss functions used are CrossEntropyLoss and LovaszLoss.

5. The system according to claim 4, wherein During the training process of the deep learning point cloud segmentation module, an automatic annotation method based on constructing an annotation model with a small number of labeled samples is used to annotate the point cloud. It includes the following steps: Step A1: Overall annotation model training; Collect a small number of overall point cloud samples, which are used as training samples after manual precise annotation. Use deep learning algorithms to train the overall annotation model, enabling the overall annotation model to initially identify the corresponding object regions in the large-scale point cloud and label the category labels. Step A2: Manual correction; Manually verify the initial recognition results of the model and manually correct problems such as blurred object contours and misclassification. Step A3: Local training; From the annotation results of the overall annotation model after manual correction, select local regions with poor annotation effects for a specified category; and perform manual annotation again based on the inherent attributes of the specified category as training samples, and use deep learning algorithms to retrain the overall annotation model.

6. The system according to claim 1, wherein The dynamic visualization module adopts a phased dynamic visualization processing mechanism, provides real-time visualization feedback through a unified point cloud rendering process, and generates real-time rendering results during the data parsing, data processing, deep learning point cloud segmentation, and manual correction stages. Through the extended function of the THREE.PCDLoader class in Three.js, asynchronously load and parse PCD files, convert the point cloud data into a THREE.BufferGeometry object for rendering; convert point cloud data from different sources into a standardized data dictionary format.

7. A method for constructing a visualization point cloud segmentation pipeline based on deep learning, characterized in that, By adopting the hierarchical architecture and modular design of the system described in any one of claims 1-6, the visual point cloud segmentation is set as follows: Front-end display layer setting, providing users with diverse interaction methods, including browser-based WebUI, RESTful interfaces supporting API access, WebGL-based online rendering of three-dimensional point clouds, and Open3D tool setting for the desktop. Core service layer setting, used for point cloud data processing, including data parsing module, data processing module, deep learning point cloud segmentation module, manual correction module, data storage module, and dynamic visualization module setting. Storage service layer settings for storing and managing point cloud data, including file systems, MongoDB databases, and cache middleware; Resource management layer settings for dynamically allocating and managing computing and storage resources, where the computing and storage resources include CPU resources, memory resources, storage resources, network resources, and GPU computing resources; Infrastructure layer settings for providing a running environment for computing resources, including public clouds, virtual machines, and physical machines.

8. A computer storage medium, characterized in that: A computer program is stored, and the computer program is called by a processor to implement: The steps of the method according to claim 7.

9. An electronic terminal, characterized in that: At least includes: One or more processors; And a memory storing one or more computer programs; Wherein, the processor calls the computer program to execute: The steps of the method according to claim 7.

Citation Information

Cited By

  • Shoeprint point cloud identification method based on image processing

    CN121010832A