Intelligent segmentation and CAD model generation method and system for city-level point cloud scene
By introducing voxelized downsampling, quadtree segmentation and deep learning models in point cloud processing, the problems of low efficiency and insufficient accuracy of urban-level high-density point cloud scenarios are solved, efficient and accurate CAD model generation is achieved, and intelligent missing completion and user-friendly interaction are provided.
Patent Information
- Application Number
- CN202411945553.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-16
AI Technical Summary
The existing technology cannot efficiently process urban-level high-density point cloud scenarios and accurately convert them into CAD models, which has problems such as low processing efficiency, lack of intelligence and lack of completion capabilities, dynamic objects affect modeling accuracy and inconvenient user interaction.
Through the optimization of the processing process, voxelized downsampling and quadtree segmentation technology were introduced, combined with deep learning models to achieve point cloud semantic segmentation, instance recognition and intelligent completion of missing areas, and ensure the accuracy of model splicing through global coordinates, ultimately realizing the automated conversion of large-scale point cloud data to a complete CAD model.
It realizes efficient and accurate point cloud data processing and three-dimensional modeling, improves the modeling efficiency of urban-level scenarios, solves the problems of missing completion and dynamic object filtering, and provides user-friendly visual interface and interactive support.
Smart Images

Figure CN120012217A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of three-dimensional point cloud data processing and automatic generation of computer-aided design (CAD), and in particular to an intelligent processing platform based on large-scale point cloud scenes, which is used for automatic segmentation, instance recognition and multi-scale CAD conversion of high-density three-dimensional scene point cloud data at the city level. With the assistance of deep learning models, the platform realizes intelligent segmentation and instance recognition of point cloud data, enabling users to quickly generate high-precision, visual CAD models in large-scale scenes. This technology is suitable for multiple application fields such as building information modeling (BIM), smart cities, digital twins, virtual reality (VR) and augmented reality (AR), geographic information systems (GIS), etc. Background Art
[0002] As a high-precision expression of the spatial form of the real world, three-dimensional point cloud data has been widely used in many fields, including building information modeling (BIM), smart cities, virtual reality (VR), augmented reality (AR), and geographic information systems (GIS). High-density and massive point cloud data can be obtained through laser scanning (LiDAR), drone aerial photography, etc. However, due to the unstructured, complex and large amount of point cloud data itself, traditional point cloud data processing technology is difficult to meet the needs of large-scale scenarios in terms of efficiency and accuracy.
[0003] In recent years, with the development of deep learning technology, point cloud data processing has made significant progress in feature extraction, semantic segmentation, and instance segmentation. These technologies can efficiently extract semantic and geometric features from point clouds by training neural networks, and apply them to scene analysis, 3D modeling and other scenarios. At the same time, some traditional 3D modeling tools also provide the ability to convert point clouds to CAD models, providing important support for the practical application of 3D point cloud data.
[0004] In the prior art, the solution closest to the present invention generally includes the following key modules:
[0005] 1) Point cloud data preprocessing: In order to reduce computational complexity, downsampling or blocking techniques are often used to preprocess point cloud data. Some solutions sparse point cloud data through voxelization operations, or use region partitioning algorithms to segment large-scale scenes to ensure that the data volume is within the hardware processing capacity.
[0006] 2) Semantic and instance segmentation: Deep learning models are widely used in point cloud segmentation tasks, which generate semantic and instance labels by reasoning point by point on point cloud data. This function can effectively distinguish different objects in the scene, such as buildings, roads, bridges, etc.
[0007] 3) Point cloud to CAD model conversion: point cloud data is converted into CAD models through traditional 3D surface reconstruction algorithms or regularized modeling techniques. These methods can generate visual 3D models for architectural design, engineering analysis and other fields.
[0008] Existing technical solutions are relatively mature in small-scale point cloud scenarios, and some commercial tools and research prototypes can achieve a certain degree of automated modeling. However, the processing range and capabilities of these technologies are still greatly limited in large-scale point cloud scenarios at the city level. The existing technologies mainly have the following technical problems:
[0009] 1) Unable to adapt to large scenes. Currently, existing technologies and related platforms mostly focus on converting point cloud data of a single building or a single part into a CAD model, and the modeling of small-scale, independent objects is relatively mature. However, in city-level point cloud scenes, due to the huge amount of data, high scene complexity, and close relationships between instances, existing technologies cannot yet achieve the automatic conversion of the overall point cloud scene into a CAD model.
[0010] 2) Low processing efficiency. When faced with large-scale, high-density point cloud data at the city level, existing technologies have high computational complexity and are often unable to process all scene data at once. In particular, deep learning inference models are limited by the graphics card memory capacity and require manual block operations on point cloud data. This manual block division method is not only inefficient, but also easily leads to inconsistent segmentation problems.
[0011] 3) Lack of intelligent missing completion capabilities. In the process of converting point clouds to CAD models, there are often missing areas in the point cloud data, such as the bottom surface of a building or details in the structure. Existing surface reconstruction methods have limited processing capabilities for these missing information, making it difficult to generate a continuous and complete CAD model, which affects the practicality and accuracy of the model.
[0012] 4) Dynamic objects affect modeling accuracy. Urban point cloud data usually contains dynamic objects (such as pedestrians and vehicles), which interfere with the modeling process of fixed objects (such as buildings and roads). Existing technologies are not capable of automatically filtering dynamic objects, resulting in the generated CAD models containing redundant information or impaired accuracy.
[0013] 5) Inconvenient user interaction. Existing point cloud processing platforms usually lack efficient visualization tools, making it difficult for users to intuitively select specific instances or categories for modeling. At the same time, manual correction of segmentation results is relatively complicated and cannot meet the convenience requirements in practical applications. Summary of the invention
[0014] The present invention aims to solve the technical problem that the existing technology cannot efficiently process urban-level high-density point cloud scenes and accurately convert them into CAD models, and provide an intelligent segmentation and CAD model generation method and system for urban-level point cloud scenes, and realize efficient and accurate point cloud data processing and three-dimensional modeling by optimizing the processing flow and introducing deep learning models. The present invention optimizes the data processing flow through voxel downsampling and quadtree segmentation technology, combines deep learning models to realize point cloud semantic segmentation, instance recognition and intelligent completion of missing areas, and further ensures the accuracy of model splicing through global coordinates, and finally efficiently completes the automatic conversion of large-scale point cloud data to complete CAD models, meeting the high-precision modeling requirements of urban-level scenes.
[0015] The technical solution adopted by the present invention is as follows:
[0016] A method for intelligent segmentation and CAD model generation for city-level point cloud scenes, comprising the following steps:
[0017] Pre-process the point cloud data uploaded by the user, and divide the point cloud scene into several sub-regions by performing voxelization and region segmentation;
[0018] Use the deep learning model to segment the preprocessed sub-areas and generate point cloud data with semantic labels and instance labels;
[0019] Using the point cloud data with semantic labels and instance labels generated after point cloud segmentation, point cloud visualization is performed in the visualization interface according to the instance or category selected by the user;
[0020] Perform CAD conversion operations based on the instance or category selected by the user to convert the point cloud file into a CAD file;
[0021] Based on the CAD file obtained by the conversion operation, a three-dimensional model of the entire scene is displayed in the CAD visualization interface.
[0022] Furthermore, the preprocessing of the point cloud data uploaded by the user includes:
[0023] Receiving raw point cloud data and a voxel size parameter specified by a user as input, wherein the raw point cloud data contains three-dimensional point information of a city level or other large-scale scenes, and the voxel size is used to control the accuracy of downsampling;
[0024] Downsample the original point cloud data based on the voxel size;
[0025] The number of points in the downsampled point cloud data is judged. If the number of points in the downsampled point cloud data does not exceed the set threshold, the point cloud segmentation step is entered; if the number of points exceeds the threshold, block processing is performed, and the number of points in each sub-block is within the threshold range;
[0026] After the block processing is completed, each sub-block is recorded as an independent sub-scene in the point cloud sub-scene list, which contains the location information and point cloud data of each sub-block.
[0027] Furthermore, the use of the deep learning model to perform point cloud segmentation on the preprocessed sub-regions includes:
[0028] Receive the point cloud sub-scene list obtained after preprocessing, and perform inference processing of the deep learning model on each sub-scene data in the list in turn. During the inference process, choose block-by-block inference or multi-graphics card parallel inference according to available resources;
[0029] After each inference, the inference result is temporarily stored in the memory, and the processed sub-scenes are deleted from the point cloud sub-scene list until the list is empty;
[0030] When the list is empty, the sub-scene merging phase begins. During the merging process, the segmentation results of each sub-scene are unified through the common area alignment technology to ensure that the semantics and instance labels of each sub-scene are consistent and generate a complete global segmentation result.
[0031] Based on the shortest distance principle, for each point in the original point cloud data, the nearest point in the segmented point cloud scene is searched in a cube space with a side length of twice the voxel centered on it, and each point in the original point cloud data is matched with the label in the segmentation result to generate large-scale high-density point cloud data with semantic and instance labels;
[0032] The segmentation results are downsampled according to the voxel size set by the user to generate low-density point cloud data suitable for visualization.
[0033] Furthermore, performing point cloud visualization in the visualization interface according to the instance or category selected by the user includes:
[0034] Receive voxelized downsampled scene point cloud data with semantic and instance labels, perform LOD processing on it, and achieve efficient rendering on the user's browser through a multi-resolution point cloud data generation mechanism;
[0035] The user can intuitively view and filter point cloud data of different instances and categories in the visual interface, and select a specific modeling object type or select one or several instance objects for modeling;
[0036] Users set the voxel size to control modeling accuracy and file data volume, ensuring that the generated CAD model meets the actual needs of the scene while maintaining high efficiency of file conversion.
[0037] Furthermore, performing a CAD conversion operation according to the instance or category selected by the user to convert the point cloud file into a CAD file includes:
[0038] Receive instance labels, category labels and voxel size parameters, and obtain large-scale scene high-density point cloud data with semantic and instance labels;
[0039] Filter and process point cloud data according to category labels and instance labels, and retain category or instance point cloud data specified by the user;
[0040] Perform voxel downsampling processing on the filtered high-density point cloud data;
[0041] The downsampled point cloud data is processed one by one according to the instance label, and the data of each independent instance is extracted and converted to CAD in turn. During the conversion process, the deep learning model is used to complete the missing structure in each instance.
[0042] After the point cloud data of each instance is converted into a CAD instance model, the conversion result is saved as a separate CAD file. When the CAD conversion of all instances is completed, the CAD files of the individual instances are merged one by one into a complete scene CAD file.
[0043] Furthermore, the CAD visualization interface supports real-time observation and zooming in and out from multiple angles, allowing users to examine in detail the geometry, structural integrity and relative position of each instance in the scene; the CAD visualization interface provides interactive tools for users to fine-tune the model, including correction of geometric errors, edge alignment, and structural improvement, to ensure the accuracy and integrity of the final model.
[0044] An intelligent segmentation and CAD model generation system for city-level point cloud scenes, comprising:
[0045] The preprocessing module is used to preprocess the point cloud data uploaded by the user and divide the point cloud scene into several sub-regions by performing voxelization and region segmentation;
[0046] Point cloud segmentation module, which is used to segment the preprocessed sub-areas using a deep learning model to generate point cloud data with semantic labels and instance labels;
[0047] The point cloud visualization module is used to use the point cloud data with semantic labels and instance labels generated after point cloud segmentation to visualize the point cloud according to the instance or category selected by the user in the visualization interface;
[0048] A point cloud CAD conversion module is used to perform CAD conversion operations according to the instance or category selected by the user, and convert the point cloud file into a CAD file;
[0049] The CAD visualization module is used to display the three-dimensional model of the entire scene in the CAD visualization interface according to the CAD file obtained by the conversion operation.
[0050] The beneficial effects of the present invention are as follows:
[0051] Key point 1: Point cloud instance segmentation and semantic labeling based on deep learning models. The present invention uses deep learning models to perform instance segmentation and semantic labeling on large-scale point cloud data at the city level, and can accurately extract instance information such as buildings, roads, bridges, etc., providing a clear input data foundation for instance-by-instance CAD modeling. In response to the problem that large-scale point cloud scenes may exceed the video memory limit of the graphics card, the present invention divides the scene into video memory-friendly sub-areas through a quadtree segmentation algorithm, so that each deep learning inference can run efficiently under limited hardware resources. After the segmentation is completed, a unique common area alignment and shortest distance matching technology is used to splice the segmented sub-areas into a complete scene, and ensure the consistency of global semantics and instance labels, providing an accurate global view for overall modeling.
[0052] Key point 2: Novel structural design for instance-by-instance CAD modeling. The innovative instance-by-instance modeling architecture proposed in the present invention enables users to generate high-precision CAD models for specific instances (such as a single building or bridge) according to their needs without the need for overall scene modeling, fundamentally reducing modeling complexity and resource consumption. After generation, each instance model is spliced into a CAD model of the complete scene through global coordinate calibration technology, which not only ensures the spatial consistency between instances, but also supports the flexibility of modular design. Compared with the prior art, the present invention greatly improves the modeling efficiency of city-level scenes and provides an efficient and practical solution for smart city planning and building information modeling (BIM).
[0053] Key point 3: Intelligent completion of missing data driven by deep learning. In the case of common acquisition blind spots in point cloud data (such as the bottom surface of a building or complex structural details), traditional methods cannot generate a complete model. The present invention uses a deep learning model to intelligently predict and complete the missing areas, and the generated CAD model is more complete and continuous in geometry. This completion technology significantly improves the applicability of the model, allowing the generated 3D model to more realistically reflect the actual scene.
[0054] Key point 4: Intelligent filtering of dynamic objects. The deep learning model of the present invention can automatically identify and filter the point cloud data of dynamic objects (such as pedestrians and vehicles), avoiding the interference of dynamic data on the modeling of fixed facilities, and ensuring that the generated CAD model is more accurate and the data is purer. Compared with the cumbersome process of manually removing dynamic objects in traditional technology, the present invention significantly simplifies the data processing steps.
[0055] Key point 5: User-friendly visualization interface and interactive support. The visualization interface based on LOD (level of detail) technology enables users to preview segmentation results in real time, select specific instances for modeling, or dynamically adjust modeling parameters to meet specific needs. While improving operational efficiency, it provides convenient support for multi-scenario and refined modeling. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 : The architecture design diagram of the automated processing and modeling platform for large-scale, high-density urban point cloud data of the present invention.
[0057] Figure 2 :Flowchart for converting large urban scene point cloud files into CAD models.
[0058] Figure 3 : Flowchart of the preprocessing module.
[0059] Figure 4 : Flowchart of the point cloud segmentation module.
[0060] Figure 5 : Flowchart of the point cloud visualization module.
[0061] Figure 6 : Flowchart of point cloud CAD conversion module. DETAILED DESCRIPTION
[0062] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below through specific embodiments and drawings.
[0063] 1. Analysis of key points and invention points
[0064] 1) Point cloud instance segmentation based on deep learning model. The present invention uses deep learning technology to achieve accurate segmentation of each instance (such as a single building, bridge or road) in a large-scale point cloud scene. Compared with the traditional overall modeling method, this method can decompose complex city-level scenes into multiple independent instances, greatly reducing the modeling difficulty and computational complexity.
[0065] 2) A novel structural design for instance-by-instance CAD modeling. This invention innovatively proposes a technical architecture for instance-by-instance modeling. Users do not need to directly model the entire point cloud scene, but instead generate a CAD model for each independent instance based on the instance segmentation results. This architecture not only improves the flexibility of model generation, but also simplifies the organization and splicing of the entire scene through instance management.
[0066] 3) Global consistency guarantee of instance splicing. After generating the CAD model of a single instance, the present invention uses global coordinate calibration and common area alignment technology to reassemble the CAD models of each instance into a complete urban scene, while ensuring that the geometric structure and spatial relationship between the models are accurately consistent.
[0067] 4) Dynamic object filtering and instance selection optimization. The deep learning model of the present invention can automatically filter dynamic objects (such as pedestrians and vehicles) in the point cloud to ensure that instance segmentation focuses on fixed facilities. Combined with the visual interface, users can flexibly select the instance category or range to be modeled, further improving modeling efficiency.
[0068] 5) Technical architecture that efficiently adapts to multi-instance scenarios. The instance-by-instance modeling architecture of the present invention adapts to the requirements of multi-instance scenarios, supports both high-precision modeling of a single instance and batch modeling and integration of instantiation results, and meets the requirements of efficient and automated modeling of city-level scenarios.
[0069] The present invention provides an automated processing and modeling platform for large-scale, high-density urban point cloud data. The platform uses technologies such as voxel downsampling, sub-region division, deep learning reasoning, and automatic completion to achieve intelligent segmentation, semantic annotation, instance modeling, and CAD model generation of urban point cloud scenes, providing users with a high-precision, customizable 3D modeling solution.
[0070] The innovation of the present invention lies in that it proposes a point cloud data to CAD conversion mechanism combined with a deep learning model, which can cope with the difficulties caused by computing resource limitations in the overall processing of large-scale point cloud data.
[0071] First, to ensure data processing efficiency and reasonable resource allocation, the platform performs voxel downsampling on point cloud data during data preprocessing, making the point cloud data more sparse and reducing the computational load. Next, the point cloud scene is divided into multiple sub-areas through the quadtree segmentation algorithm so that they can be loaded into the deep learning model one by one for processing in a single inference. The upper limit of the number of points in the sub-area division is automatically adjusted according to the graphics card memory capacity to ensure that the system can efficiently process large-scale point cloud data.
[0072] In the point cloud segmentation and instance recognition stage, the platform infers each sub-region of the point cloud based on the deep learning model, thereby accurately segmenting and labeling fixed objects in the scene, such as buildings, roads, bridges, etc. Through this sub-region step-by-step splicing reasoning mechanism, the platform further performs minimum distance matching on the original point cloud after the scene is spliced, and associates each point cloud point with the model's prediction result, thereby achieving accurate semantic segmentation.
[0073] In the process of converting point cloud to CAD model, the present invention specifically introduces automatic completion technology based on deep learning to improve the integrity and accuracy of CAD model. Traditional point cloud to CAD conversion methods may lead to incomplete modeling in the case of missing information, while the deep learning model adopted by this platform can intelligently complete the missing points based on the existing spatial distribution characteristics and neighborhood information in the point cloud. For example, users can choose to perform CAD modeling only on fixed objects such as buildings, roads, and bridges, and the system will automatically remove the point data of dynamic objects such as pedestrians and vehicles in the point cloud data before conversion. At this time, the ground or other related structures may produce discontinuous areas due to the lack of point data, but the deep learning model can automatically complete the surrounding data in the process of generating the CAD model, making the modeling results more coherent.
[0074] In addition, since most city-level point cloud data is collected from a bird's-eye view, the bottom surface information of individual buildings, bridges, and other instances is often missing, resulting in incomplete structures when modeling one instance at a time. The deep learning model used in the present invention can automatically generate the missing bottom surface based on the overall structural information when restoring the shape of the object, thereby restoring the complete three-dimensional form. This intelligent completion capability enables the CAD model generated by the system to more accurately reflect the actual scene, improving the applicability of the final modeling result.
[0075] In the model output stage, the system of the present invention can generate CAD files for segmented instances or entire scenes according to user selection, and calibrate the position of each CAD model based on the global coordinate system. The platform's visual interface supports users to freely select specific instance categories to generate CAD files, and can filter out unnecessary dynamic objects during the modeling process. In addition, the generated CAD files can be downloaded and previewed, which is convenient for subsequent BIM modeling, smart city planning, digital twins and other multi-scenario applications.
[0076] In summary, the present invention provides an efficient and flexible city-level point cloud data modeling platform by introducing technologies such as voxel downsampling, quadtree segmentation, and deep learning automatic completion, solving the efficiency bottleneck and incomplete shape problems in large-scale point cloud data processing. Through intelligent segmentation, automatic completion, and three-dimensional modeling, the present invention significantly improves the processing speed and modeling accuracy of point cloud data, and provides reliable technical support for three-dimensional city reconstruction, virtual reality, building information modeling (BIM), and other fields.
[0077] The following is an introduction to the intelligent segmentation and CAD model generation method and system for city-level point cloud scenes provided by the present invention in conjunction with the accompanying drawings and embodiments:
[0078] like Figure 1As shown, the overall system architecture of the present invention is divided into two parts: the front-end and the back-end. The front-end part mainly runs in the user's browser and provides an interactive interface for the user. The front-end module includes: 1. File upload / download, allowing users to upload point cloud files and download the generated CAD files after processing. 2. File management / selection, users can manage and select the uploaded point cloud files to facilitate subsequent processing. 3. Parameter input, providing an input interface for user-defined settings, including the voxel size for point cloud segmentation and the voxel size for point cloud visualization. 4. Visualization, used to display the visualization results of point cloud semantics and instance segmentation, different semantics and instance labels are displayed in different colors, which is convenient for users to select specific instances for CAD conversion. The back-end part consists of programs, databases, asynchronous processing tasks and data storage to realize the core functions of data processing. 1. Program module, mainly for the rapid response of front-end requests, so it is only responsible for processing simple logical tasks, including data transmission and storage, database writing, asynchronous task calls and user authority management. For complex data processing tasks, the back-end will execute this task by calling a subroutine, and immediately send a response to the front-end that the data is being processed after the call is successful. This process is called asynchronous task call. 2. Database, which stores information about point cloud files and generated CAD files, including file path, file status, file permissions, etc., as well as user information, to facilitate data management and permission control for a large number of point cloud files and multiple users. 3. Asynchronous processing tasks, including point cloud file preprocessing, deep learning segmentation, visual point cloud file generation, and point cloud to CAD conversion tasks. Each task module works together to realize the processing and conversion of point cloud data.
[0079] Figure 2 The detailed interaction process between the user and the background (backend) processing module and the task flow of each module are shown. The processing flow of the present invention mainly includes the following steps:
[0080] 1) The user uploads the point cloud file through the interactive interface on the browser side, and after the upload is completed, the data is transferred to the backend server for storage. After the upload process is completed, the user can enter the voxel size for point cloud segmentation and the voxel size for point cloud visualization, and request the system to preprocess the file. Among them, the point cloud is a three-dimensional coordinate data set obtained by a three-dimensional scanning device or other means, usually composed of a large number of irregularly distributed points, which is used to represent the spatial shape and characteristics of an object or scene.
[0081] 2) In the preprocessing stage, after receiving the user's request, the system's preprocessing module performs voxelization and region segmentation on the uploaded point cloud data, dividing the point cloud scene into several sub-regions to ensure that the number of points in each sub-region does not exceed the set video memory threshold. The purpose of this segmentation step is to ensure that the subsequent deep learning model will not cause video memory overload due to excessive input points during processing, thereby ensuring the normal operation of the segmentation model. Among them, voxelization is the process of dividing the continuous three-dimensional space into discrete voxels (Voxel, i.e. three-dimensional pixels), which is used to downsample the point cloud data to reduce computational complexity while retaining the main geometric features.
[0082] 3) In the point cloud segmentation module, the system uses a deep learning model to perform semantic and instance segmentation on the preprocessed sub-region point cloud to generate point cloud data with semantic labels and instance labels. When the segmentation tasks of all sub-regions are completed, the system reorganizes the segmentation results of these sub-regions, and forms a corresponding mapping between the semantic labels and instance labels of each point in the original point cloud and the reorganized point cloud after segmentation to ensure the consistency of the model output results with the original data. Based on the voxel size previously entered by the user for visualization, the system further generates a downsampled point cloud file for visualization. The downsampled point cloud file is used for instance selection and category filtering on the front end to ensure that users can obtain a sufficiently clear instance outline display in the browser without causing excessive consumption of browser hardware resources due to loading too much point data. Semantic segmentation is to assign each point in the point cloud data to a semantic category, such as a building, road or bridge, through deep learning technology, so as to achieve semantic understanding of the scene; instance segmentation is a further processing based on semantic segmentation, which can distinguish different objects in the same semantic category as independent instances, such as distinguishing different buildings in the same scene.
[0083] 4) In the visualization interface of the point cloud visualization module, users can select specific instances or categories for CAD modeling, such as selecting buildings or roads and ignoring dynamic objects (such as pedestrians and vehicles). Or model a specific building. The visualization module can also be replaced with an efficient annotation platform based on multi-resolution level of detail (LOD) technology, and the segmentation results of the deep learning model can be manually adjusted to further improve the accuracy of modeling each instance.
[0084] 5) When the user completes the instance or category selection, the selected parameters will be passed to the point cloud CAD conversion module. According to the user's selection, the module performs CAD conversion operations on each instance within the specified instance range one by one, or generates CAD models for all instances in a specific category in sequence, and temporarily stores the generated models in the storage module. During the conversion process, the system uses the deep learning model to complete the missing areas to ensure the geometric integrity and structural continuity of the model. The point cloud CAD conversion module can also be replaced by a traditional 3D point cloud surface reconstruction algorithm. Compared with the end-to-end direct conversion of point cloud files to CAD files by the deep learning model, the surface reconstruction algorithm is more complex to implement and cannot automatically associate and complete the exact area. However, the traditional surface reconstruction algorithm has better hardware resource requirements and computational efficiency than the deep learning method. When all selected instances have completed the CAD conversion, the system merges the CAD models of each instance into a CAD file of a complete scene, and transmits the merged data to the front-end browser interface.
[0085] 6) In the CAD visualization interface of the CAD visualization module, users can preview the model effect of the entire scene to ensure that the modeling results meet the requirements, and download the final complete CAD file in this interface. After downloading, users can fine-tune the model according to the actual size.
[0086] The above introduces the overall architecture of the system. Next, the functions of each module and its technical implementation will be introduced one by one. Figure 3 The preprocessing module in the present invention is shown. This module is designed to downsample and block large-scale point cloud data to ensure the normal reasoning process of the subsequent deep learning model and prevent the model processing from failing due to video memory overload. The workflow of the preprocessing module is as follows: Figure 3As shown in the figure, the original point cloud data and the voxel size parameter specified by the user are first received as input, where the original point cloud data contains three-dimensional point information at the city level or other large-scale scenes, and the voxel size is used to control the accuracy of downsampling. In the downsampling step, the system performs sparse processing on the original point cloud data based on the voxel size, adjusts the data space density through the voxel grid, reduces the data volume while retaining the main geometric structure, thereby reducing the computational load. The specific operation of downsampling is to divide the x, y, and z coordinates of each point in the point cloud data by the voxel size, round down, and finally remove the points with the same x, y, and z. After this step, one point is retained in each voxel median. The downsampled point cloud data will enter the point count judgment link, which judges whether the data volume exceeds the threshold set by the system by counting the number of points. The threshold is set according to the video memory capacity of the graphics card to ensure the safe reasoning of the subsequent deep learning model. If the number of points in the downsampled point cloud data does not exceed the threshold, it will be directly passed to the next module; if the number of points exceeds the threshold, it will enter the block process. The block processing adopts the quadtree algorithm. First, the point cloud scene is divided into four blocks on the x and y planes. Then, the number of points in each block is determined one by one. If it still exceeds the threshold, the recursive block division is continued until the number of points in each sub-block is within the threshold range. During the block division process, a 5% common area will be reserved for adjacent sub-blocks for subsequent splicing. After the block processing is completed, the system records each sub-block as an independent sub-scene in the point cloud sub-scene list, which contains the location information and point cloud data of each sub-block. These sub-scenes will be loaded and processed one by one in the subsequent point cloud segmentation module. The pre-processing module combines voxel downsampling and quadtree block technology to decompose large-scale high-density point cloud data into multiple memory-friendly sub-areas, thereby ensuring the normal progress of deep learning reasoning. The combination of voxel downsampling and quadtree block technology ensures that each sub-scene can guarantee sufficient receptive field without losing geometric structure, effectively improving the efficiency and accuracy of subsequent semantic and instance segmentation models in processing large-scale high-density point cloud scenes.
[0087] Figure 4The point cloud segmentation module in the present invention is shown. The function of this module is to segment the preprocessed point cloud data to generate a complete point cloud scene containing semantic and instance labels. The module receives the point cloud sub-scene list from the preprocessing module and performs inference processing of the deep learning model on each sub-scene data in the list in turn. In this process, the point cloud segmentation module will choose block-by-block inference or multi-graphics card parallel inference according to the available resources to ensure processing efficiency. After each inference, the system temporarily stores the inference results in the memory and deletes the processed sub-scenes from the point cloud sub-scene list until the list is empty. When the list is empty, it indicates that all sub-scenes have been segmented, and the system immediately enters the sub-scene merging stage. During the merging process, the module unifies the segmentation results of each sub-scene through the common area alignment technology, ensures that the semantics and instance labels between each sub-scene are consistent, and generates a complete global segmentation result. Next, based on the shortest distance principle, for each point in the original point cloud data, the system searches for the nearest point in the segmented point cloud scene in a cube space with a side length of two times the voxel centered on it. Each point in the original point cloud data is matched with the label in the segmentation result, adding semantic and instance labels to the original data. Finally, the system generates large-scale high-density point cloud data with semantic and instance labels. In addition, the module downsamples the segmentation results according to the voxel size set by the user to generate low-density point cloud data suitable for visualization, so that users can intuitively select specific instances or categories in the subsequent visualization module.
[0088] Figure 5The point cloud visualization module of the present invention is shown. After the point cloud segmentation module is processed, the visualization module receives the voxelized downsampled scene point cloud data with semantic and instance labels, and performs LOD (level of detail) processing on it. Through the multi-resolution point cloud data generation mechanism, efficient rendering is achieved on the user's browser side. LOD processing enables the system to automatically load the adapted resolution data according to the viewing angle and zoom level adjusted by the user in real time in the browser, thereby balancing the rendering performance and display effect, and ensuring the smooth visualization of large-scale point cloud data in the browser. In the visualization interface, the user can intuitively view and filter point cloud data of different instances and categories, and select a specific modeling object type (such as fixed facilities such as buildings and roads) or select one or several instance objects for modeling (such as a certain building, several sections of roads, etc.). In addition, the user can set the voxelization size to control the modeling accuracy and file data volume, ensuring that the generated CAD model meets the actual needs of the scene and maintains the efficiency of file conversion. The user's selection and setting parameters will be further transmitted to the point cloud CAD conversion module for processing. The point cloud visualization front-end interface of the visualization module in the present invention can also be replaced by a large scene point cloud annotation interface. In this way, users can modify the semantic labels or instance labels predicted by the deep learning model according to their needs. This modification operation is first performed in the front-end interface. After the user confirms, the updated label information is synchronized to the database through the background program to ensure the consistency and accuracy of the data.
[0089] Figure 6The detailed processing flow of the point cloud CAD conversion module in the present invention is shown. The module first receives the instance label, category label and voxel size parameter from the visualization module, and obtains the high-density point cloud data of large-scale scenes with semantic and instance labels. The system filters the point cloud data according to the category label and instance label, and retains the category or instance point cloud data specified by the user to ensure that the subsequent conversion operation focuses on the object of user concern. Then, the system performs voxel downsampling processing on the filtered high-density point cloud data, thereby reducing the amount of data while maintaining the main structural information to improve the conversion efficiency. In order to adapt to the diverse application requirements, the system allows users to customize the voxel size parameters for conversion. The flexibility of setting the voxel size enables the system to adapt to different modeling requirements. For example, when the user needs to model the entire city-level scene as a whole, the system pays more attention to the global positional relationship between the instances, and has lower requirements for the details of a single instance. At this time, the user can set the voxel size to 0.1 meters to reduce the number of points in each instance point cloud, thereby significantly improving the CAD conversion efficiency of the subsequent deep learning model. On the contrary, when the user wants to model a specific instance (such as a building) with high precision, the voxel size can be set to 0.01 meters to retain more detailed information. Although a smaller voxel size will reduce the conversion speed accordingly, it can ensure that the generated CAD model accurately displays the instance details. In summary, users can flexibly adjust the voxel size parameters according to different application scenarios to balance model accuracy and conversion efficiency.
[0090] The downsampled point cloud data is processed one by one according to the instance label. The system extracts the data of each independent instance in turn and performs CAD conversion on it. During the conversion process, the system uses the deep learning model to complete the missing structures in each instance to ensure the geometric integrity of the CAD model. After the point cloud data of each instance is converted into a CAD instance model, the system saves the conversion result as a separate CAD file. When the CAD conversion of all instances is completed, the system merges the single instance CAD files one by one into a complete scene CAD file to ensure that different instances are accurately positioned in the global coordinate system. Finally, the merged CAD file is passed to the CAD visualization module.
[0091] In the CAD visualization interface, users can intuitively browse the 3D model of the entire scene. The interface supports real-time observation and zooming in and out from multiple angles, allowing users to examine in detail the geometry, structural integrity, and relative position of each instance in the scene. Users can also use the interactive tools in the interface to fine-tune the model, including correction of geometric errors, edge alignment, structural improvement, etc., to ensure the accuracy and integrity of the final model.
[0092] After fine-tuning, users can export the final CAD model to a variety of common file formats for continued use in other software platforms. The generated CAD files can be used in a variety of application scenarios, such as detailed building and facility modeling in building information models (BIM); as basic geographic data support in smart city systems; for scene reconstruction in virtual reality (VR) or augmented reality (AR) environments, or for spatial analysis and management in geographic information systems (GIS). This variety of applications not only expands the scope of application of the system, but also provides strong support for users' different business needs.
[0093] In summary, the present invention provides users with a flexible and accurate 3D modeling platform through intelligent point cloud segmentation and efficient CAD conversion functions. The system's multi-resolution processing, deep learning intelligent completion, and visual fine-tuning functions combine to make the modeling process of large-scale urban scenes more efficient and convenient. Whether it is used for urban planning, digital twins, or intelligent scene reconstruction, the system can provide users with reliable technical support and accurate 3D model output.
[0094] Other embodiments of the present invention:
[0095] 1. In terms of voxel downsampling of point cloud data, in order to further expand the scope of technical protection and implementation methods, the following alternatives can be considered:
[0096] Random point deletion: Instead of voxel downsampling technology, random point deletion can be used to reduce the density of point cloud data. Random deletion technology reduces the amount of point data while retaining global geometric features by setting random sampling rates or deletion rules. This method is suitable for scenes with relatively uniform geometric distribution or scenes that do not require high precision. For example, points in the point cloud are randomly selected at a preset ratio, and the remaining points are deleted to form sparse point cloud data for subsequent block and segmentation operations.
[0097] Sampling techniques based on regional importance: When replacing voxelization techniques, sampling methods based on regional importance can be introduced, such as retaining more points in feature-dense areas (such as edges or surfaces) in the point cloud and downsampling more in flat areas (such as the ground), thereby taking into account both sparseness and detail preservation. This method can combine the normal features, curvature or distance thresholds of the point cloud for importance assessment.
[0098] Clustering-based downsampling technology: Instead of voxelization methods, point cloud clustering-based downsampling technology can also be used, such as K-means clustering or DBSCAN algorithm. By clustering, adjacent points are grouped together, and each group retains a representative point to form a sparse point cloud, which not only reduces the amount of data but also retains the main spatial structure of the point cloud.
[0099] Resolution adaptive downsampling: Dynamically adjust the sampling method according to the spatial resolution and point cloud density of the scene, such as using a higher sampling rate for denser areas and a lower sampling rate for low-density areas to form an adaptively sparse point cloud. This method can be implemented through distance thresholds or point cloud distribution histogram analysis as a flexible alternative to voxelization.
[0100] Optimization based on sampling rules: Adopt a rule-based sampling method, such as sparse sampling of the point cloud on the top of a building with dense distribution in the Z-axis direction, or optimize the distribution of point cloud data according to the scanning trajectory of the acquisition equipment to achieve the purpose of downsampling.
[0101] 2. In terms of semantic instance segmentation processing of point cloud data, in order to further expand the scope of technical protection and implementation methods, the following alternatives can be considered:
[0102] Traditional segmentation algorithms: In addition to using deep learning models for point cloud semantic segmentation and instance recognition, segmentation methods based on traditional algorithms can be used, such as point cloud segmentation techniques based on region growing, hierarchical clustering, or density clustering. These methods can be used as a supplement or alternative to deep learning models, especially in some scenarios with small data volumes or low precision requirements.
[0103] Rule-based feature extraction algorithms: directly perform instance segmentation by identifying the characteristics of geometric shapes (such as planes, cylinders, or edges).
[0104] 3. In terms of point cloud data segmentation, the following alternatives may be considered to further expand the scope of technical protection and implementation methods:
[0105] Density-based segmentation algorithm: Instead of the quadtree segmentation algorithm, dynamic block technology based on grid partitioning or point cloud density can be used to adapt to different data structures and processing requirements.
[0106] 4. Alternatives to point cloud visualization interfaces:
[0107] It can replace the existing LOD visualization interface and adopt an efficient rendering method based on layered cache to ensure the smoothness of real-time user operations. It can enhance the interactive function and replace the traditional mouse interaction mode through gesture recognition, voice control and other methods to further improve the user experience.
[0108] 5. Flexible organization of multi-instance scenarios:
[0109] On the basis of instance-by-instance modeling, it can support nested hierarchical management of multi-instance modeling. For example, sub-instances (such as doors, windows, roofs) can be further divided within the same building instance and modeled separately, expanding the flexibility of hierarchical modeling.
[0110] 6. Hardware architecture replacement:
[0111] In terms of hardware implementation, point cloud segmentation and modeling can be realized based on a distributed computing environment, replacing the single graphics card or stand-alone mode. Distributed processing of point cloud data through cloud computing clusters or edge computing devices can further improve processing scale and efficiency.
[0112] It should be understood that the methods and systems disclosed in the above embodiments provided by the present invention can be implemented in other ways. For example, the division of the above modules can have other division methods in actual implementation, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. The various modules in the present invention can be implemented in the form of software functional units, which can be stored in a computer-readable storage medium, including several instructions for a computer device to perform some or all steps of the method described in the present invention. For example, one embodiment of the present invention provides a computer device (computer, server, smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention. For example, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk, etc.), the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the various steps of the method of the present invention are implemented.
[0113] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. It can be understood by those skilled in the art that various replacements, changes and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the contents disclosed in the embodiments of this specification, and the scope of protection of the present invention shall be subject to the scope defined in the claims.
Claims
1. A method for intelligent segmentation and CAD model generation for city-level point cloud scenes, characterized in that: The following steps are involved: Pre-process the point cloud data uploaded by the user, and divide the point cloud scene into several sub-regions by performing voxelization and region segmentation; Use the deep learning model to segment the preprocessed sub-areas and generate point cloud data with semantic labels and instance labels; Using the point cloud data with semantic labels and instance labels generated after point cloud segmentation, point cloud visualization is performed in the visualization interface according to the instance or category selected by the user; Perform CAD conversion operations based on the instance or category selected by the user to convert the point cloud file into a CAD file; Based on the CAD file obtained by the conversion operation, a three-dimensional model of the entire scene is displayed in the CAD visualization interface.
2. The method according to claim 1, characterized in that The preprocessing of the point cloud data uploaded by the user includes: Receiving raw point cloud data and a voxel size parameter specified by a user as input, wherein the raw point cloud data contains three-dimensional point information of a city level or other large-scale scenes, and the voxel size is used to control the accuracy of downsampling; Downsample the original point cloud data based on the voxel size; The number of points in the downsampled point cloud data is judged. If the number of points in the downsampled point cloud data does not exceed the set threshold, the point cloud segmentation step is entered; if the number of points exceeds the threshold, block processing is performed, and the number of points in each sub-block is within the threshold range; After the block processing is completed, each sub-block is recorded as an independent sub-scene in the point cloud sub-scene list, which contains the location information and point cloud data of each sub-block.
3. The method according to claim 2, characterized in that The block processing adopts a quadtree algorithm, first dividing the point cloud scene into four blocks on the x and y planes, and then judging whether the number of points in each block meets the requirements one by one. If it still exceeds the threshold, recursive block division continues until the number of points in each sub-block is within the threshold range.
4. The method according to claim 2, characterized in that The method of performing point cloud segmentation on the preprocessed sub-regions using a deep learning model includes: Receive the point cloud sub-scene list obtained after preprocessing, and perform inference processing of the deep learning model on each sub-scene data in the list in turn. During the inference process, choose block-by-block inference or multi-graphics card parallel inference according to available resources; After each inference, the inference result is temporarily stored in the memory, and the processed sub-scenes are deleted from the point cloud sub-scene list until the list is empty; When the list is empty, the sub-scene merging phase begins. During the merging process, the segmentation results of each sub-scene are unified through the common area alignment technology to ensure that the semantics and instance labels of each sub-scene are consistent and generate a complete global segmentation result. Based on the shortest distance principle, for each point in the original point cloud data, the nearest point in the segmented point cloud scene is searched in a cube space with a side length of twice the voxel centered on it, and each point in the original point cloud data is matched with the label in the segmentation result to generate large-scale high-density point cloud data with semantic and instance labels; The segmentation results are downsampled according to the voxel size set by the user to generate low-density point cloud data suitable for visualization.
5. The method according to claim 1, characterized in that: The point cloud visualization is performed in the visualization interface according to the instance or category selected by the user, including: Receive voxelized downsampled scene point cloud data with semantic and instance labels, perform LOD processing on it, and achieve efficient rendering on the user's browser through a multi-resolution point cloud data generation mechanism; The user can intuitively view and filter point cloud data of different instances and categories in the visual interface, and select a specific modeling object type or select one or several instance objects for modeling; Users set the voxel size to control modeling accuracy and file data volume, ensuring that the generated CAD model meets the actual needs of the scene while maintaining high efficiency of file conversion.
6. The method according to claim 1, characterized in that The performing of the CAD conversion operation according to the instance or category selected by the user to convert the point cloud file into a CAD file includes: Receive instance labels, category labels and voxel size parameters, and obtain large-scale scene high-density point cloud data with semantic and instance labels; Filter and process point cloud data according to category labels and instance labels, and retain category or instance point cloud data specified by the user; Perform voxel downsampling processing on the filtered high-density point cloud data; The downsampled point cloud data is processed one by one according to the instance label, and the data of each independent instance is extracted and converted to CAD in turn. During the conversion process, the deep learning model is used to complete the missing structure in each instance. After the point cloud data of each instance is converted into a CAD instance model, the conversion result is saved as a separate CAD file. When the CAD conversion of all instances is completed, the CAD files of the individual instances are merged one by one into a complete scene CAD file.
7. The method according to claim 1, characterized in that The CAD visualization interface supports real-time observation and zooming in and out from multiple angles, allowing users to examine in detail the geometry, structural integrity and relative position of each instance in the scene; The CAD visualization interface provides interactive tools for users to fine-tune the model, including correction of geometric errors, edge alignment, and structural improvement, to ensure the accuracy and completeness of the final model.
8. An intelligent segmentation and CAD model generation system for city-level point cloud scenes, characterized in that: include: The preprocessing module is used to preprocess the point cloud data uploaded by the user and divide the point cloud scene into several sub-regions by performing voxelization and region segmentation; Point cloud segmentation module, which is used to segment the preprocessed sub-areas using a deep learning model to generate point cloud data with semantic labels and instance labels; The point cloud visualization module is used to use the point cloud data with semantic labels and instance labels generated after point cloud segmentation to visualize the point cloud according to the instance or category selected by the user in the visualization interface; A point cloud CAD conversion module is used to perform CAD conversion operations according to the instance or category selected by the user, and convert the point cloud file into a CAD file; The CAD visualization module is used to display the three-dimensional model of the entire scene in the CAD visualization interface according to the CAD file obtained by the conversion operation.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Dynamic and static combined multi-scale point cloud adaptive characterization method and system
CN120852764A