Device and method for automated, three-dimensional building data modelling

EP4659137A1Pending Publication Date: 2025-12-10FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024701692
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-31
Filing Date
2024-01-24
Publication Date
2025-12-10

AI Technical Summary

Technical Problem

Traditional methods for creating Building Information Models (BIM) are inefficient and error-prone, particularly when dealing with existing buildings, as they rely on manual data collection and processing of point clouds, leading to timeliness and accuracy issues due to the complexity and detail of scenes.

Method used

A device and method utilizing an artificial intelligence module trained with machine learning to process point cloud data, enabling automated conversion of point clouds into geometric models, which includes semantic segmentation to remove ambient noise and instance segmentation for accurate object-level modeling.

Benefits of technology

Significantly improves the efficiency and accuracy of BIM model creation by automating the conversion of point clouds into geometric models, ensuring precise and timely generation of BIM data for existing buildings, enhancing the combination of 3D scanning and BIM technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024051592_08082024_PF_FP
    Figure EP2024051592_08082024_PF_FP
Patent Text Reader

Abstract

The invention relates to a device for creating a digital model of an existing building, or an existing industrial plant, or an existing infrastructure, according to one embodiment. The device comprises an input unit (110) for providing point cloud data representing the building or the industrial plant or the infrastructure. The device also comprises a processing unit (120) having an artificial intelligence module (125) for creating the model according to the point cloud data, wherein the artificial intelligence module (125) is trained using machine learning.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Device and method for automated, three-dimensional building data modelling

[0002] Description

[0003] The application relates to building data modelling, in particular to a device and a method for automated building data modelling.

[0004] Building Information Modeling (BIM) describes a working method for the networked planning, construction, and management of buildings and other structures using software. It is primarily used for building visualization, cost estimation, construction planning, conflict, fault, and collision detection, and facility management (see [1], [2]). BIM can therefore avoid planning errors, optimize material usage, reduce construction and maintenance costs, and improve the efficiency of resource utilization throughout the entire construction industry chain (see [2]). Especially in industrialized countries with low new construction rates, the activities of the construction sector are increasingly shifting toward conversions, retrofitting, and the demolition of existing buildings (see [3]). The use of BIM is therefore essential for the cost-effective maintenance of existing buildings.

[0005] The traditional methods for creating BIM models are divided into the following two main types: First, there is computer-aided 3D modeling, which mainly relies on basic measuring instruments such as laser rangefinders, digital cameras, and tape measures to capture the dimensions of all building formations through manual measurements and uses commercial software (such as Autodesk Revit, AutoCAD, 3DS MAX, etc.) to perform modeling based on the measurement parameters. The second method is the generation of BIM models based on early industrial design drawings, for example, using the original 2D design drawings of the building as a reference and generating BIM models based on the drawing parameters (see [4]).These two traditional methods can provide acceptable BIM models in some situations, but suffer from difficulties in data acquisition, low automation, inefficiency, or lack of timeliness due to missing drawings or changes during renovation of building facilities. With the development of computer, sensor, and image processing technology, a 3D point cloud can be quickly generated, which consists of numerous and scattered points with coordinates (X, Y, Z) and attributes (e.g., RGB colors, intensity). This form of data provides effective support for 3D modeling because it effectively represents the complex real world. The difficulties of data acquisition are becoming increasingly less, and this can improve the timeliness and accuracy of the data compared to traditional methods. Therefore, 3D scanning sensors are increasingly being used to generate BIM for existing buildings.These generate 3D point clouds from which building parameters can be derived. In industrial contexts, employees are employed to manually convert point clouds into BIM models. This process requires software interaction, is time-consuming, and error-prone. As a result, the advantages of 3D scanning technology and BIM technology are not utilized efficiently.

[0006] This approach has several problems: On the one hand, the system can become extremely slow or even fail due to the large amount of point cloud data involved in modeling projects (see [5]). On the other hand, manual 3D modeling of scenes with many details can be time-consuming and error-prone (see [6]-[8]). Furthermore, if the scenes are too complex, walls and other areas of interest may be missing, which presents many challenges (see [8], [9]).

[0007] The traditional method for creating a BIM model involves measuring building data using measuring instruments or using data from existing 2D drawing files to obtain building parameters and manually modeling the building using this data, as shown in Fig. 2. Most renovation and real estate companies still use these traditional methods to solve the problem. Thus, Fig. 2 illustrates traditional BIM construction methods.

[0008] When converting 2D construction drawings to BIM, several research institutions have published methods for improving modeling efficiency in 3D geometric modeling. On the one hand, semi-automatic reconstruction of virtual 3D architectural models from 2D architectural drawings can improve the efficiency of the previous manual reconstruction process (see

[0010] ). On the other hand, improvement can be achieved by combining the geometry and semantics from multiple construction drawings and tables to integrate and normalize building data (see

[0011] ). One approach is to combine an interactive sketching framework with perception-based techniques, using the hinge angle scheme of interactive sketching to achieve stable, accurate, and more automated 3D model construction (see

[0012] ).Using matching, mapping, and classification methods, interaction during 3D reconstruction can be reduced, allowing the basic structure of the building model to be reconstructed easily and quickly (see

[0013] ). Although several approaches based on 2D construction drawings exist, it has not been possible to fully automate the generation of 3D building models from 2D construction drawings, as the representation of building drawings varies somewhat from region to region, and the drawings do not necessarily reflect the current construction status.

[0009] With the advancement of hardware and software technologies, advanced methods have been developed. Using laser scanners and various cameras (monocular cameras, binocular cameras, and drone cameras), it is possible to efficiently collect point cloud data of a building's current interior and exterior scenes. Then, using point cloud data processing and appropriate design drawing software, the building parameters can be extracted and modeled from the existing scenes.

[0010] Fig. 3 illustrates advanced BIM construction methods in a point cloud workflow. Specifically, Fig. 3 shows the workflow for today's point cloud generation and processing. The novel approach to the BIM creation process is highlighted with a dotted line in Fig. 3. Meshing is a well-known standard process.

[0011] The point cloud generation is described below.

[0012] Point cloud generation is essential in this advanced BIM design method. A distinction is made between two main types of point cloud data: the first type is obtained indirectly through the processing of photo / video images, and the second type is obtained directly through the sensors that generate the point cloud data. Because direct point cloud data sensors are not easy to use and hardware costs are high, the majority of current research focuses on acquiring point cloud data from 2D images through photo / video image processing, initially based on the "Structure from Motion" (SfM) technique to obtain sparse point cloud data, which can then be processed to produce dense point clouds.

[0013] The main steps for generating point clouds are shown on the left side of Fig. 3.

[0014] The following section explains the algorithms and methods currently used for each of the main steps of point cloud generation.

[0015] First, feature extraction is described. This is described, for example, in references

[0014] to

[0029] . The algorithms used in the feature extraction step can be divided into two categories: feature point detectors, which determine the position of feature points in the image, and feature point descriptors, which generate vectors or character strings to describe the feature points detected by the feature point detectors. Feature points are some special points in the image data that have certain unique properties and more information than ordinary points and can be used to describe the key information in the image based on these feature points. Once the feature points and their descriptions are known, the correspondence of these points in different images can be used for subsequent algorithms.

[0016] Next, feature matching is described. This is described, for example, in the references

[0017] ,

[0021] and

[0030] to

[0037] . Thus, after the feature points of each image have been extracted, the feature points in each image should be matched, i.e., the corresponding feature points of one image should be found in the other image. If two points in different images have the same description, then these points can be considered identical in terms of their appearance in the scene; if two images have a common set of points, then they can be said to depict a common part of the scene. Various strategies can be used to efficiently compute correspondences between images. The result of this phase is a set of images that overlap at least pairwise, and a set of correspondences between the features. Motion estimation is now described.This is illustrated, for example, in

[0021] ,

[0024] ,

[0029] ,

[0032] , and

[0038] to

[0042] . The goal of this step is to find the camera parameters for each pair of images using the feature points of each image pair, both intrinsic (i.e., focal length and radial aberration) and extrinsic (i.e., rotation and translation) parameters. In the polar geometry of an image pair, two matrices can be used to describe the correspondence between the matching feature points of the image pair, namely the essential matrix and fundamental matrix

[0053] . The essential matrix

[0053] contains the extrinsic parameters that describe the correspondence between matching feature points in camera coordinates, while the fundamental matrix

[0053] contains the intrinsic and extrinsic parameters that describe the correspondence between matching feature points in pixel coordinates.Typically, the five-point algorithm is used to calculate extrinsic parameters, while the eight-point algorithm is used to calculate both intrinsic and extrinsic parameters. In other words, the five-point algorithm is more suitable when the intrinsic parameters are known and the goal is to calculate extrinsic parameters, while the eight-point algorithm is more suitable when the goal is to calculate both extrinsic and intrinsic parameters. Direct linear transformation (DLT): Direct Linear Transform (DLT) uses feature points in pixel coordinates and corresponding 3D points in absolute coordinates in the scene to calculate intrinsic and extrinsic parameters based on the least squares method, while the five- and eight-point algorithms do not require 3D points in absolute coordinates in the scene.

[0017] Next, sparse 3D reconstruction is described. This is presented, for example, in references

[0015] ,

[0017] ,

[0026] , and

[0043] . The goal of this step is to compute the 3D points by using the feature points with correspondences obtained from the "Feature Matching" section and the camera parameters of each image obtained from the "Motion Estimation" section to generate a point cloud of the scene.

[0018] Parameter correction is now described. This is presented in references

[0015] -

[0017] ,

[0021] ,

[0035] , and

[0044] . The goal of this step is to correct the camera parameters of each image to reduce the reprojection error. To do this, the intrinsic and extrinsic camera parameters for each image from the "Motion Estimation" section and the 3D positions of the points in the point cloud from the previous step are used to generate the precise camera parameters of each image and the exact 3D positions of the points in the point cloud.

[0019] The next step is dense 3D reconstruction. This step is described in references

[0015] ,

[0017] ,

[0021] ,

[0043] , and

[0045] -

[0052] . The goal of this step is to recover the details of the scene by using the images of a given scene, the intrinsic and extrinsic camera parameters for each image obtained from the "Parametric Correction" section, and the sparse point cloud obtained from the previous step to generate a dense point cloud.

[0020] The point cloud processing is described below.

[0021] The above-mentioned steps in point cloud generation result in a complete, dense point cloud from the raw data of the underlying scene. This can then be processed to obtain the final 3D model. 3D models are divided into three main types: tessellated (meshed) models, point cloud models, and geometric models. The requirements for 3D models vary in different scenarios. In industry, geometric models are primarily used due to the high demands on model accuracy. Geometric models are more versatile and can be easily converted into the other two 3D models. 3D BIM is a typical format for geometric models.

[0022] Due to the limitations of the dense point cloud generated in the point cloud generation step, it is not possible to directly meet the application requirements. For example, the generated point cloud contains not only the target object information but also other point cloud data from the surrounding environment, which can be an obstacle to further application of the point cloud. Therefore, according to current research, point cloud preprocessing, point cloud segmentation, and point cloud modeling are essential to meet the application requirements of converting point clouds into geometric models.

[0023] The algorithms and methods used in the various steps of processing point clouds into geometric models are explained below.

[0024] First, point cloud preprocessing is described. This is shown in

[0015] ,

[0017] , and

[0054] to

[0057] . The purpose of this step is to create a point cloud model or generate a high-quality point cloud for subsequent steps. As shown in Fig. 3, this step consists of four substeps: alignment, noise filtering, outlier removal, and downsampling. Which of the four substeps is used depends on the application requirements and the quality of the point cloud.

[0025] Point cloud segmentation is now described. This is presented in references

[0054] and

[0058] -

[0062] . The goal of this step is to segment the point cloud to obtain the points of the object of interest. Although the algorithms used in the publications differ slightly, the traditional algorithms mainly used for point cloud segmentation today can be divided into two categories: feature-based segmentation algorithms and model-based segmentation algorithms. Feature-based segmentation algorithms segment a point cloud based on the features of each point in the cloud, grouping points with the same features into a subset.Model-based segmentation algorithms segment a point cloud based on a mathematical model of the object under consideration, where the mathematical model used in model-based segmentation is usually a plane equation, which can usually be mathematically parameterizable geometries such as plane, cylinder, torus, etc.

[0026] Finally, point cloud modeling is described. This is illustrated in

[0056] and

[0063] . The goal of this step is to create a geometric model of the object under consideration using the segmented point cloud segments from the previous step. To create a geometric model of the object in question, the model parameters should be determined based on the results of point cloud segmentation. The method for determining the model parameters generally involves determining the dimensions of the object.For example, the radius of a cylindrical pipe is one of its model parameters and can be determined by calculating the shortest distance between the centerline and the pipe's point cloud obtained from point cloud segmentation

[0056] . Alternatively, by projecting the pipeline's point cloud onto a plane perpendicular to the pipeline's centerline to generate a resulting circle and calculate the circle's radius (see

[0046] and

[0064] ). Using the model parameters, a geometric model of the target object can be easily created.

[0027] An apparatus according to claim 1, a method according to claim 26 and a computer program according to claim 27 are provided.

[0028] A device for creating a digital model of an existing building or an existing industrial facility or an existing infrastructure according to one embodiment is provided. The device comprises an input unit for providing point cloud data representing the building, industrial facility, or infrastructure. Furthermore, the device comprises a processing unit comprising an artificial intelligence module for creating the model based on the point cloud data, wherein the artificial intelligence module is trained using machine learning.

[0029] Furthermore, a method for creating a digital model of an existing building or an existing industrial facility or an existing infrastructure is provided according to one embodiment. The method comprises:

[0030] Providing point cloud data representing the building, industrial facility, or infrastructure. And:

[0031] Creating the model depending on the point cloud data using a processing unit that includes an artificial intelligence module, wherein the artificial intelligence module is trained using machine learning.

[0032] Furthermore, a computer program with program code for implementing the method described above is provided. Embodiments aim to solve the above-mentioned problems, improve the automation and accuracy of existing BIM design, and achieve a perfect combination of the advantages of 3D scanning technology and BIM technology.

[0033] Embodiments mainly concern the point cloud processing phase, while the prior art concerns the point cloud generation and application phase.

[0034] In particular, embodiments relate to the process of converting point clouds into geometric models.

[0035] According to embodiments, the efficiency and accuracy of the process of converting point clouds into geometric models will be significantly improved and it will be possible to automate the process.

[0036] Embodiments provide semantic segmentation to remove the influence of ambient noise from the point cloud data.

[0037] In embodiments, segmentation of point cloud targets at the object level is performed by instance segmentation to obtain geometric and positional parameters of the target object, e.g., rooms, furniture, walls.

[0038] According to one embodiment, a database is extended by converting existing geometric models into point cloud data to achieve better segmentation accuracy.

[0039] Preferred embodiments of the invention are described below with reference to the drawings.

[0040] The drawings show:

[0041] Fig. 1 shows an apparatus for creating a digital model of an existing building, industrial facility, or infrastructure according to one embodiment. Fig. 2 shows traditional BIM construction methods.

[0042] Fig. 3 shows advanced BIM construction methods in a point cloud workflow.

[0043] Fig. 4 a target building scene with point cloud noise containing environmental information after point cloud registration.

[0044] Fig. 5 shows different methods and groups of approaches for semantic segmentation of point clouds.

[0045] Fig. 6 shows a state after removing ambient noise information.

[0046] Fig. 7 shows the target object without any other buildings.

[0047] Fig. 8 shows different methods for instance segmentation of point clouds.

[0048] Fig. 9 shows a flow of a feedback process.

[0049] Fig. 10 shows an extraction of boxes from a point cloud according to a

[0050] Embodiment to obtain training data.

[0051] Fig. 11 shows a visualization of the data augmentation according to a first embodiment.

[0052] Fig. 12 shows a visualization of the data augmentation according to a second embodiment.

[0053] Fig. 1 shows an apparatus for creating a digital model of an existing building or an existing industrial facility or an existing infrastructure according to one embodiment.

[0054] The device comprises an input unit 110 for providing point cloud data representing the building, industrial facility, or infrastructure. Furthermore, the device comprises a processing unit 120, which includes an artificial intelligence module 125 for creating the model based on the point cloud data, wherein the artificial intelligence module 125 is trained using machine learning.

[0055] An infrastructure can be, for example, a train station, a platform, a bridge, or another type of infrastructure.

[0056] According to one embodiment, the input unit 110 can be configured, for example, to receive two or more point cloud partial data sets, each representing a partial area of ​​the building, a partial area of ​​the industrial facility, or a partial area of ​​the infrastructure. The input unit 110 can be configured, for example, to generate the point cloud data representing the building, the industrial facility, or the infrastructure from the two or more point cloud data sets.

[0057] In one embodiment, the input unit 110 can be configured, for example, to receive sensor data from one or more cameras and / or from one or more laser scanners and / or from one or more depth cameras. The input unit 110 can be configured, for example, to determine the point cloud data and a portion of the point cloud data depending on the sensor data; or wherein the sensor data comprises the point cloud data or a portion of the point cloud data.

[0058] According to one embodiment, the processing unit 120 may, for example, be configured to detect data in the point cloud data that does not represent the building, industrial facility, or infrastructure, and / or that is located outside the building, industrial facility, or infrastructure. For example, these detected data that do not represent the building, industrial facility, or infrastructure may subsequently be removed (e.g., if necessary).

[0059] In one embodiment, the processing unit 120 may, for example, be configured to determine boundaries of the building or industrial plant or infrastructure in order to detect the data in the point cloud data that lie outside the building or industrial plant or infrastructure.

[0060] According to one embodiment, the point cloud data may include information about surfaces of the building or industrial plant or infrastructure and information about the interior of the building or industrial plant or infrastructure.

[0061] In one embodiment, the processing unit 120 may, for example, be configured to process the point cloud data such that one or more subsets of the point cloud data are classified such that the one or more subsets are each assigned to one of one or more object categories.

[0062] According to one embodiment, the one or more object categories may, for example, include one or more building component types. The processing unit 120 may, for example, be configured to perform the classification in such a way that it includes a first instance segmentation that classifies at least one subset of the one or more subsets of the point cloud data such that it is classified as one of the one or more building component types of the building, industrial facility, or infrastructure.

[0063] According to one embodiment, the one or more building component types can be freely defined, for example, by a developer. A developer is, for example, a user who configures the device.

[0064] In one embodiment, the one or more building part types may include, for example, one or more of the following building part types: a roof, an exterior wall, an exterior door, an exterior window, an interior space, a stairwell, an elevator shaft, or another type of building part type.

[0065] According to one embodiment, the one or more object categories may, for example, comprise one or more spatial structure types. The processing unit 120 may, for example, be configured to perform the classification in such a way that it comprises a second instance segmentation that classifies at least one subset of the one or more subsets of the point cloud data in such a way that it is classified as a spatial structure from the one or more spatial structure types of a room of the building, industrial facility, or infrastructure.

[0066] In one embodiment, the one or more spatial structure types can be freely definable, for example, by a developer. A developer is, for example, a user who configures the device. In one embodiment, the one or more spatial structure types can include, for example, one or more of the following spatial structure types: a wall, a ceiling, a floor, a door, a window, furniture, or another type of spatial structure type.

[0067] According to one embodiment, the processing unit 120 may, for example, be configured to classify the one or more subsets of the point cloud data in such a way that classification is performed by using the artificial intelligence module 125, which is trained by means of machine learning and which is configured to receive the point cloud data as input data.

[0068] In one embodiment, the artificial intelligence module 125 may be trained using supervised deep learning, for example.

[0069] According to one embodiment, the artificial intelligence module 125 can be designed, for example, to perform the classification in a projection-based and / or discretization-based and / or point-based manner.

[0070] In one embodiment, the artificial intelligence module 125 may, for example, comprise a deep learning neural network having at least two hidden layers (e.g., at least two inner layers, i.e., layers that do not have input and output nodes).

[0071] According to one embodiment, the artificial intelligence module 125 may, for example, comprise a recurrent neural network for determining one or more features of a portion of the point cloud data.

[0072] In one embodiment, the artificial intelligence module 125 may be trained, for example, by means of machine learning by assigning a first point cloud to an object category of the one or more object categories, and wherein the first point cloud and said object category were used as a first training data set for the artificial intelligence module 125.

[0073] According to one embodiment, object classes or categories are open or freely definable.

[0074] According to one embodiment, the artificial intelligence module 125 may be trained, for example, by means of machine learning by modifying the first point cloud one or more times to obtain one or more modified point clouds, wherein each of the one or more modified point clouds and the said object category are used as a further training data set from one or more further training data sets.

[0075] In one embodiment, the first point cloud may, for example, have been modified one or more times by applying translation dithering and / or rotation and / or random scaling and / or elastic deformation of the first point cloud.

[0076] According to one embodiment, the point cloud data may, for example, be three-dimensional point cloud data. The processing unit 120 may, for example, be configured to map the three-dimensional point cloud data to two-dimensional point cloud data and to classify the one or more subsets of the point cloud data using the two-dimensional point cloud data.

[0077] In one embodiment, the point cloud data may, for example, be three-dimensional point cloud data. Processing unit 120 may, for example, be configured to perform an assignment of the three-dimensional point cloud data to a plurality of voxels and to assign the point cloud data to the plurality of voxels depending on the assignment.

[0078] According to one embodiment, the processing unit 120 may, for example, be configured to create a virtual construction plan depending on the classification of the one or more subsets of the point cloud data.

[0079] In one embodiment, the processing unit 120 may, for example, be configured to create a three-dimensional building data model depending on the classification of the one or more subsets of the point cloud data.

[0080] Specific embodiments of the invention are presented below.

[0081] Embodiments relate to a device and a method for the automated creation of digital parametric geometric models for existing buildings, industrial facilities, and infrastructures. Embodiments can be applied to building information modeling (BIM models), digital factories, digital twins, and others. The input data for a device according to the invention is point cloud data of a complete scene containing the target object.

[0082] The input data is then processed by various artificial intelligence models tailored to the user's situation and based on different devices, resulting in the automatic creation of a digital, parameterized, geometric model.

[0083] The entered point cloud data can be in any format.

[0084] A focus of some embodiments is the processing of point cloud data in real-time. This point cloud data can be acquired using a variety of environment-specific sensor devices, including, but not limited to, high-resolution cameras, laser scanners, or depth cameras, among others. These sensors can also be operated via various media such as mobile robots and drones, and data can be acquired efficiently. When acquiring point cloud data, it is not possible to acquire all point clouds of an object at once due to the limited number of sensors and the target object being a large scene such as a building or a factory. It is therefore necessary to acquire sub-scenes of the target object according to sub-areas of the target building. The sub-scenes should then be registered multiple times (point cloud registration) to obtain the complete scene.However, these complete point cloud data contain much more information than the target object, e.g., trees, vehicles, pedestrians, facades of surrounding houses, and park and street scenes, as shown in Fig. 4.

[0085] In particular, Fig. 4 shows a target building scene with point cloud noise containing environmental information after point cloud registration.

[0086] The point cloud data other than the target building are redundant, meaning point cloud noise may also occur.

[0087] According to the above description, all redundant point cloud information outside the target object must be eliminated to construct the parameterized geometric model, as otherwise it will significantly affect the accuracy of subsequent steps. The case of the input scene in Fig. 4 intuitively shows that the complete scene is a complex scene with resolution variations, noise disturbances, object occlusions, background effects, etc. However, the semantic information within the scene is limited to certain categories, namely buildings, trees, roads, cars, pedestrians, etc. The target object, a building, is determined through multiple point cloud registrations and contains information not only about the surface of the building but also about its interior.

[0088] An advantageous embodiment consists in achieving the elimination of point cloud information that does not correspond to the category of the target object by means of a scene-level 3D point cloud segmentation algorithm, i.e., semantic segmentation.

[0089] Point cloud semantic segmentation (POSS) generates semantic information for each point in the point cloud data using supervised learning methods. This allows for the classification of point cloud data into subsets based on the semantic information of the points and the classification of different target object categories. Point cloud segmentation encompasses both regular supervised machine learning and modern deep learning, with regular supervised machine learning referring to algorithms that are not part of deep supervised learning.

[0090] According to previous research results, point cloud semantic segmentation (PCSS) for regular supervised machine learning can be divided into two types

[0065] . One type is an individual PCSS, which classifies each point or point cluster only based on its individual features, e.g. B. Maximum Likelihood classifiers based on Gaussian Mixture Models

[0066] , Random Forests

[0067] , AdaBoost [68,69], Support Vector Machines [70,71], a cascade of binary classifiers

[0072] , and Bayesian Discriminant Classifiers

[0073] , The other type is PCSS statistical context model, such as Associative and Non-Associative Markov Networks [74-76], Conditional Random Fields [77-81], Simplified Markov Random Fields

[0082] , Multi-stage inference method using point cloud statistics and learning relational information over fine and coarse scales

[0083] , and a framework for spatial inference engines,which models mid-range and long-range dependencies in the data

[0084] . The general procedure of individual classification for PCSS is divided into four steps: neighborhood selection, feature extraction, feature selection, and semantic segmentation

[0085] . Since individual PCSS does not consider contextual features (with neighboring points) of points, individual classifiers, although they can work effectively, will inevitably generate noise, leading to uneven PCSS results. Statistical context models can alleviate this problem.

[0091] Deep learning-based PCSS can be used to obtain high-dimensional features from the training data using more than two hidden layers and then process them, while traditional handcrafted features are designed using domain-specific knowledge. There are four paradigms of deep learning-based PCSS for semantic segmentation: projection-based approaches, discretization-based approaches, point-based approaches, and hybrid approaches.

[0086] In these different approaches, there are different methods under different methods. In the different methods, the features of the target object point cloud can be identified in different ways and aligned for relatively accurate segmentation, mainly by changing the structure of the deep learning (neural) network.

[0092] Projection-based approaches typically project a 3D point cloud onto a 2D image, including multi-view images and spherical images, i.e., the multi-view representation method and the spherical representation method. There are several methods for multi-view representation, one of which allows for first projecting a 3D point cloud from multiple virtual camera views onto a 2D plane, and then using a multi-stream FCN (multi-stream fully convolutional network) to predict pixel-by-pixel scores on synthetic images.The final semantic label for each point is obtained by fusing the reprojected scores across different views

[0087] . In another method, multiple RGB and depth snapshots of a point cloud are first created using multiple camera locations

[0088] , and then pixel-wise labeling of these snapshots is performed using 2D segmentation networks. The scores predicted from the RGB and depth images are fused using residual correction

[0089] . In the third method, tangent convolution can be introduced for the segmentation of dense point clouds, which is based on the assumption that the point cloud is sampled from a local Euclidean surface. This method first projects the local surface geometry around each point onto a virtual tangent plane. The tangent convolution is then applied directly to the surface geometry.This method exhibits good scalability and allows the processing of large sets of point clouds containing millions of points

[0090] . In general, the performance of multi-view segmentation methods depends on the selection of the viewpoint and the occlusion of the viewpoint. Furthermore, these methods do not fully exploit the underlying geometric and structural information, as the projection step inevitably leads to information loss. There are also various methods for the spherical representation method, one of which is an end-to-end network based on SqueezeNet

[0091] and Conditional Random Fields (ORF).

[0093]

[0092] To further improve the segmentation accuracy, SqueezeSegV2

[0094]

[0093] to solve the domain transfer problem by using an unsupervised domain adaptation pipeline. Another method first transfers the semantic labels from 2D range images to 3D point clouds, using an efficient GPU-enabled ANN-based post-processing step to mitigate the problem of discretization errors and fuzzy inference results.

[0095]

[0094] Compared to the single-view projection, the spherical projection contains more information and is suitable for labeling point clouds. However, this intermediate representation inevitably brings with it several problems, such as discretization errors and occlusions.

[0096] Discretization-based approaches typically convert a point cloud into a dense / sparse discrete

[0097] Representation methods, e.g., into volumetric and sparse permutohedral grids. In dense discretization representation, one method consists of first dividing the point cloud into a series of occupancy voxels and then feeding this cached data into a full 3D CNN for voxel-wise segmentation

[0095] . Finally, all point distributions within each voxel are aggregated into a compact latent region. Finally, all points within a voxel are assigned the same semantic label as the voxel. (The performance of this method is severely limited by the granularity of the voxels and boundary artifacts due to the point cloud partitioning.)Another method is to introduce deterministic trilinear interpolation to map the coarse voxel predictions generated by 3D fully convolutional networks (3D-FCNN)

[0096] back to the point cloud, and then use a fully connected conditional random field (CRF) or fully connected CRF (FCCRF) to enforce the spatial consistency of these derived perpoint labels, enabling fine-grained and globally consistent semantic segmentation

[0097] . The third approach is to encode the local geometric structures within each voxel via a kernel-based interpolated variational autoencoder architecture.

[0098] Instead of a binary occupancy representation, RBFs are used for each voxel to obtain a continuous representation and capture the distribution of points within each voxel. Variational Auto Encoders (VAE) are further used to map the point distribution within each voxel to a compact latent region. Both symmetry groups and an equivalence CNN are then used to achieve robust feature learning. A fourth method consists of first hierarchically inferring geometric relationships at different levels of point clouds and then using 3D convolution and weighted mean pooling to extract features and incorporate long-range dependencies. This method can handle large point clouds and has good scalability in the inference process (it is called a fully convolutional point network (FCPN).

[0099] The fifth method takes advantage of the scalability of the fourth method (FCPN) and can adapt to different input data sizes during training and testing. A coarse-to-fine strategy was used to hierarchically improve the resolution of the predicted results, allowing the 3D scans to be completed and semantic labels to be applied per voxel.

[0100] The first three methods voxelize the point cloud into a dense grid and then use standard 3D convolution. The last two methods are volume-based 3D convolutional neural networks (3D-CNNs), which also use volumetric representations. (However, the voxelization step inevitably leads to discretization artifacts and information loss. High resolution results in high storage and computational costs, while low resolution results in a loss of detail.)

[0098] In sparse discretized representation methods, the volumetric representation is inherently sparse because the number of non-zero values ​​is only a small percentage, making it inefficient to apply dense convolutional neural networks to spatially sparse data. One method is a sparse submanifold convolutional network based on an indexing structure. This method significantly reduces storage and computational costs by limiting the convolution output to the occupied voxels. In addition, sparse convolution allows control over the sparsity of the extracted features. This sparse submanifold convolution is suitable for the efficient processing of high-dimensional and spatially sparse data.

[0101] , Another method is a 4D spatio-temporal neural network for 3D video perception, which uses a generalized sparse convolution to efficiently process high-dimensional data and also uses a trilateral-stationary conditional random field to enforce consistency

[0102] A third method uses Sparse Lattice Networks (SPLATNet) based on Bilateral Convolution Layers (BCLs), in which the original point cloud is first interpolated into a permutohedral sparse grid, then BCL is applied to convolve the populated part of the sparse grid, and the filtered output is then interpolated back to the original point cloud, which enables flexible joint processing of images and point clouds with multiple views.

[0103] The fourth method uses fast grid convolution by embedding local geometry into a permutohedral sparse grid, keeping the memory footprint low and introducing a novel interpolation module as data is learned to project grid features back into the point cloud

[0104] ,

[0099] Hybrid approaches are used to learn multimodal features from 3D scans using a number of methods. One method uses a 3D CNN stream and multiple 2D streams for feature extraction, and a differentiable backprojection layer is used to jointly fuse the learned 2D embeddings and 3D geometric features. This is a joint 3D multi-view network to combine RGB features and geometric features.

[0105] , Another method is a unified point-based framework that learns 2D texture features, 3D structures, and global context features from point clouds and directly applies a point-based network to extract local geometric features and global context from a sparsely sampled point set without voxelization

[0106] , Another method is a Multi-View PointNet to aggregate appearance features from 2D multi-view images and spatial geometric features in the canonical point cloud space

[0107] ,

[0100] In point-based approaches, point clouds are disordered and unstructured, so that the direct application of standard CNNs is not possible. Therefore, the pioneering work PointNet

[0108] developed using shared multilayer perceptrons (MLPs) to learn features per point and using symmetric pooling functions to learn global features. A number of point-based networks are developed on the basis of PointNet. In general, these methods can be roughly divided into pointwise MLP methods, point convolution methods, RNN-based methods, and graph-based methods. Point-based pointwise MLP methods use shared MLPs as the basic units of their networks to achieve high efficiency (point-like features extracted by a shared MLP, however, do not capture the local geometry and interactions between points in a point cloud

[0108] ).To capture a broader context for each point and learn richer local structures, several specialized networks were introduced, including methods based on neighbor feature pooling, attention-based aggregation, and local-global feature aggregation.

[0101] Point-based pointwise MLP methods based on attention-based aggregation have several methods for capturing local geometric patterns, which learn a feature for each point by aggregating the information from local neighboring points. One of them, PointNet++, groups points hierarchically at multiple scales and at multiple resolutions to overcome the problems caused by the non-uniformity and different density of the point cloud, and progressively learns from larger local regions.

[0109] , In addition, there is a method that stacks and encodes information from eight spatial orientations through a three-stage ordered convolution. Multiscale features are concatenated to achieve adaptability to different scales for orientation encoding and scale awareness.

[0110] Another method, different from the clustering technique used in PointNet++ (i.e., sphere query), uses K-means clustering and KNN to define two neighborhoods separately in world space and feature space. Based on the assumption that points of the same class are expected to be closer to each other in the feature space, a pairwise distance loss and a centroid loss are introduced to further regularize the feature learning.

[0111] The fourth method uses an Adaptive Feature Adjustment (AFA) module to achieve information exchange and feature refinement. This aggregation process helps the network learn a discriminatory feature representation and model the interactions between different points, as well as explore the relationships between all point pairs in a local region through the dense construction of a locally fully connected network.

[0112] The fifth method is a permutation-invariant convolution based on the statistics of concentric sphere shells. In this method, a set of concentric spheres is first queried using multiscale, then the max-pooling operation within different shells is used to summarize the statistics, and MLPs and 1D convolution are used to obtain the final convolution output.

[0113] The sixth method is an efficient, lightweight network for segmenting large point clouds. The network uses random point sampling and achieves significant efficiency in terms of memory and computation. Furthermore, this network has a local feature aggregation module to capture and preserve geometric features.

[0114] ,

[0102] To further improve the segmentation accuracy, the point-based pointwise MLP methods based on attention-based aggregation

[0115] One method is to model the relationship between the points using group shuffle attention, where a permutation-invariant, task-independent, and distinguishable Gumbel Subset Sampling (GSS) replaces the widely used FPS approach. This module is less sensitive to outliers and allows the selection of a representative subset of points.

[0116] , Another method is to better capture the spatial distribution of point clouds by learning spatial awareness weights through a Local Spatial Aware (LSA) layer based on the spatial arrangement and local structure of the point clouds

[0117] Another method, similar to Conditional Random Fields (CRF), uses an attention-based score refinement (ASR) module to post-process the segmentation results generated by the network. The original segmentation result is refined by merging the scores of neighboring points with learned attention weights. This module can be easily integrated into existing deep networks to improve segmentation performance.

[0118] , Third point-based pointwise MLP methods based on local-global features have a method that uses four repeatedly stacked encoders, each encoder has two basic components Edgeconv

[0119] and NetVLAD

[0120] to capture local information and global features at the scene level to incorporate the local structure and global context of the point cloud

[0121] ,

[0103] Point-based point convolution methods tend to propose effective convolution operators for point clouds. One method may use a pointwise convolution operator, which allows neighboring points to be divided into kernel cells and then convolved with kernel weights

[0122] , Another approach can be achieved by using parametric kernel functions that cover the entire continuous vector space and enable parametric continuous convolution (PCCN) that can be learned on any data structure whose support relations are computable

[0123] , Another method is a Kernel Point Fully Convolutional Network based on Kernel Point Convolution (KPConv)

[0124] Specifically, the convolution weights of KPConv are determined by the Euclidean distances to the kernel points, and the number of kernel points is not fixed. The positions of the kernel points are formulated as a best coverage optimization problem in a spherical space. Note that the radius neighborhood is used to obtain a consistent receptive field, while subsampling the grid in each layer is used to achieve high robustness under varying point cloud densities. The next method aggregates dilated neighboring features instead of k-nearest neighbors using a dilated point convolution (DPC) operation. This method proved to be very effective in increasing the receptive field and can be easily integrated into existing aggregation-based networks.

[0125] ,

[0104] To capture inherent context features of point clouds, Recurrent Neural Networks (RNN) are also used for semantic segmentation of point clouds. There are several RNN-based methods based on points. One such method is based on PointNet

[0108] , where a point block is first transformed into multi-scale blocks and grid blocks to obtain contexts at the input level. The block-wise features extracted by PointNet are then sequentially fed into Consolidation Units (CU) or Recurrent Consolidation Units (RCU) to obtain the context at the output level, which can improve the segmentation performance by incorporating the spatial context through these methods.

[0126] , There is also a method for converting an unordered point feature set into an ordered sequence of feature vectors by using a slice pooling layer and a lightweight module for modeling local dependencies

[0127] Another approach is to capture the local structure from coarse to fine structure through a Pointwise Pyramid Pooling (3P) module and then use hierarchical RNNs in two directions to obtain long-range spatial dependencies. RNNs are then deployed to achieve end-to-end learning.

[0128] (However, these methods lose the rich geometric features and density distribution of the point cloud when local neighborhood features are pooled with global structural features.) To mitigate the problems associated with rigid and static pooling operations, there is also a dynamic aggregation network method that considers both the global complexity of the scene and local geometric features by using a self-adapted receptive field and node weights to aggregate inter-medium features.

[0129] Another method involves learning the spatial distribution and color features using a 3D CNN network, and then using DQN (Deep Q-Learning) to localize objects belonging to a specific class. The final concatenated feature vector is fed into a residual RNN to obtain the final segmentation result.

[0130] ,

[0105] Finally, there are some methods, the graph-based method in the point-based approach, to capture the underlying shapes and geometric structures of 3D point clouds. One such method consists of representing the point cloud as a set of interconnected simple shapes and superpoints and using an attributive directed graph (i.e., the superpoint graph) to capture the structural and contextual information. The problem of segmenting large point clouds is then divided into three subproblems, namely geometrically homogeneous segmentation, superpoint embedding, and contextual segmentation.

[0131] Another method optimizes the partitioning steps of the previous methods by using a supervised framework to over-segment the point cloud into pure superpoints. This problem is formulated as a deep metric learning problem structured by an adjacency graph. Furthermore, a graph-structured contrastive loss is proposed to support the detection of boundaries between objects.

[0132] , The third method enables a better capture of local geometric relationships in high-dimensional space by a

[0106] PyramNet, which is based on the Graph Embedding Module (GEM) and the Pyramid Attention Network (PAN)

[0133] The GEM module formulates the point cloud as a directed acyclic graph and uses a covariance matrix to replace the Euclidean distance for constructing the neighbor similarity matrix. In the PAN module, four different convolution kernels are used to extract features with different semantic intensities.

[0133] In the fourth method, the relevant features are selectively learned from a local neighborhood set by Graph Attention Convolution (GAC). This process is achieved by dynamically assigning attention weights to different neighboring points and feature channels based on the spatial positions and feature differences. GAC can learn to capture discriminatory features for segmentation and has similar properties to the commonly used CRF (Conditional Random Field) model.

[0134] The fifth method uses a Point Global Context Reasoning (PointGCR) module, which captures global context information along the channel dimension using an undirected graph representation. PointGCR is a plug-and-play and end-to-end trainable module and can be easily integrated into an existing segmentation network to improve performance.

[0135] , There are also some methods for implementing semantic segmentation of point clouds under weak supervision, such as a two-stage approach to train a segmentation network with subcloud level labels

[0136] and a network that can be trained with partially labeled points (e.g., 10%) for an inaccurate supervised method for semantic segmentation of point clouds

[0137] ,

[0107] Fig. 5 shows different methods of approaches and groups for semantic segmentation of point clouds (Point Cloud Semantic Segmentation, PCSS).

[0108] In particular, Fig. 5 shows the various methods and groups of approaches for semantic segmentation of point clouds. Embodiments include, but are not limited to, the above-mentioned methods and techniques to achieve the first step of removing ambient noise information from the 3D point clouds that does not match the category of the target object. Fig. 6 shows the result of performing this step according to embodiments.

[0109] Thus, Fig. 6 shows the result after removing environmental noise information that is not of the target object category using PCSS. Point clouds of the same category (buildings) can be determined using the method described above. Some embodiments allow automatic noise reduction based on the density and spacing of the point clouds using a clustering algorithm at specific distances or a histogram statistic of the x-, y-, and z-axis coordinates of the point clouds. After this step, the results can be seen in Fig. 7, where the extraction of the target building is achieved. Thus, Fig. 7 shows the target object without any other buildings. After the extraction of the target buildings, there is a very small accuracy error due to the semantic segmentation.To avoid this accuracy error, another effective implementation of some embodiments is to extract the boundaries of the different views of the target building and then use the boundary information to segment the original point cloud data so that the point cloud data of the target building can be obtained without errors.

[0110] A further advantageous embodiment consists in the instance segmentation of point clouds (Point Cloud Instance Segmentation, PCIS) of the entire target building or factory, wherein the targets of the instance segmentation are primarily stairwells, elevator shafts, interior rooms and exterior walls of the target building, as well as the doors and windows on the exterior walls of the target building. This is also the third step of an embodiment. This method first eliminates the influence of stairwells and elevator shafts on the determination of the floor heights. Then, from the heights of the interior rooms, the floor heights and roof heights of each floor of the entire building or factory are determined. Subsequently, the point clouds of the windows and doors on the exterior walls are used to obtain the geometric parameters and positions of the windows and doors.

[0111] To be able to model with high accuracy, another advantageous embodiment is the PCIS of the various spaces of a house or factory, such as walls, various pieces of furniture (chairs, tables, cabinets, PCs, etc.) and machine tools, materials, etc. in the factory. This is the main step of the fourth step of an embodiment, which then allows the determination of the geometric parameters and the position of these objects using our algorithms and later, using the algorithms of this embodiment, the reconstruction of a highly detailed geometric model, such as BIM, CAD, etc.

[0112] The two embodiments mentioned above are both based on instance segmentation of point clouds. As an object-level segmentation method, this also means that it is more challenging than semantic segmentation, as it should draw more accurate and fine-grained reasoning on points. PCIS should distinguish not only between points with different semantics, but also between instances with the same semantic information. There are two main types of instance segmentation: proposal-based methods and proposal-free methods. In the proposal-based method, the category of a group of objects is first predicted, i.e., proposals are generated, which has a similar effect to object detection, and then foreground-background segmentation is performed in each bounding box.With the proposal-free method, the step of creating proposals is omitted.

[0113] As mentioned above, the proposal-based method transforms the problem of instance segmentation into two subtasks: 3D object detection and instance mask prediction. One method is to use a 3D fully convolutional Semantic Instance Segmentation (3D-SIS) to perform semantic instance segmentation for RGB-D scans. This network learns from both color and geometric features. Similar to 3D object detection, a 3D Region Proposal Network (3DRPN) and a 3D Region of Interesting (3D-Rol) layer are used to predict bounding box positions, object class labels, and instance masks.

[0138] Another method uses the analysis-by-synthesis strategy to generate 3D suggestions with high objectivity using a Generative Shape Suggestion Network (GSPN) via a synthetic analysis.

[0139] These proposals are further refined by a Region-based PointNet (R-PointNet). The final label is obtained by predicting a binary mask per point for each class label. Unlike direct regression of 3D bounding boxes from point clouds, this method removes a large amount of meaningless proposals by enforcing geometric understanding. Another method is to implement an online 3D volumetric mapping system by extending 2D panoptic segmentation to 3D mapping to jointly achieve large-scale 3D reconstruction, semantic labeling, and instance segmentation. The method first uses 2D semantic and instance segmentation networks to obtain pixel-wise panoptic labels and then integrates these labels into the volumetric map.A fully linked Conditional Random Fields (CRF) is further utilized to achieve accurate segmentation. This semantic mapping system can achieve high-quality semantic mapping and discriminatory object detection.

[0140] A fourth method proposes a one-stage, anchor-free, and end-to-end trainable network called 3D-BoNet to achieve instance segmentation on point clouds. This method directly regresses coarse 3D bounding boxes for all potential instances and then uses a point-level binary classifier to obtain instance labels. Specifically, the task of generating bounding boxes is formulated as an optimal matching problem. Furthermore, a multi-objective loss function is proposed to regularize the generated bounding boxes. This method requires no post-processing and is computationally efficient.

[0141] The fifth method is a network for instance segmentation of large-scale outdoor LiDAR point clouds. This method learns a bird's-eye view feature representation of point clouds using self-attention blocks. The final instance labels are determined based on the predicted horizontal center and elevation boundaries.

[0142] The sixth method is the prediction of the layout of a 3D indoor space using a hierarchy-aware Variational Denoising Recursive AutoEncoder (VDRAE). The object predictions are generated iteratively and refined by recursive context aggregation and propagation.

[0143] The proposed methods [138,139,141,144] are intuitive and simple, and the instance segmentation results are generally of good objectivity. However, these proposed methods require multi-stage training and filtering out redundant proposals. Therefore, they are generally time-consuming and computationally intensive.

[0114] Proposal-free methods do not have proposal generation, i.e., they do not have a module for object prediction. Instead, they usually consider instance segmentation as a subsequent clustering step after semantic segmentation. In particular, most existing methods are based on the assumption that points belonging to the same instance should have very similar features. Therefore, these methods mainly focus on learning discriminatory features and grouping points. In particular, most existing methods are based on the assumption that points belonging to the same instance should have very similar features. Therefore, these methods mainly focus on learning discriminatory features and grouping points. One such method is the Similarity Group Proposal Network (SGPN).This method first learns a feature and a semantic map for each point and then introduces a similarity matrix to represent the similarity between each pair of features. To learn further discriminatory features, a double-hinge loss is used to mutually adapt the similarity matrix and the semantic segmentation results. Finally, a heuristic and non-maximum suppression method is employed to cluster similar points into instances. Since the construction of a similarity matrix requires a large amount of memory, the scalability of this method is limited

[0145] . The second method uses sparse convolution

[0101] of submanifolds to predict semantic values ​​for each voxel and the affinity between neighboring voxels.A clustering algorithm is then used to group points into instances based on the predicted affinities and network topology

[0146] . Another method uses a structure-aware loss for learning discriminatory embeddings. This loss considers both feature similarity and geometric relationships between points. An attention-based graph CNN was further used to adaptively refine the learned features by aggregating various information from neighbors

[0147] . Since the semantic category and the instance label of a point are usually interdependent, several methods have been proposed to combine these two tasks into a single task. One method involves integrating the two tasks by introducing an end-to-end and adaptive associative instance semantic segmentation (ASIS) module.Experiments show that semantic features and instance features can support each other to achieve improved performance through this ASIS module

[0148] , Another novel combined method for implementing instance and semantic segmentation is JSNet

[0149] , An alternative method is the MultiTask Pointwise Network (MT-PNet) to assign a label to each point and regularize the embeddings in the feature space by introducing a discriminatory loss

[0150] , Then, it fused the predicted semantic labels and embeddings with an MV-CRF (Multi-Value Conditional Random Field) model for joint optimization.Finally, variational mean-field inference is used to generate semantic labels and instance labels

[0151] . Another method is to use dynamic region growing (DRG) to divide the point cloud into a series of disjointed patches, and then use an unsupervised Kmeans++ algorithm to cluster all these patches. Multi-stage segmentation of the patches is then performed using context information between the patches. Finally, these labeled patches are merged at the object level to obtain final semantic labels and instance labels

[0152] . Another method to achieve instance segmentation in full 3D scenes is to use a hybrid 2D-3D network to jointly learn globally consistent instance features from a BEV representation and local geometric features from point clouds.The learned features are then combined to achieve semantic and instance segmentation

[0153] . Another method consists of learning the abstract feature embedding of each instance and the direction information of the instance center for each voxel. The feature embedding loss and direction loss are proposed to adjust the learned feature embeddings in the latent feature space. Mean-shift clustering and non-maximum suppression are employed to group voxels into instances

[0154] . The other method consists of a semantic segmentation branch and an offset prediction branch.A dual-set clustering algorithm and ScoreNet are further used to achieve better clustering results

[0155] . In another method, a semantic superpoint tree (SST) is constructed from the semantic features of the learned superpoints, which is then traversed and segmented at the nodes of the intermediate tree to propose object instances. This method is complemented by a refinement module, CliqueNet, to clean up superpoints that might be incorrectly classified as instance proposals

[0156] .

[0115] Figure 8 shows various methods for point cloud instance segmentation (PCIS). Embodiments include the above-mentioned methods and techniques in the context of algorithms intended to implement object-level instance segmentation.

[0116] All of the above-mentioned semantic segmentations of point clouds and the instance segmentation of point cloud instances based on deep learning and supervised learning require good database support to achieve excellent segmentation accuracy and results for industrial use through training. Since our database is constantly expanding and can be manually corrected, it becomes increasingly useful as the number of segmentations and applications increases. Furthermore, according to one embodiment, the database of the segmentation algorithm can also be expanded in various steps by building a point cloud database from existing geometric models, which can be used to improve the segmentation accuracy of the algorithm.

[0117] After these steps, the algorithm according to the invention is capable of determining the geometric parameters, the position and orientation of an object in the interior of the building as well as the individual walls, windows and doors of the building, e.g., the thickness of the walls, the start and end points, etc. With these parameters, it is possible to establish interfaces to various drawing software programs and then to automate the modeling according to the drawing software methods.

[0118] Regression based purely on the geometric properties of the objects contained in the point cloud is highly maintenance-intensive, and the initial effort of describing all objects rule-based in order to then be able to recognize them automatically is not cost-effective. Therefore, it is necessary to extend the geometry-based algorithms with AI-based algorithms and create a hybrid process. The sequence of the hybrid process is shown in Fig. 9. Thus, Fig. 9 shows a sequence of the regression process.

[0119] For example, there should be a clear mapping between the parameters required in the authoring system for modeling the 3D models and the "real" parameters of the objects under consideration, present as a virtualized point cloud. This is the only way to ensure that the necessary parameters are extracted from the point cloud prior to transfer. This should make it possible to ensure a neutral "blueprint" for transfer between the point cloud and the authoring systems.

[0120] The software should, for example, be accessible to end users through easy installation. The constraints, libraries, and utilities should be able to be set or integrated without user input. The optional interfaces to external software APIs enable the integration of the framework into existing software environments. To meet the needs of individual technical applications, various implementations are conceivable. Depending on the given conditions, it can either be open-source-based or implemented in proprietary software using a suitable programming language. For the present use case, for example, connections to the proprietary software products Revit and ArchiCAD are required, which represent the desired front end for the end user.

[0121] The virtual construction plan is the basis for the automated modeling of 3D BIM models in the authoring systems. It contains all necessary parameters and is transferred to the authoring system via an interface. Modeling proceeds according to the scheme stored in the construction plan, taking into account the stored parameters. The result is a 3D BIM model. This represents the virtual representation of the real initial scene.

[0122] Training data augmentation according to one embodiment is described below.

[0123] Thus, the point cloud of the specified box size (e.g., cuboid size) in the original scene is extracted sequentially during training (the box size can be changed based on the hardware). Online data augmentation of the point clouds in the boxes is performed. For example, the point cloud box can be randomly extracted 100 times from the entire scene. This provides a solution to the problem of training large scenes on limited hardware.

[0124] Fig. 10 shows such an extraction of boxes (cuboid-shaped sub-regions) of the point cloud according to one embodiment in order to obtain training data.

[0125] For online data augmentation according to one embodiment, downsampling can be performed based on different categories in different buildings with different voxel sizes. Larger voxels can be defined, for example, for categories with more points, and smaller voxels for categories with fewer points, for example.

[0126] For example, (e.g., to generate a larger number of training data sets), translation dithering for the entire building and / or a 360-degree rotation around the entire building and / or random scaling for the entire building and / or elastic deformation of the entire building may be used. And / or the data augmentation and / or cross-combining described above may be used.

[0127] Fig. 11 shows a visualization of the data augmentation according to a first embodiment.

[0128] Fig. 12 shows a visualization of the data augmentation according to a second embodiment.

[0129] Architects are currently being commissioned, and will be increasingly commissioned, to remodel existing buildings in the coming years. They are currently faced with two problems. First, they should create the most accurate as-built plan possible of the building to be remodeled. Second, since 2020, the Building Information Model (BIM) has been the required standard for all public construction projects in Germany.

[0130] This transition, along with precise building surveying and the associated data processing, repeatedly presents architects with major challenges. This is often the case because many of them still work with outdated techniques and principles that are being superseded by the new BIM standard. This is confirmed by a study by the North Rhine-Westphalia Chamber of Architects, which states that in 2018, only 10% of all architects worked with the Building Information Model.

[0131] Executive forms offer architects the opportunity to solve these problems. Automated building modeling allows architects and project developers to save hours of work at the push of a button, thereby reducing the major problems of insufficient time and budget. This allows them to begin directly with the actual value creation, the planning of the construction measures. In addition, embodiments should offer the opportunity to check, quantify, and validate the accuracy of the regressed model. In the current manual process, the positioning of building elements in the virtual model is done by eye. This represents a potential source of error in the quality of the model, which can be minimized or even eliminated using algorithms.A quick visual and quantitative review of the results creates a high level of user confidence in the quality of the model, which is essential for further planning processes.

[0132] Technical application areas of embodiments include the generation of 3D BIM for buildings, industrial facilities, and infrastructures for serial renovation, renovation in general, construction planning, construction processes in general, insurance services, damage analyses and predictions, visualizations (XR), etc. Further application areas include the generation of 3D BIM of the current status / construction progress of buildings, industrial facilities, and infrastructures, as well as the generation of 3D BIM of infrastructure, and the generation of smart city models. Further use cases include the generation of 3D BIM of individual rooms and the generation of interior models for houses or apartments to facilitate interior design for designers.Further fields of application include the generation of 3D BIMs of production plants, lines, and factories, as well as the production and operating resources contained therein, the generation of corresponding FE meshes as input for FE analyses of the aforementioned models, and the generation of corresponding digital twins of the aforementioned models. Another field of application is the generation of current 3D models of the condition of buildings, industrial facilities, or infrastructure. On this basis, algorithms for network coverage and losses, antenna prediction and installation simulation, IoT implementation, and energy management can be implemented.

[0133] Although some aspects have been described in connection with a device, it is understood that these aspects also represent a description of the corresponding method, so that a block or component of a device can also be understood as a corresponding method step or as a feature of a method step. Analogously, aspects described in connection with or as a method step also represent a description of a corresponding block, detail, or feature of a corresponding device. Some or all of the method steps may be performed by (or using) a hardware device, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, some or more of the key method steps may be performed by such an device.

[0134] Depending on specific implementation requirements, embodiments of the invention may be implemented in hardware or in software, or at least partially in hardware or at least partially in software. The implementation may be carried out using a digital storage medium, for example a floppy disk, a DVD, a Blu-ray disc, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a FLASH memory, a hard disk, or other magnetic or optical storage device storing electronically readable control signals that can interact or interact with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer-readable.

[0135] Some embodiments according to the invention thus comprise a data carrier having electronically readable control signals capable of interacting with a programmable computer system such that one of the methods described herein is carried out.

[0136] In general, embodiments of the present invention may be implemented as a computer program product having a program code, wherein the program code is effective to perform one of the methods when the computer program product is run on a computer.

[0137] The program code can, for example, also be stored on a machine-readable medium.

[0138] Other embodiments comprise the computer program for performing one of the methods described herein, wherein the computer program is stored on a machine-readable medium. In other words, one embodiment of the method according to the invention is thus a computer program that has program code for performing one of the methods described herein when the computer program runs on a computer. Another embodiment of the method according to the invention is thus a data carrier (or a digital storage medium or a computer-readable medium) on which the computer program for performing one of the methods described herein is recorded. The data carrier or the digital storage medium or the computer-readable medium is typically tangible and / or non-transitory.

[0139] A further embodiment of the method according to the invention is thus a data stream or a sequence of signals that represents the computer program for carrying out one of the methods described herein. The data stream or the sequence of signals can be configured, for example, to be transferred via a data communication connection, for example, via the Internet.

[0140] A further embodiment comprises a processing device, for example a computer or a programmable logic device, which is configured or adapted to carry out one of the methods described herein.

[0141] A further embodiment comprises a computer on which the computer program for performing one of the methods described herein is installed.

[0142] A further embodiment according to the invention comprises a device or a system designed to transmit a computer program for performing at least one of the methods described herein to a recipient. The transmission may, for example, be electronic or optical. The recipient may, for example, be a computer, a mobile device, a storage device, or a similar device. The device or system may, for example, comprise a file server for transmitting the computer program to the recipient.

[0143] In some embodiments, a programmable logic device (e.g., a field-programmable gate array, an FPGA) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field-programmable gate array may cooperate with a microprocessor to perform any of the methods described herein. Generally, in some embodiments, the methods are performed by any hardware device. This may be general-purpose hardware such as a computer processor (CPU) or hardware specific to the method, such as an ASIC. The embodiments described above are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art.Therefore, it is intended that the invention be limited only by the scope of the following claims and not by the specific details presented in the description and explanation of the embodiments herein.

[0144] Literature:

[0145] [1] B. Becerik-Gerber, F. Jazizadeh, N. Li, G. Calis, Application Areas and Data Requirements for BIM-Enabled Facilities Management, J. Constr. Closely. Manage. 138 (2012) 431-442.

[0146] [2] S. Azhar, Building Information Modeling (BIM): Trends, Benefits, Risks, and Challenges for the AEC Industry, Leadership Manage. Closely. 11 (2011) 241-252.

[0147] [3] R. Volk, J. Stengel, F. Schultmann, Building Information Modeling (BIM) for existing buildings — Literature review and future needs, Automation in Construction 38 (2014) 109-127.

[0148] [4] Y. Deng, J.C. Cheng, C. Anumba, Mapping between BIM and 3D GIS in different levels of detail using schema mediation and instance comparison, Automation in Construction 67 (2016) 1-21.

[0149] [5] X. Xiong, A. Adan, B. Akinci, D. Huber, Automatic creation of semantically rich 3D building models from laser scanner data, Automation in Construction 31 (2013) 325-337.

[0150] [6] S.M. Sepasgozar, S. Lim, S. Shirowzhan, Y.M. Kim, Implementation of As-Built Information Modelling Using Mobile and Terrestrial Lidar Systems, in: Q. Ha, X. Shen, A. Akbarnezhad (Eds.), Proceedings of the 31st International Symposium on Automation and Robotics in Construction and Mining (ISARC), International Association for Automation and Robotics in Construction (IAARC), 2014.

[0151] [7] S. Ochmann, R. Vock, R. Wessel, R. Klein, Automatic reconstruction of parametric building models from indoor point clouds, Computers & Graphics 54 (2016) 94- 103.

[0152] [8] C. Mura, O. Mattausch, A. Jaspe Villanueva, E. Gobbetti, R. Pajarola, Automatic room detection and reconstruction in cluttered indoor environments with complex room layouts, Computers & Graphics 44 (2014) 20-32. [9] C. Mura, O. Mattausch, R. Pajarola, Piecewise-planar Reconstruction of Multiroom Interiors with Arbitrary Wall Arrangements, Computer Graphics Forum 35 (2016) 179-188.

[0153]

[0010] C. So, G. Baciu, H. Sun, Reconstruction of 3D virtual buildings from 2D architectural floor plans, in: J.M. Shieh, S.-N. Yang (Eds.), Proceedings of the ACM symposium on Virtual reality software and technology 1998 -VRST '98, ACM Press, New York, New York, USA, 1998, pp. 17-23.

[0154]

[0011] T. Lu, C.-L. Tai, L. Bao, F. Su, S. Cai, 3D Reconstruction of Detailed Buildings from Architectural Drawings, Computer-Aided Design and Applications 2 (2005) 527-536.

[0155]

[0012] S. Lee, D. Feng, C. Grimm, B. Gooch, A Sketch-Based User Interface for Reconstructing Architectural Drawings, Computer Graphics Forum 27 (2008) SI- 90.

[0156]

[0013] T. Li, B. Shu, X. Qiu, Z. Wang, Efficient reconstruction from architectural drawings, IJCAT 38 (2010) 177.

[0157]

[0014] D.G. Lowe, Distinctive Image Features from Scale-Invariant Keypoints, International Journal of Computer Vision 60 (2004) 91-110.

[0158]

[0015] Y.-F. Liu, S. Cho, B.F. Spencer, J.-S. Fan, Concrete Crack Assessment Using Digital Image Processing and 3D Scene Reconstruction, J. Comput. Civ. Eng. 30 (2016) 4014124.

[0159]

[0016] M. Golparvar-Fard, F. Peha-Mora, S. Savarese, Automated Progress Monitoring Using Unordered Daily Construction Photographs and I FC-Based Building Information Models, J. Comput. Civ. Eng. 29 (2015) 4014025.

[0160]

[0017] A. Khaloo, D. Lattanzi, Hierarchical Dense Structure-from-Motion Reconstructions for Infrastructure Condition Assessment, J. Comput. Civ. Eng. 31 (2017) 4016047.

[0161]

[0018] P. Alcantarilla, J. Nuevo, A. Bartoli, Fast Explicit Diffusion for Accelerated Features in Nonlinear Scale Spaces, in: T. Burghardt, D. Damen, W. Mayol-Cuevas, M. Mirmehdi (Eds.), Procedings of the British Machine Vision Conference 2013, British Machine Vision Association, 2013, 13.1-13.11.

[0162]

[0019] F. Huang, Z. Mao, W. Shi, ICA-ASIFT-Based Multi-Temporal Matching of High- Resolution Remote Sensing Urban Images, Cybernetics and Information Technologies 16 (2016) 34-49.

[0163]

[0020] J.-M. Morel, G. Yu, ASIFT: A New Framework for Fully Affine Invariant Image Comparison, SIAM J. Imaging Sei. 2 (2009) 438-469.

[0164]

[0021] P. Rodriguez-Gonzalvez, D. Gonzalez-Aguilera, G. Lopez-Jimenez, I. Picon- Cabrera, Image-based modeling of built environment from an unmanned aerial system, Automation in Construction 48 (2014) 44-52.

[0165]

[0022] H. Bay, T. Tuytelaars, L. van Gool, SURF: Speeded Up Robust Features, in: A. Leonardis, H. Bischof, A. Pinz (Eds.), Computer Vision -ECCV 2006, Springer Berlin Heidelberg, Berlin, Heidelberg, 2006, pp. 404-417.

[0166]

[0023] H. Bay, A. Ess, T. Tuytelaars, L. van Gool, Speeded-Up Robust Features (SURF), Computer Vision and Image Understanding 110 (2008) 346-359.

[0167]

[0024] G.M. Jog, H. Fathi, I. Brilakis, Automated computation of the fundamental matrix for vision based construction site applications, Advanced Engineering Informatics 25 (2011) 725-735.

[0168]

[0025] C. Harris, M. Stephens, A Combined Corner and Edge Detector, in: C.J. Taylor (Ed.), Procedings of the Alvey Vision Conference 1988, Alvey Vision Club, 1988, 23.1-23.6.

[0169]

[0026] C. Sung, P.Y. Kim, 3D terrain reconstruction of construction sites using a stereo camera, Automation in Construction 64 (2016) 65-77.

[0170]

[0027] E. Rosten, T. Drummond, Machine Learning for High-Speed Corner Detection, in: A. Leonardis, H. Bischof, A. Pinz (Eds.), Computer Vision -ECCV 2006, Springer Berlin Heidelberg, Berlin, Heidelberg, 2006, pp. 430-443.

[0028] A. Alahi, R. Ortiz, P. Vandergheynst, FREAK: Fast Retina Keypoint, in: 2012 IEEE Conference on Computer Vision and Pattern Recognition, IEEE, 2012, pp. 510- 517.

[0171]

[0029] H. Bae, M. Golparvar-Fard, J. White, High-precision vision-based mobile augmented reality system for context-aware architectural, engineering, construction and facility management (AEC / FM) applications, Vis. in Eng. 1 (2013).

[0172]

[0030] X. Du, Y. Zhu, A flexible method for 3D reconstruction of urban building, in: 2008 9th International Conference on Signal Processing, IEEE, 2008, pp. 1355-1359.

[0173]

[0031] M. Golparvar-Fard, Y. Ham, Automated Diagnostics and Visualization of Potential Energy Performance Problems in Existing Buildings Using Energy Performance Augmented Reality Models, J. Comput. Civ. Eng. 28 (2014) 17-29.

[0174]

[0032] Y. Ham, M. Golparvar-Fard, EPAR: Energy Performance Augmented Reality models for identification of building energy performance deviations between actual measurements and simulation results, Energy and Buildings 63 (2013) 15-28.

[0175]

[0033] A. Rashidi, F. Dai, I. Brilakis, P. Vela, Comparison of Camera Motion Estimation Methods for 3D Reconstruction of Infrastructure, in: Y. Zhu, R.R. Issa (Eds.), Computing in Civil Engineering (2011), American Society of Civil Engineers, Reston, VA, 2011, pp. 363-371.

[0176]

[0034] HYBRID 4-DIMENSIONAL AUGMENTED REALITY - A High-precision Approach to Mobile Augmented Reality, in: Proceedings of the 2nd International Conference on Pervasive Embedded Computing and Communication Systems, SciTePress - Science and and Technology Publications, 2012, pp. 156-161.

[0177]

[0035] H. Bae, M. Golparvar-Fard, J. White, High-Precision and Infrastructure- Independent Mobile Augmented Reality System for Context-Aware Construction and Facility Management Applications, in: I. Brilakis, S. Lee, B. Becerik-Gerber (Eds.), Computing in Civil Engineering, American Society of Civil Engineers, Reston, VA, 2013, pp. 637-644.

[0036] M. Muja, D.G. Lowe, Scalable Nearest Neighbor Algorithms for High Dimensional Data, IEEE transactions on pattern analysis and machine intelligence 36 (2014) 2227-2240.

[0178]

[0037] J. Cheng, C.Leng, J. Wu, H. Cui, H. Lu, Fast and Accurate Image Matching with Cascade Hashing for 3D Reconstruction, in: 2014 IEEE Conference on Computer Vision and Pattern Recognition, IEEE, 2014, pp. 1-8.

[0179]

[0038] M. Golparvar-Fard, V. Balali, J.M. de La Garza, Segmentation and Recognition of Highway Assets Using Image-Based 3D Point Clouds and Semantic Texton Forests, J. Comput. Civ. Eng. 29 (2015) 4014023.

[0180]

[0039] R. Hartley, A. Zisserman, Multiple View Geometry in Computer Vision, Cambridge University Press, 2011.

[0181]

[0040] D. Nister, An efficient solution to the five-point relative pose problem, IEEE transactions on pattern analysis and machine intelligence 26 (2004) 756-777.

[0182]

[0041] R.l. Hartley, In defense of the eight-point algorithm, IEEE Trans. Pattern Anal. Machine Intell. 19 (1997) 580-593.

[0183]

[0042] W. Zhang, Y. Zhang, Z. Li, Z. Zheng, R. Zhang, J. Chen, A rapid evaluation method of existing building applied photovoltaic (BAPV) potential, Energy and Buildings 135 (2017) 39-49.

[0184]

[0043] M.-D. Yang, C.-F. Chao, K.-S. Huang, L.-Y. Lu, Y.-P. Chen, Image-based 3D scene reconstruction and exploration in augmented reality, Automation in Construction 33 (2013) 48-60.

[0185]

[0044] B. Triggs, P.F. McLauchlan, R.l. Hartley, A.W. Fitzgibbon, Bundle Adjustment — A Modern Synthesis, in: G. Goos, J. Hartmanis, J. van Leeuwen, B. Triggs, A. Zisserman, R. Szeliski (Eds.), Vision Algorithms: Theory and Practice, Springer Berlin Heidelberg, Berlin, Heidelberg, 2000, pp. 298-372.

[0045] Y. Furukawa, B. Curless, S.M. Seitz, R. Szeliski, Towards Internet-scale multi-view stereo, in: 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE, 2010, pp. 1434-1441.

[0186]

[0046] M.M. Torok, M. Golparvar-Fard, K.B. Kochersberger, Image-Based Automated 3D Crack Detection for Post-disaster Building Assessment, J. Comput. Civ. Eng. 28 (2014).

[0187]

[0047] Y. Ham, M. Golparvar-Fard, Identification of Potential Areas for Building Retrofit Using Thermal Digital Imagery and CFD Models, in: R.R. Issa, I. Flood (Eds.), Computing in Civil Engineering (2012), American Society of Civil Engineers, Reston, VA, 2012, pp. 642-649.

[0188]

[0048] Y. Ham, M. Golparvar-Fard, Mapping actual thermal properties to building elements in gbXML-based BIM for reliable building energy performance modeling, Automation in Construction 49 (2015) 214-224.

[0189]

[0049] C. Koch, S.G. Paal, A. Rashidi, Z. Zhu, M. König, I. Brilakis, Achievements and Challenges in Machine Vision-Based Inspection of Large Concrete Structures, Advances in Structural Engineering 17 (2014) 303-318.

[0190]

[0050] A. Rashidi, I. Brilakis, P. Vela, Generating Absolute-Scale Point Cloud Data of Built Infrastructure Scenes Using a Monocular Camera Setting, J. Comput. Civ. Eng. 29 (2015) 4014089.

[0191]

[0051] W. Yuan, S. Chen, Y. Zhang, J. Gong, R. Shibasaki, AN AERIAL-IMAGE DENSE MATCHING APPROACH BASED ON OPTICAL FLOW FIELD, Int. Arch. Photogramm. Remote Sens. Spatial Inf. Sei. XLI-B3 (2016) 543-548.

[0192]

[0052] Z. Zhou, J. Gong, M. Guo, Image-Based 3D Reconstruction for Posthurricane Residential Building Damage Assessment, J. Comput. Civ. Eng. 30 (2016) 4015015.

[0193]

[0053] S. Birchfield, Essential and fundamental matrices, 1999, http: / / ai.stanford.edu / ~birch / projective / node20.html, accessed 18 February 2022.

[0054] M. Nahangi, C.T. Haas, Skeleton-based discrepancy feedback for automated realignment of industrial assemblies, Automation in Construction 61 (2016) 147- 161.

[0194]

[0055] S. Zollmann, C. Hoppe, S. Kluckner, C. Poglitsch, H. Bischof, G. Reitmayr, Augmented Reality for Construction Site Monitoring and Documentation, Proc. IEEE 102 (2014) 137-154.

[0195]

[0056] J. Lee, H. Son, C. Kim, C. Kim, Skeleton-based 3D reconstruction of as-built pipelines from laser-scan data, Automation in Construction 35 (2013) 199-207.

[0196]

[0057] H. Son, C. Kim, C. Kim, Fully Automated As-Built 3D Pipeline Extraction Method from Laser-Scanned Data Based on Curvature Computation, J. Comput. Civ. Eng. 29 (2015).

[0197]

[0058] A. Dimitrov, M. Golparvar-Fard, Segmentation of building point cloud models including detailed architectural / structural features and MEP systems, Automation in Construction 51 (2015) 32-45.

[0198]

[0059] S. Wang, Q. Gou, M. Sun, Simple Building Reconstruction from Lidar Data and Aerial Imagery, in: 2012 2nd International Conference on Remote Sensing, Environment and Transportation Engineering, IEEE, 2012, pp. 1-5.

[0199]

[0060] L. Cheng, Y. Wu, Y. Wang, L. Zhong, Y. Chen, M. Li, Three-Dimensional Reconstruction of Large Multilayer Interchange Bridge Using Airborne LiDAR Data, IEEE J. Sei. Top. Appl. Earth Observations Remote Sensing 8 (2015) 691-708.

[0200]

[0061] A. Adan, D. Huber, 3D Reconstruction of Interior Wall Surfaces under Occlusion and Clutter, in: 2011 International Conference on 3D Imaging, Modeling, Processing, Visualization and Transmission, IEEE, 2011, pp. 275-281.

[0201]

[0062] E. Valero, A. Adän, F. Bosche, Semantic 3D Reconstruction of Furnished Interiors Using Laser Scanning and RFID Technology, J. Comput. Civ. Eng. 30 (2016) 4015053.

[0063] A.K. Patil, P. Holi, S.K. Lee, Y.H. Chai, An adaptive approach for the reconstruction and modeling of as-built 3D pipelines from point clouds, Automation in Construction 75 (2017) 65-78.

[0202]

[0064] M.M. Torok, M.G. Fard, K.B. Kochersberger, Post-Disaster Robotic Building Assessment: Automated 3D Crack Detection from Image-Based Reconstructions, in: R.R. Issa, I. Flood (Eds.), Computing in Civil Engineering (2012), American Society of Civil Engineers, Reston, VA, 2012, pp. 397-404.

[0203]

[0065] M. Weinmann, A. Schmidt, C. Mallet, S. Hinz, F. Rottensteiner, B. Jutzi, Contextual classification of point cloud data by exploiting individual 3d neigbourhoods, Göttingen Copernicus GmbH, 2015.

[0204]

[0066] J. Lalonde, R. Unnikrishnan, N. Vandapel, M. Hebert, Scale Selection for Classification of Point-Sampled 3-D Surfaces, in: Fifth International Conference on 3-D Digital Imaging and Modeling (3DIM'O5), IEEE, 2005, pp. 285-292.

[0205]

[0067] Nesrine Chehata, Li Guo, Clement Mallet, Airborne lidar feature selection for urban classification using random forests, in: International Archives of Photogrammetry. Remote Sensing and Spatial Information Sciences, pp. 207-212.

[0206]

[0068] S.K. Lodha, D.M. Fitzpatrick, D.P. Helmbold, Aerial Lidar Data Classification using Ada-Boost, in: Sixth International Conference on 3-D Digital Imaging and Modeling (3DIM 2007), IEEE, 2007, pp. 435-442.

[0207]

[0069] Z. Wang, L. Zhang, T. Fang, P.T. Mathiopoulos, X. Tong, H. Qu, Z. Xiao, F. Li, D. Chen, A Multiscaleand Hierarchical Feature Extraction Method for Terrestrial Laser Scanning Point Cloud Classification, IEEE Trans. Geosci. Remote Sensing 53 (2015) 2409-2425.

[0208]

[0070] J. Zhang, X. Lin, X. Ning, SVM-Based Classification of Segmented Airborne LiDAR Point Cloudsin Urban Areas, Remote Sensing 5 (2013) 3749-3775.

[0209]

[0071] Z. Li, L. Zhang, X. Tong, B. Du, Y. Wang, L. Zhang, Z. Zhang, H. Liu, J. Mei, X. Xing, P.T. Mathiopoulos, A Three-Step Approach for TLS Point Cloud Classification, IEEE Trans. Geosci. Remote Sensing 54 (2016) 5412-5424.

[0072] M. Carlberg, P. Gao, G. Chen, A. Zakhor, Classifying urban landscape in aerial LiDAR using 3D shape analysis, in: 2009 16th IEEE International Conference on Image Processing (ICIP), IEEE, 2009, pp. 1701-1704.

[0210]

[0073] K. Khoshelham, S.O. Elberink, Accuracy and resolution of Kinect depth data for indoor mapping applications, Sensors (Basel, Switzerland) 12 (2012) 1437-1454.

[0211]

[0074] Roman Shapovalov, Er Velizhev, Olga Barinova, Nonassociative markov networks for 3d point cloud classification. The, 2010.

[0212]

[0075] M. Najafi, S. Taghavi Namin, M. Salzmann, L. Petersson, Non-associative Higher- Order Markov Networks for Point Cloud Classification, in: D. Fleet, T. Pajdla, B. Schiele, T. Tuytelaars (Eds.), Computer Vision -ECCV 2014, Springer International Publishing, Cham, 2014, pp. 500-515.

[0213]

[0076] D. Munoz, J. A. Bagnell, N. Vandapel, M. Hebert, Contextual classification with functional Max-Margin Markov Networks, in: 2009 IEEE Conference on Computer Vision and Pattern Recognition, IEEE, 2009, pp. 975-982.

[0214]

[0077] J. Niemeyer, F. Rottensteiner, U. Soergel, CONDITIONAL RANDOM FIELDS FOR LIDAR POINT CLOUD CLASSIFICATION IN COMPLEX URBAN AREAS, ISPRS Ann. Photogramm. Remote Sens. Spatial Inf. Sei. I-3 (2012) 263-268.

[0215]

[0078] J. Niemeyer, F. Rottensteiner, U. Soergel, Contextual classification of lidar data and building object detection in urban areas, ISPRS Journal of Photogrammetry and Remote Sensing 87 (2014) 152-165.

[0216]

[0079] G. Vosselman, M. Coenen, F. Rottensteiner, Contextual segment-based classification of airborne laser scanner data, ISPRS Journal of Photogrammetry and Remote Sensing 128 (2017) 354-371.

[0217]

[0080] E.H. Lim, D. Suter, 3D terrestrial LIDAR classifications with super-voxels and multiscale Conditional Random Fields, Computer-Aided Design 41 (2009) 701-710.

[0081] A. Schmidt, F. Rottensteiner, II. Sörgel, Classification of airborne laser scanning data in wadden sea areas using conditional random fields, Göttingen Copernicus GmbH, 2012.

[0218]

[0082] Y. Lu, C. Rasmussen, Simplified markov random fields for efficient semantic labeling of 3D point clouds, in: 2012 IEEE / RSJ International Conference on Intelligent Robots and Systems, IEEE, 2012, pp. 2690-2697.

[0219]

[0083] X. Xiong, D. Munoz, J. A. Bagnell, M. Hebert, 3-D scene analysis via sequenced predictions over points and regions, in: 2011 IEEE International Conference on Robotics and Automation, IEEE, 2011, pp. 2609-2616.

[0220]

[0084] R. Shapovalov, D. Vetrov, P. Kohli, Spatial Inference Machines, in: 2013 IEEE Conference on Computer Vision and Pattern Recognition, IEEE, 2013, pp. 2985- 2992.

[0221]

[0085] M. Weinmann, B. Jutzi, S. Hinz, C. Mallet, Semantic point cloud interpretation based on optimal neighborhoods, relevant features and efficient classifiers, ISPRS Journal of Photogrammetry and Remote Sensing 105 (2015) 286-304.

[0222]

[0086] Y. Guo, H. Wang, Q. Hu, H. Liu, L. Liu, M. Bennamoun, Deep Learning for 3D Point Clouds: A Survey, arXiv, 2019.

[0223]

[0087] F.J. Lawin, M. Danelljan, P. Tosteberg, G. Bhat, F.S. Khan, M. Felsberg, Deep Projective 3D Semantic Segmentation, in: M. Felsberg, A. Heyden, N. Krüger (Eds.), Computer Analysis of Images and Patterns, Springer International Publishing, Cham, 2017, pp. 95-107.

[0224]

[0088] A. Boulch, B. Le Saux, N. Audebert, Unstructured Point Cloud Semantic Labeling Using Deep Segmentation Networks, The Eurographics Association, 2017.

[0225]

[0089] N. Audebert, B. Le Saux, S. Lefevre, Semantic Segmentation of Earth Observation Data Using Multimodal and Multiscale Deep Networks, in: S.-H. Lai, V. Lepetit, K. Nishino, Y. Sato (Eds.), Computer Vision -ACCV 2016, Springer International Publishing, Cham, 2017, pp. 180-196.

[0090] M. Tatarchenko, J. Park, V. Koltun, Q.-Y. Zhou, Tangent Convolutions for Dense Prediction in 3D, arXiv, 2018.

[0226]

[0091] F.N. landola, S. Han, M.W. Moskewicz, K. Ashraf, W.J. Dally, K. Keutzer, SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size, arXiv, 2016.

[0227]

[0092] B. Wu, A. Wan, X. Yue, K. Keutzer, SqueezeSeg: Convolutional Neural Nets with Recurrent CRF for Real-Time Road-Object Segmentation from 3D LiDAR Point Cloud, in: 2018 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 2018, pp. 1887-1893.

[0228]

[0093] B. Wu, X. Zhou, S. Zhao, X. Yue, K. Keutzer, SqueezeSegV2: Improved Model Structure and Unsupervised Domain Adaptation for Road-Object Segmentation from a LiDAR Point Cloud, in: 2019 International Conference on Robotics and Automation (ICRA), IEEE, 2019, pp. 4376-4382.

[0229]

[0094] A. Milioto, I. Vizzo, J. Behley, C. Stachniss, RangeNet ++: Fast and Accurate LiDAR Semantic Segmentation, in: 2019 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2019, pp. 4213-4220.

[0230]

[0095] J. Huang, S. You, Point cloud labeling using 3D Convolutional Neural Network, in: 2016 23rd International Conference on Pattern Recognition (ICPR), IEEE, 2016, pp. 2670-2675.

[0231]

[0096] E. Shelhamer, J. Long, T. Darrell, Fully Convolutional Networks for Semantic Segmentation, IEEE Trans. Pattern Anal. Machine Intell. 39 (2017) 640-651.

[0232]

[0097] L. Tchapmi, C. Choy, I. Armeni, J. Gwak, S. Savarese, SEGCloud: Semantic Segmentation of 3D Point Clouds, in: 2017 International Conference on 3D Vision (3DV), IEEE, 2017, pp. 537-547.

[0233]

[0098] H.-Y. Meng, L. Gao, Y. Lai, D. Manocha, VV-Net: Voxel VAE Net with Group Convolutions for Point Cloud Segmentation, arXiv, 2018.

[0099] D. Rethage, J. Wald, J. Sturm, N. Navab, F. Tombari, Fully-Convolutional Point Networks for Large-Scale Point Clouds, arXiv, 2018.

[0234]

[0100] A. Dai, D. Ritchie, M. Bokeloh, S. Reed, J. Sturm, M. Nießner, ScanComplete: Large-Scale Scene Completion and Semantic Segmentation for 3D Scans, arXiv, 2017.

[0235]

[0101] B. Graham, M. Engelcke, L. van der Maaten, 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks, 2017.

[0236]

[0102] C. Choy, J. Gwak, S. Savarese, 4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks, 2019.

[0237]

[0103] H. Su, V. Jampani, D. Sun, S. Maji, E. Kalogerakis, M.-H. Yang, J. Kautz, SPLATNet: Sparse Lattice Networks for Point Cloud Processing, 2018.

[0238]

[0104] R.A. Rosu, P. Schütt, J. Quenzel, S. Behnke, LatticeNet: Fast Point Cloud Segmentation Using Permutohedral Lattices, 2019.

[0239]

[0105] A. Dai, M. Nießner, 3DMV: Joint 3D-Multi-View Prediction for 3D Semantic Scene Segmentation, 2018.

[0240]

[0106] H.-Y. Chiang, Y.-L. Lin, Y.-C. Liu, W.H. Hsu, A Unified Point-Based Framework for 3D Segmentation, in: 2019 International Conference on 3D Vision (3DV), IEEE, 2019, pp. 155-163.

[0241]

[0107] M. Jaritz, J. Gu, H. Su, Multi-view PointNet for 3D Scene Understanding, 2019.

[0242]

[0108] C.R. Qi, H. Su, K. Mo, L.J. Guibas, PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, 2016.

[0243]

[0109] C.R. Qi, L. Yi, H. Su, L.J. Guibas, PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, 2017.

[0244]

[0110] M. Jiang, Y. Wu, T. Zhao, Z. Zhao, C. Lu, PointSIFT: A SIFT-like Network Module for 3D Point Cloud Semantic Segmentation, 2018.

[0111] F. Engelmann, T. Kontogianni, J. Schult, B. Leibe, Know What Your Neighbors Do: 3D Semantic Segmentation of Point Clouds 11131 (2019) 395-409.

[0245]

[0112] H. Zhao, L. Jiang, C.-W. Fu, J. Jia, PointWeb: Enhancing Local Neighborhood Features for Point Cloud Processing, in: 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2019, pp. 5560-5568.

[0246]

[0113] Z. Zhang, B.-S. Hua, S.-K. Yeung, ShellNet: Efficient Point Cloud Convolutional Neural Networks using Concentric ShellsStatistics, 2019.

[0247]

[0114] Q. Hu, B. Yang, L. Xie, S. Rosa, Y. Guo, Z. Wang, N. Trigoni, A. Markham, RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds, 2019.

[0248]

[0115] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser, I. Polosukhin, Attention Is All You Need, 2017.

[0249]

[0116] J. Yang, Q. Zhang, B. Ni, L. Li, J. Liu, M. Zhou, Q. Tian, Modeling Point Clouds with Self-Attention and Gumbel Subset Sampling, 2019.

[0250]

[0117] L.-Z. Chen, X.-Y. Li, D.-P. Fan, K. Wang, S.-P. Lu, M.-M. Cheng, LSANet: Feature Learning on Point Sets by Local Spatial Aware Layer, 2019.

[0251]

[0118] C. Zhao, W. Zhou, L. Lu, Q. Zhao, Pooling Scores of Neighboring Points for Improved 3D Point Cloud Segmentation, in: 2019 IEEE International Conference on Image Processing (ICIP), IEEE, 2019, pp. 1475-1479.

[0252]

[0119] Y. Wang, Y. Sun, Z. Liu, S.E. Sarma, M.M. Bronstein, J.M. Solomon, Dynamic Graph CNN for Learning on Point Clouds, 2018.

[0253]

[0120] R. Arandjelovic , P. Gronat, A. Torii, T. Pajdla, J. Sivic, NetVLAD: CNN architecture for weakly supervised place recognition, 2015.

[0254]

[0121] N. Zhao, T.-S. Chua, G.H. Lee, PS A 2-Net: A Locally and Globally Aware Network for Point-Based Semantic Segmentation 12667 (2021).

[0122] B.-S. Hua, M.-K. Tran, S.-K. Yeung, Pointwise Convolutional Neural Networks, 2017.

[0255]

[0123] S. Wang, S. Suo, W.-C. Ma, A. Pokrovsky, R. Urtasun, Deep Parametric Continuous Convolutional Neural Networks (2018) 2589-2597.

[0256]

[0124] H. Thomas, C.R. Qi, J.-E. Deschaud, B. Marcotegui, F. Goulette, L.J. Guibas, KPConv: Flexible and Deformable Convolution for Point Clouds, 2019.

[0257]

[0125] F. Engelmann, T. Kontogianni, B. Leibe, Dilated Point Convolutions: On the Receptive Field Size of Point Convolutions on 3D Point Clouds, 2019.

[0258]

[0126] F. Engelmann, T. Kontogianni, A. Hermans, B. Leibe, Exploring Spatial Context for 3D Semantic Segmentation of Point Clouds (2017) 716-724.

[0259]

[0127] Q. Huang, W. Wang, U. Neumann, Recurrent Slice Networks for 3D Segmentation of Point Clouds, 2018.

[0260]

[0128] X.Ye, J. Li, H. Huang, L. Du, X. Zhang, 3D Recurrent Neural Networks with Context Fusion for Point Cloud Semantic Segmentation, in: V. Ferrari, M. Hebert, C. Sminchisescu, Y. Weiss (Eds.), Computer Vision -ECCV 2018, Springer International Publishing, Cham, 2018, pp. 415-430.

[0261]

[0129] Z. Zhao, M. Liu, K. Ramani, DAR-Net: Dynamic Aggregation Network for Semantic Scene Segmentation, 2019.

[0262]

[0130] F. Liu, S. Li, L. Zhang, C. Zhou, R. Ye, Y. Wang, J. Lu, 3DCNN-DQN-RNN: A Deep Reinforcement Learning Framework for Semantic Parsing of Large-scale 3D Point Clouds, 2017.

[0263]

[0131] L. Landrieu, M. Simonovsky, Large-scale Point Cloud Semantic Segmentation with Superpoint Graphs, 2017.

[0264]

[0132] L. Landrieu, M. Boussaha, Point Cloud Oversegmentation with Graph-Structured Deep MetricLearning, 2019.

[0133] K. Zhiheng, L. Ning, PyramNet: Point Cloud Pyramid Attention Network and Graph Embedding Module for Classification and Segmentation, 2019.

[0265]

[0134] L. Wang, Y. Huang, Y. Hou, S. Zhang, J. Shan, Graph Attention Convolution for Point Cloud Semantic Segmentation, in: 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2019, pp. 10288-10297.

[0266]

[0135] Y. Ma, Y. Guo, H. Liu, Y. Lei, G. Wen, Global Context Reasoning for Semantic Segmentation of 3D Point Clouds, in: 2020 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE, 2020, pp. 2920-2929.

[0267]

[0136] J. Wei, G. Lin, K.-H. Yap, T.-Y. Hung, L. Xie, Multi-Path Region Mining For Weakly Supervised 3D Semantic Segmentation on Point Clouds, 2020.

[0268]

[0137] X. Xu, G.H. Lee, Weakly Supervised Semantic Point Cloud Segmentation: Towards 10X Fewer Labels, 2020.

[0269]

[0138] J. Hou, A. Dai, M. Nießner, 3D-SIS: 3D Semantic Instance Segmentation of RGB- D Scans, 2018.

[0270]

[0139] L. Yi, W. Zhao, H. Wang, M. Sung, L. Guibas, GSPN: Generative Shape Proposal Network for 3D Instance Segmentation in Point Cloud, 2018.

[0271]

[0140] G. Narita, T. Seno, T. Ishikawa, Y. Kaji, PanopticFusion: Online Volumetric Semantic Mapping at the Level of Stuff and Things, 2019.

[0272]

[0141] B. Yang, J. Wang, R. Clark, Q. Hu, S. Wang, A. Markham, N. Trigoni, Learning Object Bounding Boxes for 3D Instance Segmentation on Point Clouds, 2019.

[0273]

[0142] F. Zhang, C. Guan, J. Fang, S. Bai, R. Yang, P.H. Torr, V. Prisacariu, Instance Segmentation of LiDAR Point Clouds, in: 2020 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 52020, pp. 9448-9455.

[0274]

[0143] Y. Shi, A.X. Chang, Z. Wu, M. Savva, K. Xu, Hierarchy Denoising Recursive Autoencoders for 3D Scene Layout Prediction, 2019.

[0144] F. Engelmann, M. Bokeloh, A. Fathi, B. Leibe, M. Nießner, 3D-MPA: Multi Proposal Aggregation for 3D Semantic Instance Segmentation, 2020.

[0275]

[0145] W. Wang, R. Yu, Q. Huang, II. Neumann, SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance Segmentation, 2017.

[0276]

[0146] C. Liu, Y. Furukawa, MASC: Multi-scale Affinity with Sparse Convolution for 3D Instance Segmentation, 2019.

[0277]

[0147] Z. Liang, M. Yang, C. Wang, 3D Graph Embedding Learning with a Structure- aware Loss Function for Point Cloud Semantic Instance Segmentation (2019).

[0278]

[0148] X. Wang, S. Liu, X. Shen, C. Shen, J. Jia, Associatively Segmenting Instances and Semantics in Point Clouds, 2019.

[0279]

[0149] L. Zhao, W. Tao, JSNet: Joint Instance and Semantic Segmentation of 3D Point Clouds, 2019.

[0280]

[0150] B.D. Brabandere, D. Neven, L. van Gool, Semantic Instance Segmentation with a Discriminative Loss Function, 2017.

[0281]

[0151] Q.-H. Pham, D.T. Nguyen, B.-S. Hua, G. Roig, S.-K. Yeung, JSIS3D: Joint Semantic-Instance Segmentation of 3D Point Clouds with Multi-Task Pointwise Networks and Multi-Value Conditional Random Fields, 2019.

[0282]

[0152] S.-M. Hu, J.-X. Cai, Y.-K. Lai, Semantic Labeling and Instance Segmentation of 3D Point Clouds Using Patch Context Analysis and Multiscale Processing, IEEE transactions on visualization and computer graphics 26 (2020) 2485-2498.

[0283]

[0153] C. Elich, F. Engelmann, T. Kontogianni, B. Leibe, 3D-BEVIS: Bird's-Eye- iew Instance Segmentation 11824 (2019) 48-61.

[0284]

[0154] J. Lahoud, B. Ghanem, M. Pollefeys, M.R. Oswald, 3D Instance Segmentation via Multi-Task Metric Learning, 2019.

[0155] L. Jiang, H. Zhao, S. Shi, S. Liu, C.-W. Fu, J. Jia, PointGroup: Dual-Set Point Grouping for 3D Instance Segmentation, 2020.

[0285]

[0156] Z. Liang, Z. Li, S. Xu, M. Tan, K. Jia, Instance Segmentation in 3D Scenes using Semantic Superpoint Tree Networks, 2021.

Claims

Patent claims 1. A device for creating a digital model of an existing building or an existing industrial plant or an existing infrastructure, the device comprising: an input unit (110) for providing point cloud data representing the building or the industrial plant or the infrastructure, a processing unit (120) comprising an artificial intelligence module (125) for creating the model depending on the point cloud data, the artificial intelligence module (125) being trained by means of machine learning.

2. Device according to claim 1, wherein the input unit (110) is designed to receive two or more point cloud partial data sets, each representing a partial area of ​​the building or a partial area of ​​the industrial plant or a partial area of ​​the infrastructure, and wherein the input unit (110) is designed to generate the point cloud data representing the building or the industrial plant or the infrastructure from the two or more point cloud data.

3. The device according to claim 1 or 2, wherein the input unit (110) is configured to receive sensor data from one or more cameras and / or from one or more laser scanners and / or from one or more depth cameras; and wherein the input unit (110) is configured to determine the point cloud data and a portion of the point cloud data depending on the sensor data; or wherein the sensor data comprises the point cloud data or a portion of the point cloud data.

4. Device according to one of the preceding claims, wherein the processing unit (120) is configured to detect the data in the point cloud data that do not represent the building or the industrial plant or the infrastructure, and / or that lie outside the building or the industrial plant or the infrastructure.

5. The apparatus of claim 4, wherein the processing unit (120) is configured to determine boundaries of the building or industrial plant or infrastructure in order to detect the data in the point cloud data that lie outside the building or industrial plant or infrastructure.

6. Apparatus according to any one of the preceding claims, wherein the point cloud data comprises information about surfaces of the building or industrial plant or infrastructure and information about the interior of the building or industrial plant or infrastructure.

7. Device according to one of the preceding claims, wherein the processing unit (120) is designed to process the point cloud data in such a way that one or more subsets of the point cloud data are classified in such a way that the one or more subsets are each assigned to one of one or more object categories.

8. The apparatus of claim 7, wherein the one or more object categories comprise one or more building part types, and wherein the processing unit (120) is configured to perform the classification in such a way that it comprises a first instance segmentation that classifies at least a subset of the one or more subsets of the point cloud data in such a way that it is classified as one of the one or more building part types of the building or the industrial plant or the infrastructure.

9. Device according to claim 8, where the one or more building component types can be freely defined by a developer.

10. The apparatus of claim 8, wherein the one or more building part types comprise one or more of the following building part types: a roof, an exterior wall, an exterior door, an exterior window, an interior space, a stairwell, an elevator shaft.

11. Device according to one of claims 8 to 10, wherein the one or more object categories comprise one or more spatial structure types, and wherein the processing unit (120) is designed to carry out the classification in such a way that this comprises a second instance segmentation which classifies at least a subset of the one or more subsets of the point cloud data in such a way that it is classified as a spatial structure from the one or more spatial structure types of a room of the building or the industrial plant or the infrastructure.

12. The device according to claim 11, wherein the one or more spatial structure types are freely definable by a developer.

13. The device according to claim 11, wherein the one or more spatial structure types comprise one or more of the following spatial structure types: a wall, a ceiling, a floor, a door, a window, a piece of furniture.

14. Device according to one of claims 7 to 13, wherein the processing unit (120) is configured to classify the one or more subsets of the point cloud data in such a way that classification is carried out by using the artificial intelligence module (125) which is trained by means of machine learning and which is configured to receive the point cloud data as input data.

15. The apparatus of claim 14, wherein the artificial intelligence module (125) is trained by supervised deep learning.

16. The device according to claim 15, wherein the artificial intelligence module (125) is designed to perform the classification in a projection-based and / or discretization-based and / or point-based manner.

17. The device according to any one of claims 14 to 16, wherein the artificial intelligence module (125) comprises a deep learning neural network having at least two hidden layers.

18. Device according to one of claims 14 to 17, wherein the artificial intelligence module (125) comprises a recurrent neural network for determining one or more features of a sub-area of ​​the point cloud data.

19. Device according to one of claims 14 to 18, wherein the artificial intelligence module (125) is trained by means of machine learning by assigning a first point cloud to an object category of the one or more object categories, and wherein the first point cloud and said object category were used as a first training data set for the artificial intelligence module (125).

20. Device according to claim 19, wherein the artificial intelligence module (125) is trained by means of machine learning by modifying the first point cloud one or more times to obtain one or more modified point clouds, wherein in each case one of the one or more modified point clouds and the said object category are used as a further training data set from one or more further training data sets.

21. The apparatus of claim 20, wherein the first point cloud has been modified one or more times by applying translation dithering, rotation, random scaling, and / or elastic deformation of the first point cloud.

22. Apparatus according to any one of claims 14 to 21, wherein the point cloud data is three-dimensional point cloud data, and wherein the processing unit (120) is configured to map the three-dimensional point cloud data to two-dimensional point cloud data and to classify the one or more subsets of the point cloud data using the two-dimensional point cloud data.

23. Device according to one of claims 7 to 22, wherein the point cloud data is three-dimensional point cloud data, and wherein the processing unit (120) is designed to perform an assignment of the three-dimensional point cloud data to a plurality of voxels and to perform the point cloud data depending on the assignment to the plurality of voxels.

24. Device according to one of claims 7 to 23, wherein the processing unit (120) is designed to create a virtual construction plan depending on the classification of the one or more subsets of the point cloud data.

25. Device according to one of claims 7 to 24, wherein the processing unit (120) is designed to create a three-dimensional building data model depending on the classification of the one or more subsets of the point cloud data.

26. A method for creating a digital model of an existing building or industrial facility or infrastructure, the method comprising: Providing point cloud data representing the building or industrial facility or infrastructure, and Creating the model depending on the point cloud data using a processing unit (120) comprising an artificial intelligence module (125), wherein the artificial intelligence module (125) is trained by means of machine learning.

27. A computer program with a program code for carrying out the method according to claim 26.