Building indoor scene deep learning reconstruction method and system based on point cloud 2D-3D fusion
By adopting a deep learning method of point cloud 2D-3D fusion in architectural indoor scenarios, and combining the classification results of YOLO and PointNet series models, semantic label updates are performed, which solves the problem of low segmentation accuracy of point cloud reconstruction in the existing technology, achieves higher accuracy and efficiency, and supports intelligent construction and maintenance of BIM.
Patent Information
- Application Number
- CN202411816349.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-11
AI Technical Summary
The prior art has low segmentation accuracy when dealing with point cloud reconstruction in building indoor scenes, especially in point cloud data processing with similar shapes or sparse shapes.
The deep learning reconstruction method based on point cloud 2D-3D fusion is adopted to semantic segmentation of point clouds through deep learning algorithms, and a virtual camera is used to generate two-dimensional images, combining the classification results of YOLO and PointNet series models to update semantic labels, and finally build a building information model BIM.
The accuracy of point cloud semantic segmentation is significantly improved, especially in point cloud data processing with similar shapes or sparse shapes, providing higher accuracy and efficiency than existing methods, supporting intelligent construction and maintenance of building information model (BIM).
Smart Images

Figure CN119991930A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of civil engineering and artificial intelligence interaction, and specifically relates to a deep learning reconstruction method and system for building indoor scenes based on point cloud 2D-3D fusion. Background Art
[0002] Point cloud data is important information for architectural scene modeling and is usually used for scene reconstruction, object recognition and other tasks. Although traditional point cloud semantic segmentation methods have achieved certain results, due to the sparse point cloud data and complex local geometric structure, the existing deep learning algorithms still have low segmentation accuracy when dealing with categories with fewer instances or similar shapes. In recent years, 2D and 3D fusion algorithms based on deep learning have made some progress in improving the accuracy of deep learning models, but they still face the challenge of requiring pre-registered images and point cloud data. How to better fuse 2D image information to improve segmentation accuracy is still a research problem that needs to be solved urgently in this field. Summary of the invention
[0003] The present invention aims to solve the existing problems of point cloud reconstruction of building indoor scenes in the prior art, and provides a method and system for deep learning reconstruction of building indoor scenes based on point cloud 2D-3D fusion, aiming to improve the reconstruction accuracy of point cloud of building indoor scenes by combining point cloud semantic segmentation and image semantic segmentation technology. Unlike traditional methods, the present invention does not require 2D images registered with point clouds, but directly generates corresponding 2D images through point clouds, and semantically segments and recognizes point clouds and images at the same time, and finally fuses the segmentation results of the two, which can improve the accuracy of point cloud semantic segmentation.
[0004] In order to solve the above technical problems, the present invention provides the following technical solutions: a deep learning reconstruction method for building indoor scenes based on point cloud 2D-3D fusion, comprising the following steps:
[0005] S1. Use deep learning algorithms to perform semantic segmentation on the original 3D point cloud data of indoor scenes to obtain 3D semantic segmentation results;
[0006] S2, generate several two-dimensional images from the original three-dimensional point cloud data through a virtual camera, then use the YOLO model to perform target detection and semantic segmentation on the two-dimensional images, generate a mask of a two-dimensional graphic, and then obtain the two-dimensional semantic segmentation result of voxel reconstruction through posture adjustment;
[0007] S3, updating the semantic label according to the two-dimensional semantic segmentation result obtained in step S2 and the three-dimensional semantic segmentation result in step S1, selecting the category with the highest comprehensive confidence as the final semantic label of the point, and obtaining the final semantic segmentation result;
[0008] S4. Based on the final semantic segmentation results, a building information model (BIM) is constructed to reconstruct the building interior scene.
[0009] Then the final scene reconstruction result.
[0010] Furthermore, the aforementioned step S2 generates a plurality of two-dimensional images from the original three-dimensional point cloud data by using a virtual camera, and includes the following sub-steps:
[0011] S2.1, setting multiple virtual camera positions in the scene of 3D point cloud data, and taking images at different angles;
[0012] S2.2, record the depth information of each image and the internal and external parameter data of the camera, the depth information is used to calculate the three-dimensional point cloud position corresponding to each image pixel;
[0013] S2.3. Map the pixels of each image to a three-dimensional point cloud and generate a correspondence between the image and the point cloud.
[0014] Furthermore, the aforementioned step S2.2 includes the following sub-steps:
[0015] S2.2-1. When generating a two-dimensional image, obtain the intrinsic parameter matrix and extrinsic parameter matrix of the virtual camera, establish the corresponding relationship between the point cloud data and the image, and the intrinsic parameter matrix of the virtual camera maps the three-dimensional points in the camera coordinate system to the two-dimensional image coordinate system; the intrinsic parameter matrix of the camera is:
[0016]
[0017] Among them, f x and c y Respectively represent the focal length of the camera in the x and y directions, c x and c y are the coordinates of the principal point;
[0018] S2.2-2. Use the camera's extrinsic matrix E to transform the three-dimensional point from the world coordinate system to the camera coordinate system. The extrinsic matrix consists of a rotation matrix and a translation vector:
[0019]
[0020] Where R is the rotation matrix and t is the translation vector.
[0021] Furthermore, the aforementioned step S2.3 includes the following sub-steps:
[0022] S2.3-1. By using the intrinsic parameter matrix K and the extrinsic parameter matrix E, as well as the captured image and depth information, the position of each pixel in the three-dimensional space is calculated. For any pixel (u, v) in the image and its corresponding depth value D(u, v), the three-dimensional point (X) in the camera coordinate system is calculated by the following formula:cam ,Y cam ,Z cam ),
[0023]
[0024] Z cam =D(u,v) (5)
[0025] The three-dimensional point P in the camera coordinate system cam It is expressed as:
[0026]
[0027] S2.3-2 After generating the point cloud, a voxel grid is constructed to aggregate points from different viewpoints by converting the three-dimensional points in the camera coordinate system into the three-dimensional points in the world coordinate system as shown in the following formula;
[0028] P world =E -1 ·P cam (7)
[0029] Among them, P world is a three-dimensional point in the world coordinate system;
[0030] S2.3-3. Based on the three-dimensional point cloud data generated by the image, a plurality of three-dimensional points from different perspectives are aggregated into corresponding voxels using a voxel grid method, and the category of the voxel is determined according to the category of the most points in the voxel.
[0031] Furthermore, the aforementioned step S3 includes the following sub-steps:
[0032] S3.1. For each point λ in the point cloud, the point-by-point classification probability obtained by the PointNet model is expressed as:
[0033] P PointNet family (λ)=[p1(λ), p2(λ),…,p n (λ)] (8)
[0034] Among them, p i (λ) is the predicted probability that point λ belongs to the i-th class, and the sum of all probabilities is 1, that is:
[0035] S3.2, YOLO prediction results, when projected into a 3D point cloud, provide a confidence score for each predicted object category. For a point λ in 3D space corresponding to a YOLO 2D detection result, YOLO is based on the confidence score P YOLO (λ)
[0036] It is expressed as follows:
[0037] P YOLO (λ)=[q1(λ), q2(λ),…,q n (λ)] (9)
[0038] Among them, q i (λ) is the confidence score that point λ belongs to the i-th class based on YOLO prediction;
[0039] S3.3. For each point λ in the point cloud, its final category confidence is obtained by adding the confidence values of each category of the PointNet series model and YOLO, as shown in the following formula:
[0040] P combined (λ)=P PointNet family (λ)+P YOLO (λ) (10)
[0041] S3.4, determine the label of point λ, first in P PointNetfamily (λ) and P YOLO (λ) to find the category with the largest confidence:
[0042] In P PointNetfamily The category with the largest confidence in (λ) is:
[0043] C PointMet family (λ)=argmax(P PointNet family (λ)) (11)
[0044] In P combined The category with the largest confidence in (λ) is:
[0045] C combined (λ)=argmax(P combined (λ)) (12)
[0046] S3.5. The decision to update the semantic label of point λ is based on the comparison of the confidence scores of the two categories in step S3.4. The specific update rules are as follows:
[0047] P combined (λ,C combined (λ))and P PointNet family (λ,C PointNet family (λ)) (13) If the confidence of the fused label is greater than the confidence of the original PointNet series label, the label is updated to
[0048] C combined (λ);
[0049] Otherwise, keep the original label C PointNet family (λ),
[0050] The formula for label update is:
[0051]
[0052] Another aspect of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the present invention when executing the computer program.
[0053] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is used by a processor to execute the steps of any method described in the present invention.
[0054] Compared with the prior art, the beneficial technical effects of the above technical solutions adopted by the present invention are as follows: by directly generating images from point clouds without pre-registered two-dimensional images and combining 2D and 3D data, the accuracy of semantic segmentation of point clouds is improved. This method effectively processes complex and sparse point cloud data, combines the classification results of YOLO and PointNet series models, and improves classification accuracy through confidence update, thereby achieving more accurate reconstruction of building interior scenes and supporting intelligent construction and maintenance of building information models (BIM). BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flow chart of the method of the present invention.
[0056] Figure 2 This is an example diagram of the present invention using majority voting to assign semantic labels to point clouds; in the figure, (a) is a classification diagram of consistent voxel p, and (b) is an example diagram of majority voting for voxel Q.
[0057] Figure 3 It is a visualization of the point cloud scene in Autodesk Recap; (a) is the visualization of scene 1, and (b) is the visualization of scene 2.
[0058] Figure 4 : is an example diagram of the process of updating semantic information using the Point-YOLO method; in the figure, (a) is a schematic diagram before the semantic information is updated, (b) is a schematic diagram after the semantic information is updated, (c) is a diagram of the semantic update process of point 1, and (d) is a diagram of the semantic update process of point 2.
[0059] Figure 5 It is a schematic diagram of the confusion matrix comparison results of scene 1; in the figure, (a) is a schematic diagram of the PointNet++ confusion matrix results, and (b) is a schematic diagram of the PointNet++-YOLO confusion matrix results.
[0060] Figure 6 It is a schematic diagram of the confusion matrix comparison results of scene 2; in the figure, (a) is a schematic diagram of the PointNet++ confusion matrix results, and (b) is a schematic diagram of the PointNet++-YOLO confusion matrix results. DETAILED DESCRIPTION
[0061] In order to better understand the technical content of the present invention, specific embodiments are given and described as follows in conjunction with the accompanying drawings.
[0062] Various aspects of the invention are described herein with reference to the accompanying drawings, in which many illustrative embodiments are shown. The embodiments of the invention are not limited to those described in the accompanying drawings. It should be understood that the invention is implemented by any of the various concepts and embodiments described above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed in the invention are not limited to any implementation. In addition, some aspects disclosed in the invention may be used alone or in any appropriate combination with other aspects disclosed in the invention.
[0063] refer to Figure 1 The present invention provides a deep learning reconstruction method for building indoor scenes based on point cloud 2D-3D fusion. First, a point-based deep learning algorithm, such as PointNet or PointNet++, is used to semantically segment the point cloud. Secondly, a simulated camera method is used to generate virtual images from the point cloud scene, and a deep learning algorithm such as YOLO is used to semantically segment these images. Finally, based on the 2D segmentation results, a 3D voxel representation with detected object labels is reconstructed. The 3D voxel representation is used to optimize and update the semantic segmentation of the point cloud according to the image segmentation results.
[0064] This implementation adopts two different scenes for experimental verification, namely scene 1 (building indoor scene) and scene 2 (underground garage dataset UGD). The building indoor dataset contains six areas (Area 1 to Area 6) and defines 13 semantic categories: ceiling, floor, wall, beam, column, window, door, table, chair, sofa, bookshelf, plate and cluttered objects. The UGD dataset was scanned in the underground garage of the Humanities Building of Southeast University by Leica RTC360 scanner, containing XYZ coordinates, RGB color values and density information, but does not contain registered image data. The UGD dataset is divided into five areas (Area 1 to Area 5) and includes 10 component categories: ducts, cable trays, ceilings, fire extinguishers, floors, lamps, columns, pipes, walls and doors. Since the dataset combines large building components (such as ceilings and floors) with smaller and complex electromechanical system (MEP) components, its scene segmentation task is highly challenging.
[0065] The datasets for scene 1 and scene 2 are divided into training set and test set respectively. In scene 1, the point cloud data of Area 1, Area 2, Area 3, Area 4 and Area 6 are used for training set, and the 42 office scenes of Area 5 are used as test set. For scene 2, Area 1, Area 2 and Area 5 are used for training set, and Area 3 and Area 4 are used as test set. The legend of the scene is shown in Autodesk Recap as follows Figure 3 The visualization of the point cloud scene in Autodesk Recap is shown below; (a) is the visualization of scene 1, and (b) is the visualization of scene 2.
[0066] The specific steps of the present invention are as follows:
[0067] S1. The original three-dimensional point cloud data of the indoor scene is semantically segmented using a deep learning algorithm to obtain a three-dimensional semantic segmentation result. The present invention uses a deep learning algorithm based on PointNet and PointNet++ to semantically segment point cloud data. PointNet processes each point through a shared multi-layer perceptron (MLPs), extracts features, and aggregates features through symmetric functions (such as maximum pooling) to obtain global features, which is suitable for processing unordered point cloud data. However, PointNet has the problem of insufficient capture of local geometric structures. To this end, PointNet++ introduces a hierarchical learning mechanism, which selects the center point and forms a local area through farthest point sampling and radius search, thereby capturing multi-scale local structural information.
[0068] S2, generate several two-dimensional images of the original three-dimensional point cloud data through a virtual camera, then use the YOLO model to perform target detection, semantic segmentation, generate a mask of a two-dimensional graphic, and then obtain the two-dimensional semantic segmentation result of voxel reconstruction through posture adjustment. The present invention uses YOLOv8 to perform semantic segmentation of images. YOLOv8 is an image-based object detection and segmentation algorithm that combines efficient speed and accuracy. This method is based on a convolutional neural network (CNN), which can directly predict bounding boxes and category probabilities, divide the image into multiple grids, and each grid predicts an object and its confidence score.
[0069] Step S2 generates a number of two-dimensional images of the original three-dimensional point cloud data through a virtual camera, including the following sub-steps:
[0070] S2.1, setting multiple virtual camera positions in the scene of 3D point cloud data, and taking images at different angles;
[0071] S2.2, recording the depth information of each image and the internal and external parameter data of the camera, the depth information is used to calculate the three-dimensional point cloud position corresponding to each image pixel; specifically, step S2.2 includes the following sub-steps:
[0072] S2.2-1. When generating a two-dimensional image, obtain the intrinsic parameter matrix and extrinsic parameter matrix of the virtual camera, establish the corresponding relationship between the point cloud data and the image, and the intrinsic parameter matrix of the virtual camera maps the three-dimensional points in the camera coordinate system to the two-dimensional image coordinate system; the intrinsic parameter matrix of the camera is:
[0073]
[0074] Among them, f x and c y Respectively represent the focal length of the camera in the x and y directions, c x and c y are the coordinates of the principal point;
[0075] S2.2-2. Use the camera's extrinsic matrix E to transform the three-dimensional point from the world coordinate system to the camera coordinate system. The extrinsic matrix consists of a rotation matrix and a translation vector:
[0076]
[0077] Wherein, R is the rotation matrix and t is the translation vector. S2.3, mapping the pixels of each image to the three-dimensional point cloud, generating the corresponding relationship between the image and the point cloud, specifically, step S2.3 includes the following sub-steps:
[0078] S2.3-1. By using the intrinsic parameter matrix K and the extrinsic parameter matrix E, as well as the captured image and depth information, the position of each pixel in the three-dimensional space is calculated. For any pixel (u, v) in the image and its corresponding depth value D(u, v), the three-dimensional point (X) in the camera coordinate system is calculated by the following formula: cam ,Y cam ,Z cam )
[0079]
[0080] Z cam =D(u,v)(5)
[0081] The three-dimensional point P in the camera coordinate system cam It is expressed as:
[0082]
[0083] S2.3-2 After generating the point cloud, a voxel grid is constructed to aggregate points from different viewpoints by converting the three-dimensional points in the camera coordinate system into the three-dimensional points in the world coordinate system as shown in the following formula;
[0084] P world =E -1 ·P cam (7)
[0085] Among them, P world is a three-dimensional point in the world coordinate system;
[0086] S2.3-3. Based on the three-dimensional point cloud data generated by the image, a voxel grid method is used to aggregate multiple three-dimensional points from different perspectives into corresponding voxels. The category of the voxel is determined according to the category with the most points in the voxel. The voxel grid helps to resolve the conflict problem when multiple points from different perspectives correspond to the same voxel. In this case, a majority voting mechanism is used to assign voxel labels to ensure that the category with the most points in the voxel is selected. Figure 2 As shown, Figure 2 This is an example diagram of the present invention using majority voting to assign semantic labels to point clouds; in the figure, (a) is a schematic diagram of the consistent classification of voxel p, and (b) is an example diagram of the majority voting of voxel Q. As shown in the figure, four pixels from different images (AD) correspond to voxel P or voxel Q, respectively. For voxel P, the four pixels from different images are all classified as category 1. Therefore, the category of voxel P is updated to category 1. On the other hand, for voxel Q, two pixels are classified as category 2, and the other two pixels are classified as category 1 and category 3, respectively. In this case, the category of voxel Q is updated to the majority category, i.e., category 2.
[0087] S3. Update the semantic label according to the two-dimensional semantic segmentation result obtained in step S2 and the three-dimensional semantic segmentation result in step S1, select the category with the highest comprehensive confidence as the final semantic label of the point, and obtain the final semantic segmentation result. The present invention uses a point-based deep learning algorithm (such as the PointNet series) to semantically classify each point to obtain the classification probability of each point; uses the YOLO model to detect objects in the image and generate a confidence score for each object, and projects these confidence scores into the three-dimensional point cloud; calculates the final classification confidence of each point in the point cloud by combining the confidence of the YOLO model with the classification probability of the PointNet series model; selects the category with the largest classification probability as the final label of the point according to the final confidence value,
[0088] The embodiment provides a method for directly generating images from a three-dimensional point cloud, avoiding the need for external 2D-3D data registration and the need for two-dimensional images that lack registration of corresponding point clouds. The method uses a virtual camera to capture multiple images from a point cloud scene and records the corresponding depth and position information. Specifically, Open3D is used to visualize the entire scene, and images are generated at different positions and angles by a virtual camera. During the generation of each image, depth data is collected, and the camera's internal and external parameters are recorded to establish a corresponding relationship between the point cloud and the image. Specifically, the following sub-steps are included:
[0089] S3.1. For each point λ in the point cloud, the point-by-point classification probability obtained by the PointNet model is expressed as:
[0090] P PointNet family (λ)=[p1(λ), p2(λ),…,p n (λ)] (8)
[0091] Among them, p i (λ) is the predicted probability that point λ belongs to the i-th class, and the sum of all probabilities is 1, that is:
[0092] S3.2, YOLO prediction results, when projected into a 3D point cloud, provide a confidence score for each predicted object category. For a point λ in 3D space corresponding to a YOLO 2D detection result, YOLO is based on the confidence score P YOLO (λ) is expressed as follows:
[0093] P YOLO (λ)=[q1(λ), q2(λ),…,q n (λ)] (9)
[0094] Among them, q i (λ) is the confidence score that point λ belongs to the i-th class based on YOLO prediction;
[0095] S3.3. For each point λ in the point cloud, its final category confidence is obtained by adding the confidence values of each category of the PointNet series model and YOLO, as shown in the following formula:
[0096] P combined (λ)=P PointNet family (λ)+P YOLO (λ) (10)
[0097] S3.4, determine the label of point λ, first in P PointNetfamily (λ) and P YOLO (λ) to find the category with the largest confidence: PointNetfamilyThe category with the largest confidence in (λ) is:
[0098] C PointNet family (λ)=argmax(P PointNet family (λ)) (11)
[0099] In P combined The category with the largest confidence in (λ) is:
[0100] C combined (λ)=argmax(P combined (λ)) (12)
[0101] S3.5. The decision to update the semantic label of point λ is based on the comparison of the confidence scores of the two categories in step S3.4. The specific update rules are as follows:
[0102] P combined (λ,C combined (λ))and P PointNet family (λ,C PointNet family (λ)) (13)
[0103] If the confidence of the fused label is greater than the confidence of the original PointNet series label, the label is updated to C combined (λ); otherwise, keep the original label C PointNet family (λ),
[0104] The formula for label update is:
[0105]
[0106] S4. Based on the final semantic segmentation results, a building information model (BIM) is constructed to reconstruct the building interior scene.
[0107] Then the final scene reconstruction result.
[0108] The method of the present invention is applied to point cloud reconstruction of building interior scenes, which can significantly improve the accuracy of point cloud semantic segmentation, especially in the processing of point cloud data with similar or sparse shapes, providing higher accuracy and efficiency than existing methods.
[0109] This example first selects PointNet and PointNet++ as the benchmark algorithms for experimental verification. Subsequently, PointNext and PointVector are used as comparison algorithms to further evaluate the performance of the proposed algorithm. In the initial experiments, deep learning algorithms (such as PointNet and PointNet++) were used, and subsequent Point-YOLO algorithm tests were performed. Specifically, PointNet-YOLO refers to the combination of the YOLO algorithm and the PointNet algorithm, and PointNet++-YOLO is the combination of the YOLO algorithm and the PointNet++ algorithm.
[0110] During the image annotation and training process, 415 images were randomly selected from the training set of scene 1, and 1 to 3 images were selected for each room. These images were annotated and covered eight categories: beams, columns, windows, doors, tables, chairs, sofas, and plates. Since these categories are relatively rare in the S3DIS dataset, they are more likely to be misclassified as more common categories such as floors, walls, and ceilings. In scene 2, 90 images were randomly selected from the training set of each area for shooting, and these images included seven categories: air ducts, cable trays, fire extinguishers, lamps, columns, pipes, and doors. Similar to scene 1, these components are also relatively rare in scene 2 and are prone to misclassification. After the image annotation is completed, the YOLOv8 model is used for training. Among the annotated images, 90% of the data is used for training, and the remaining 10% is used as a validation set. The initial training model uses a pre-trained model based on the COCO dataset.
[0111] During the image semantic segmentation process, for scene 1, each room in the test set was photographed using a virtual camera. The virtual camera was set in the center of the room and took an image every 45 degrees to complete a 360-degree panoramic shot. Then, the camera moved 0.5 meters in four directions (front, back, left, and right) each time and took 32 images. In this way, a total of 40 images were taken in each room, effectively covering the entire scene. For scene 2, due to the large venue, the central point method was not sufficient to cover all areas, so Open3D was used for manual shooting, and a total of 106 images were taken across the two test areas.
[0112] The image training process lasts for 200 epochs, while the point cloud training process lasts for 100 epochs. All experiments are performed on a single NVIDIA L40 GPU, and the training parameters use the default settings of the model.
[0113] The present invention uses evaluation indicators such as confusion matrix (CM), overall accuracy (OA), precision (Precision), recall (Recall), F1 value, intersection over union (IoU), mean intersection over union (mIoU) and mean category accuracy (mAcc) to comprehensively evaluate the performance of deep learning algorithms. The confusion matrix is a basic tool for evaluating the performance of classification models. By comparing model predictions with true labels, it is refined into four key parts: true positives, false positives, true negatives and false negatives. mAcc calculates the accuracy of each category and takes the average to ensure that uncommon categories are not masked by frequent categories, thereby comprehensively reflecting the performance of the model. The overall accuracy (OA) represents the proportion of correctly classified samples to the total number of samples. The precision and recall rate reflect the proportion of positive examples correctly predicted by the model to the total predicted positive examples, and the proportion of correctly predicted positive examples to actual positive examples, respectively. The F1 value comprehensively considers the precision and recall rates, balancing the weights of the two. The intersection over union (IoU) measures the overlap between the predicted segmentation and the true label, while the mean intersection over union (mIoU) provides the average of the IoUs for all categories as an evaluation metric for the overall segmentation performance.
[0114] Figure 4 An example of the semantic label update process is shown. In the figure, (a) is a schematic diagram before the semantic information is updated, (b) is a schematic diagram after the semantic information is updated, (c) is a diagram of the semantic update process of point 1, and (d) is a diagram of the semantic update process of point 2. For point 1 from the door, when using PointNet for semantic segmentation, this point may be misclassified as floor because the semantic score of the plate (0.303) is the highest among all categories. When using the PointNet-YOLO method, YOLO will classify the point as door and give a confidence score of 0.65. Therefore, adding the semantic score of PointNet to the confidence score of YOLO, the total score of door is 0.284+0.65=0.934. In this way, this point will be correctly classified as door.
[0115] For point 2 from the wall, PointNet correctly classifies the point as wall, and its semantic score is 0.641, which is the highest. However, YOLO classifies the point as chair with a confidence of 0.51. In this case, the total score of chair is 0.056+0.51=0.566, which is still lower than wall (0.641). Therefore, the final classification of this point is still wall.
[0116] Figure 5 and Figure 6 The confusion matrix comparison results of PointNet++ and PointNet++-YOLO in scene 1 and scene 2 are shown respectively. Figure 5Figure, (a) is a schematic diagram of the PointNet++ confusion matrix results, and (b) is a schematic diagram of the PointNet++-YOLO confusion matrix results. Figure 6 In the figure, (a) is a schematic diagram of the confusion matrix results of PointNet++, and (b) is a schematic diagram of the confusion matrix results of PointNet++-YOLO. The results show that the accuracy of various components of PointNet++-YOLO is higher than that of PointNet++ in both scene 1 and scene 2.
[0117] Another aspect of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the present invention when executing the computer program.
[0118] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of any one of the methods described in the present invention when executed by a processor.
[0119] Although the present invention has been described above with preferred embodiments, it is not intended to limit the present invention. A person skilled in the art of the present invention may make various modifications and improvements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the definition of the claims.
Claims
1. A deep learning reconstruction method for building indoor scenes based on point cloud 2D-3D fusion, characterized in that: The following steps are involved: S1. Use deep learning algorithms to perform semantic segmentation on the original 3D point cloud data of indoor scenes to obtain 3D semantic segmentation results; S2, generate several two-dimensional images from the original three-dimensional point cloud data through a virtual camera, then use the YOLO model to perform target detection and semantic segmentation on the two-dimensional images, generate a mask of a two-dimensional graphic, and then obtain the two-dimensional semantic segmentation result of voxel reconstruction through posture adjustment; S3, updating the semantic label according to the two-dimensional semantic segmentation result obtained in step S2 and the three-dimensional semantic segmentation result in step S1, selecting the category with the highest comprehensive confidence as the final semantic label of the point, and obtaining the final semantic segmentation result; S4. Based on the final semantic segmentation results, a building information model (BIM) is constructed to reconstruct the building interior scene. Then the final scene reconstruction result.
2. The method for deep learning reconstruction of indoor building scenes based on point cloud 2D-3D fusion according to claim 1, characterized in that: Step S2 generates a number of two-dimensional images of the original three-dimensional point cloud data through a virtual camera, including the following sub-steps: S2.1, setting multiple virtual camera positions in the scene of 3D point cloud data, and taking images at different angles; S2.2, record the depth information of each image and the internal and external parameter data of the camera, the depth information is used to calculate the three-dimensional point cloud position corresponding to each image pixel; S2.
3. Map the pixels of each image to a three-dimensional point cloud and generate a correspondence between the image and the point cloud.
3. The method for deep learning reconstruction of indoor building scenes based on point cloud 2D-3D fusion according to claim 1, characterized in that: Step S2.2 includes the following sub-steps: S2.2-1. When generating a two-dimensional image, obtain the intrinsic parameter matrix and extrinsic parameter matrix of the virtual camera, establish the corresponding relationship between the point cloud data and the image, and the intrinsic parameter matrix of the virtual camera maps the three-dimensional points in the camera coordinate system to the two-dimensional image coordinate system; the intrinsic parameter matrix of the camera is: Among them, f x and c y Respectively represent the focal length of the camera in the x and y directions, c x and c y are the coordinates of the principal point; S2.2-2. Use the camera's extrinsic matrix E to transform the three-dimensional point from the world coordinate system to the camera coordinate system. The extrinsic matrix consists of a rotation matrix and a translation vector: Where R is the rotation matrix and t is the translation vector.
4. The method for deep learning reconstruction of indoor building scenes based on point cloud 2D-3D fusion according to claim 1, characterized in that: Step S2.3 includes the following sub-steps: S2.3-1. By using the intrinsic parameter matrix K and the extrinsic parameter matrix E, as well as the captured image and depth information, the position of each pixel in the three-dimensional space is calculated. For any pixel (u, v) in the image and its corresponding depth value D(u, v), the three-dimensional point (X) in the camera coordinate system is calculated by the following formula: cam ,Y cam ,Z cam ), Z cam =D(u,v) (5) The three-dimensional point P in the camera coordinate system cam It is expressed as: S2.3-2 After generating the point cloud, a voxel grid is constructed to aggregate points from different viewpoints by converting the three-dimensional points in the camera coordinate system into three-dimensional points in the world coordinate system as follows; P world =E -1 ·P cam (7) Among them, P world is a three-dimensional point in the world coordinate system; S2.3-3. Based on the three-dimensional point cloud data generated by the image, a voxel grid method is used to aggregate multiple three-dimensional points from different perspectives into corresponding voxels, and the category of the voxel is determined according to the category of the most points in the voxel.
5. The method for deep learning reconstruction of indoor building scenes based on point cloud 2D-3D fusion according to claim 1, characterized in that: Step S3 includes the following sub-steps: S3.
1. For each point λ in the point cloud, the point-by-point classification probability obtained by the PointNet model is expressed as: P PointNet family (λ)=[p1(λ),p2(λ),…,p n (l)] (8) Among them, p i (λ) is the predicted probability that point λ belongs to the i-th class, and the sum of all probabilities is 1, that is: S3.2, YOLO prediction results, when projected into a 3D point cloud, provide a confidence score for each predicted object category. For a point λ in 3D space corresponding to a YOLO 2D detection result, YOLO is based on the confidence score P YOLO (λ) is expressed as follows: P YOLO (λ)=[q1(λ),q2(λ),…,q n (l)] (9) Among them, q i (λ) is the confidence score that point λ belongs to the i-th class based on YOLO prediction; S3.
3. For each point λ in the point cloud, its final category confidence is obtained by adding the confidence values of each category of the PointNet series model and YOLO, as shown in the following formula: P combined (λ)=P PointNet family (λ)+P YOLO (l) (10) S3.4, determine the label of point λ, first in P PointNet family (λ) and P YOLO (λ) to find the category with the largest confidence: In P PointNet family The category with the largest confidence in (λ) is: C PointNet family (λ)=argmax(P PointNet family (l)) (11) In P combined The category with the largest confidence in (λ) is: C combined (λ)=argmax(P combined (l)) (12) S3.5, the decision to update the semantic label of point λ is based on the comparison of the confidence scores of the two categories in step S3.4, The specific update rules are as follows: P combined (λ, C combined (λ))and P PointNet family (λ, C PointNet family (l)) (13) If the confidence of the fused label is greater than the confidence of the original PointNet series label, update the label C combined (λ); otherwise, keep the original label C PointNet family (λ), The formula for label update is:
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 5 when executing the computer program.
7. A computer-readable storage medium having a computer program stored thereon, wherein the computer program is used by a processor to execute the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Scene three-dimensional intelligent reconstruction system and method based on BIM and deep learning
CN115147545A
Online carbon semantic map construction method based on sparse fusion
CN115496900A
Indoor building structure point cloud semantic segmentation method and system based on deep learning
CN117710975A
Substation equipment identification method and system based on two-dimensional image and three-dimensional data
CN118608905A
Cited By
Stacked material mass identification and generation method fusing single-view 3D reconstruction and BIM calibration
CN120374885A
Through-wall and cross-border compliance judgment method based on scene graph and trajectory analysis
CN120913081A
Full-scene real-time energy consumption simulation optimization method and system based on AR (Augmented Reality) equipment
CN120976435A