Scene-level synthetic point cloud enhanced semantic segmentation method and system based on diffusion model
By adopting a scene-level synthetic point cloud enhanced semantic segmentation method based on diffusion model in point cloud data processing, the time-consuming and labor-intensive problem of point cloud data acquisition and semantic segmentation in the existing technology is solved, and efficient and accurate three-dimensional model generation is achieved, supporting the effective management and maintenance of underground garages.
Patent Information
- Application Number
- CN202411963037.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The prior art has time-consuming and labor-intensive problems in the acquisition and semantic segmentation of point cloud data, and it is difficult to effectively improve the semantic segmentation accuracy and efficiency of point cloud data in underground garages.
Using a scene-level synthetic point cloud enhancement semantic segmentation method based on diffusion model, a high-fidelity panoramic map and depth map are generated by constructing and training a dual-branch diffusion model and a depth prediction model based on vision transformer, and then synthesize three-dimensional point cloud data, and use a deep learning semantic segmentation network for model training and segmentation.
It significantly improves the semantic segmentation accuracy and efficiency of point cloud data in underground garages, reduces manual intervention and artificial operation time, can quickly and efficiently process large-scale point cloud data, generate accurate three-dimensional models, and support the construction, monitoring and maintenance of underground garages.
Smart Images

Figure CN119992082A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of civil engineering and artificial intelligence interaction technology, and specifically relates to a semantic segmentation method and system for scene-level synthetic point cloud enhancement based on a diffusion model. Background Art
[0002] The reconstruction of 3D models based on point clouds can obtain the BIM model (Building Information Model) of the building, which can effectively promote the operation and maintenance of the building. The first step in 3D model reconstruction is semantic segmentation. Deep learning semantic segmentation significantly reduces the reliance on manually designed features and achieves better performance by autonomously learning and extracting complex data features from a large amount of labeled training data. However, deep learning usually requires a large amount of sample data for training. Compared with images, the collection of point cloud data relies on multi-station scanning laser scanners, which is very time-consuming and labor-intensive. Point cloud generation methods are urgently needed. Summary of the invention
[0003] In view of the problems existing in the prior art, the present invention provides a semantic segmentation method and system for scene-level synthetic point cloud enhancement based on a diffusion model, which can improve the semantic segmentation accuracy and efficiency of underground garage point cloud data.
[0004] In order to solve the above technical problems, the present invention provides the following technical solution: a semantic segmentation method for scene-level synthetic point cloud enhancement based on a diffusion model, comprising the following steps:
[0005] S1. Scanning a real point cloud scene to be reconstructed to obtain a real point cloud and preprocessing the real point cloud, wherein the real three-dimensional point cloud includes XYZ coordinates and RGB color information of each point;
[0006] S2. Construct and train a diffusion model, specifically: a dual-branch method based on a panoramic branch and a perspective branch, with the panoramic branch as the input of the diffusion model; with the noisy panoramic latent map as input, the perspective branch generates a noisy perspective latent map corresponding to different parts of the scene, and the two branches use the equidistant perspective projection attention EPPA mechanism to achieve feature fusion of the panoramic image and the perspective image, and use the low-rank adaptation LoRA layer to adaptively adjust the model weights. At the same time, in the two branches, a U-Net structure with jump connections is used, and the diffusion model finally outputs a high-fidelity panoramic image;
[0007] S3, using the diffusion model to obtain a high-fidelity panorama, performing semantic segmentation and annotation on the panorama, and generating a semantic label of the panorama;
[0008] S4. Build and train a depth prediction model based on the visual transformer ViT, specifically: take the panorama as input and the corresponding depth map as output, and train to obtain the depth prediction model; during the training process, the input panorama is divided into small grids, each small block is flattened and linearly transformed into an embedding through the learned weight matrix and bias parameters, and the embedded tags are processed by the self-attention mechanism. After several self-attention layers, the final depth prediction is mapped to the output depth map of the transformer through the regression head;
[0009] S5, using the depth prediction model to obtain a depth map, and using a cylindrical projection method to map the depth map and semantic labels onto the panoramic image to generate synthetic 3D point cloud data including spatial coordinates, RGB color values, and semantic labels;
[0010] S6. Add the synthetic 3D point cloud to the real 3D point cloud to obtain hybrid point cloud data, and construct a deep learning semantic segmentation network at the same time. Train the deep learning semantic segmentation network with the hybrid point cloud data as input and the semantic segmentation result as output to obtain a deep learning semantic segmentation model. Use the deep learning semantic segmentation model to segment the real point cloud scene to be reconstructed, and build a BIM model based on the segmentation result.
[0011] Furthermore, in the aforementioned step S2, the diffusion model is trained by minimizing the following loss function L:
[0012] L=Eε (x) ,t,∈,y[||∈-∈ θ (z t , t, τ(y))||2] (1)
[0013] Among them, Eε (x) ,t,∈,y respectively represent the expectation of input image x, time step t, noise ∈ and conditional variable y,∈ θ (z t , t, τ(y)) is a denoising function that predicts the noise in the latent space through the U-Net neural network, [||∈-∈ θ (z t , t, τ(y))||2 represents the L2 norm, which is used to calculate the prediction noise ∈ θ (z t , t, the Euclidean distance between τ(y) and the actual noise ∈.
[0014] Furthermore, in the aforementioned step S3, the Roboflow platform is used for semantic annotation, using an intelligent annotation method based on the SegmentAnything Model.
[0015] Furthermore, the aforementioned step S4 includes the following sub-steps:
[0016] S4.1. Divide the input image I into a small grid I_patch. Each small grid is flattened and linearly transformed into an embedding Em through the learned weight matrix and bias parameters, as shown below:
[0017] Em=Flatten(I _patch )·W E +b E (2)
[0018] Among them, W E and b E are the learned embedding weights and bias parameters respectively
[0019] S4.2, the embedded tags are processed through the self-attention mechanism to evaluate the correlation between each small grid and other small grids, and dynamically adjust the correlation of each small grid to capture local and global spatial dependencies;
[0020] S4.3. After several self-attention layers, the final depth prediction is mapped to the output depth map of the transformer through the regression head:
[0021] D=Em·W D +b D (3)
[0022] Where D is the predicted depth map, W D and b D are the learned parameters in the regression head.
[0023] Furthermore, the aforementioned step S5 includes the following sub-steps:
[0024] S5.1. Calculate the spherical coordinates of each pixel (u, v) of the panoramic image of size H×W, as follows:
[0025]
[0026] Among them, θ is the azimuth angle and φ is the polar angle;
[0027] S5.2. Calculate the three-dimensional coordinates (X, Y, Z) of each pixel in the world coordinate system using the following formula:
[0028] X u,v =D u,v ·cos(φ)·cos(θ) (5)
[0029] Y u,v =D u,v ·cos(φ)·sin(θ) (6)
[0030] Z u,v =D u,v·sin(φ) (7)
[0031] Among them, D u,v is the depth value of the pixel;
[0032] S5.3. Each pixel (u, v) of the panoramic image I is assigned a semantic label L according to the predefined scene category. u,v ; Label L u,v Associated with the corresponding 3D point, each point in the generated synthetic point cloud has a corresponding label. Finally, the generated point cloud P consists of spatial coordinates (X, Y, Z), RGB color values and semantic labels, which are expressed as follows:
[0033] P=(X u,v , Y u,v , Z u,v , RGB u,v .L u,v )|(u,v)∈I (8)
[0034] Among them, RGB represents the pixel color extracted from the panoramic image I, L u,v Is a semantic tag.
[0035] Furthermore, in the aforementioned step S6, a deep learning semantic segmentation network is constructed based on the PointNet or PointNet++ network, and three key indicators, namely, overall accuracy OA, intersection over union (IoU) and average intersection over union (mIoU), are used to evaluate the deep learning semantic segmentation model.
[0036] Another aspect of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the present invention when executing the computer program.
[0037] The present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of any one of the methods described in the present invention when executed by a processor.
[0038] Compared with the prior art, the beneficial technical effects of the present invention using the above technical solution are as follows:
[0039] By using deep learning methods, the present invention greatly improves the semantic segmentation accuracy and efficiency of underground garage point cloud data, and significantly reduces manual intervention and manual operation time. Using the trained model, large-scale point cloud data can be processed quickly and efficiently to generate accurate three-dimensional models to support the construction, monitoring and maintenance of underground garages. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a flow chart of the method of the present invention.
[0041] Figure 2 It is a flow chart of the double-branch method of the present invention.
[0042] Figure 3 This is an example of annotating a panorama.
[0043] Figure 4 This is an example of panoramic depth prediction. In the figure, (a) is an RGB panoramic image and (b) is the predicted depth map.
[0044] Figure 5 This is an example of the process of generating a synthetic point cloud. DETAILED DESCRIPTION
[0045] In order to better understand the technical content of the present invention, specific embodiments are given and described as follows in conjunction with the accompanying drawings.
[0046] Various aspects of the invention are described herein with reference to the accompanying drawings, in which many illustrative embodiments are shown. The embodiments of the invention are not limited to those described in the accompanying drawings. It should be understood that the invention is implemented by any of the various concepts and embodiments described above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed in the invention are not limited to any implementation. In addition, some aspects disclosed in the invention may be used alone or in any appropriate combination with other aspects disclosed in the invention.
[0047] refer to Figure 1 The present invention provides a semantic segmentation method for scene-level synthetic point cloud enhancement based on a diffusion model, comprising the following steps:
[0048] S1. Scan the real point cloud scene to be reconstructed to obtain the real point cloud and perform preprocessing. The real three-dimensional point cloud includes the XYZ coordinates and RGB color information of each point.
[0049] S2. Construct and train a diffusion model, specifically: a dual-branch method based on a panoramic branch and a perspective branch, with the panoramic branch as the input of the diffusion model; with the noisy panoramic latent map as input, the perspective branch generates a noisy perspective latent map corresponding to different parts of the scene, and the two branches use the equidistant perspective projection attention (EPPA) mechanism to achieve feature fusion of panoramic and perspective images, and use a low-rank adaptation LoRA layer to adaptively adjust the model weights. At the same time, in the two branches, a U-Net structure with jump connections is used, and the diffusion model finally outputs a high-fidelity panoramic image.
[0050] S3, using the diffusion model to obtain a high-fidelity panorama, performing semantic segmentation and annotation on the panorama, and generating a semantic label of the panorama;
[0051] S4. Build and train a depth prediction model based on the visual transformer ViT, specifically: take the panorama as input and the corresponding depth map as output, and train to obtain the depth prediction model; during the training process, the input panorama is divided into small grids, each small block is flattened and linearly transformed into an embedding through the learned weight matrix and bias parameters, and the embedded tags are processed by the self-attention mechanism. After several self-attention layers, the final depth prediction is mapped to the output depth map of the transformer through the regression head;
[0052] S5, using the depth prediction model to obtain a depth map, and using a cylindrical projection method to map the depth map and semantic labels onto the panoramic image to generate synthetic 3D point cloud data including spatial coordinates, RGB color values, and semantic labels;
[0053] S6. Add the synthetic 3D point cloud to the real 3D point cloud to obtain hybrid point cloud data, and construct a deep learning semantic segmentation network at the same time. Train the deep learning semantic segmentation network with the hybrid point cloud data as input and the semantic segmentation result as output to obtain a deep learning semantic segmentation model. Use the deep learning semantic segmentation model to segment the real point cloud scene to be reconstructed, and build a BIM model based on the segmentation result.
[0054] As a preferred embodiment of the present invention, step S1 uses a laser scanner such as Leica RTC360 to obtain point cloud data in the building interior. The collected point cloud data is preprocessed, including noise point removal and abnormal point removal. In addition, the point cloud data is downsampled, and the voxel downsampling method is used to simplify the point cloud data to a reasonable scale, while retaining key geometric features to facilitate subsequent semantic segmentation and model reconstruction. The preprocessed point cloud data is stored in a standard format such as .pcd or .las for subsequent processing.
[0055] As a preferred embodiment of the present invention, step S2 performs diffusion model training. In the model training stage, real panoramic images and their corresponding text descriptions are used. This study uses the Matterport3D dataset, which contains 10,800 panoramic images and 2,295 room layout annotations. However, the original Matterport3D dataset does not contain text descriptions. To solve this problem, the BLIP-2 model is used, which is an advanced visual-language model for generating image captions and understanding visual-language relationships. Each image begins with a prompt of "a 360-degree view of" to generate a concise text description related to the panoramic image for text conditional generation.
[0056] For example, for the generation of panoramic images of building interiors, descriptive languages such as "An office with awindow and a bookcase", "A study room with a bright window" can be used to generate synthetic point cloud scenes.
[0057] In step S2, the goal of the diffusion model is to reconstruct a high-fidelity image by gradually optimizing the noise latent variables. The diffusion model is trained by minimizing the following loss function L:
[0058] L=Eε (x) ,t,∈,y[||∈-∈ θ (z t , t, τ(y))||2] (1)
[0059] Among them, Eε (x) ,t,∈,y respectively represent the expectation of input image x, time step t, noise ∈ and conditional variable y,∈ θ (z, t, τ(y)) is a denoising function that uses a U-Net neural network to predict the noise in the latent space. This function will be based on the latent variable z t and conditional encoding τ(y) to predict noise. Through this denoising process, the model can gradually correct z in the latent space t , and finally reconstruct the image. θ (z t , t, τ(y))||2 represents the L2 norm, which is used to calculate the prediction noise ∈ θ (z t , t,τ(y) and the Euclidean distance between the actual noise ∈. The closer the value is to zero, the better the denoising effect of the model.
[0060] Traditional diffusion-based image generation methods are limited to generating regular images and cannot generate panoramic images. To address this limitation, the present invention adopts a dual-branch approach, including a panoramic branch and a perspective branch, such as Figure 2 The panoramic branch is responsible for generating panoramic views, while the perspective branch generates perspective projections to facilitate cross-format information exchange.
[0061] In the model architecture, the generation process starts with a noisy panoramic latent map, which is initially created based on the input text prompt. This latent map serves as the basis for generating panoramic views, capturing a wide field of view. Meanwhile, the perspective branch generates noisy perspective latent maps corresponding to different parts of the scene, which are then projected onto the panorama to create a seamlessly fused panoramic representation.
[0062] To facilitate the integration of information between the two branches, the Equidistant Perspective Projection Attention (EPPA) mechanism is adopted. EPPA achieves efficient cross-branch attention by aligning and merging the features of the panorama and perspective latent maps, thereby enhancing spatial consistency and providing a unified scene representation. In addition, a Low Rank Adaptation (LoRA) layer is used to adaptively adjust the weights of the model to optimize its performance in the panorama generation task.
[0063] Within each branch, a U-Net structure with skip connections is used to ensure high-resolution output and preserve details, ultimately generating high-quality panoramic images. By leveraging these design choices, the dual-branch approach achieves panoramic image generation that maintains spatial consistency and visual fidelity across the entire field of view.
[0064] As a preferred embodiment of the present invention, step S3 performs semantic segmentation and annotation on the panoramic image. Since the point cloud is composed of a set of points in three-dimensional space, directly annotating these points may take a lot of time and manpower. One advantage of generating a point cloud through a panoramic image is that semantic annotation can be performed on the image and then these annotations can be projected onto the three-dimensional point cloud. Compared with three-dimensional point clouds, the annotation of two-dimensional images is relatively simple and has more mature technical support. The present invention performs annotation on the Roboflow image annotation platform, which adopts an intelligent annotation method based on the Segment Anything Model (SAM). The method allows the user to automatically select similar pixels near the manually clicked area, thereby significantly improving the annotation efficiency. In addition, only the required semantic categories need to be annotated, and any unannotated areas will be automatically set to the "cluttered" category. In addition, the algorithm can also be customized according to the annotation preference. For example, if there is a window embedded in the wall, and the pixels in the window area are labeled as both wall and window categories, a custom algorithm can be designed to retain only the window category. Annotating the panoramic image in this way significantly improves the efficiency of point cloud annotation, Figure 3 Shows an example of completed annotation.
[0065] Since the synthetic image generated by the diffusion model lacks depth information, the present invention uses a deep learning method, the visual transformer (ViT), to directly predict depth from the image data. As a preferred embodiment of the present invention, step S4 includes the following sub-steps:
[0066] S4.1. The depth prediction process of ViT starts by dividing the input image into small patches, each of which represents a local area in the image. By converting these patches into labels, the model is able to process the entire image while preserving the spatial structure. Specifically, the input image I is divided into a small patch grid I_patch, each of which is flattened and linearly transformed into an embedding Em through the learned weight matrix and bias parameters, as follows
[0067] Em=Flatten(I _patch )·W E +b E (2)
[0068] Among them, W E and b E are the learned embedding weights and bias parameters respectively
[0069] S4.2, the embedded tags are processed through the self-attention mechanism to evaluate the correlation between each small grid and other small grids, and dynamically adjust the correlation of each small grid to capture local and global spatial dependencies;
[0070] S4.3. After several self-attention layers, the final depth prediction is mapped to the output depth map of the transformer through the regression head:
[0071] D=Em·W D +b D (3)
[0072] Where D is the predicted depth map, W D and b D are the learned parameters in the regression head.
[0073] The depth prediction model based on the visual transformer ViT is built using the Omnidata pre-trained model. The Omnidata pre-trained model has been trained on a wide range of depth prediction tasks on various datasets, which enhances the generalization ability of the model. In addition, the predicted depth values are adjusted. For example, in indoor panoramic scenes, the maximum depth range is limited to the typical depth range of the indoor environment to avoid the predicted value exceeding the actual range. An example of depth prediction is Figure 4 As shown in the figure, (a) shows the RGB panoramic image and (b) shows the predicted depth map.
[0074] As a preferred embodiment of the present invention, step S5 includes the following sub-steps:
[0075] S5.1. After obtaining depth information and semantic labels from the panoramic image, this study uses the spherical projection method to generate point cloud scenes. For a panoramic image I of size H×W and its corresponding depth map D, the two-dimensional pixel values are converted into three-dimensional coordinates by spherical projection. For each pixel (u, v), where u and v represent the horizontal and vertical indices respectively, the spherical coordinates are calculated as follows:
[0076]
[0077] Among them, θ is the azimuth angle and φ is the polar angle;
[0078] S5.2. Calculate the three-dimensional coordinates (X, Y, Z) of each pixel in the world coordinate system using the following formula:
[0079] X u,v =D u,v·cos (φ)·cos(θ) (5)
[0080] Y u,v =D u,v ·cos(φ)·sin(θ) (6)
[0081] Z u,v =D u,v ·sin(φ) (7)
[0082] Among them, D u,v is the depth value of the pixel;
[0083] S5.3. Each pixel (u, v) of the panoramic image I is assigned a semantic label L according to the predefined scene category. u,v ; Label L u,v Associated with the corresponding 3D point, each point in the generated synthetic point cloud has a corresponding label. Finally, the generated point cloud P consists of spatial coordinates (X, Y, Z), RGB color values and semantic labels, which are expressed as follows:
[0084] P=(X u,v , Y u,v , Z u,v , RGB u,v .L u,v )|(u,v)∈I (8)
[0085] Among them, RGB represents the pixel color extracted from the panoramic image I, L u,v Is a semantic tag.
[0086] Through this method, the present invention is able to convert the generated panoramic image into an accurate three-dimensional point cloud and assign the correct semantic category to each point.
[0087] An example of projecting a panorama into a 3D point cloud, from text description to generating a panorama, and finally generating a synthetic point cloud is as follows: Figure 5 shown.
[0088] Step S6 specifically adds the generated synthetic point cloud to the dataset of the real scene for training. The present invention selects PointNet++ as the main model because its architecture is fundamental and widely used. In addition, other deep learning models (such as PointNet and PointNeXt) are also used to further evaluate the effectiveness of the proposed method. In order to evaluate the performance of semantic segmentation based on deep learning, this study uses the following three key indicators: overall accuracy (OA), intersection over union (IoU), and mean intersection over union (mIoU).
[0089] Model optimization and evaluation. After training, the validation set is used to tune the model, adjust the network structure and hyperparameters, and further optimize the model performance. The best model configuration is selected through cross-validation and multiple experiments. The evaluation process includes calculating the mean intersection over union (mIoU) and overall accuracy (OA) of the model on the test set to ensure that the accuracy and robustness of the model meet the needs of practical applications. The optimized model is used for point cloud semantic segmentation tasks in actual scenes. Actual scene application. The trained and optimized deep learning model is applied to the actual scanning point cloud data processing. First, the point cloud data of the new scene is collected using a laser scanner, and the same preprocessing as step S1 is performed. Then, the preprocessed point cloud data is input into the trained model for automatic semantic segmentation, identification and classification of various objects in the point cloud scene, and the segmentation results are used to reconstruct the 3D model in BIM software such as Revit, and generate a 3D digital twin model of the point cloud scene, which can be used for subsequent facility management, monitoring and maintenance work.
[0090] The laboratory conditions required for training include: high-performance GPU (such as NVIDIA RTX 4090), running Linux system, using Python programming language and Pytorch deep learning framework.
[0091] Another aspect of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods described in the present invention when executing the computer program.
[0092] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any method described in the present invention.
[0093] Although the present invention has been described above with preferred embodiments, it is not intended to limit the present invention. A person skilled in the art of the present invention may make various modifications and improvements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the definition of the claims.
Claims
1. A semantic segmentation method for scene-level synthetic point cloud enhancement based on diffusion model, characterized in that: The following steps are involved: S1. Scanning a real point cloud scene to be reconstructed to obtain a real point cloud and preprocessing the real point cloud, wherein the real three-dimensional point cloud includes XYZ coordinates and RGB color information of each point; S2. Construct and train a diffusion model, specifically: a dual-branch method based on a panoramic branch and a perspective branch, with the panoramic branch as the input of the diffusion model; with the noisy panoramic latent map as input, the perspective branch generates a noisy perspective latent map corresponding to different parts of the scene, and the two branches use the equidistant perspective projection attention EPPA mechanism to achieve feature fusion of the panoramic image and the perspective image, and use the low-rank adaptation LoRA layer to adaptively adjust the model weights. At the same time, in the two branches, a U-Net structure with jump connections is used, and the diffusion model finally outputs a high-fidelity panoramic image; S3, using the diffusion model to obtain a high-fidelity panorama, performing semantic segmentation and annotation on the panorama, and generating a semantic label of the panorama; S4. Build and train a depth prediction model based on the visual transformer ViT, specifically: take the panorama as input and the corresponding depth map as output, and train to obtain the depth prediction model; during the training process, the input panorama is divided into small grids, each small block is flattened and linearly transformed into an embedding through the learned weight matrix and bias parameters, and the embedded tags are processed by the self-attention mechanism. After several self-attention layers, the final depth prediction is mapped to the output depth map of the transformer through the regression head; S5, using the depth prediction model to obtain a depth map, and using a cylindrical projection method to map the depth map and semantic labels onto the panoramic image to generate synthetic 3D point cloud data including spatial coordinates, RGB color values, and semantic labels; S6. Add the synthetic 3D point cloud to the real 3D point cloud to obtain hybrid point cloud data, and construct a deep learning semantic segmentation network at the same time. Train the deep learning semantic segmentation network with the hybrid point cloud data as input and the semantic segmentation result as output to obtain a deep learning semantic segmentation model. Use the deep learning semantic segmentation model to segment the real point cloud scene to be reconstructed, and build a BIM model based on the segmentation result.
2. The semantic segmentation method for scene-level synthetic point cloud enhancement based on diffusion model according to claim 1, characterized in that: In step S2, the diffusion model is trained by minimizing the following loss function L: L=Eε (x) ,t,∈,y[||∈-∈ θ (z t ,t,τ(y))||2] (1) Among them, Eε (x) ,t,∈,y respectively represent the expectation of input image x, time step t, noise ∈ and conditional variable y,∈ θ (z t , t, τ(y)) is a denoising function that predicts the noise in the latent space through the U-Net neural network, [||∈-∈ θ (z t , t, τ(y))||2 represents the L2 norm, which is used to calculate the prediction noise ∈ θ (z t , t, the Euclidean distance between τ(y) and the actual noise ∈.
3. The semantic segmentation method for scene-level synthetic point cloud enhancement based on diffusion model according to claim 1, characterized in that: In step S3, the Roboflow platform is used for semantic annotation, using an intelligent annotation method based on the Segment AnythingModel.
4. The semantic segmentation method for scene-level synthetic point cloud enhancement based on diffusion model according to claim 1, characterized in that: Step S4 includes the following sub-steps: S4.
1. Divide the input image I into a small grid I_patch. Each small grid is flattened and linearly transformed into an embedding Em through the learned weight matrix and bias parameters, as shown below: Em=Flatten(I _patch )·W E +b E (2) Among them, W E and b E are the learned embedding weights and bias parameters respectively S4.2, the embedded tags are processed through the self-attention mechanism to evaluate the correlation between each small grid and other small grids, and dynamically adjust the correlation of each small grid to capture local and global spatial dependencies; S4.
3. After several self-attention layers, the final depth prediction is mapped to the output depth map of the transformer through the regression head: D=Em·W D +b D (3) Where D is the predicted depth map, W D and b D are the learned parameters in the regression head.
5. The semantic segmentation method for scene-level synthetic point cloud enhancement based on diffusion model according to claim 1, characterized in that: Step S5 includes the following sub-steps: S5.
1. Calculate the spherical coordinates of each pixel (u, v) of the panoramic image of size H×W, as follows: Among them, θ is the azimuth angle and φ is the polar angle; S5.
2. Calculate the three-dimensional coordinates (X, Y, Z) of each pixel in the world coordinate system using the following formula: X u,v =D u,v ·cos (φ)·cos (θ) (5) Y u,v =D u,v ·cos (φ)·sin (θ) (6) WITH u,v =D u,v sin (φ) (7) Among them, D u,v is the depth value of the pixel; S5.
3. Each pixel (u, v) of the panoramic image I is assigned a semantic label L according to the predefined scene category. u,v ; Label L u,v Associated with the corresponding 3D point, each point in the generated synthetic point cloud has a corresponding label. Finally, the generated point cloud P consists of spatial coordinates (X, Y, Z), RGB color values and semantic labels, which are expressed as follows: P=(X u,v , Y u,v , Z u,v ,RGB u,v . L u,v )∣(u,v)∈I (8) Among them, RGB represents the pixel color extracted from the panoramic image I, L u,v Is a semantic tag.
6. The semantic segmentation method for scene-level synthetic point cloud enhancement based on diffusion model according to claim 1, characterized in that: In step S6z, a deep learning semantic segmentation network is constructed based on the PointNet or PointNet++ network, and three key indicators, namely, overall accuracy OA, intersection over union (IoU), and average intersection over union (mIoU), are used to evaluate the deep learning semantic segmentation model.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Image semantic segmentation method and device, equipment and storage medium
CN114004972A
Indoor building structure point cloud semantic segmentation method and system based on deep learning
CN117710975A
Point cloud semantic segmentation method based on global feature enhancement
CN118247511A
High-precision point cloud completion method based on deep learning and device thereof
US20230206603A1
Systems and methods for generating a road surface semantic segmentation map from a sequence of point clouds
US20230267615A1
Cited By
Indoor scene reconstruction method and device, storage medium and equipment
CN120182509A
Method, device, storage medium and equipment for indoor scene reconstruction
CN120182509B
Building identification method and system based on three-dimensional point cloud
CN120526319A
Building recognition method and system based on three-dimensional point cloud
CN120526319B
Wide-viewing-angle dynamic scene reconstruction system and method
CN120599143A