Point cloud semantic segmentation model and method for object and indoor scene

By introducing the channel space attention module and edge information perception module into the point cloud semantic segmentation model, the point cloud segmentation accuracy is optimized using multi-scale and edge information, and the problem of difficult or unsmooth segmentation of objects in the prior art is solved, and more efficient and accurate point cloud segmentation is achieved.

CN119963834AActive Publication Date: 2025-05-09CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510034778.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-09
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The prior art is difficult to fully consider multi-scale and edge information in point cloud semantic segmentation, resulting in difficult or unsmoothing edges of some objects.

Method used

A point cloud semantic segmentation model is proposed. By introducing a channel space attention module and edge information perception module, the point cloud segmentation accuracy is optimized using space channel information and edge information. The channel space attention module combines channel and spatial attention to extract more comprehensive point cloud features, while the edge information perception module integrates and reconstructs features through dynamic weighting to enhance the model's ability to capture edge information.

Benefits of technology

Improve point cloud segmentation accuracy, especially when dealing with complex scenes and edge details, enhance the performance and generalization capabilities of the model, and reduce the sensitivity to noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963834A_ABST
    Figure CN119963834A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud semantic segmentation model for an object and an indoor scene, and the model comprises an obtaining module which is used for obtaining target point cloud data; the coding module is used for carrying out down-sampling on the point cloud data, introducing a channel space attention module and extracting point cloud data features by utilizing channel and space information at the same time; and the decoding module is used for carrying out up-sampling on the point cloud data features processed by the coding module, introducing an edge information sensing module, and fusing and reconstructing the features in a dynamic weighting mode. By improving the feature extraction capability when the model processes the point cloud data, the understanding and processing of the model on the input data are optimized, and the accuracy of point cloud data processing is improved. Therefore, the efficiency of extracting valuable features from the point cloud data is improved; the method is beneficial for leaving edge information of tasks, meanwhile, irrelevant edge information is restrained, sensitivity of the model to noise is reduced, original point cloud resolution is better recovered, and model performance and generalization ability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and deep learning technology, and in particular to a point cloud semantic segmentation model and method for objects and indoor scenes. Background Art

[0002] With the rapid development of autonomous driving, augmented reality and medical image analysis, people are paying more and more attention to 3D scene segmentation technology. Point cloud, as a common form of 3D representation, is widely used in 3D computer vision research. Compared with traditional 2D image segmentation, 3D point cloud segmentation can handle more complex scenes, understand scenes more comprehensively, provide richer geometric and spatial information, and better handle occlusion problems. Traditional point cloud segmentation methods, such as model fitting-based methods, region growing-based methods, edge-based methods, etc., mostly rely on manually designed geometric features and are limited by prior knowledge, resulting in unstable segmentation results. In order to have a deeper understanding of 3D scenes, researchers have proposed point cloud semantic segmentation methods based on machine learning, such as support vector machines and conditional random fields. Although these methods have improved the accuracy and efficiency of segmentation to a certain extent, they still have limitations in the face of complex scenes. As deep neural networks have greatly promoted the advancement of computer vision technology, more and more studies have begun to use deep learning methods to achieve point cloud semantic segmentation.

[0003] However, 3D deep learning methods are still challenging because point clouds are unevenly distributed and unstructured in 3D space, which makes it impossible to use convolutional neural networks directly on point clouds. Early studies would convert point clouds into regular structures, such as multi-view projection-based methods and voxel-based methods, but these methods would cause geometric information loss and high computational storage overhead. In order to solve these problems, technicians in this field have proposed Pointnet, the first deep learning framework directly on unstructured point clouds. It uses a shared multi-layer perceptron model (MLP) to learn point features and a maximum pooling function to directly process point clouds. Pointnet++ proposes a hierarchical downsampling strategy to capture local geometric detail information and multi-scale grouping, which significantly improves the accuracy of point cloud segmentation. It also proposes a PointSIFT network to stack and encode information from 8 key spatial directions, so that the model can capture information in different directions and achieve multi-scale representation. It also proposes a RandLA-Net network, which first introduces a local spatial encoding unit to compensate for the local structural information lost by random sampling, and then uses attention pooling learning and aggregating neighborhood point features to strengthen important information, and uses dilated residual blocks to increase the receptive field of each point to extract richer features. It also proposes DGCNN to capture the distance information between each point and its neighboring points through EdgeConv to more effectively extract local features. In order to capture the local contextual features of point cloud data, recent research has explored the direct application of explicit convolution kernels in point space. Among them, KPconv uses deformable convolution to capture local information more accurately; PointNeXt is also proposed to re-examine Pointnet++, fully explore its potential by improving the training strategy, and capture more detailed local information to improve segmentation accuracy. Although most of the above methods have advantages in local feature learning, they lack the ability to capture long-range correlations, cannot understand the global context well, and are difficult to adapt to complex scenes.

[0004] With the revolutionary progress of Transformer in the field of natural language processing, many recent studies have also applied the attention mechanism to the analysis of point clouds. The use of the attention mechanism enables the model to dynamically focus on important features, effectively overcoming the disorder and irregularity of point cloud data. PointASNL proposed by technicians in this field adopts adaptive sampling block AS and local-non-local L-NL block to capture the long-range correlation of sampling points, and more effectively processes noisy point cloud data; it also proposes an attention-based group relationship aggregator module RPNet to more effectively capture local semantics and positional relationships; it also proposes Point Cloud Transformer (PCT) to add a vector mapping structure and offset attention structure of domain points in the encoder to enhance the model's ability to learn local context features; it also proposes PointTransformer to capture local and global features using a self-attention mechanism to more effectively capture geometric information; it also proposes a hierarchical Transformer (Stratified Transformer) that captures long-range contextual relationships, and the hierarchical strategy more effectively enhances the receptive field; it also proposes a dual-branch Transformer (PointCAT) based on a cross-attention mechanism to enhance the model's ability to capture long-range relationships; it also proposes a self-positioning point Transformer (SPoTr), which uses a set of self-positioning points to calculate attention weights and improve the scalability of global self-attention. Although these works are very effective in point cloud segmentation, most of them fail to fully consider the importance of multi-scale and edge information, and cannot make good use of spatial channel information and edge information to optimize the point cloud segmentation accuracy, resulting in the edges of some objects being difficult to segment or the segmentation being uneven.

[0005] Therefore, technicians in this field are committed to developing a point cloud semantic segmentation model and method for objects and indoor scenes. Summary of the invention

[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is: fully consider the importance of multi-scale information and edge information, use spatial channel information and edge information to optimize the point cloud segmentation accuracy, and solve the problem that the edges of some objects are difficult to segment or the segmentation is not smooth.

[0007] To achieve the above-mentioned object, the present invention provides a point cloud semantic segmentation model for objects and indoor scenes, and an acquisition module for acquiring target point cloud data;

[0008] An encoding module, used for down-sampling the target point cloud data, introducing a channel space attention module and utilizing the channel and spatial information of the channel space attention module to extract features of the target point cloud data;

[0009] A decoding module, used for up-sampling the target point cloud data features processed by the encoding module, introducing an edge information perception module, and fusing and reconstructing the sampled target point cloud data features by a dynamic weighting method;

[0010] Furthermore, the channel space attention module is a combination of channel attention and spatial attention, and the output formula of the channel space attention module is:

[0011] F CSPA =F CA +F SA =CA(x,f)+SA(F CA )

[0012] Among them, F CSPA is the output result of the channel space attention module, F CA is the result of the channel attention output, F SA is the result of the spatial attention output, CA is the channel attention, (x, f) is the three-dimensional coordinates and feature vector of the target point cloud data, and SA is the spatial attention;

[0013] Furthermore, the downsampling adopts farthest point sampling;

[0014] Furthermore, the result of the channel attention output is expressed as:

[0015]

[0016] in, is x q , x k The corresponding point feature, Ω key is the set of key points, A q,k,c is the channel attention weight, φ qk To obtain the query point x through normalization q Calculate the key point x k The relative position encoding, α is a query point feature f q With the key point feature f k Linear transformation, α = f q -f k ;

[0017] Furthermore, the channel attention weight is calculated by calculating the attention weight between the query point and the key point of each channel, the query point is selected as a point obtained by randomly sampling the original point cloud data and enhancing it with a random vector generated by a learnable parameter to form an identifiable point; the key point is a point retained after sampling the farthest point;

[0018] Furthermore, the channel attention weight is expressed as:

[0019]

[0020] Among them, c is the channel index, τ is the smoothness parameter used to control the SoftMax function, γ is the mapping function, and is a multilayer perceptron model consisting of 2 linear layers and ReLu. ψ and is a linear change;

[0021] Furthermore, the spatial attention is applied to the F by average pooling and maximum pooling operations. CA Process again and merge the results of the channel attention and the spatial attention. The result of the spatial attention output is expressed as: SA =SA(F CA )=sigmoid(concat[Avgpool(F CA ),Maxpool(F CA )]);

[0022] Furthermore, the upsampled features are fused in a dynamic weighted manner, feature selection is performed through a nonlinear activation function Relu and a convolutional layer, a weight coefficient is obtained through a Sigmoid activation function, and the features are dynamically weighted using the weight coefficient;

[0023] Furthermore, the edge information perception module is a new module that acts on edge information in the dynamic weighted manner based on a diffusion unit, wherein the diffusion unit is represented by:

[0024]

[0025] in, is a point set, Represented as the center point and the neighboring point, respectively, n , f s Represent the features of the center point and the neighboring points respectively, σ is ReLu, normalization, and convolution operation;

[0026] The result output by the edge information perception module is expressed as follows:

[0027] F EIA =Concat[Sigmoid[σ[Avgpool(F C ' SPA )+F C ' SPA ]],Avgpool(F C ' SPA )]+f

[0028] Among them, F C ' SPA is the F after downsampling CSPA Features, F EIA The result output by the edge information perception module EIA;

[0029] The present invention provides a point cloud semantic segmentation method for objects and indoor scenes, which obtains target point cloud data; inputs the point cloud data into a point cloud semantic segmentation model as described in any one of claims 1 to 9, and outputs the segmented point cloud data.

[0030] The point cloud semantic segmentation model and method for objects and indoor scenes proposed in the present invention improve the feature extraction capability of the model when processing point cloud data by introducing a channel spatial attention module, and optimize the model's understanding and processing of input data, thereby improving the efficiency of extracting valuable features from point cloud data; the introduction of an edge information awareness (EIA) module is beneficial to retaining edge information of the task, while suppressing irrelevant edge information, reducing the model's sensitivity to noise, better restoring the original point cloud resolution, and improving the model's performance and generalization capability. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Schematic diagram of the architecture of the semantic segmentation model of the present invention;

[0032] Figure 2 Schematic diagram of the structure of the channel space attention module of the present invention;

[0033] Figure 3 It is a structural schematic diagram of the edge information perception module of the present invention;

[0034] Figure 4 It is a schematic diagram of visualization results of partial segmentation of ShapeNetPart dataset according to an embodiment of the present invention;

[0035] Figure 5 This is a schematic diagram of the visualization results of the semantic segmentation of the S3DIS dataset according to an embodiment of the present invention. DETAILED DESCRIPTION

[0036] The following describes several preferred embodiments of the present invention with reference to the drawings in the specification, so that the technical content is clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0037] In the drawings, components with the same structure are indicated by the same numerical reference numerals, and components with similar structures or functions are indicated by similar numerical reference numerals. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. In order to make the illustration clearer, the thickness of the components is appropriately exaggerated in some places in the drawings.

[0038] In the field of point cloud segmentation research, this paper proposes a point cloud semantic segmentation model (CSEANet) for objects and indoor scenes, which can effectively handle the complexity of point cloud data; the design of CSEANet adopts the most common U-net encoder-decoder structure in current point cloud semantic segmentation, which effectively realizes hierarchical feature learning. Figure 1 As shown, the architectural schematic diagram of the semantic segmentation model of the present invention includes an acquisition module for acquiring target point cloud data; an encoding module for down-sampling the point cloud data, introducing a channel spatial attention module (CSPA), and extracting features of the point cloud data using channel and spatial information; a decoding module for up-sampling the point cloud data features processed by the encoding module, introducing an edge information perception module (EIA), and fusing and reconstructing the features by dynamic weighting.

[0039] In a specific embodiment, the encoder part is responsible for extracting and abstracting features from the input point cloud. In the encoder, the point cloud data undergoes a series of downsampling operations to reduce the data dimension by downsampling. While ensuring the effective encoding of multi-scale features, a more discriminative feature representation is extracted. The farthest point sampling (FPS) method is used for downsampling to effectively maintain the geometric features and spatial distribution features of the point cloud and reduce information loss. However, due to the structural complexity of the point cloud data, downsampling may lose local detail information, resulting in insufficient feature representation and extraction. In order to improve the integrity of feature extraction and segmentation accuracy, we add a channel spatial attention module (CSPA) after each downsampling, so that the model can understand and utilize feature information more comprehensively and improve the model's ability to represent complex data. The point features processed by the spatial channel module are passed to the next processing stage.

[0040] In the decoder part, the point cloud is gradually restored to its original resolution through multiple upsampling and the edge information perception module. The skip connection mechanism is used to alleviate the problem of lost spatial information, which is conducive to feature fusion and reconstruction. In order to better understand contextual information and adapt to data changes, and improve segmentation accuracy, the edge information perception module (EIA) is introduced. EIA acts on the features after upsampling. This module is conducive to retaining the edge information of the task, while suppressing irrelevant edge information, reducing the model's sensitivity to noise, better restoring the original point cloud resolution, and improving model performance and generalization ability.

[0041] In some embodiments, in order to better utilize spatial channel information to optimize the accuracy of point cloud segmentation, the channel attention CA is combined with the spatial attention SA to propose an efficient global information attention module, called the channel spatial attention module (CSPA), whose structural diagram is shown in Figure 2 As shown. CSPA can simultaneously utilize channel and spatial information to better understand the intrinsic structure and spatial relationship of point cloud data, thereby achieving more comprehensive point cloud feature extraction. The CSPA output formula is shown in formula (1):

[0042] F CSPA =F CA +F SA =CA(x,f)+SA(F CA ) (1)

[0043] Among them, F CSPA is the output of the CSPA module, F CA is the result of channel attention output, F SA It is the result of spatial attention output. CA is the attention on the point based on the channel. (x, f) is the three-dimensional coordinate and feature vector of the point, and SA is the spatial attention.

[0044] It is worth noting that the way CSPA calculates weights between channels is different from standard channel attention. Based on channel attention CA, the attention weights between the query point and the key point of each channel are calculated. The query point is selected as the point after random sampling of the original point cloud, and the random vector generated by the learnable parameters is enhanced to form a recognizable point. The key point is the point retained after the farthest point is sampled. The result of the channel attention CA output is shown in formula (2):

[0045]

[0046] in, is x q , x k The corresponding point feature, Ω key is the set of key points, A q,k,cis the channel attention weight. In order to better integrate local and global context information, the relative position encoding function φ is introduced qk ,φ qk is obtained by normalization based on the query point x q Calculate the key point x k The relative position encoding, α is a query point feature f q With the key point feature f k Linear transformation, α = f q -f k .

[0047] A q,k,c is the channel attention weight, defined as shown in formula (3):

[0048]

[0049] Among them, c is the channel index, τ is the smoothness parameter used to control the SoftMax function, γ is the mapping function, and is a multi-layer perceptron model (MLP) consisting of 2 linear layers and ReLu, ψ and It is a linear change.

[0050] In some embodiments, in order to more fully process the local details in the point cloud data, spatial attention is added after the channel attention to help the model capture the detailed geometric structure of the local point cloud. In the spatial attention part, F is firstly averaged and max pooled. CA Process again and merge the results of channel attention and spatial attention to obtain the final feature. The result of spatial attention SA output is shown in formula (4):

[0051] F SA =SA(F CA )=sigmoid(concat[Avgpool(F CA ),Maxpool(F CA )]) (4)

[0052] CSPA improves the model's feature extraction capabilities when processing point cloud data and optimizes the model's understanding and processing of input data, thereby improving the efficiency of extracting valuable features from point cloud data.

[0053] In some embodiments, in order to better understand the context information and adapt to the changes in point cloud feature data, the present invention proposes a dynamic edge information perception module (EIA), the structural diagram of which is shown in FIG. Figure 3As shown. After each upsampling in the decoder stage, the EIA module not only dynamically strengthens the edge features that are beneficial to the task or suppresses irrelevant edge information, but also reduces the model's sensitivity to noise and better restores the original point cloud resolution, thereby improving the performance and generalization ability of the model. EIA is a new module based on the idea of ​​the diffusion unit (DU) that acts on edge information in a dynamically weighted manner. The diffusion unit DU is an edge perception unit extended from the classical diffusion theory. The principle of the diffusion unit DU is shown in formula (5):

[0054]

[0055] in, is a point set, Represented as the center point and the neighboring point, respectively, n 、f s They represent the features of the center point and the neighboring points respectively, and σ is the ReLu, normalization, and convolution operation.

[0056] In some embodiments, the present invention generates new feature representations by convolution operations for the upsampled local features and the features after global average pooling, and fuses the two feature representations. Feature selection is performed through the nonlinear activation function Relu and the convolution layer, and the weight coefficient is obtained through the Sigmoid activation function, and the feature is weighted using the weight coefficient. This dynamic weighting method further optimizes the fusion process, helps to restore local details that may be lost during the downsampling process, and retains global information. The EIA module assigns different weights to different features, strengthens the features that are important to the task, and suppresses relatively unimportant information, which helps to segment the edge details of the object point cloud. Finally, edge optimization is further achieved through DU to make the information smoother. The principle of the EIA module is shown in formula (6):

[0057] F EIA =Concat[Sigmoid[σ[Avgpool(F C ' SPA )+F C ' SPA ]],Avgpool(F C ' SPA )]+f (6)

[0058] Among them, F C ' SPA is the F after downsampling CSPA Features, F EIA The output of the EIA.

[0059] In some embodiments, a point cloud semantic segmentation method for objects and indoor scenes is also provided, including: acquiring target point cloud data; inputting the point cloud data into a point cloud semantic segmentation model such as in the above embodiments, and outputting the segmented point cloud data.

[0060] In some embodiments, the performance of the point cloud semantic segmentation model of the present invention is evaluated on benchmark datasets: including the point cloud part segmentation dataset ShapeNetPart of objects and the point cloud semantic segmentation dataset S3DIS for indoor scenes:

[0061] 1) Object part segmentation:

[0062] Dataset: ShapeNetPart is an object-level dataset for part segmentation. It consists of 16,880 shapes from 16 different shape categories, each with 2-6 parts and 50 part segmentation labels. In our experiments, we followed the dataset division method of predecessors, with 14,006 samples as training set and 2,847 samples as validation set. On each shape, 2,048 points are randomly sampled.

[0063] Performance comparison: The model performance of the point cloud semantic segmentation model of the present invention is evaluated using the average of instance IoU and category IoU, and the IoU of each category. Table 1 lists some segmentation results on ShapeNetPart. Compared with some recent studies, including PointNet++, the results show that the present invention achieves 83.5 and 86.1 in class mIoU and instance mIoU, respectively, which is competitive performance.

[0064] Table 1 Results of partial segmentation of ShapeNetPart dataset

[0065]

[0066] Visualization: The visualization results of the partial segmentation of the ShapeNetPart dataset are shown in Figure 4. By comparing the visualization results, it is found that the point cloud semantic segmentation model of the present invention is closer to the true value. In terms of equal categories, the present invention has certain advantages.

[0067] 2) Indoor scene segmentation:

[0068] Dataset: S3DIS (Stanford Large-Scale 3D Indoor Spaces Dataset) is a large-scale indoor 3D point cloud dataset provided by Stanford University. The dataset contains 6 large indoor areas with a total of 271 rooms, including classrooms and offices, with a total coverage area of ​​more than 6000m 2,Each room has a point cloud file and a label file, and the annotated objects are divided into a total of 13 semantic classes (such as ceiling, floor, wall, door, etc.) for training and evaluating the semantic segmentation model. In this embodiment, following the settings of previous work, Area5 is used as the test set.

[0069] Performance comparison: The overall accuracy OA, the average mAc of class accuracy and the average IoU of instances are used to evaluate the model performance of the point cloud semantic segmentation model of the present invention, and the results are shown in Table 2. The point cloud semantic segmentation model and method of the present invention were tested in the most difficult segmentation area 5 of S3DIS. The CSEANet of the present invention achieved mIoU / mAcc / OA results of 69.8 / 75.4 / 90.5 respectively, achieving certain advantages.

[0070] Table 2 Results of semantic segmentation on the S3DIS dataset

[0071]

[0072] Visualization: The semantic segmentation visualization results of the S3DIS dataset are as follows Figure 5 As shown in the figure, the first two columns represent the input point cloud and the true annotation value, respectively, and the last two columns represent the prediction results of PointNet++ and our CSEANet. The comparison of the visualization results shows that the point cloud semantic segmentation model of the present invention is closer to the true value. In terms of equal categories, the present invention has certain advantages.

[0073] The preferred specific embodiments of the present invention are described in detail above. It should be understood that ordinary technicians in the field can make many modifications and changes based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by technicians in the technical field based on the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A point cloud semantic segmentation model for objects and indoor scenes, characterized in that: include: An acquisition module is used to acquire target point cloud data; An encoding module, used for down-sampling the target point cloud data, introducing a channel space attention module and utilizing the channel and spatial information of the channel space attention module to extract features of the target point cloud data; The decoding module is used to upsample the target point cloud data features processed by the encoding module, introduce an edge information perception module, and fuse and reconstruct the sampled target point cloud data features through a dynamic weighting method.

2. The point cloud semantic segmentation model for objects and indoor scenes according to claim 1, characterized in that: The channel space attention module is a combination of channel attention and spatial attention. The output formula of the channel space attention module is: F CSPA =F CA +F SA =CA(x,f)+SA(F CA ) Among them, F CSPA is the output result of the channel space attention module, F CA is the result of the channel attention output, F SA is the result of the spatial attention output, CA is the channel attention, (x, f) is the three-dimensional coordinates and feature vector of the target point cloud data, and SA is the spatial attention.

3. The point cloud semantic segmentation model for objects and indoor scenes according to claim 2, characterized in that: The down sampling adopts farthest point sampling.

4. The point cloud semantic segmentation model for objects and indoor scenes according to claim 3, characterized in that: The result of the channel attention output is expressed as: Among them, f q is x q , x k The corresponding point feature, Ω key is the set of key points, A q,k,c is the channel attention weight, φ qk To obtain the query point x through normalization q Calculate the key point x k The relative position encoding, α is a query point feature f q With the key point feature f k Linear transformation, α = f q -f k .

5. The point cloud semantic segmentation model for objects and indoor scenes according to claim 4, characterized in that: The calculation of the channel attention weight is to calculate the attention weight between the query point and the key point of each channel, the query point is selected as a point after random sampling of the original point cloud data, and enhanced by a random vector generated by a learnable parameter to form a recognizable point; the key point is the point retained after sampling the farthest point.

6. The point cloud semantic segmentation model for objects and indoor scenes according to claim 5, characterized in that: The channel attention weight is expressed as: Among them, c is the channel index, τ is the smoothness parameter used to control the SoftMax function, γ is the mapping function, and is a multilayer perceptron model consisting of 2 linear layers and ReLu. ψ and It is a linear change.

7. The point cloud semantic segmentation model for objects and indoor scenes according to claim 6, characterized in that: The spatial attention is applied to the F by average pooling and maximum pooling operations. CA Process again and merge the results of the channel attention and the spatial attention. The result of the spatial attention output is expressed as: SA =SA(F CA )=sigmoid(concat[Avgpool(F CA ),Maxpool(F CA )]).

8. The point cloud semantic segmentation model for objects and indoor scenes according to claim 1, characterized in that: The upsampled features are fused in a dynamic weighted manner, feature selection is performed through a nonlinear activation function Relu and a convolutional layer, a weight coefficient is obtained through a Sigmoid activation function, and the features are dynamically weighted using the weight coefficient.

9. The point cloud semantic segmentation model for objects and indoor scenes according to claim 8, characterized in that: The edge information perception module is a new module based on a diffusion unit that acts on edge information in a dynamic weighted manner, wherein the diffusion unit is represented by: in, is a point set, s, Represented as the center point and the neighboring point, respectively, n 、f s Represent the features of the center point and the neighboring points respectively, σ is ReLu, normalization, and convolution operation; The result output by the edge information perception module is expressed as follows: F EIA =Concat[Sigmoid[σ[Avgpool(F C ' SPA )+F C ' SPA ]],Avgpool(F C ' SPA )]+f Among them, F C ' SPA is the F after downsampling CSPA Features, F EIA It is the result output by the edge information perception module EIA.

10. A point cloud semantic segmentation method for objects and indoor scenes, characterized in that: include: Obtain target point cloud data; The point cloud data is input into the point cloud semantic segmentation model as described in any one of claims 1-9, and the segmented point cloud data is output.

Citation Information

Patent Citations

  • Three-dimensional semantic segmentation method based on channel attention and multi-scale fusion

    CN114743007A

  • Three-dimensional point cloud semantic segmentation method based on local and global context awareness

    CN117218351A

  • Point cloud segmentation network based on combination of attention mechanism and double-graph convolution

    CN117541795A

  • Method for semantic segmentation using correlations and regional associations of multi-scale features, and computer program recorded on record-medium for executing method thereof

    KR102546206B1

  • Semantic segmentation of point clouds using confidence-based adaptive feature aggregation and 3D neighborhood feature augmentation

    WO2024243725A1