A 3D point cloud semantic segmentation method and system based on visual assistance and feature enhancement

By introducing visual assistance and feature enhancement modules into the three-dimensional point cloud semantic segmentation method, using self-attention and channel attention mechanisms, the shortcomings of the existing methods in large-scale data processing and semantic boundary point segmentation are solved, and higher segmentation accuracy is achieved.

CN116229079BActive Publication Date: 2025-05-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310324023.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2025-05-06
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

The existing three-dimensional point cloud semantic segmentation method has shortcomings in processing large-scale data and segmenting boundary points of different semantic classes. It is unable to effectively utilize the information of point cloud data, resulting in low segmentation accuracy.

Method used

A three-dimensional point cloud semantic segmentation method based on visual assistance and feature enhancement is adopted. By building a deep learning model, visual assisted tasks and feature enhancement modules are introduced, and self-attention mechanisms and channel attention mechanisms are used to enhance semantic segmentation performance.

Benefits of technology

The accuracy of three-dimensional point cloud semantic segmentation is improved, especially in segmenting the boundary points of different semantic classes, and the model's processing ability of large-scale point cloud data is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229079B_ABST
    Figure CN116229079B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer vision graphics, and specifically relates to a method and system for three-dimensional point cloud semantic segmentation based on visual assistance and feature enhancement. The method comprises constructing and training a deep learning model for three-dimensional point cloud semantic segmentation, inputting three-dimensional point cloud data to be segmented into the trained point cloud semantic segmentation model, explicitly extracting visual color features by designing a reconstruction auxiliary network, introducing a channel attention mechanism in a backbone segmentation network to fully utilize the features, and constructing a point feature enhancement module in a decoding layer to further improve the model's ability to segment points at the boundaries of different semantic classes. The present invention can effectively aggregate the local neighborhood of a point, improve the effect of a deep learning model on the semantic segmentation of three-dimensional point clouds, and promote the development of related technical fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision graphics, and specifically to a three-dimensional point cloud semantic segmentation method and system based on visual assistance and feature enhancement. Background Art

[0002] The goal of 3D point cloud semantic segmentation is to divide the point cloud into several subsets based on the semantics of the points. 3D point cloud semantic segmentation is an important research topic in the field of computer vision, and many technologies are built on this basis, such as target detection, classification, and recognition based on 3D point clouds, which are the main technologies currently used to solve scene understanding. Traditional point cloud segmentation methods use information such as the position and shape of point clouds to segment different area boundaries. The segmentation results require manual semantic annotation of the results, which cannot achieve acceptable accuracy and cannot be generalized to large-scale data. Therefore, it is necessary to develop an efficient point cloud semantic segmentation algorithm to automatically obtain semantic annotations for large-scale point clouds.

[0003] Driven by the great success of deep learning technology in 2D computer vision, it has also become the tool of choice for 3D segmentation tasks. On this basis, three types of semantic segmentation methods have emerged: projection-based, discretization-based, and point-based. Due to some common problems, projection-based and discretization-based methods are not optimal for practical applications. On the one hand, they require several time-consuming pre- and post-processing steps to make predictions, and on the other hand, the generated intermediate representations may partially lose the contextual information of the surrounding environment. Point-based methods have gradually become the mainstream method because such methods can directly process irregular point clouds. However, the point-by-point features extracted by MLP cannot capture the local geometry of points and the interactions between points. Although some researchers have also proposed related solutions, such as using a hierarchical sampling mechanism in PointNet++ to obtain sampled points that fuse the feature information of local neighborhood points. However, these methods do not make full use of the existing information of point cloud data to improve the semantic segmentation effect, and have limited ability to segment boundary points of different semantic classes. Summary of the invention

[0004] In view of this, the present invention provides a 3D point cloud semantic segmentation method based on visual assistance and feature enhancement, including constructing and training a 3D point cloud semantic segmentation deep learning model, inputting the 3D point cloud data to be segmented into the trained point cloud semantic segmentation model, introducing a visual assistance task and a feature enhancement module to enhance the semantic segmentation performance, and calculating the segmentation result; the training process of the 3D point cloud semantic segmentation model specifically includes the following steps:

[0005] S1, obtaining the three-dimensional point cloud to be segmented and preprocessing it;

[0006] S2, inputting the preprocessed point cloud data into the segmentation network and the reconstruction network respectively, wherein the data input into the segmentation network includes XYZ space coordinates and RGB color information, and the data input into the reconstruction network includes RGB color information;

[0007] S3, the segmentation network and the reconstruction network both obtain the features of each point. In the segmentation network, the neighborhood index and its features of each point in the point cloud are obtained; the reconstruction network shares the neighborhood index calculated in the segmentation network;

[0008] S4, for the encoder of the segmentation network, geometric encoding and feature encoding are performed on each local neighborhood; the encoder of the reconstruction network performs visual color feature extraction;

[0009] S5, the segmentation network splices the features obtained by geometric coding and feature coding, and obtains accurate point-by-point features through weighted aggregation by the self-attention mechanism; the reconstruction network extracts the most significant features of visual color from its neighborhood and splices the weighted average of its neighborhood as the point-by-point features of the reconstruction network;

[0010] S6. In each encoding layer of the segmentation network, the spatial position and volume ratio are used to learn the global context of the 3D point cloud.

[0011] S7. Use the channel attention mechanism in the segmentation network to fuse the top-level pyramid features from the segmentation network and the reconstruction network, and use the fused features as the input of the decoder in the segmentation network; the reconstruction network inputs the top-level pyramid features of the reconstruction network into the decoder in the reconstruction network;

[0012] S8. In the decoder, the segmentation network performs feature enhancement on the upsampled point cloud; the reconstruction network performs nearest neighbor interpolation upsampling and MLP to extract point-by-point features;

[0013] S9, the cross entropy loss function is used for supervision of the segmentation network, and the mean square error is used for supervision of the reconstruction network;

[0014] S10, start the gradient back propagation mechanism, optimize the loss function, update the network parameters, and save the model when the model converges or reaches the set number of epochs.

[0015] Furthermore, the process of obtaining the features of each point by the segmentation network and the reconstruction network includes using a fully connected layer to increase the dimension of the features of the input point cloud, and then using MLP to extract the features of each point in the entire three-dimensional point cloud.

[0016] Furthermore, in the segmentation network, the process of obtaining the neighborhood of each point in the point cloud includes: taking each point in the point cloud as the center point and using the nearest neighbor algorithm to find the indexes of its K neighbor points, and obtaining the xyz three-dimensional coordinates and corresponding point features of each neighbor point according to the index.

[0017] Furthermore, the encoder of the segmentation network performs geometric and feature encoding on each local neighborhood by:

[0018] Calculate the relative coordinates of each neighbor point and the center point, and concatenate the relative coordinates, center point coordinates, and neighbor point coordinates as the geometric context information of the entire neighborhood;

[0019] Use MLP to extract the entire local geometric context features;

[0020] Calculate the absolute average of the feature differences between the center point and the neighboring points, and concatenate the neighboring point features and the negative exponential of the feature distance as the new neighboring point features.

[0021] Furthermore, using spatial position and volume ratio to learn the global context of three-dimensional point clouds includes: using the distance between the center point and its farthest neighbor in the point cloud neighborhood to calculate the local volume, using the farthest distance from the coordinate origin in the point cloud to calculate the global volume, and using MLP learning based on the ratio of the local volume to the global volume to obtain the global context of the three-dimensional point cloud.

[0022] Furthermore, the process of feature enhancement of the upsampled point cloud by the segmentation network includes the following steps: using the nearest neighbor algorithm to obtain K neighbor points of each point, subtracting the corresponding center point feature from the neighbor point feature to obtain the absolute feature difference, and summing the absolute feature differences of all neighbor points; after extracting the features through MLP, adding them element by element with the center point feature to obtain the enhanced features as the input of the next decoding layer.

[0023] Furthermore, the visual assistance and feature enhancement module is introduced into the 3D point cloud semantic segmentation model, including: taking each point in the 3D point cloud data input to the segmentation network as the center point, using the K nearest neighbor algorithm to find its corresponding K neighbor points, and geometrically encoding the neighbor points in terms of local geometric context to obtain In terms of feature space, feature encoding of local neighborhood is obtained and The attention weights of neighboring points are calculated through the self-attention mechanism, and then the weighted sum is used to obtain an accurate local context representation containing spatial geometric information and feature distance information.

[0024] Furthermore, the segmentation network performs feature enhancement on the upsampled point cloud, and the decoding layer is expressed as follows after point-by-point feature enhancement:

[0025]

[0026] Among them, f i u is the feature of the i-th point after upsampling; For f i u Features after point-by-point enhancement; is the corresponding neighbor point feature; K is the number of points in the neighborhood; |·| means taking the absolute value.

[0027] Furthermore, the loss function of the entire end-to-end model training phase is expressed as:

[0028] L total =L ce +L mse ;

[0029] Among them, L ce represents the cross entropy loss of semantic segmentation results, L mse Represents the mean square error loss for color reconstruction.

[0030] The present invention also provides a three-dimensional point cloud semantic segmentation system based on visual assistance and feature enhancement, which is used to implement a three-dimensional point cloud semantic segmentation method based on visual assistance and feature enhancement, including a data preprocessing module, a shared sampling module, a local neighborhood search module, a local context encoding module, a shared MLP module, a splicing module, a self-attention aggregation module, a pooling module, a global feature acquisition module, a channel attention module, an upsampling module and a feature enhancement module, wherein:

[0031] The data preprocessing module is used to preprocess the 3D point cloud and reduce the number of input 3D point cloud points;

[0032] The shared sampling module uses the farthest point sampling algorithm to filter out uniform sample points and input them into the next layer;

[0033] The local neighborhood search module is used to search for neighboring points of a point and build the local neighborhood of the point;

[0034] The local context encoding module performs geometric encoding and feature encoding on the local neighborhood of the point;

[0035] The shared MLP module is used to extract local context features of points and perform point-by-point feature extraction on the entire 3D point cloud;

[0036] The splicing module is used to fuse the feature information of points and splice the local geometric context and semantic context of points together;

[0037] The self-attention aggregation module is used to aggregate the local context of each point to obtain the precise local context representation of each point;

[0038] The pooling module is used to perform maximum pooling and average pooling on the local features in the reconstruction network to obtain representative visual features;

[0039] The global feature acquisition module is used to extract the global feature representation of the points in the segmentation network;

[0040] A channel attention module to fuse visual features from the reconstruction network;

[0041] The upsampling module uses the nearest neighbor trilinear interpolation to upsample high-dimensional features.

[0042] The feature enhancement module is used to enhance the features of the decoding layer, so as to increase the feature gap between different semantic classes and improve the segmentation accuracy at the boundary points of semantic classes.

[0043] The present invention uses the self-attention mechanism to collect accurate local context representation of points by geometrically and feature encoding the local neighborhood of point clouds, and uses spatial position and volume ratio to extract global features. In addition, the method uses the channel attention mechanism to fuse visual information from the reconstruction network and enhances point-by-point features at the decoding layer. This can increase the feature gap between the local neighborhood points of the point and the center point belonging to different semantic classes, thereby improving the accuracy of semantic segmentation of three-dimensional point clouds. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 This is the overall structure diagram of the 3D point cloud semantic segmentation based on visual assistance and feature enhancement of the present invention;

[0045] Figure 2 is a point feature enhancement module diagram of the present invention;

[0046] Figure 3 This is a semantic segmentation result diagram according to an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0048] The present invention proposes a 3D point cloud semantic segmentation method based on visual assistance and feature enhancement, including constructing and training a 3D point cloud semantic segmentation deep learning model, inputting the 3D point cloud data to be segmented into the trained point cloud semantic segmentation model, introducing a visual assistance task and a feature enhancement module to enhance the semantic segmentation performance, and calculating the segmentation result; the training process of the 3D point cloud semantic segmentation model specifically includes the following steps:

[0049] S1, obtaining the three-dimensional point cloud to be segmented and preprocessing it;

[0050] S2, inputting the preprocessed point cloud data into the segmentation network and the reconstruction network respectively, wherein the data input into the segmentation network includes XYZ space coordinates and RGB color information, and the data input into the reconstruction network includes RGB color information;

[0051] S3, the segmentation network and the reconstruction network both obtain the features of each point. In the segmentation network, the neighborhood index and its features of each point in the point cloud are obtained; the reconstruction network shares the neighborhood index calculated in the segmentation network;

[0052] S4, for the encoder of the segmentation network, geometric encoding and feature encoding are performed on each local neighborhood; the encoder of the reconstruction network performs visual color feature extraction;

[0053] S5, the segmentation network splices the features obtained by geometric coding and feature coding, and obtains accurate point-by-point features through weighted aggregation by the self-attention mechanism; the reconstruction network extracts the most significant features of visual color from its neighborhood and splices the weighted average of its neighborhood as the point-by-point features of the reconstruction network;

[0054] S6. In each encoding layer of the segmentation network, the spatial position and volume ratio are used to learn the global context of the 3D point cloud.

[0055] S7. Use the channel attention mechanism in the segmentation network to fuse the top-level pyramid features from the segmentation network and the reconstruction network, and use the fused features as the input of the decoder in the segmentation network; the reconstruction network inputs the top-level pyramid features into the decoder in the reconstruction network;

[0056] S8. In the decoder, the segmentation network performs feature enhancement on the upsampled point cloud; the reconstruction network performs nearest neighbor interpolation upsampling and MLP to extract point-by-point features;

[0057] S9, the cross entropy loss function is used for supervision of the segmentation network, and the mean square error is used for supervision of the reconstruction network;

[0058] S10, start the gradient back propagation mechanism, optimize the loss function, update the network parameters, and save the model when the model converges or reaches the set number of epochs.

[0059] This embodiment proposes a specific implementation method of a 3D point cloud semantic segmentation method based on hybrid local aggregation, such as Figure 1 As shown, including:

[0060] S1. In order to quickly obtain point cloud data, first perform grid sampling on the entire point cloud scene. The sampling idea is to evenly divide the entire scene into multiple small cubes, sample the points in each cube, and use the category with the largest proportion as the sampled category. After that, in order to quickly find the nearest neighbor, a KDTree and a mapping file from the sampled point cloud to the original point cloud are established for each subxyz of the scene. This example sets the cube length to 0.04 meters, the maximum epoch number to 100, and the initial learning rate to 0.01;

[0061] S2. The preprocessed point cloud data is input into the segmentation network and the reconstruction network respectively, wherein the data input into the segmentation network contains XYZ space coordinates and RGB color information, while the reconstruction network only contains RGB color information.

[0062] S3. Both networks use a fully connected layer to increase the initial features from 6 channels to 8 channels, and then use MLP to extract the features of each point in the entire 3D point cloud. In addition, in the segmentation network, based on the 3D coordinates of the point, the nearest neighbor algorithm is used to find the index of the K neighbor points of each point in space according to the Euclidean distance, and then the coordinates and point features of all neighbor points corresponding to each point are obtained according to the index. For example, the coordinates and features of the i-th point are p respectively. i and f i , then the corresponding neighborhood The coordinates and features of the jth neighborhood point in are p j and f j The reconstructed network shares these neighborhood indices and also uses the indices to obtain neighbor point features.

[0063] S4. For the segmentation network, in order to simultaneously capture global and local information about the local neighborhood of a point, each local neighborhood is geometrically and feature encoded. First, the spatial coordinates of each point and its neighboring points are geometrically encoded. The relative coordinates of each neighboring point with the center point are calculated, and the relative coordinates, center point coordinates, and neighboring point coordinates are concatenated as the entire local geometric context, so that the corresponding point features always know their relative spatial positions.

[0064] Specifically, for the center point p i and its corresponding K neighbor points For each point in the space geometry encoding, the formula can be expressed as:

[0065]

[0066] Among them, p i represents the xyz three-dimensional coordinates of the i-th point, represents the xyz three-dimensional coordinates of the jth point in the neighborhood of the i-th point, [,,] represents a concatenation operation; MLP is a multi-layer perceptron, which is used to extract point-by-point features; Represents the local geometric features obtained after geometric encoding of the points. The positions are encoded from redundant points, but in practice, this often helps the deep learning network model learn local geometric features and achieve good performance.

[0067] Feature encoding only calculates the absolute average of the feature differences between the center point and the neighboring points, and concatenates the neighboring point features and the negative exponential of the feature distance as the new neighboring point features, which are specifically expressed as follows:

[0068]

[0069] Here, since the features are automatically learned by the network, the feature difference is an unstable feature. For this reason, we introduce a hyperparameter λ to adjust the weight of the feature difference.

[0070] The reconstruction network does not contain geometric encoding, but only feature encoding operations in feature space.

[0071] S5. The segmentation network concatenates the two encoded features and obtains accurate point-by-point features through weighted merging by the self-attention mechanism. The reconstruction network directly collects the most significant (largest) features from the K neighbors to represent the visual overview of the entire neighborhood. On the other hand, it extracts and obtains more neighborhood details by learning the weighted average of the entire neighborhood.

[0072] Specifically, the segmentation network uses the self-attention mechanism to aggregate local context features, which can be expressed as:

[0073]

[0074]

[0075] The reconstructed network aggregated local context features can be expressed as:

[0076]

[0077] Among them, max K (f′ j ) means taking the maximum value of the feature among K local neighborhood points, θ i is a set of learnable weights for the K neighbors, It represents the weighted average of the features of K local neighborhood points (that is, the softmax function is used to calculate the score of each point in the local neighborhood, which should be between 0 and 1, and then the value obtained by weighted summation is the weighted average).

[0078] S6. In each encoding layer of the segmentation network, the spatial position and volume ratio are used to learn the global context of the 3D point cloud:

[0079] v i =max(||p j -p i || 3 )

[0080] v g =max(‖PO‖ 3 )

[0081]

[0082] f iG =MLP([p i ,r i ])

[0083] Where, the local volume v i It is obtained by the cube of the maximum distance between the neighbor point and the center point, and the global volume v i It is obtained by the cube of the maximum distance from the origin O among all points in the point cloud P. It is worth noting that objects of the same type (such as chairs, tables) in different scenes usually have different styles, and their geometric structures are not exactly the same. Therefore, considering that the volume ratio is insensitive to the position of the internal points within the local and global bounding spheres, it is used to represent that slight geometric deformations of objects of the same category can be tolerated.

[0084] After that, the output of each encoding layer of the segmentation network is:

[0085]

[0086] S7. For the top-level features of the pyramid encoded by the two networks, they are concatenated through the concatenate operation, and then the channel attention mechanism is used in the segmentation network to fuse the visual feature information from the reconstruction network, and then input to the decoder. The reconstruction network has no fusion mechanism, and only the original top-level visual features are input to the decoder.

[0087] S8. In the decoding layer, the segmentation network performs feature enhancement on the upsampled point cloud. First, the nearest neighbor algorithm (KNN) is used to obtain the K neighbor points of each point. The absolute feature difference is obtained by subtracting the corresponding center point feature from the neighbor point feature, and then summing them. Finally, after the MLP extracts the features, the enhanced features are added element by element to the center point feature to obtain the input of the next decoding layer. The reconstruction network only includes the nearest neighbor interpolation upsampling and the MLP extraction of point-by-point features. The specific process is as follows Figure 2 As shown, including:

[0088] For the i-th point fi u , find its k neighbor points through the nearest neighbor algorithm, k is the sum of the number of points including the i-th point and its neighbor points, and d is the dimension of each point;

[0089] Calculate the absolute value of the difference between the eigenvalues ​​of each point and the i-th point, add all the absolute values, extract point-by-point features using a multilayer perceptron, and add them to the eigenvalue of the i-th point to obtain the eigenvalue of the i-th point after point-by-point feature enhancement at the decoding layer.

[0090] S9. The cross entropy loss function is used for supervision of the segmentation network, and the mean square error is used for supervision of the reconstruction network.

[0091] S10. Start the gradient back-propagation mechanism, optimize the loss function, update the network parameters, and save the model when the model converges or reaches the set number of epochs.

[0092] The 3D point cloud semantic segmentation model in the present invention is an end-to-end model, that is, the input of the model is the original 3D point cloud data, and the output of the model is the semantic segmentation result we want.

[0093] In the process of semantic segmentation of 3D point clouds, point-by-point feature representation is crucial for the semantic segmentation task. Although non-parametric symmetric functions can effectively summarize the local information of points, they cannot explicitly show the local uniqueness, especially for neighboring points that share similar local contexts. To solve this problem, this paper collects accurate neighborhood representations based on the attention mechanism, introduces visual reconstruction auxiliary tasks to make full use of the existing color information of the point cloud, and enhances the multi-scale features of the decoding layer to improve the semantic segmentation effect, including:

[0094] Taking each point in the 3D point cloud data input to the segmentation network as the center point, the K nearest neighbor algorithm is used to find its corresponding K neighbor points. In terms of local geometric context, the neighbor points are geometrically encoded to obtain G(p i ), in the feature space, feature encoding is performed on the local neighborhood to obtain G(f i );

[0095] G(p i ) and G(f i ) are spliced ​​together, the attention weights of neighboring points are calculated through the self-attention mechanism, and then the weighted sum is performed to obtain the accurate local context representation G(i) containing spatial geometric information and feature distance information.

[0096] The point-by-point features are enhanced in the decoding layer of the segmentation network to increase the feature gap between the boundary points of different semantic classes and improve the segmentation accuracy at the boundary points of the semantic classes. The point-by-point feature enhancement of the decoding layer is expressed as:

[0097]

[0098] Among them, f i u is the feature of the i-th point after upsampling; For f i u Features after point-by-point enhancement; f j u is the corresponding neighbor point feature; K is the number of points in the neighborhood; |·| is the absolute value operation.

[0099] Furthermore, the loss function of the entire end-to-end model training phase is expressed as:

[0100] L total =L ce +L mse ;

[0101] Among them, L ce represents the cross entropy loss of semantic segmentation results, L mse Represents the mean square error loss for color reconstruction.

[0102] In this embodiment, a data set (S3DIS) is used for experiments. The data set (S3DIS) is collected from an indoor working environment and is widely used for semantic segmentation tasks. There are six sub-areas in the data set, each of which contains 50 different rooms. The number of points in most rooms ranges from 500,000 to 2.5 million, depending on the size of the room. All points have three-dimensional coordinates and color information and are marked as one of 13 semantic categories. In the experiment, area five is used for testing and other areas are used for training. As a convention, each room is first grid sampled with a grid size of 4 cm for training and testing. The input point cloud is formed by obtaining a maximum of 40,960 points from a room, and 3D coordinates and color information are used as input features. The average intersection-over-union ratio (mIou), global accuracy (OA) and average accuracy (mAcc) are used to measure the quality of semantic segmentation. Table 1 shows the experimental results of semantic segmentation of three-dimensional point clouds.

[0103] Table 1 Test results on S3DIS dataset Area 5

[0104]

[0105] It can be seen from Table 1 that the global accuracy (OA), average accuracy (mAcc) and average intersection-over-union (mIou) obtained by this patent are better than those of other algorithms through testing with different algorithms. Figure 3As shown, the first row is the input 3D point cloud of the indoor scene, the second row is the real semantic segmentation label (referred to as the real label), the third row is the prediction result of the large-scale point cloud (referred to as RandLA-Net) method, and the fourth row is the prediction result of the 3D point cloud semantic segmentation method based on visual assistance and feature enhancement of the present invention. The result shows that the present invention can effectively improve the semantic segmentation effect based on visual assistance and feature enhancement.

[0106] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A 3D point cloud semantic segmentation method based on visual assistance and feature enhancement, characterized in that: Construct and train a 3D point cloud semantic segmentation deep learning model, input the 3D point cloud data to be segmented into the trained point cloud semantic segmentation model, introduce visual auxiliary tasks and feature enhancement modules to enhance the semantic segmentation performance, and calculate the segmentation results; the training process of the 3D point cloud semantic segmentation model specifically includes the following steps: S1, obtaining the three-dimensional point cloud to be segmented and preprocessing it; S2. Input the preprocessed point cloud data into the segmentation network and the reconstruction network respectively, wherein the data input into the segmentation network includes XYZ space coordinates and RGB color information, and the data input into the reconstruction network includes RGB color information; S3, the segmentation network and the reconstruction network both obtain the features of each point. In the segmentation network, the neighborhood index and its features of each point in the point cloud are obtained; the reconstruction network shares the neighborhood index calculated in the segmentation network; S4, for the encoder of the segmentation network, geometric encoding and feature encoding are performed on each local neighborhood; the encoder of the reconstruction network performs visual color feature extraction; S5, the segmentation network splices the features obtained by geometric coding and feature coding, and obtains accurate point-by-point features through weighted aggregation by the self-attention mechanism; the reconstruction network extracts the most significant features of visual color from its neighborhood and splices the weighted average of its neighborhood as the point-by-point features of the reconstruction network; S6. In each encoding layer of the segmentation network, the spatial position and volume ratio are used to learn the global context of the 3D point cloud. S7. Use the channel attention mechanism in the segmentation network to fuse the top-level pyramid features from the segmentation network and the reconstruction network, and use the fused features as the input of the decoder in the segmentation network; the reconstruction network inputs the top-level pyramid features of the reconstruction network into the decoder in the reconstruction network; S8. In the decoder, the segmentation network performs feature enhancement on the upsampled point cloud; the reconstruction network performs nearest neighbor interpolation upsampling and MLP to extract point-by-point features; S9, the cross entropy loss function is used for supervision of the segmentation network, and the mean square error is used for supervision of the reconstruction network; S10, start the gradient back propagation mechanism, optimize the loss function, update the network parameters, and save the model when the model converges or reaches the set number of epochs.

2. The method for semantic segmentation of three-dimensional point clouds based on visual assistance and feature enhancement according to claim 1, characterized in that: The process of segmenting the network and reconstructing the network to obtain the features of each point includes using a fully connected layer to increase the dimension of the input point cloud features, and then using MLP to extract the features of each point in the entire three-dimensional point cloud.

3. The method for semantic segmentation of three-dimensional point clouds based on visual assistance and feature enhancement according to claim 1, characterized in that: In the segmentation network, the process of obtaining the neighborhood of each point in the point cloud includes: taking each point in the point cloud as the center point and using the nearest neighbor algorithm to find the indexes of its K neighbor points, and obtaining the xyz three-dimensional coordinates and corresponding point features of each neighbor point according to the index.

4. The method for semantic segmentation of three-dimensional point cloud based on visual assistance and feature enhancement according to claim 1, characterized in that: The encoder of the segmentation network encodes the geometry and features of each local neighborhood through the following steps: Calculate the relative coordinates of each neighbor point and the center point, and concatenate the relative coordinates, center point coordinates, and neighbor point coordinates as the geometric context of the entire neighborhood; Use MLP to extract the entire local geometric context features; Calculate the absolute average of the feature differences between the center point and the neighboring points, and concatenate the neighboring point features and the negative exponential of the feature distance as the new neighboring point features.

5. The method for semantic segmentation of three-dimensional point cloud based on visual assistance and feature enhancement according to claim 1, characterized in that: Using spatial position and volume ratio to learn the global context of three-dimensional point clouds includes: using the distance between the center point and its farthest neighbor in the point cloud neighborhood to calculate the local volume, using the farthest distance from the coordinate origin in the point cloud to calculate the global volume, and using MLP learning based on the ratio of local volume to global volume to obtain the global context of the three-dimensional point cloud.

6. The method for semantic segmentation of three-dimensional point cloud based on visual assistance and feature enhancement according to claim 1, characterized in that: The process of feature enhancement of the upsampled point cloud by the segmentation network includes the following steps: using the nearest neighbor algorithm to obtain K neighbor points of each point, subtracting the corresponding center point feature from the neighbor point feature to obtain the absolute feature difference, and summing the absolute feature differences of all neighbor points; after the MLP extracts the features, the enhanced features are added element by element with the center point feature to obtain the input of the next decoding layer.

7. The method for semantic segmentation of three-dimensional point cloud based on visual assistance and feature enhancement according to claim 1, characterized in that: The introduction of visual assistance and feature enhancement modules in the 3D point cloud semantic segmentation model includes: taking each point in the 3D point cloud data input to the segmentation network as the center point, using the K nearest neighbor algorithm to find its corresponding K neighbor points, and geometrically encoding the neighbor points in terms of local geometric context to obtain In terms of feature space, feature encoding of the local neighborhood is obtained and The attention weights of neighboring points are calculated through the self-attention mechanism, and then the weighted sum is used to obtain an accurate local context representation containing spatial geometric information and feature distance information.

8. The method for semantic segmentation of three-dimensional point cloud based on visual assistance and feature enhancement according to claim 1, characterized in that: The segmentation network performs feature enhancement on the upsampled point cloud, and the decoding layer is expressed as follows after point-by-point feature enhancement: Among them, f i u is the feature of the i-th point after decoding; For f i u Features after point-by-point enhancement; is the corresponding neighbor point feature; K is the number of points in the neighborhood; |·| means taking the absolute value.

9. The method for semantic segmentation of three-dimensional point cloud based on visual assistance and feature enhancement according to claim 1, characterized in that: The loss function of the entire end-to-end model training phase is expressed as: L total =L ce +L mse ; Among them, L ce represents the cross entropy loss of semantic segmentation results, L mse Represents the mean square error loss for color reconstruction.

10. A 3D point cloud semantic segmentation system based on visual assistance and feature enhancement, characterized in that: A method for semantic segmentation of a three-dimensional point cloud based on visual assistance and feature enhancement according to claim 1 is used to implement the method, comprising a data preprocessing module, a shared sampling module, a local neighborhood search module, a local context encoding module, a shared MLP module, a splicing module, a self-attention aggregation module, a pooling module, a global feature acquisition module, a channel attention module, an upsampling module and a feature enhancement module, wherein: The data preprocessing module is used to preprocess the 3D point cloud and reduce the number of input 3D point cloud points; The shared sampling module uses the farthest point sampling algorithm to filter out uniform sample points and input them into the next layer; The local neighborhood search module is used to search for neighboring points of a point and build the local neighborhood of the point; The local context encoding module performs geometric encoding and feature encoding on the local neighborhood of the point; The shared MLP module is used to extract local context features of points and perform point-by-point feature extraction on the entire 3D point cloud; The splicing module is used to fuse the feature information of points and splice the local geometric context and semantic context of points together; The self-attention aggregation module is used to aggregate the local context of each point to obtain the precise local context representation of each point; The pooling module is used to perform maximum pooling and average pooling on the local features in the reconstruction network to obtain representative visual features; The global feature acquisition module is used to extract the global feature representation of the points in the segmentation network; A channel attention module to fuse visual features from the reconstruction network; The upsampling module uses the nearest neighbor trilinear interpolation to upsample high-dimensional features; The feature enhancement module is used to enhance the features of the decoding layer, so as to increase the feature gap between different semantic classes and improve the segmentation accuracy at the boundary points of semantic classes.

Citation Information

Patent Citations

  • Three-dimensional point cloud semantic segmentation method based on multi-feature information enhancement coding

    CN113392841A

  • Online point cloud semantic segmentation method and device, storage medium and electronic equipment

    CN115272666A