Automatic driving LiDAR point cloud semantic segmentation method and system based on multiple scales

Through the multi-scale expansion fusion module and dynamic residual connection algorithm, the problems of receptive field limitation and feature learning weight uneven in the existing point cloud semantic segmentation method are solved, simplifying the preprocessing process and significantly improving the segmentation accuracy and efficiency.

CN120198668APending Publication Date: 2025-06-24NORTH CHINA UNIV OF WATER RESOURCES & ELECTRIC POWER +1
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510294278.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing point cloud semantic segmentation methods have problems such as restricted receptive field, uneven weight allocation during feature learning, and complex preprocessing and feature extraction processes, resulting in poor performance in dynamic scenarios and large-scale point cloud processing.

Method used

By setting up a multi-scale expansion fusion module to fuse features of different scales, the capture ability of receptive fields is enhanced, and the weight allocation in feature learning is balanced using dynamic residual connection algorithms to simplify the preprocessing and feature extraction process.

Benefits of technology

It significantly improves the feature extraction ability and learning efficiency of the model, improves the accuracy and effect of point cloud semantic segmentation, and adapts to complex scenarios and diversified point cloud data distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198668A_ABST
    Figure CN120198668A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale-based automatic driving LiDAR point cloud semantic segmentation method and system, and the method comprises the steps: 1, inputting a target point cloud feature into a multi-scale fusion module, and obtaining a plurality of features of different scales; 2, fusing the features of the plurality of different scales by using a multi-scale expansion fusion module to obtain a multi-scale fusion feature; 3, processing the target point cloud features and the multi-scale fusion features by using a dynamic residual connection algorithm to obtain residual connection features; and step 4, inputting the residual connection features into the SqueezeSegV3 network model, and completing semantic segmentation. According to the method, local and global information can be effectively captured, the weight is accurately distributed during feature learning, the preprocessing and feature extraction process is simplified, the precision and effect of laser LiDAR point cloud semantic segmentation are improved, and key support is provided for automatic driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to a multi-scale based semantic segmentation method and system for LiDAR point clouds in autonomous driving. Background Art

[0002] In existing autonomous driving technologies, LiDAR point clouds provide key environmental perception data for autonomous driving, which is an important technical support for realizing autonomous driving functions. Semantic segmentation of point clouds is a key part of autonomous driving technology.

[0003] Existing point cloud segmentation methods are mainly divided into three categories: voxel-based methods, point-based methods, and view-based methods. Voxel-based methods divide the three-dimensional space into multiple volume grids with specific sizes and discrete coordinates, and then perform feature extraction and classification on each voxel. A common method is to aggregate the points within each voxel to generate a voxel feature vector, and then use a three-dimensional convolutional neural network or a sparse convolutional neural network for processing. Due to the sparsity and disorder of point clouds, blindly discretizing the three-dimensional space into voxels at high resolutions will result in too many empty voxels, leading to a decrease in computational efficiency. To overcome the problem of high computational resource requirements for voxelization, point-based segmentation methods emerged. This type of method can better retain the fine structural information of point clouds by directly processing the original point cloud data, and is more effective than voxel-based methods when dealing with sparse point cloud data. Due to the characteristics of the disorder and sparsity of point cloud data, the network requires high computational costs for preprocessing operations on point clouds. View-based segmentation methods provide another idea. By projecting point clouds onto two-dimensional views, although some spatial information is lost, the required computational resources are significantly reduced. View-based methods project LiDAR point clouds onto a two-dimensional plane to generate a bird's-eye view, a range view, or other viewpoint images, and then perform semantic segmentation on them. This type of method uses the image segmentation technology of convolutional neural networks to process the projected images, projects three-dimensional point clouds into two-dimensional images to reduce the data dimension, and improves efficiency by using mature two-dimensional image segmentation methods.

[0004] There are differences between the depth map generated through spherical projection and a conventional image. For a conventional image, the feature distribution is uniform and not affected by spatial position; for the projected depth map, its features are transformed through spherical projection, which introduces a strong spatial prior, making the distribution of depth map features vary greatly at different positions, resulting in the possible inability to effectively capture information at different scales when dealing with complex scenes. At the same time, in a deep network, features may gradually be lost during the transmission process, especially after multiple layers of convolution and multiple activation functions, leading to gradient disappearance or information loss, thus affecting the learning effect of the model.

[0005] Generally speaking, these methods have the following common drawbacks: the limited receptive field makes it difficult for the model to effectively capture multi-scale context information; the uneven weight distribution during feature learning limits the expressive power of the model; the complex preprocessing and feature extraction processes further reduce the computational efficiency, which is not conducive to the real-time requirements in practical applications. These defects limit the performance of existing point cloud semantic segmentation methods in dynamic scenarios and large-scale point cloud processing. Summary of the Invention

[0006] In order to at least partially solve the problems in existing point cloud semantic segmentation methods, such as the limited receptive field making it difficult for the model to effectively capture multi-scale context information, the uneven weight distribution during feature learning, and the complex preprocessing and feature extraction processes, the present invention provides a multi-scale-based semantic segmentation method and system for autonomous driving LiDAR point clouds. The present invention fuses features of different scales by setting a multi-scale dilation fusion module to enhance the receptive field and effectively capture local and global information. Then, dynamic weights are calculated, and the fused features and the original input features are processed and subjected to residual connection according to the dynamic weights, realizing uniform weight distribution during feature learning. Finally, the processed data is input into the SqueezeSegV3 network to complete semantic segmentation. The present invention simplifies the overall preprocessing and feature extraction processes and improves the accuracy and effect of laser LiDAR point cloud semantic segmentation.

[0007] To achieve the above object, the technical solution of the present invention is as follows:

[0008] The first aspect of the present invention proposes a multi-scale-based semantic segmentation method for autonomous driving LiDAR point clouds, including:

[0009] Step 1: Input the target point cloud features into a multi-scale fusion module to obtain features of multiple different scales;

[0010] Step 2: Use the multi-scale dilation fusion module to fuse the features of multiple different scales to obtain multi-scale fusion features, which is convenient for capturing rich local and global information;

[0011] Step 3: Use the dynamic residual connection algorithm to process the target point cloud features and the multi-scale fusion features to obtain residual connection features, which are used to improve the comprehensive learning ability of detailed information and overall context;

[0012] Step 4: Input the residual connection features into the SqueezeSegV3 network model to complete semantic segmentation.

[0013] Further, the specific content of Step 1 includes:

[0014] The multi-scale fusion module includes multiple convolutional layers, and the multiple convolutional layers respectively correspond to different dilation rates, and each dilation rate corresponds to one channel, which is convenient for capturing features of different scales and subsequent feature fusion.

[0015] Furthermore, the multi-scale dilation fusion module includes a splicing sub-module, a convolutional layer, and a batch normalization sub-module;

[0016] The splicing sub-module is used to splice features of multiple different scales to obtain spliced features;

[0017] The convolutional layer is used to reduce the dimension of the spliced features to obtain the dimension-reduced features; among them, the convolutional layer includes a 1×1 convolutional layer; the number of channels is unified through the convolutional layer;

[0018] The batch normalization sub-module is used to perform batch normalization processing on the dimension-reduced features to obtain multi-scale fusion features, which is convenient for stable training, accelerating convergence, and at the same time avoiding gradient explosion or disappearance.

[0019] Furthermore, the batch normalization sub-module is expressed by the following formula:

[0020]

[0021] Among them, is the multi-scale fusion feature, X n,c,h,w is the dimension-reduced feature, μ fusion is the average value of the dimension-reduced features, is the variance of the dimension-reduced features, N is the batch size, C1, C2, C3, and C4 are different channels respectively, H is the height, and W is the width.

[0022] Furthermore, the specific steps of step three include:

[0023] Store the target point cloud features and the multi-scale fusion features, and perform 1×1 convolution operations on the target point cloud features and the multi-scale fusion features to obtain the output features after 1×1 convolution of the target point cloud features and the output features after 1×1 convolution of the multi-scale fusion features, which is convenient for keeping the dimensions of the target point cloud features and the fusion features consistent;

[0024] Calculate the dynamic weights of the output features after 1×1 convolution of the target point cloud features and the output features after 1×1 convolution of the multi-scale fusion features, which is convenient for uniform distribution of weights during feature learning;

[0025] Perform dynamic residual connection on the output features after 1×1 convolution of the target point cloud features and the output features after 1×1 convolution of the multi-scale fusion features according to the dynamic weights.

[0026] Further, the 1×1 convolution operation on the target point cloud features and the multi-scale fusion features is expressed by the following formula:

[0027]

[0028] where Y b,c′,h,w is the output feature after 1×1 convolution, X b,c,h,w is the input feature, W c′,c is the weight matrix of the 1×1 convolution kernel, b c′ is the bias value corresponding to each output channel C′, c′ is the output channel, c is the input channel, h is the height, w is the width, and N is the batch size.

[0029] Further, the dynamic weight is expressed by the following formula:

[0030]

[0031] where α is the dynamic weight, σ singmoid is the Sigmoid activation function, Wx + b is the output feature after 1×1 convolution of the multi-scale fusion feature, μ is the average value of the normalized feature, σ 2 is the variance after normalization, and ∈ is a positive number close to 0.

[0032] Further, the dynamic residual connection of the output feature after 1×1 convolution of the target point cloud features and the output feature after 1×1 convolution of the multi-scale fusion features according to the dynamic weight is specifically expressed by the following formula:

[0033]

[0034] where F is the residual connection feature, is the output feature after 1×1 convolution of the multi-scale fusion feature, is the output feature after 1×1 convolution of the target point cloud features.

[0035] In the second aspect of the present invention, a multi-scale-based autonomous driving LiDAR point cloud semantic segmentation system is proposed, including:

[0036] A feature extraction module for inputting target point cloud features into a multi-scale fusion module to obtain features of multiple different scales;

[0037] A fusion module for fusing features of multiple different scales by using a multi-scale dilation fusion module to obtain multi-scale fusion features, facilitating the capture of rich local and global information;

[0038] A residual connection module is used to process the target point cloud features and multi-scale fusion features by using a dynamic residual connection algorithm to obtain residual connection features, which is used to improve the comprehensive learning ability of detailed information and overall context.

[0039] A segmentation module is used to input the residual connection features into the SqueezeSegV3 network model to complete semantic segmentation.

[0040] Advantages of the present invention:

[0041] (1) Aiming at the problems of missing regional information and uneven distribution of feature learning weights in existing deep network models in the field of autonomous driving, the present invention proposes a multi-scale based semantic segmentation method for LiDAR point clouds in autonomous driving. Through the innovatively designed multi-scale dilation fusion module and dynamic residual connection algorithm, the present invention significantly improves the feature extraction ability and learning efficiency of the model in the point cloud semantic segmentation task. The present invention designs a multi-scale dilation fusion module MSDFusion, which captures rich local and global information through different receptive field ranges, and solves the problem of missing regional information caused by a single receptive field in traditional methods. The present invention improves the traditional residual connection algorithm, proposes a dynamic residual connection algorithm AeightRC, introduces a learnable weight mechanism, effectively balances the relationship between the target point cloud features and the multi-scale fusion features, and thus improves the model's comprehensive learning ability of detailed information and overall context. Finally, the present invention fuses the MSDFusion module with the AeightRC algorithm to form a complete network model for training. Experiments show that this solution performs well in complex scenarios and can adapt to diverse point cloud data distributions. Generally speaking, the present invention simplifies the preprocessing and feature extraction processes, improves the accuracy and effect of LiDAR point cloud semantic segmentation, and provides key support for autonomous driving.

[0042] (2) The present invention proposes a multi-scale dilation fusion module (MSDFusion), which uses convolutional layers with different dilation rates to capture local and global features at different scales, and reduces the dimensionality through convolutional operations after splicing the features, reducing the computational amount and improving the feature expression ability, and combining batch normalization methods to stabilize training and avoid gradient problems.

[0043] (3) The present invention proposes a dynamic residual connection algorithm (AeightRC), which improves the traditional residual connection algorithm, dynamically adjusts the importance of input and output features through weight learning, uses an activation function to calculate feature weights, and dynamically balances the learning of local and global features; combines multi-scale dilation features to achieve dynamic residual connection, improving the robustness and segmentation accuracy of the model. Description of the drawings

[0044] Figure 1Flowchart of a multi-scale based semantic segmentation method for autonomous driving LiDAR point clouds provided by an embodiment of the present invention.

[0045] Figure 2 Schematic diagram of the segmentation result of a multi-scale based semantic segmentation method for autonomous driving LiDAR point clouds provided by an embodiment of the present invention.

[0046] Figure 3 Schematic diagram of the multi-scale parallel branch structure of a multi-scale dilation fusion module provided by an embodiment of the present invention.

[0047] Figure 4 Schematic diagram of the dynamic residual connection algorithm provided by an embodiment of the present invention.

[0048] Figure 5 Flowchart of a multi-scale based semantic segmentation system for autonomous driving LiDAR point clouds provided by an embodiment of the present invention. Detailed implementation manners

[0049] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0050] Embodiment 1

[0051] As Figure 1 shown, a multi-scale based semantic segmentation method for autonomous driving LiDAR point clouds includes:

[0052] S101: Input the target point cloud features into a multi-scale fusion module to obtain features of multiple different scales.

[0053] Specifically, the multi-scale fusion module includes multiple convolutional layers, and the multiple convolutional layers respectively correspond to different dilation rates, and each dilation rate corresponds to one channel. By setting convolutional layers with different dilation rates for the projection image pixel range, different receptive field ranges can be obtained according to the dilation rate size, and the number of channels can be divided according to the number of dilation rates. While capturing features of different scales, the dimensions of two features can be aligned during the subsequent feature fusion process, avoiding the problem of channel number mismatch after the fusion process.

[0054] S102: Use a multi-scale dilation fusion module to fuse the features of multiple different scales to obtain multi-scale fusion features.

[0055] S103: Process the target point cloud features and multi-scale fusion features using the dynamic residual connection algorithm to obtain the residual connection features.

[0056] S104: Input the residual connection features into the SqueezeSegV3 network model to complete semantic segmentation.

[0057] Specifically, combine the multi-scale dilation fusion module (MSDFusion) with the dynamic residual connection algorithm (AeightRC) and apply them to the SqueezeSegV3 network model to obtain a new network model, and complete semantic segmentation according to the new network model.

[0058] Preferably, use the autonomous driving point cloud dataset SemanticKITTI to train the new network model, and verify the semantic segmentation accuracy and effect of the model.

[0059] In the present invention, features of different scales are extracted through the multi-scale fusion module, and then the features of different scales are fused using the multi-scale dilation fusion module to obtain multi-scale fusion features. Then, the dynamic residual connection algorithm is used to perform residual connection between the target point cloud features (original input features) and the multi-scale fusion features. Finally, the residual connection features are input into the SqueezeSegV3 network model to complete semantic segmentation. The present invention improves the learning ability of local features and global context, enhances the segmentation effect of the network model in real complex scenarios, and the results are as Figure 2 shown. Among them, Figure 2 (a) is the segmentation result diagram of the basic model; Figure 2 (b) is the segmentation result diagram of the present invention; Figure 2 (c) On the right side is the original point cloud data, and on the left side is the point cloud segmentation effect diagram.

[0060] Embodiment 2

[0061] Based on the above embodiments, the present invention proposes the structure of the multi-scale dilation fusion module (MSDFusion), which specifically includes:

[0062] The multi-scale dilation fusion module includes a splicing sub-module, a convolutional layer, and a batch normalization sub-module. The splicing sub-module is used to splice features of multiple different scales to obtain the spliced features. The convolutional layer is used to reduce the dimension of the spliced features to obtain the dimension-reduced features; among them, the convolutional layer includes a 1×1 convolutional layer. The batch normalization sub-module is used to perform batch normalization processing on the dimension-reduced features to obtain the multi-scale fusion features.

[0063] As Figure 3As shown, different from the feature pyramid used in traditional methods, MSDFusion uses convolutional layers with different dilation rates for parallel connection, and multiple convolutional layers can process the input feature map simultaneously. This method allows the model to capture local and global features at the same level, enhancing the model's ability to understand complex scenes.

[0064] For convolutional layers with different dilation rates, they respectively correspond to different receptive field ranges, and the receptive field range calculation formula is as follows:

[0065] R i =(K - 1)·D i + 1

[0066] Among them, R i is the receptive field size of the i-th convolutional layer, K is the convolutional kernel size of the i-th convolutional layer, and D i is the dilation rate of the i-th convolutional layer.

[0067] After obtaining the output feature tensors of each convolutional layer, a concatenation operation is performed on these multiple tensors to obtain the concatenated features. The batch size N, height H, and width W in each tensor are fixed values, so the number of channels C needs to be consistent with the total number of channels during the concatenation process. The shape of the concatenated tensor can be expressed as where Y is the concatenated tensor, R is a real number, and C1, C2, C3, and C4 are different channels respectively.

[0068] Pass the concatenated features through a 1×1 convolutional layer to reduce the dimension and transform the feature map, so that the dimension of the concatenated feature map is consistent with the features of different scales. After the dimension reduction operation, batch normalization is performed on the concatenated features to help stabilize training, accelerate convergence, and avoid gradient explosion or disappearance. MSDFusion applies batch normalization to the concatenated feature map to calculate the mean and variance of the overall features, rather than normalizing the output of each convolutional layer separately.

[0069] First, calculate the average value of the concatenated features, that is:

[0070]

[0071] where μ fusion is the average value of the dimension-reduced features, N is the batch size, C1, C2, C3, and C4 are different channels respectively, H is the height, W is the width, and X n,c,h,w is the dimension-reduced feature.

[0072] Calculate the variance of the concatenated features after dimension reduction (the dimension-reduced features), which is specifically expressed by the following formula:

[0073]

[0074] Among them, is the variance of the features after dimensionality reduction.

[0075] A normalization operation is performed, which is specifically expressed by the following formula:

[0076]

[0077] Among them, is the multi-scale fusion feature.

[0078] After the above three steps, the multi-scale feature fusion process is completed, and the fused features are output for the residual connection operation.

[0079] Embodiment 3

[0080] Based on the above embodiment, as Figure 4 shown, the present invention proposes the specific process of the dynamic residual connection algorithm (AeightRC), which specifically includes:

[0081] To solve the degradation problem in the deep neural network, it is necessary to perform a residual connection between the fused features and the target point cloud features. The present invention improves the traditional residual connection algorithm, learns the weights of the multi-scale fusion features obtained by the multi-scale dilation fusion module and the target point cloud features, and obtains dynamic weights, so that the network can dynamically balance the importance of input and output features.

[0082] First, the target point cloud features and multi-scale fusion features are stored. The dimension of the target point cloud features is (I, S, S), representing a feature map with I channels and a resolution of S×S. Then, the feature map is subjected to a 1×1 convolution operation to linearly combine the features of different channels and generate a new feature representation:

[0083]

[0084] Among them, Y b,c′,h,w is the output feature after the 1×1 convolution, X b,c,h,w is the input feature, W c′,c is the weight matrix of the 1×1 convolution kernel, b c′ is the bias value corresponding to each output channel C′, c′ is the output channel, c is the input channel, h is the height, w is the width, and N is the batch size.

[0085] Dynamic weight calculation is performed according to the output feature after the 1×1 convolution of the scale fusion feature:

[0086]

[0087] Among them, α is the dynamic weight, σ singmoidis the Sigmoid activation function, W·x + b is the output feature after 1×1 convolution of the multi-scale fusion feature, μ is the average value of the normalized feature, σ 2 is the variance after normalization, ∈ is a positive number close to 0, and x is the input of the activation function.

[0088] Finally, perform dynamic residual connection on the output feature after 1×1 convolution of the target point cloud feature and the output feature after 1×1 convolution of the multi-scale fusion feature. The specific implementation is as shown in the formula:

[0089]

[0090] where F is the residual connection feature, is the output feature after 1×1 convolution of the multi-scale fusion feature, is the output feature after 1×1 convolution of the target point cloud feature.

[0091] Embodiment 4

[0092] Based on the above embodiments, as Figure 5 shown, an embodiment of the present invention provides a multi-scale based LiDAR point cloud semantic segmentation system for autonomous driving, including:

[0093] A feature extraction module for inputting the target point cloud feature into the multi-scale fusion module to obtain features of multiple different scales.

[0094] A fusion module for fusing features of multiple different scales by using a multi-scale dilation fusion module to obtain multi-scale fusion features.

[0095] A residual connection module for processing the target point cloud feature and the multi-scale fusion feature by using a dynamic residual connection algorithm to obtain a residual connection feature.

[0096] A segmentation module for inputting the residual connection feature into the SqueezeSegV3 network model to complete semantic segmentation.

[0097] It should be noted that the multi-scale based LiDAR point cloud semantic segmentation system for autonomous driving provided by the embodiment of the present invention is to implement the above-mentioned multi-scale based LiDAR point cloud semantic segmentation method for autonomous driving. Its functions can be specifically referred to the above method embodiments and will not be elaborated here.

[0098] In summary, in view of the problems of missing regional information and uneven distribution of feature learning weights in existing deep network models in the field of autonomous driving, the present invention proposes a multi-scale semantic segmentation method for LiDAR point clouds in autonomous driving. Through the innovatively designed multi-scale dilation fusion module and dynamic residual connection algorithm, the present invention significantly improves the feature extraction ability and learning efficiency of the model in the point cloud semantic segmentation task. The present invention designs a multi-scale dilation fusion module MSDFusion, which captures rich local and global information through different receptive field ranges, and solves the problem of missing regional information caused by a single receptive field in traditional methods. The present invention improves the traditional residual connection algorithm, proposes a dynamic residual connection algorithm AeightRC, introduces a learnable weight mechanism, effectively balances the relationship between the target point cloud features and the multi-scale fusion features, and thus improves the model's comprehensive learning ability of detailed information and overall context. Finally, the present invention fuses the MSDFusion module with the AeightRC algorithm to form a complete network model for training. Experiments show that this solution performs excellently in complex scenarios and can adapt to diverse point cloud data distributions. Generally speaking, the present invention simplifies the preprocessing and feature extraction processes, improves the accuracy and effect of LiDAR point cloud semantic segmentation, and provides key support for autonomous driving. The present invention proposes a multi-scale dilation fusion module (MSDFusion), which uses convolutional layers with different dilation rates to capture local and global features at different scales. After splicing the features, dimensionality reduction is performed through convolutional operations, reducing the computational amount and improving the feature expression ability. Combining with the batch normalization method, the training is stabilized and the gradient problem is avoided. The present invention proposes a dynamic residual connection algorithm (AeightRC), which improves the traditional residual connection algorithm, dynamically adjusts the importance of input and output features through weight learning, uses an activation function to calculate feature weights, and dynamically balances the learning of local and global features; combines multi-scale dilation features to achieve dynamic residual connection, improving the robustness and segmentation accuracy of the model.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multi-scale semantic segmentation method for LiDAR point cloud for autonomous driving, characterized in that: include: Step 1: Input the target point cloud features into the multi-scale fusion module to obtain features of multiple different scales; Step 2: Use the multi-scale expansion fusion module to fuse features of multiple different scales to obtain multi-scale fusion features; Step 3: Use the dynamic residual connection algorithm to process the target point cloud features and multi-scale fusion features to obtain residual connection features; Step 4: Input the residual connection features into the SqueezeSegV3 network model to complete semantic segmentation.

2. The multi-scale autonomous driving LiDAR point cloud semantic segmentation method according to claim 1, characterized in that: The step 1 specifically includes: The multi-scale fusion module includes multiple convolutional layers, each of which corresponds to a different expansion rate, and each expansion rate corresponds to a channel.

3. The multi-scale autonomous driving LiDAR point cloud semantic segmentation method according to claim 1, characterized in that: The multi-scale expansion fusion module includes a splicing submodule, a convolutional layer, and a batch normalization submodule; The splicing submodule is used to splice features of multiple different scales to obtain splicing features; The convolution layer is used to reduce the dimension of the spliced ​​features to obtain the reduced-dimensional features; wherein the convolution layer includes a 1×1 convolution layer; The batch normalization submodule is used to perform batch normalization processing on the features after dimensionality reduction to obtain multi-scale fusion features.

4. The multi-scale autonomous driving LiDAR point cloud semantic segmentation method according to claim 3, characterized in that: The batch normalization submodule is expressed as follows: in, is the multi-scale fusion feature, X n,c,h,w is the feature after dimensionality reduction, μ fusion is the average value of the feature after dimensionality reduction, is the variance of the feature after dimensionality reduction, N is the batch size, C1, C2, C3 and C4 are different channels, H is the height, and W is the width.

5. The multi-scale autonomous driving LiDAR point cloud semantic segmentation method according to claim 1, characterized in that: The step three specifically includes: The target point cloud features and the multi-scale fusion features are stored, and a 1×1 convolution operation is performed on the target point cloud features and the multi-scale fusion features to obtain the output features of the target point cloud features after the 1×1 convolution and the output features of the multi-scale fusion features after the 1×1 convolution; Calculate the dynamic weights of the output features of the target point cloud features after 1×1 convolution and the output features of the multi-scale fusion features after 1×1 convolution; According to the dynamic weights, dynamic residual connections are performed on the output features of the target point cloud features after 1×1 convolution and the output features of the multi-scale fusion features after 1×1 convolution.

6. The multi-scale autonomous driving LiDAR point cloud semantic segmentation method according to claim 5, characterized in that: The 1×1 convolution operation on the target point cloud features and the multi-scale fusion features is expressed by the following formula: Among them, Y b,c′,h,w is the output feature after 1×1 convolution, X b,c,h,w is the input feature, W c′,c is the weight matrix of the 1×1 convolution kernel, b c′ is the bias value corresponding to each output channel C′, c′ is the output channel, c is the input channel, h is the height, w is the width, and N is the batch size.

7. The multi-scale autonomous driving LiDAR point cloud semantic segmentation method according to claim 5, characterized in that: The dynamic weight is expressed by the following formula: Among them, α is the dynamic weight, σ singmoid is the Sigmoid activation function, Wx+b is the output feature after 1×1 convolution of multi-scale fusion features, μ is the normalized feature average, σ 2 is the normalized variance, ∈ is a positive number close to 0.

8. The multi-scale autonomous driving LiDAR point cloud semantic segmentation method according to claim 7, characterized in that: The output features of the target point cloud features after 1×1 convolution and the output features of the multi-scale fusion features after 1×1 convolution are dynamically connected according to the dynamic weights, which is specifically expressed by the following formula: Among them, F is the residual connection feature, It is the output feature after 1×1 convolution of multi-scale fusion features. It is the output feature of the target point cloud feature after 1×1 convolution.

9. A multi-scale autonomous driving LiDAR point cloud semantic segmentation system, characterized in that: include: The feature extraction module is used to input the target point cloud features into the multi-scale fusion module to obtain features of multiple different scales; A fusion module is used to fuse features of multiple scales using a multi-scale expansion fusion module to obtain multi-scale fusion features; The residual connection module is used to process the target point cloud features and the multi-scale fusion features using a dynamic residual connection algorithm to obtain residual connection features; The segmentation module is used to input the residual connection features into the SqueezeSegV3 network model to complete semantic segmentation.

Citation Information

Patent Citations

  • Multi-threat target reconstruction and situation awareness method based on generative adversarial network

    CN110969637A

  • Image segmentation method, system and equipment

    CN116129124A

  • Connected double-attention multi-scale fusion semantic segmentation network

    CN116630626A

  • Aviation laser point cloud semantic segmentation method and device based on multistage context feature fusion network

    CN116824585A

  • Three-dimensional point cloud semantic segmentation method based on multi-scale feature jump fusion

    CN117456172A