A LiDAR point cloud semantic segmentation method based on asymmetric convolution

By adopting asymmetric convolutional backbone network and context feature enhancement module in the semantic segmentation model of lidar point cloud, the shortcomings of existing models in feature extraction and detection accuracy are solved, and higher semantic segmentation accuracy and robustness of rotation targets are achieved.

CN115937850BActive Publication Date: 2025-06-24SHANGHAI SECOND POLYTECHNIC UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211425576.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-15
Publication Date
2025-06-24
Estimated Expiration
2042-11-15

AI Technical Summary

Technical Problem

The existing lidar point cloud semantic segmentation model has shortcomings in inference speed and detection accuracy, especially in the feature extraction stage, where the feature information in the point cloud is not fully discovered and utilized, resulting in limited detection accuracy.

Method used

A semantic segmentation method of lidar point cloud based on asymmetric convolution is proposed, using asymmetric convolution backbone network and context feature enhancement module. In the feature extraction stage, the interference of the rotation target on the segmentation is reduced by asymmetric convolution, and the higher-order context feature information is extracted through the context feature enhancement module.

Benefits of technology

Without increasing the inference time, the performance and generalization ability of the model are improved, the semantic segmentation accuracy of the point cloud is improved, and the robustness of the rotation target is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115937850B_ABST
    Figure CN115937850B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for lidar point cloud semantic segmentation based on asymmetric convolution, which specifically relates to the field of computer vision technology, and includes the following steps: Step 1: Input the lidar point cloud data into a projection-based point cloud encoder to obtain the encoded point cloud; Step 2: Input the encoded point cloud into an asymmetric convolution backbone network to extract features of the point cloud; Step 3: Input the extracted point cloud features into a context feature enhancement module to extract context feature information in the point cloud features; Step 4: Input the feature information of the point cloud into a semantic classifier to identify the semantic category of each point cloud, and finally obtain the semantic segmentation result of the entire point cloud data. The asymmetric convolution backbone network and context feature enhancement module proposed by the present invention can effectively improve the ability to extract point cloud features and improve the semantic segmentation accuracy of the point cloud.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology. More specifically, the present invention relates to a method for lidar point cloud semantic segmentation based on asymmetric convolution. Background Art

[0002] Achieving scene understanding of the surrounding environment is one of the most important tasks in the field of autonomous driving. Although semantic segmentation of two-dimensional images is an important step towards achieving scene understanding, pure vision sensors have some limitations, such as being unable to effectively obtain information under low-light conditions, lacking depth information, and having a limited field of view. In contrast, lidar can obtain accurate depth information with high density and wide viewing angle regardless of the lighting conditions, making it a more reliable information source for the environmental perception part of autonomous driving. Therefore, in recent years, research on lidar point cloud semantic segmentation has become a hot topic.

[0003] On the premise of ensuring real-time model inference, it is of great practical significance to improve the detection accuracy of the model. Current lidar point cloud semantic segmentation models can be divided into three major categories according to different point cloud encoding methods: point-based methods, voxel-based methods, and projection-based methods. In terms of inference speed, the large computational cost and memory consumption make it difficult for point-based and voxel-based methods to achieve real-time inference results, and difficulties will be encountered in actual deployment and application. In contrast, projection-based methods have a significant advantage in inference speed and can achieve real-time inference results during deployment. In terms of detection accuracy, although projection-based methods have achieved certain results, since the feature information in the point cloud is not fully explored and utilized during the feature extraction stage, there is still room for improvement in the detection accuracy of lidar point cloud semantic segmentation models. Summary of the Invention

[0004] To overcome the shortcomings of the prior art, the present invention proposes a method for lidar point cloud semantic segmentation based on asymmetric convolution. The method includes an asymmetric convolution backbone network and a context feature enhancement module. In the feature extraction stage, the backbone network is constructed by replacing square convolution with asymmetric convolution to enhance the backbone part in the convolution, thereby reducing the interference caused by the rotation of the target to segmentation. By designing the context feature enhancement module, the decomposition and re-aggregation of features are performed to achieve the full extraction of context feature information. These methods further improve the performance and generalization ability of the model without increasing the inference time.

[0005] To achieve the above object, the present invention provides the following technical solution: A method for lidar point cloud semantic segmentation based on asymmetric convolution, comprising the following steps:

[0006] Step 1: Input lidar point cloud data into a projection-based point cloud encoder to obtain encoded point cloud;

[0007] Step 2: Input the encoded point cloud into the asymmetric convolutional backbone network to extract features of the point cloud;

[0008] Step 3: Input the extracted point cloud features into the context feature enhancement module to extract the context feature information in the point cloud features;

[0009] Step 4: Input the feature information of the point cloud into the semantic classifier to identify the semantic category of each point cloud, and finally obtain the semantic segmentation result of the entire point cloud data;

[0010] Among them, in Step 2, the asymmetric convolutional backbone network is composed of four downsampling asymmetric convolutional modules and four upsampling asymmetric convolutional modules; three skip connections are also used to cascade the upsampling results (low-level features) with the corresponding downsampling, effectively fusing the network's low-level features and high-level features, and improving the model's learning ability for detailed information;

[0011] In one downsampling process, first perform a square convolution operation on the features with a stride of 2, and then perform convolution operations using two groups of asymmetric convolution combinations respectively, and then add the operation results and output; these two groups of asymmetric convolution combinations are respectively composed of an asymmetric convolution kernel group of 3×1 and 1×3 and an asymmetric convolution kernel group of 1×3 and 3×1;

[0012] The formula for the downsampling asymmetric convolutional module is:

[0013] F out =C 3×1 (C 1×3 (C 3×3 (F in )))+C 1×3 (C 3×1 (C 3×3 (F in ))) (1)

[0014] In the formula, F in and F out are the input feature and the output feature respectively, and C 3×3 , C 1×3 and C 3×1 are the convolution operations of 3×3, 1×3, and 3×1 respectively;

[0015] In one upsampling process, first perform bilinear interpolation on the features, then splice them with the low-order features of the skip connection, and finally perform convolution operations on the features using a group of asymmetric convolutions, and this asymmetric convolution group is composed of asymmetric convolution kernels of 1×3 and 3×1;

[0016] The formula for the upsampling asymmetric convolutional module is:

[0017] F out = C 3×1 (C 1×3 (Δ(B(F in ), F low ))) (2)

[0018] In the formula, F in , F out and F low are the input feature, output feature and low-order feature respectively, C 3×3 , C 1×3 and C 3×1 are the convolution operations of 3×3, 1×3 and 3×1 respectively, Δ is the feature concatenation operation, and B is the bilinear interpolation operation.

[0019] In a preferred embodiment, the context feature enhancement module in the third step is composed of the following steps:

[0020] Step 3-1: Use a low-rank convolution kernel to generate low-rank encodings by dimension decomposition of the high-rank context features respectively;

[0021] Step 3-2: Use the Sigmoid function to activate the convolution results respectively and then add them;

[0022] Step 3-3: Multiply with the context features before processing to obtain the enhanced context features;

[0023] The calculation process of the context feature enhancement module is shown in the following formula:

[0024] F out = F in ·(Sig(C 3×1 (F in )) + Sig(C 1×3 (F in ))) (3)

[0025] In the formula, F in and F out are the input feature and output feature respectively, Sig is the Sigmoid function, C 3×1 and C 1×3 are the convolution operations of 3×1 and 1×3 respectively.

[0026] Compared with the prior art, the technical effects and advantages of the present invention:

[0027] 1. It can be directly applied to the original point cloud data, and can achieve high semantic segmentation accuracy of the point cloud;

[0028] 2. The proposed asymmetric convolution backbone network can improve the robustness to rotating targets;

[0029] 3. The proposed context feature enhancement module can avoid the difficulty of directly extracting high-rank information by further extracting features in different dimensions using low-rank convolutional kernels, effectively extract the feature information in high-order contexts, and improve the segmentation accuracy of point clouds. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic diagram of the overall framework of the present invention;

[0031] Figure 2 It is a schematic diagram of the overall process of the present invention;

[0032] Figure 3 It is a schematic diagram of the asymmetric convolutional backbone network of the present invention;

[0033] Figure 4 It is a schematic diagram of the asymmetric downsampling module of the present invention;

[0034] Figure 5 It is a schematic diagram of the asymmetric upsampling module of the present invention;

[0035] Figure 6 It is a schematic diagram of the context feature enhancement module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0037] According to Figure 1-2 A lidar point cloud semantic segmentation method based on asymmetric convolution shown below includes the following steps:

[0038] Step 1: Input the lidar point cloud data into the projection-based point cloud encoder to obtain the encoded point cloud;

[0039] Step 2: Input the encoded point cloud into the asymmetric convolutional backbone network to extract features from the point cloud;

[0040] Step 3: Input the extracted point cloud features into the context feature enhancement module to extract the context feature information in the point cloud features;

[0041] Step 4: Input the feature information of the point cloud into the semantic classifier to identify the semantic category of each point cloud, and finally obtain the semantic segmentation result of the entire point cloud data;

[0042] As Figure 3 shown:Figure 3 Schematic diagram of the asymmetric convolution backbone network of the present invention; among them, in step two, the asymmetric convolution backbone network is composed of four downsampling asymmetric convolution modules and four upsampling asymmetric convolution modules; three skip connections are also used to cascade the upsampling results (low-level features) with the corresponding downsampling, effectively fusing the low-level features and high-level features of the network and improving the model's learning ability for detailed information;

[0043] As shown in Figure 4 the figure, Figure 4 Schematic diagram of the asymmetric downsampling module of the present invention; during one downsampling process, first perform a square convolution operation on the feature with a stride of 2, and then perform convolution operations with two groups of asymmetric convolution combinations respectively, and then add the operation results and output; these two groups of asymmetric convolution combinations are respectively composed of an asymmetric convolution kernel group of 3×1 and 1×3 and an asymmetric convolution kernel group of 1×3 and 3×1;

[0044] The formula for the downsampling asymmetric convolution module is:

[0045] F out = C 3×1 (C 1×3 (C 3×3 (F in ))) + C 1×3 (C 3×1 (C 3×3 (F in ))) (1)

[0046] In the formula, F in and F out are the input feature and the output feature respectively, and C 3×3 , C 1×3 and C 3×1 are the convolution operations of 3×3, 1×3 and 3×1 respectively;

[0047] As shown in Figure 5 the figure, Figure 5 Schematic diagram of the asymmetric upsampling module of the present invention; during one upsampling process, first perform bilinear interpolation on the feature, then splice it with the low-order feature of the skip connection, and finally perform convolution operation on the feature with a group of asymmetric convolutions, and this group of asymmetric convolutions is composed of asymmetric convolution kernels of 1×3 and 3×1;

[0048] The formula for the upsampling asymmetric convolution module is:

[0049] F out = C 3×1 (C 1×3 (Δ(B(F in ), F low ))) (2)

[0050] Where F in , F out and F low are the input feature, output feature, and low-order feature respectively, C 3×3 , C 1×3 and C 3×1 are the convolution operations of 3×3, 1×3, and 3×1 respectively, Δ is the feature concatenation operation, and B is the bilinear interpolation operation.

[0051] As Figure 6 shown, Figure 6 is the schematic diagram of the context feature enhancement module of the present invention; the context feature enhancement module in step 3 is composed of the following steps:

[0052] Step 3-1: Use a low-rank convolution kernel to generate low-rank encodings by dimension decomposition of the high-rank context features respectively;

[0053] Step 3-2: Use the Sigmoid function to activate the convolution results respectively and then add them;

[0054] Step 3-3: Multiply with the context features before processing to obtain enhanced context features;

[0055] The calculation process of the context feature enhancement module is shown in the following formula:

[0056] F out = F in · (Sig(C 3×1 (F in )) + Sig(C 1×3 (F in ))) (3)

[0057] Where F in and F out are the input feature and output feature respectively, Sig is the Sigmoid function, C 3×1 and C 1×3 are the convolution operations of 3×1 and 1×3 respectively.

[0058] In summary, the present invention provides a lidar point cloud semantic segmentation method based on asymmetric convolution, which can be directly applied to the original point cloud data and can achieve high accuracy of point cloud semantic segmentation; the proposed asymmetric convolution backbone network can improve the robustness to rotating targets; the proposed context feature enhancement module can avoid the difficulty of directly extracting high-rank information by further extracting features in different dimensions with a low-rank convolution kernel, effectively extract the feature information in high-order contexts, and improve the segmentation accuracy of the point cloud. Finally, the following points should be noted:

[0059] First, in the description of the present application, it should be noted that unless otherwise specified and defined, the terms "installed", "connected", and "coupled" should be understood in a broad sense, which can be a mechanical connection or an electrical connection, or can be the communication inside two components, and can be directly connected. The terms "upper", "lower", "left", "right", etc. are only used to represent the relative position relationship. When the absolute position of the object being described changes, the relative position relationship may change;

[0060] Secondly, in the drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments are involved. Other structures can refer to the general design. Without conflict, the same embodiment and different embodiments of the present invention can be combined with each other;

[0061] Finally, the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for lidar point cloud semantic segmentation based on asymmetric convolution, characterized in that: It includes the following steps: Step 1: Input the lidar point cloud data into a projection-based point cloud encoder to obtain the encoded point cloud; Step 2: Input the encoded point cloud into an asymmetric convolutional backbone network to extract features of the point cloud; Step 3: Input the extracted point cloud features into a context feature enhancement module to extract the context feature information in the point cloud features; Step 4: Input the feature information of the point cloud into a semantic classifier to identify the semantic category of each point cloud, and finally obtain the semantic segmentation result of the entire point cloud data; Among them, the asymmetric convolutional backbone network in Step 2 is composed of four downsampling asymmetric convolutional modules and four upsampling asymmetric convolutional modules; three skip connections are also used to cascade the upsampling result with the corresponding downsampling. The upsampling result is the low-level feature, effectively fusing the low-level and high-level features of the network and improving the model's learning ability for detailed information; In one downsampling process, first perform a square convolution operation on the feature with a stride of 2, and then perform convolution operations with two groups of asymmetric convolution combinations respectively, and then add the operation results and output; these two groups of asymmetric convolution combinations are respectively composed of an asymmetric convolution kernel group of 3×1 and 1×3 and an asymmetric convolution kernel group of 1×3 and 3×1; The formula for the downsampling asymmetric convolutional module is: F out = C 3×1 (C 1×3 (C 3×3 (F in )))+ C 1×3 (C 3×1 (C 3×3 (F in ))) (1) where F in and F out are the input feature and the output feature respectively, and C 3×3 , C 1×3 and C 3×1 are the convolution operations of 3×3, 1×3, and 3×1 respectively; In one upsampling process, first perform bilinear interpolation on the feature, then splice it with the low-order feature of the skip connection, and finally perform a convolution operation on the feature with a group of asymmetric convolutions. This asymmetric convolution group is composed of asymmetric convolution kernels of 1×3 and 3×1; The formula for the upsampling asymmetric convolutional module is: F out = C 3×1 (C 1×3 (Δ(B(F in ), F low ))) (2) where F in , F out and F low are the input feature, output feature, and low-order feature respectively, C 3×3 , C 1×3 and C 3×1 are the convolution operations of 3×3, 1×3, and 3×1 respectively, Δ is the feature concatenation operation, and B is the bilinear interpolation operation.

2. The method for lidar point cloud semantic segmentation based on asymmetric convolution according to claim 1, wherein: The context feature enhancement module in Step 3 is composed of the following steps: Step 3-1: Use a low-rank convolution kernel to generate low-rank encodings for the high-rank context features respectively based on dimension decomposition; Step 3-2: Use the Sigmoid function to activate the convolution results respectively and then add them; Step 3-3: Multiply with the context features before processing to obtain the enhanced context features; The calculation process of the context feature enhancement module is shown in the following formula: F out = F in ·(Sig(C 3×1 (F in )) + Sig(C 1×3 (F in ))) (3) where F in and F out are the input feature and the output feature respectively, Sig is the Sigmoid function, C 3×1 and C 1×3 are the convolution operations of 3×1 and 1×3 respectively.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method, system and equipment and storage medium

    CN114022785A

  • Three-dimensional point cloud semantic segmentation method and apparatus, and device and medium

    WO2022088676A1