A Semantic Segmentation Method for LiDAR Point Cloud Data in Adverse Weather
Through the dual-branch design and robust multi-level feature collaboration module, the semantic segmentation problem in severe weather conditions is solved, and efficient segmentation performance and generalization ability in severe weather is achieved.
Patent Information
- Application Number
- CN202510195353.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-02-21
AI Technical Summary
The existing semantic segmentation method of lidar point cloud data is insufficient in robustness and generalization capabilities under severe weather conditions, and cannot effectively deal with the domain offset between standard conditions and harsh conditions.
The dual-branch design is adopted to process spatial geometric structures and reflection intensity features respectively. The features can be extracted through sparse three-dimensional convolution and depth separation convolution, and local and global fusion is combined with the robust multi-level feature collaboration module to optimize the network loss function to generate semantic segmentation results.
Under severe weather conditions, the semantic segmentation performance of lidar point cloud data is significantly improved, effective generalization from the source domain to the target domain, and the robustness and generalization ability of the model are improved.
Smart Images

Figure CN119832251B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of semantic segmentation of lidar point cloud data. More specifically, it relates to a method for semantic segmentation of lidar point cloud data for adverse weather conditions. Background Art
[0002] Current research on lidar point cloud data is usually carried out under idealized conditions: using standard lidar point cloud datasets for training and testing. These lidar point cloud datasets usually exclude various interference factors that may be encountered in the real world, such as adverse weather conditions. The resulting idealized lidar point cloud data semantic segmentation system performs well in standard environments, but it is often difficult to generalize under adverse weather conditions. The reason is that the spatial geometric structure features and reflection intensity features are affected by the degradation caused by the weather, especially in adverse weather conditions such as fog, rain, snow, etc., with poor robustness and adaptability. Common feature unified processing methods cannot effectively cope with the domain shift between standard conditions and adverse conditions. Therefore, there is an urgent need for a domain generalization lidar point cloud data semantic segmentation method that can effectively separate and fuse geometric structure features and reflection intensity features while reducing the mutual degradation effect under adverse weather conditions. Summary of the Invention
[0003] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for semantic segmentation of lidar point cloud data for adverse weather conditions, which can be generalized to the target domain after source domain training and can still efficiently complete the semantic segmentation task of lidar point cloud data under adverse weather conditions.
[0004] To achieve the above invention purpose, the method for semantic segmentation of lidar point cloud data for adverse weather conditions of the present invention is characterized by including the following steps:
[0005] (1) Extraction of spatial geometric structure features
[0006] First, exclude the reflection intensity information from the obtained lidar point cloud data to obtain point cloud data containing only geometric coordinates, that is, geometric point cloud, and discretize the geometric point cloud into regular voxel grids. Then, use an encoder network composed of multiple cascaded sparse three-dimensional convolutional blocks to process the voxel grids, layer by layer to obtain spatial resolution and geometric context information at different scales. Finally, aggregate the spatial resolution and geometric context information at different scales to obtain spatial geometric structure features in a sparse representation form;
[0007] (2) Extraction of reflection intensity features
[0008] First, the acquired lidar point cloud data is converted into a two-dimensional view image using spherical projection, where the reflection intensity information of each point is used as the input feature. Second, an encoder network composed of two-dimensional depthwise separable convolutional blocks is used to process the reflection intensity information, and features of different scales are obtained layer by layer to form the reflection intensity feature;
[0009] (3) Robust Multi-level Feature Collaboration
[0010] 3.1) Calculate the complementary information constraint loss L for enhancing feature complementarity and robustness cic :
[0011] L cic = KL(p(m geo |f geo ) || r(m geo )) + KL(p(m ref |f ref ) || r(m ref )) - KL(p(m geo |f geo ) || p(m ref |f ref )) - KL(p(m ref |f ref ) || p(m geo |f geo ))
[0012] where KL(· || ·) is the maximum between-class scatter, f geo is the feature vector extracted from the spatial geometric structure feature, f ref is the feature vector extracted from the reflection intensity feature, m geo is the mean μ geo and standard deviation σ geo calculated based on the feature vector f geo through the addition of additive noise to generate the enhanced spatial geometric structure feature, m ref is the mean μ ref and standard deviation σ ref calculated based on the feature vector f ref , through the addition of additive noise to generate the enhanced reflection intensity feature:
[0013] m geo = μ geo + ε · σ geo
[0014] m ref = μ ref + ε · σ ref
[0015] r(m geo ) is to generate the enhanced feature m of the spatial geometric structure geo The formula is converted into a standard Gaussian prior distribution with additive noise. r(m ref ) is to generate the enhanced feature m of the reflection intensity ref The formula is converted into a standard Gaussian prior distribution with additive noise. p(m geo ∣f geo ) is the conditional distribution for measuring the spatial geometric structure feature, and p(m ref ∣f ref ) is the conditional distribution for measuring the reflection intensity feature;
[0016] 3.2), Local-level feature fusion
[0017] The mean μ geo calculated based on the feature vector f geo and the standard deviation σ geo and the mean μ ref calculated based on the feature vector f ref and the standard deviation σ ref are dynamically fused to obtain the local fusion feature F local :
[0018] F local =α·μ geo +(1 - α)·μ ref
[0019] where the fusion weight α is dynamically calculated according to the standard deviation σ geo and the standard deviation σ ref dynamically:
[0020]
[0021] where, and are respectively the means of the standard deviation σ geo and the standard deviation σ ref in the channel dimension;
[0022] 3.3), Global-level feature fusion
[0023] First, introduce a set of learnable global query values Q. The enhanced feature m of the spatial geometric structure geo is used as the key value K, value V of the cross-attention mechanism and the global query value Q are aggregated to obtain a temporary global feature value, and then the temporary global feature value is used as the key value K, value V of the cross-attention mechanism and the enhanced feature m of the reflection intensity ref is further aggregated to obtain the global fusion feature F global ;
[0024] 3.4), Concatenated into a joint representation F
[0025] The local fusion feature F local and the global fusion feature F global are concatenated to form the joint representation F;
[0026] (4) Optimize the network
[0027] The joint representation F is fed into the 3D voxel decoding module to obtain the semantic segmentation result, and the network is optimized by combining the cross-entropy loss and the complementary information constraint loss. The loss L of the optimized network is:
[0028] L = L CE + βL cic
[0029] where L CE is the cross-entropy loss calculated from the semantic segmentation result obtained by the 3D voxel decoding module and the label, and β is the combination coefficient;
[0030] (5) Semantic segmentation of lidar point cloud data
[0031] The joint representation F is obtained according to steps (1) to (3) and fed into the 3D voxel decoding module to obtain the semantic segmentation result.
[0032] The object of the present invention is achieved in this way.
[0033] The semantic segmentation method for lidar point cloud data in adverse weather of the present invention adopts a dual-branch design. The spatial geometric structure branch is responsible for capturing the spatial geometric form to obtain the spatial geometric structure feature, and the reflection intensity branch focuses on the reflection intensity of the lidar signal to obtain the reflection intensity feature. Then, a robust multi-level feature collaboration module is used for feature fusion at different scales, thereby reducing the impact of adverse weather on the segmentation performance. The spatial geometric structure feature is processed by voxelization and extracted using a sparse 3D convolutional network; the reflection intensity information is transformed into a 2D image through spherical projection and extracted through depthwise separable convolution. The fused spatial geometric structure feature and reflection intensity feature are further optimized through the local and global feature fusion mechanism, and finally an accurate semantic segmentation result is generated through the 3D voxel decoding module. The present invention can be generalized to the target domain after training in the source domain, effectively cope with the domain shift between the standard condition and the adverse condition, and show excellent performance under various adverse weather conditions.
[0034] The present invention adopts a dual-branch design, independently encoding geometric and reflection information, avoiding interference between them, and improving the effectiveness of features; through a multi-level feature collaboration mechanism, feature fusion is performed at both the local and global levels, significantly enhancing the robustness and complementarity of the fused features; an introduced information constraint mechanism reduces redundancy between geometric and reflection features. Through this innovative design, the present invention can better learn from source domain data and effectively transfer it to the target domain, thereby maintaining excellent performance under different weather conditions and solving the problem of lidar semantic segmentation under adverse weather conditions.
[0035] Compared with the prior art, the present invention significantly improves the performance of the model under harsh weather conditions through innovative design when processing lidar data. Many existing methods uniformly process geometric structure information and reflection intensity information, often ignoring their degradation characteristics under different weather conditions, resulting in insufficient robustness and generalization ability of the model under adverse weather. In contrast, through the dual-branch design, multi-level feature collaboration mechanism, and information constraint strategy, the present invention demonstrates stronger robustness, higher generalization ability, and more accurate semantic segmentation performance when processing lidar data under adverse weather conditions, providing a more reliable and effective solution for the lidar semantic segmentation task. Brief Description of the Drawings
[0036] Figure 1 is a flowchart of a specific implementation manner of the method for semantic segmentation of lidar point cloud data for adverse weather according to the present invention;
[0037] Figure 2 is the overall network structure diagram of the method for semantic segmentation of lidar point cloud data for adverse weather according to the present invention;
[0038] Figure 3 is a visualization diagram of the reflection intensity under different weather conditions;
[0039] Figure 4 is a comparison diagram of segmentation effects. Detailed Implementation Manner
[0040] The following describes the specific implementation manner of the present invention with reference to the accompanying drawings, so that those skilled in the art can better understand the present invention. It should be particularly noted that in the following description, when the detailed description of known functions and designs may dilute the main content of the present invention, these descriptions will be omitted here.
[0041] The present invention processes the spatial geometric structure features and reflection intensity features in 3D point cloud data respectively to solve the problem of lidar semantic segmentation under adverse weather conditions. The spatial geometric structure branch focuses on capturing the geometric shape of spatial points, and the reflection intensity branch focuses on the return intensity of lidar signals. At the same time, effective fusion is achieved through a robust multi-level feature collaboration module to reduce the impact of adverse weather on the segmentation performance. Among them, after the spatial geometric information is voxelized, a sparse 3D convolutional network is used to extract the spatial geometric structure features; the reflection intensity information branch converts the spherical projection into a 2D image and uses a depthwise separable 2D convolutional network to extract the reflection features. The extracted spatial geometric structure and reflection features are efficiently fused through a robust multi-level feature collaboration module, enabling it to generalize to the target domain after training in the source domain and still efficiently complete the lidar semantic segmentation task under adverse weather conditions.
[0042] Figure 1 It is a flowchart of a specific implementation manner of the lidar point cloud data semantic segmentation method for adverse weather of the present invention.
[0043] In this embodiment, as Figure 1 shown, the lidar point cloud data semantic segmentation method for adverse weather of the present invention includes the following steps:
[0044] Step S1: Extraction of spatial geometric structure features
[0045] First, exclude the reflection intensity information from the acquired lidar point cloud data to obtain point cloud data containing only geometric coordinates, i.e., geometric point cloud, and discretize the geometric point cloud into regular voxel grids. Then, use an encoder network composed of multiple cascaded sparse 3D convolutional blocks to process the voxel grids, layer by layer to obtain spatial resolutions and geometric context information at different scales. Finally, aggregate the spatial resolutions and geometric context information at different scales to obtain spatial geometric structure features in a sparse representation form.
[0046] In this embodiment, as Figure 2 shown, the extraction of spatial geometric structure features is completed by a geometric feature extraction module. The encoder network composed of the sparse 3D convolutional blocks is MinkNet18 or MinkNet32.
[0047] Step S2: Extraction of reflection intensity features
[0048] First, use spherical projection to convert the acquired lidar point cloud data into a 2D view image, where only the reflection intensity information of each point is used as the input feature. Second, use an encoder network composed of 2D depthwise separable convolutional blocks to process the reflection intensity information, layer by layer to obtain features at different scales to form reflection intensity features.
[0049] In this embodiment, asFigure 2 As shown, the reflection intensity feature extraction is completed by the reflection feature extraction module. The encoder network composed of the two-dimensional depthwise separable convolution blocks is MobileNetV2.
[0050] Step S3: Robust multi-level feature collaboration
[0051] The robust multi-level feature collaboration includes a complementary information constraint loss for calculating enhanced feature complementarity and robustness, a local-level fusion based on the reliability of features to dynamically weight and combine geometric and reflection information, and a global-level fusion of reflection features through a two-stage cross-attention mechanism. In this embodiment, as Figure 2 shown, this step is completed in the robust multi-level feature collaboration, specifically:
[0052] Step S3.1: Calculate the complementary information constraint loss L for enhancing feature complementarity and robustness cic :
[0053] L cic = KL(p(m geo |f geo ) || r(m geo )) + KL(p(m ref |f ref ) || r(m ref )) - KL(p(m geo |f geo ) || p(m ref |f ref )) - KL(p(m ref |f ref ) || p(m geo |f geo ))
[0054] where KL(·|·) is the maximum between-class scatter, f geo is the feature vector extracted from the spatial geometric structure features, f ref is the feature vector extracted from the reflection intensity feature extraction, m geo is the mean μ geo calculated based on the feature vector f geo and the standard deviation σ geo , and the spatial geometric structure enhanced feature generated by adding additive noise , m ref is the mean μ ref calculated based on the feature vector f ref and the standard deviation σ ref , and the reflection intensity enhanced feature generated by adding additive noise :
[0055] m geo= μ geo + ε·σ geo
[0056] m ref = μ ref + ε·σ ref
[0057] r(m geo ) is to generate the enhanced feature m of the spatial geometric structure geo The formula is transformed into a standard Gaussian prior distribution with additive noise, r(m ref ) is to generate the enhanced feature m of the reflection intensity ref The formula is transformed into a standard Gaussian prior distribution with additive noise, p(m geo | f geo ) is the conditional distribution measuring the spatial geometric structure feature, p(m ref | f ref ) is the conditional distribution measuring the reflection intensity feature.
[0058] In the present invention, the first term KL(p(m geo | f geo ) || r(m geo )) measures the KL divergence between the conditional distribution p(m geo | f geo ) of the geometric feature and the standard Gaussian prior distribution r(m geo )(transformed into a Gaussian distribution with additive noise through the above reflection intensity enhanced feature formula). Minimizing this term means making the enhanced feature m of the spatial geometric structure geo close to the standard Gaussian prior distribution r(m geo ) with additive noise, improving the generalization ability of the feature while suppressing overfitting, and at the same time reducing the conditional dependence of the enhanced feature m of the spatial geometric structure geo on the feature vector f geo extracted from the spatial geometric structure feature.
[0059] The second term KL(p(m ref | f ref ) || r(m ref )) is similar to the first term and is used to measure the KL divergence between the conditional distribution p(m ref | f ref ) of the reflection intensity feature and the standard Gaussian prior distribution r(m ref ). Minimizing this term means making the enhanced feature m of the reflection intensity ref close to the standard Gaussian prior distribution r(m ref ) with additive noise, improving the generalization ability of the feature while suppressing overfitting, and at the same time reducing the conditional dependence of the enhanced feature m of the reflection intensity ref on the feature vector fgeo Conditional dependence.
[0060] The negative KL divergence in the third and fourth terms is used to increase the difference between p(m geo |f geo ) and p(m ref |f ref ) and reduce their similarity. Minimizing the negative value of this binomial (i.e., maximizing the KL divergence) means encouraging the distributions to be as different as possible, thereby reducing information redundancy and enhancing the complementarity of features.
[0061] The Kullback-Leibler (KL) divergence, i.e., the maximum between-class divergence, is a measure of the difference between two probability distributions P and Q, defined as follows:
[0062]
[0063] The KL divergence is asymmetric, i.e., usually KL(P||Q) ≠ KL(Q||P), which means it has a directional dependence on the deviation between the two distributions. Due to the asymmetry of the KL divergence, using KL divergences in both directions simultaneously can more evenly separate m geo and m ref to ensure their complementarity.
[0064] Step S3.2: Local-level feature fusion
[0065] Dynamically fuse the mean μ geo calculated based on the feature vector f geo and the standard deviation σ geo and the mean μ ref calculated based on the feature vector f ref and the standard deviation σ ref to obtain the local fusion feature F local :
[0066] F local = α·μ geo + (1 - α)·μ ref
[0067] where the fusion weight α is dynamically calculated according to the standard deviation σ geo and the standard deviation σ ref :
[0068]
[0069] where and are the means of the standard deviation σ geo and the standard deviation σ ref in the channel dimension, respectively.
[0070] Step S3.3: Global-level feature fusion
[0071] First, introduce a group of learnable global query values Q, and use the spatially geometric structure enhanced feature m geo as the key value K, value V of the cross-attention mechanism and the global query value Q to aggregate and obtain a temporary global feature value. Then, use the temporary global feature value as the key value K, value V of the cross-attention mechanism and the reflection intensity enhanced feature m ref to further aggregate and obtain the global fusion feature F global :
[0072] F global = CA(M geo , CA(Q, M ref ))
[0073] where CA(·,·) represents the cross-attention operation. In the parentheses of CA(·,·), the latter item is used as the key value K, value V of the cross-attention mechanism, and the former item is used as the query value Q for aggregation.
[0074] Step S3.4: Concatenate into a joint representation F
[0075] The joint representation F is formed by concatenating the local fusion feature F local and the global fusion feature F global .
[0076] Step S4: Optimize the network
[0077] In this embodiment, as shown in Figure 2 , the joint representation F is fed into the 3D voxel decoding module to obtain the semantic segmentation result. The network is optimized by combining the cross-entropy loss and the complementary information constraint loss. The loss L of the optimized network is:[[]]
[0078] L = L CE + βL cic
[0079] where L CE is the cross-entropy loss calculated from the semantic segmentation result obtained by the 3D voxel decoding module and the label, and β is the combination coefficient.
[0080] In this embodiment, the 3D voxel decoding module is composed of a MinkNet32 network opposite to the sparse 3D convolutional network structure and a linear layer in series. The output result is a 19-dimensional vector, and the index corresponding to the maximum value is the final result predicted by the network.
[0081] The optimized network is to update the network weights. In this embodiment, the publicly available dataset SemanticKITTI is selected. It contains 64 LiDAR sequences and annotates more than 19 semantic categories. According to the official protocol, we use sequences 00 - 07 and 09 - 10 as the training set, and sequence 08 as the validation set. We applied standard data augmentation, including random dropout, rotation, flipping, scaling, etc. We use the momentum SGD optimizer with a momentum of 0.9, a weight decay of 0.0001, and a batch size of 6, which is suitable for SemanticKITTI. The learning rate is initially set to 0.24 and is adjusted using the OneCycleLR strategy over a total of 50 epochs.
[0082] Step S5: Semantic segmentation of LiDAR point cloud data
[0083] The joint representation F obtained according to steps S1 - S3 is fed into the 3D voxel decoding module to obtain the semantic segmentation result.
[0084] Figure 2 It is the overall network structure diagram of the LiDAR point cloud data semantic segmentation method for adverse weather of the present invention.
[0085] In this embodiment, the present invention is achieved through Figure 2 the overall network structure shown.
[0086] The present invention is based on a dual - branch network architecture and realizes the semantic segmentation of LiDAR point cloud data by constructing a LiDAR segmentation network containing a geometric feature extraction module, a reflection feature extraction module, a robust multi - level feature collaboration module, and a 3D voxel decoding module in the offline stage. After the collected LiDAR signals are voxelized, the geometric feature extraction module is used to extract the spatial geometric structure features; the reflection intensity information is transformed into a 2D image through spherical projection, and the reflection intensity features are extracted by the reflection feature extraction module. The extracted spatial geometric structure features and reflection intensity features are efficiently fused through the robust multi - level feature collaboration module, performing local fusion and global fusion. These two fusion methods work together to process the feature information at different scales respectively, improving the performance of the model in complex environments, especially under adverse weather conditions. Among them, local fusion focuses on the local region of the features and dynamically fuses geometric and reflection features by calculating the weighted standard deviation; global fusion uses the cross - attention mechanism to extract useful information from the global context and effectively fuse it into the geometric features. The fused local and global features are concatenated, and the semantic segmentation result is generated through the 3D voxel decoding module. Finally, the network is optimized by combining the cross - entropy loss and the complementary information constraint loss, enabling it to generalize from the source domain to the target domain after training in the source domain and still efficiently complete the LiDAR semantic segmentation task under adverse weather conditions.
[0087] Test Performance of LiDAR Semantic Segmentation Network
[0088] The publicly available datasets SemanticSTF and SemanticKITTI-C are selected. SemanticSTF is a LiDAR segmentation dataset for adverse weather, containing per-point annotations for 21 semantic classes. This dataset includes four common adverse weather conditions: thick fog, light fog, snow, and rain, as Figure 3 shown. We follow the official protocol and use its validation set (containing 19 semantic classes) for testing. SemanticKITTI-C is a corrupted LiDAR segmentation dataset, with eight types of corruptions imposed on the validation set of SemanticKITTI, divided into three severity levels. The experiment uses the mean intersection over union (mIoU) as the evaluation metric. The generalization performance of the present invention on the SemanticSTF dataset is significantly improved. Compared with the widely adopted baseline MinkNet, the present invention achieves a significant increase of +18.1 in mIoU. In addition, the present invention exceeds the state-of-the-art method RDA in mIoU by an improvement of +3.0. The present invention performs best under all tested adverse weather conditions on the SemanticKITTI-C dataset, demonstrating its strong generalization ability from good weather to adverse conditions. The mIoU is improved by +2.1 under thick fog conditions, +2.6 under light fog conditions, +0.5 under rainy conditions, and +1.0 under snowy conditions, compared with the state-of-the-art method. In comparison with the MinkNet baseline, our method achieves a significant increase of +8.6 mIoU in thick fog, +14.1 mIoU in light fog, +9.7 mIoU in rainy conditions, and +12.7 mIoU in snowy conditions. In addition to performing well under adverse weather conditions, our method also achieves excellent results in most semantic classes, as Figure 4 shown. Compared with the baseline, we achieve an improvement of more than +10 mIoU in 13 classes. Notably, our method achieves an improvement of more than +5 in mIoU compared with the state-of-the-art method in classes such as trucks, other vehicles, motorcyclists, roads, parking, and terrain. At the same time, the previous state-of-the-art method relies on reinforcement learning for data augmentation and requires four A6000 GPUs for training. In contrast, our method achieves better performance through a simpler framework, without such complex augmentation techniques, and can be efficiently trained on a single RTX4090 GPU.
[0089] In summary, the method proposed in the present invention surpasses the previous methods and establishes a new state-of-the-art performance benchmark.
[0090] Although the above-described illustrative specific embodiments of the present invention have been described to facilitate understanding of the present invention by those skilled in the art, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
Claims
1. A semantic segmentation method for lidar point cloud data for adverse weather, characterized in that, It includes the following steps: (1) Spatial geometric structure feature extraction First, exclude the reflection intensity information from the acquired lidar point cloud data to obtain point cloud data containing only geometric coordinates, i.e., geometric point cloud, and discretize the geometric point cloud into regular voxel grids. Then, use an encoder network composed of multiple cascaded sparse 3D convolutional blocks to process the voxel grids, layer by layer to obtain spatial resolution and geometric context information at different scales. Finally, aggregate the spatial resolution and geometric context information at different scales to obtain the spatial geometric structure features in a sparse representation form; (2) Reflection intensity feature extraction First, use spherical projection to convert the acquired lidar point cloud data into a 2D view image, where the reflection intensity information of each point is used as the input feature. Second, use an encoder network composed of 2D depthwise separable convolutional blocks to process the reflection intensity information, layer by layer to obtain features at different scales to form the reflection intensity features; (3) Robust multi-level feature collaboration 3.1), Calculate the complementary information constraint loss L for enhancing feature complementarity and robustness cic : L cic = KL(p(m geo | f geo ) || r(m geo )) +KL(p(m ref |f ref )||r(m ref )) -KL(p(m geo |f geo )||p(m ref |f ref )) -KL(p(m ref |f ref )||p(m geo |f geo )) where KL(·||·) is the maximum between-class scatter, f geo is the feature vector extracted from the spatial geometric structure features, f ref is the feature vector extracted from the reflection intensity features, m geo is the mean μ geo calculated based on the feature vector f geo and the standard deviation σ geo The enhanced spatial geometric structure features generated by adding additive noise , m ref is the mean μ ref calculated based on the feature vector f ref and the standard deviation σ ref , and the enhanced reflection intensity features generated by adding additive noise : m geo = μ geo + ε·σ geo m ref = μ ref + ε·σ ref r(m geo ) is to generate the enhanced feature m of the spatial geometric structure geo The formula is converted into a standard Gaussian prior distribution with additive noise, and r(m ref ) is to generate the enhanced feature m of the reflection intensity ref The formula is converted into a standard Gaussian prior distribution with additive noise, p(m geo ∣f geo ) is the conditional distribution for measuring the spatial geometric structure feature, and p(m ref ∣f ref ) is the conditional distribution for measuring the reflection intensity feature; 3.2) Local-level feature fusion Based on the feature vector f geo The calculated mean μ geo And the standard deviation σ geo And based on the feature vector f ref The calculated mean μ ref And the standard deviation σ ref Perform dynamic fusion to obtain the local fusion feature F local : F local = α·μ geo + (1 - α)·μ ref Among them, the fusion weight α is calculated dynamically according to the standard deviation σ geo and the standard deviation σ ref as follows: Among them, and are the mean values of the standard deviation σ geo and the standard deviation σ ref on the channel dimension, respectively; 3.3) Global-level feature fusion First, introduce a set of learnable global query values Q, and enhance the spatial geometric structure feature m geo As the key value K, value V of the cross-attention mechanism and the global query value Q are aggregated to obtain a temporary global feature value, and then the temporary global feature value is used as the key value K, value V of the cross-attention mechanism and the reflection intensity enhancement feature m ref Further aggregation results in the global fusion feature F global ; 3.4) Concatenate into a joint representation F The locally fused feature F local and the globally fused feature F global are concatenated to form the joint representation F; (4) Optimization network Send the joint representation F into a 3D voxel decoding module to obtain the semantic segmentation result, and optimize the network by combining the cross-entropy loss and the complementary information constraint loss. The loss L of the optimization network is: L = L CE + βL cic Among them, L CE is the cross-entropy loss calculated from the semantic segmentation result obtained by the three-dimensional voxel decoding module and the label, and β is the combination coefficient; (5) Semantic segmentation of lidar point cloud data Obtain the joint representation F according to steps (1) to (3), and send it into the 3D voxel decoding module to obtain the semantic segmentation result.
2. The method for semantic segmentation of lidar point cloud data for adverse weather according to claim 1, wherein The 3D voxel decoding module is composed of a MinkNet32 network opposite to the sparse 3D convolutional network structure and a linear layer in series. The output result is a 19-dimensional vector, and the index corresponding to the maximum value is the final result predicted by the network.
Citation Information
Patent Citations
Aviation laser point cloud semantic segmentation method and device based on multistage context feature fusion network
CN116824585A
Laser radar semantic segmentation domain generalization method based on multi-layer aerial view assistance
CN119478402A