LiDAR Target Detection Method under Rain and Snow Weather Based on Deep Learning

Through deep learning-based methods, the noise of lidar in medium and heavy rain and snow scenes is extracted and filtered, and the problem of poor noise filtering effect in the existing technology is solved, achieving higher real-time and accuracy.

CN119478894BActive Publication Date: 2025-05-30青岛蚂蚁机器人有限责任公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411630362.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-05-30
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

The prior art is difficult to effectively filter lidar noise in moderate and heavy rain and snow scenarios, resulting in a decrease in real-time and accuracy of the perception system.

Method used

Using a deep learning-based method, through data acquisition and training, the features of point clouds are extracted and converted into depth maps, intensity maps and echo maps. Using convolutional neural networks and the timing attention mechanism of multi-heads, we extract and filter rain and snow noise to improve the quality of point clouds.

Benefits of technology

It significantly improves the point cloud noise filtering effect in moderate and heavy rain and snow scenes, improves the real-time and accuracy of ground target detection, and solves the technical difficulties of lidar in bad weather.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478894B_ABST
    Figure CN119478894B_ABST
Patent Text Reader

Abstract

This application proposes a lidar target detection method based on deep learning, and proposes a denoising solution for accurately filtering medium and heavy rain and snow scenarios, aiming to improve the real-time performance of ground target detection and solve the technical difficulties of lidar vehicle deployment. It includes the following implementation steps: Step 1), data collection and training environment preparation; Step 2), read data and extract original environmental features; Step 3), image processing; Step 4), feature extraction; Step 5), extract medium and heavy rain and snow point cloud noise points; Step 6), remove rain and snow noise points; Step 7), input the set of filtered normal point clouds into the centerpoint target detection network to realize the perception of surrounding obstacles in complex weather.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a target detection method for lidar noise filtering in rainy and snowy weather conditions, belonging to the fields of driverless and autonomous navigation. Background Art

[0002] To ensure the safety of the vehicle itself, road facilities, and pedestrians, driverless vehicles must have a high environmental perception ability. Environmental perception is the basic part of obtaining surrounding environmental information for autonomous decision-making and navigation control, and non-ground data detection is an important part of environmental perception ability. Existing obstacle detection means for driverless vehicles at home and abroad mainly include lidar, because it has the advantages of being less affected by light and having higher accuracy.

[0003] Lidar sensors in autonomous driving are vulnerable to adverse weather conditions such as snow, fog, and rain, which cause unnecessary noise factors and seriously reduce the cognitive performance of the perception system. The following previously published domestic patent applications, for example, the publication number is CN112213735A, and the name is a lidar point cloud noise reduction method for rainy and snowy weather. The disclosed method judges the rain and snow noise points by the number of connected components in the eight-connected domain after converting the point cloud into a depth map, and has a positive effect on light rain and light snow; however, it has poor effects on moderate to heavy rain and snow, will retain more rain and snow noise points, and has a long calculation time and poor real-time performance. Another example is the publication number CN116660934A, and the name is a vehicle, a rain / snow weather recognition method, a vehicle control method and system. Its solution can only identify rain and snow weather, but cannot solve the impact of rain and snow noise on perception.

[0004] Existing technologies generally search for neighboring point clouds based on the current point cloud, need to traverse and search the surrounding point clouds of each point, with a large amount of calculation and poor real-time performance. Suppose a 32-line lidar generates about 570,000 points per second. If the surrounding point clouds of each point are obtained and the density is calculated, it will take a quite long time. Moreover, after calculating the neighboring point clouds, it is necessary to perform point count statistics and neighboring distance calculation through a preset threshold. Since the noise point distributions of moderate to heavy rain and snow and light rain and snow are completely different, a single set of thresholds can usually only solve some of the noise points in some light rain and light snow scenarios, and the filtering effect is limited; in the face of moderate to heavy rain and snow scenarios, the noise filtering is poor, and if the filtering threshold is limited too harshly, it may cause misdeletion of the actual surrounding environmental point clouds, affecting the perception effect and causing missed detections.

[0005] In view of this, the present patent application is specifically proposed. Summary of the Invention

[0006] The method for lidar target detection in rain and snow weather based on deep learning described in this application aims to solve the problems existing in the above-mentioned prior art and proposes a denoising solution for accurately filtering medium to heavy rain and snow scenarios, with the expectation of improving the real-time performance of ground target detection and solving the technical difficulties of lidar deployment on vehicles.

[0007] To achieve the above design purpose, the method for lidar target detection in rain and snow weather based on deep learning includes the following implementation steps:

[0008] Step 1), data collection and training environment preparation;

[0009] Step 2), read the data and extract the original environmental features;

[0010] Extract the point cloud PCD in rainy weather, and obtain the point cloud features (x, y, z, i, diff) from it; where x is the coordinate value of the point cloud on the X-axis, y is the coordinate value of the point cloud on the Y-axis, z is the coordinate value on the Z-axis, i is the intensity value of the point cloud, and diff is the echo value of the point cloud;

[0011] Calculate the average echo value diff_average for each frame of input point cloud, and compare the average echo value diff_average with a pre-set threshold T to confirm whether it is light rain and snow weather or medium to heavy rain and snow weather; T is the average echo value statistically obtained according to the proportion of actual point cloud rain and snow noise points by different lidars;

[0012] Convert the Euclidean distance, reflection intensity, and echo value of the input point cloud into a depth map, intensity map, and echo map respectively, and project them to the corresponding pixel positions of the image;

[0013] Step 3), image processing;

[0014] Smoothly fill the blank values in the depth map, intensity map, and echo map;

[0015] Step 4), feature extraction;

[0016] Send the smoothed feature map into a convolutional neural network for feature extraction, and extract the sparse point cloud spatial feature map through multiple downsampling modules and multiple upsampling modules;

[0017] According to the judgment result of Step 2), for light rain and snow weather, jump to Step 6) to execute; for medium to heavy rain and snow weather, continue to execute the following Step 5);

[0018] Step 5), extraction of medium to heavy rain and snow point cloud noise;

[0019] Fuse the sparse distribution feature map of the point cloud space in the previous frame with the current frame through a multi-head temporal attention mechanism, that is, by passing in the feature map of the point cloud space information in the previous frame and the time difference between the two frames, and let the deformable attention mechanism extract the rain and snow noise features in the sparse distribution map of the point cloud space in the previous frame;

[0020] Step (6), removing rain and snow noise points;

[0021] For the obtained sparse distribution map of the current frame point cloud, predict the rain and snow noise point results through the detection head, and output a three-channel rain and snow noise point result map of depth, intensity, and echo;

[0022] Combine the point cloud height, intensity value, echo value, and the prediction result of the model to filter the rain and snow noise points;

[0023] Step (7), input the set of filtered normal point clouds into the centerpoint object detection network to realize the perception of surrounding obstacles in complex weather.

[0024] Furthermore, the said step (1) includes the following steps:

[0025] Install the lidar on the vehicle and collect data in the rain and snow weather environment; convert the data collected by the lidar into frames of point cloud PCD and save it with the corresponding time information; build a training environment and construct a training network architecture including pre-processing, single-frame network part, temporal network part, post-processing, and loss function code; prepare to extract point cloud data, conduct training, and output the model.

[0026] Furthermore, in the said step (2), the pixel position of each of the depth map, intensity map, and echo map comes from the angle conversion of the point cloud relative to the lidar coordinate system; among them, the height H of the image resolution is determined by the number of lidar beams, and the width W of the image resolution is determined by the maximum number of points scanned horizontally by the lidar; determine the angular range of α in the height direction through the maximum angular range of the up and down angles of the lidar, and determine the angular range of β in the width direction according to the maximum angular range of the left and right angles of the lidar; set the resolution of the image according to the number of lidar beams and the maximum number of points scanned horizontally; among them, the angular value α of the current point cloud point in the height direction is obtained by performing an arcsine calculation based on the z value of the current point cloud and the distance d value of the current point from the radar center; the product of the proportion of α in the angular range of α in the height direction and the total height is the row coordinate r of the current point in the image, and the calculation formula is as follows: α = arcsin(z, d), r = α / α H * H; the angular value β of the current point cloud point in the width direction is determined according to the angle of the point cloud from the X axis of the lidar coordinate system in the (X, Y) plane, and the proportion of β in the angular range of β in the width direction W is used to determine the angular range; set the resolution of the image according to the number of lidar beams and the maximum number of points scanned horizontally; among them, the angular value α of the current point cloud point in the height direction is obtained by performing an arcsine calculation based on the z value of the current point cloud and the distance d value of the current point from the radar center; the proportion of α in the angular range of α in the height direction H multiplied by the total height is the row coordinate r of the current point in the image, and the calculation formula is as follows: α = arcsin(z, d), r = α / α H * H; the angular value β of the current point cloud point in the width direction is determined according to the angle of the point cloud from the X axis of the lidar coordinate system in the (X, Y) plane, and the proportion of β in the angular range of β in the width directionW The product of the ratio and the total width is the column coordinate c of the current point in the image. The calculation formula is as follows: β = arctan(y, x), c = β / β W …W.

[0027] Further, in step (iii), the Savitzky-Golay filter is used for smoothing and filling. The length m of the window M is set to 5, which can retain more high-frequency information while paying attention to the smoothing effect. At the same time, the reflect filling is adopted for the left and right boundaries of the image, and the blank areas in the input vector (3 * H * W) of the three-channel image are filtered and smoothed along the width direction. The original values in the areas with original values remain unchanged, and the data accuracy is improved without changing the signal trend. Finally, the feature maps of the depth, intensity, and echo of the three channels after smoothing and filling (3 * H * W) are obtained;

[0028] The smoothing function expression of the Savitzky-Golay filter is as follows:

[0029]

[0030] where i is the index of the channel, j is the index in the height direction, k is the index in the width direction, M n is the weight value within the filter window, and m is the total length of the window.

[0031] Further, in step (iv), the downsampling module adopts a combination of a wavelet transform layer, a convolutional layer, a Batchnorm layer, a RELU non-linear activation layer, and a Dropout layer; each downsampling doubles the number of channels and reduces the length and width of the feature map by half; the upsampling module uses a residual network to receive the features from the downsampling, and the specific processing layers adopt a combination of a convolutional layer, a Batchnorm layer, a RELU non-linear activation layer, a Dropout layer, and an inverse wavelet transform layer; each upsampling reduces the number of channels by half and increases the length and width of the feature map by half; through the downsampling and upsampling of convolution and discrete wavelet transform, all subbands are used as inputs after each transformation, and it has a stronger ability to model the spatial context and inter-subband dependencies; finally, the feature map of the point cloud spatial distribution of the current frame (3 * H * W) is obtained;

[0032] Further, in step (v), for the point cloud spatial distribution feature map (3*H*W) obtained in step (iv), a learnable bias parameter of (N*2*H*W) and a learnable weight coefficient parameter of (N*H*W) are configured for the feature map of each channel; wherein, the learnable bias parameter includes the relative offset results of each pixel in this frame of feature map to N (x_pixel, y_pixel) around the same pixel position in the previous frame; 2 represents the x_pixel pixel values offset in the pixel width direction and the y_pixel pixel values offset in the height direction.

[0033] Further, in step (vi), the prediction of the rain, snow and noise results is based on a detection head represented by a 1x1 three-channel two-dimensional convolutional layer.

[0034] Further, step (vi) includes the following steps:

[0035] 6.1), Set all the original point clouds in this frame as P0;

[0036] All the ground points with the point cloud z value lower than the ground plane z_ground are first excluded and set as non-rain, snow and noise ground points P1, and the remaining unclassified point clouds P2 are left;

[0037] 6.2), Judge the point cloud categories in the remaining unclassified point clouds P2. According to the results predicted by the three-channel rain, snow and noise result map, the results of the three channels in each pixel are processed to predict the point cloud category in this pixel; among them, for the depth result d_result, intensity result i_result and echo result diff_result of each pixel, judge whether the result of the point cloud in the remaining unclassified point clouds P2 is greater than the threshold δ, that is, the following expression:

[0038] d_result*i_result 2 …diff_result>δ

[0039] If the result of the point cloud in the remaining unclassified point clouds P2 is less than the threshold δ, then judge these points as normal point clouds P3 excluded by the model; if the result of the point cloud in the remaining unclassified point clouds P2 is greater than the threshold δ, then judge these points as candidate rain, snow and noise points P4;

[0040] 6.3), Further distinguish the candidate rain, snow and noise points in the candidate rain, snow and noise points P4;

[0041] Set the points with the intensity value between 4 and 255 and the echo value between 0 and 3 cm as the normal point clouds P5 misdetected in the model, and the remaining part in P4 is set as the verified rain, snow and noise points P6 in this frame;

[0042] 6.4), Aggregate to obtain the set of normal point clouds;

[0043] The normal point cloud set includes the ground point cloud P1, the normal point cloud P3 excluded by the model, and the normal point cloud P5 misdetected by the model.

[0044] In summary, the advantages and beneficial effects of the present application are as follows:

[0045] 1. It can better filter out point cloud noise in light, moderate, and heavy rain and snow weather, providing a better solution for point cloud applications in harsh environments;

[0046] 2. The use of a deep neural network enables the acceleration of point cloud data calculation through hardware such as GPUs, significantly improving the running speed;

[0047] 3. The self-supervised algorithm accurately identifies noise points in medium and heavy rain and snow scenes, and does not require a large amount of annotation of rain and snow point clouds, with a small amount of calculation and high accuracy;

[0048] 4. The present application can be trained with self-generated point cloud noise to further reduce the detection cost of ground targets in subsequent real-time scenes. Description of the Drawings

[0049] The following specifically describes the implementation embodiments in conjunction with the drawings;

[0050] Figure 1 It is a flowchart of the lidar target detection method based on deep learning according to the present application;

[0051] Figure 2 It is a schematic diagram of the coordinate transformation from point cloud to image;

[0052] Figure 3 It is a schematic diagram of the rain and snow noise feature extraction network;

[0053] Figure 4 It is a schematic diagram of the downsampling structure;

[0054] Figure 5 It is a schematic diagram of the upsampling structure;

[0055] Figure 6 It is a schematic diagram of the multi-head temporal attention mechanism; Detailed Embodiments

[0056] Example 1, as Figures 1 to 6 shown, the present application proposes a lidar target detection method based on deep learning in rain and snow weather. This method converts a frame of point cloud into a depth map, an intensity map, and a double echo map, and inputs them into a neural network to extract noise points in medium and heavy rain and snow scenes using the spatial information and temporal information of rain and snow noise, and predicts the points caused by rain and snow particles. This part of the noise points is deleted from the point cloud to form a point cloud without rain and snow noise.

[0057] Spatial information means that the target signals in the real environment are sparse, while rain and snow noise points are non-sparse in space. Therefore, after rain and snow noise is converted from the time domain to the frequency domain, it will have a relatively large bandwidth. The rain and snow noise points with a large bandwidth in the frequency domain are extracted through a convolutional neural network (CNN), and the rain and snow point clouds are filtered out on the premise of maximizing the sparsity of the background real environment point cloud;

[0058] However, relying solely on spatial information, it is difficult to filter out all the noise points. Based on the above noise filtering relying on spatial information, the acquisition of temporal information is added for moderate to heavy rain and snow days to strengthen the sparse information of the current frame, and then the noise points in moderate to heavy rain and snow days are filtered out according to the prediction. Since rain and snow noise points are scattered and random and are easily affected by airflows, the noise points reflected by rain and snow particles are unpredictable. Reflected in the time series, the noise points that existed at the previous moment are very likely not in the same position at the current moment. For this reason, the spatial information obtained from the previous frame of point cloud is retained, and then through the multi-head temporal attention mechanism, the elements in each feature map of the current frame are used to find the elements at several positions around this element in the previous frame for more accurate feature fusion.

[0059] Specifically, it includes the following process steps:

[0060] Step 1), Data collection and training environment preparation

[0061] Install the lidar on the vehicle and collect data in rainy and snowy weather;

[0062] Convert the data collected by the lidar into frames of point cloud PCD and save it with the corresponding time information;

[0063] Build a training environment and construct a training network architecture including pre-processing, single-frame network part, temporal network part, post-processing and loss function code;

[0064] Prepare to extract point cloud data, train, and output the model.

[0065] Step 2), Read the data and extract the original environmental features;

[0066] Extract the point cloud PCD in rainy weather and obtain the point cloud features (x, y, z, i, diff) from it, where x is the X-axis coordinate value of the point cloud, y is the Y-axis coordinate value of the point cloud, z is the Z-axis coordinate value, i is the intensity value of the point cloud, and diff is the echo value of the point cloud;

[0067] Calculate the average echo value diff_average for each frame of input point cloud, and compare the average echo value diff_average with a pre-set threshold T to confirm whether it is light rain and snow weather or moderate to heavy rain and snow weather;

[0068] T is the average echo value statistically obtained based on whether the proportion of actual point cloud rain, snow, and noise points of different lidars is above 3 to 5%;

[0069] In this embodiment, it is set that when the number of rain, snow, and noise points is less than 3% of the current frame of point cloud, it is light rain and snow weather. Conversely, if the rain, snow, and noise points are greater than 3%, it is moderate to heavy rain and snow weather;

[0070] The calculation formula of the average echo value diff_average is where N is the number of points in the current frame;

[0071] If the average echo value diff_average is less than or equal to the threshold T, it is considered that the current frame is the point cloud of light rain and snow weather; if the average echo value diff_average is greater than the threshold T, it is considered that the current frame is the point cloud of moderate to heavy rain and snow weather, and the time series information needs to be added for filtering;

[0072] Such as Figure 2 As shown, the Euclidean distance, reflection intensity, and echo value of the input point cloud are respectively converted into a depth map, intensity map, and echo map, and projected onto the corresponding pixel positions of the image;

[0073] Each pixel position of the above three maps comes from the angle conversion of the point cloud relative to the lidar coordinate system; among them, the height H of the image resolution is determined by the number of beams of the lidar, and the width W of the image resolution is determined by the maximum number of points in the horizontal scan of the lidar;

[0074] Determine the angular range of α in the height direction through the maximum angular range of the up and down angles of the lidar H The angular range of β in the width direction is determined according to the maximum angular range of the left and right angles of the lidar W The angular range;

[0075] Set the resolution of the image according to the number of lidar beams and the maximum number of points in the horizontal scan;

[0076] Among them, the angular value α of the current point cloud point in the height direction is calculated by the arcsine calculation based on the z value of the current point cloud and the distance d value of the current point from the radar center;

[0077] The product of the proportion of α in the angular α in the height direction and the total height is the row coordinate r of the current point in the image, and the calculation formula is as follows: α = arcsin(z, d), r = α / α H *H; H *H;

[0078] The angular value β of the current point cloud point in the width direction is determined according to the angle of the point cloud from the X-axis in the (X, Y) plane. The proportion of β in the angular β in the width direction WThe product of the ratio and the total width is the column coordinate c of the current point in the image, and the calculation formula is as follows: β = arctan(y, x), c = β / β W *W;

[0079] Step 3), Image processing;

[0080] Smoothly fill the blank values in the depth map, intensity map, and echo map;

[0081] Use the Savitzky-Golay filter, where the length m of the window M is set to 5, retaining more high-frequency information while paying attention to the smoothing effect; at the same time, use reflect padding for the left and right boundaries of the image, and filter and smooth the blanks in each place in the input vector (3 * H * W) of the three-channel image along the width direction. The original values remain unchanged, and the data accuracy is improved without changing the signal trend; finally, the feature maps of the three channels of depth, intensity, and echo after smooth filling (3 * H * W) are obtained;

[0082] The smoothing function expression of the Savitzky-Golay filter is as follows:

[0083]

[0084] Among them, i is the index of the channel, j is the index in the height direction, k is the index in the width direction, M n is the weight value within the filter window, and m is the total length of the window;

[0085] Step 4), Feature extraction;

[0086] Send the feature map after smoothing processing into the convolutional neural network (CNN) as shown in Figure 3 to extract the sparse feature map of the point cloud space through multiple downsampling modules and multiple upsampling modules;

[0087] Convolutional neural networks (CNNs) usually use pooling to expand the receptive field, which has the advantage of low computational complexity. However, pooling may cause information loss, which has an adverse impact on subsequent operations such as feature extraction and analysis. Normal point clouds are sparse in space, and rain and snow noise points will affect the sparsity of the point cloud. After discrete wavelet transform, the noise points will have a large bandwidth in the frequency domain, so it is more suitable for analyzing and extracting signals with a large bandwidth.

[0088] Therefore, this application adopts a network form that combines wavelet transform and convolutional neural network to extract the features of non-sparse rain and snow noise points, that is, the way of wavelet transform plus downsampling of CNN, and inverse wavelet transform combined with upsampling of CNN, in order to better balance the receptive field size and computational efficiency of rain and snow noise points.

[0089] Specifically, the downsampling module adopts a combination of a wavelet transform layer, a convolutional layer, a Batchnorm layer, a RELU non-linear activation layer, and a Dropout layer, as Figure 4 shown; each downsampling doubles the number of channels and halves the length and width of the feature map;

[0090] The upsampling module uses a residual network to receive the features from the downsampling. The specific processing layers adopt a combination of a convolutional layer, a Batchnorm layer, a RELU non-linear activation layer, a Dropout layer, and an inverse wavelet transform layer, as Figure 5 shown; each upsampling halves the number of channels and doubles the length and width of the feature map;

[0091] Through the above series of convolutions, downsampling, and upsampling of discrete wavelet transforms, all subbands are used as inputs after each transformation, and it has a stronger ability to model spatial context and inter-subband dependencies; finally, a point cloud spatial distribution feature map (3*H*W) of the current frame is obtained;

[0092] According to the judgment result in step ii), for light rain, snow, or sleet weather, jump to step vi) for execution; for moderate to heavy rain, snow, or sleet weather, continue to execute the following step v);

[0093] Step v), Extract noise points from point clouds in moderate to heavy rain, snow, or sleet;

[0094] As Figure 6 shown, fuse the point cloud spatial sparse distribution feature map of the previous frame with the current frame through a multi-head temporal attention mechanism, that is, by passing in the point cloud spatial information feature map of the previous frame and the time difference between the two frames, and let the deformable attention mechanism extract the noise point features in the point cloud spatial sparse distribution map of the previous frame in rainy and snowy weather;

[0095] For the point cloud spatial distribution feature map (3*H*W) obtained in step iv), configure a learnable bias parameter of (N*2*H*W) and a learnable weight coefficient parameter of (N*H*W) for each channel's feature map; among them, the learnable bias parameter contains the relative offset results of each pixel in this frame's feature map to the N surrounding (x_pixel, y_pixel) at the same pixel position in the previous frame; 2 represents the x_pixel pixel values offset in the pixel width direction and the y_pixel pixel values offset in the height direction.

[0096] The learnable weight coefficient contains the specific contribution ratio of each pixel in this frame's feature map to the N surrounding pixels in the previous frame, so as to strengthen the sparsity of the moderate to heavy rain and snow noise points in the current frame through the network self-attention mechanism;

[0097] Step vi), Remove rain and snow noise points;

[0098] For the obtained current-frame point cloud sparse distribution map, predict the rain, snow, and noise point results through a detection head represented by a 1x1 three-channel two-dimensional convolutional layer, and output a three-channel rain, snow, and noise point result map of depth, intensity, and echo;

[0099] Combine the point cloud height, intensity value, echo value, and the prediction result of the model to filter the noise points in rainy and snowy days;

[0100] It includes the following steps:

[0101] 6.1), Set all the original point clouds of this frame as P0 (original input points); first, exclude all the ground points whose point cloud z value is lower than the ground plane z_ground, set them as non-rain, snow, and noise points P1 (ground points), and leave the remaining unclassified point clouds P2 (that is, P2 is the remaining unclassified points after excluding the ground points);

[0102] 6.2), Judge the point cloud categories in the remaining unclassified point clouds P2. According to the prediction results of the three-channel rain, snow, and noise point result map, process the results of the three channels within each pixel to predict the point cloud category within this pixel; among them, for the depth result d_result (cm), intensity result i_result, and echo result diff_result (cm) of each pixel, judge whether the result of the point cloud in the remaining unclassified point clouds P2 is greater than the threshold δ, that is, the following expression:

[0103] d_result * i_result 2 * diff_result > δ

[0104] If the result of the point cloud in the remaining unclassified point clouds P2 is less than the threshold δ, then judge these points as non-rain, snow, and noise points P3 (normal point clouds excluded by the model); if the result of the point cloud in the remaining unclassified point clouds P2 is greater than the threshold δ, then judge these points as candidate rain, snow, and noise points P4 (candidate rain, snow, and noise points after excluding the ground);

[0105] 6.4), Further distinguish the candidate rain, snow, and noise points within the candidate rain, snow, and noise points P4;

[0106] Set the points with intensity values between 4 and 255 and echo values between 0 and 3 cm as non-rain, snow, and noise points P5 (that is, the misdetections in the model, which are normal point clouds; the reason is that rain, snow, and noise points have very low intensity and very high echo values due to refraction and echo effects), and the remaining part in P4 is set as the rain, snow, and noise points P6 of this frame (rain, snow, and noise points after verification);

[0107] 6.4), Aggregate to obtain the set of normal point clouds

[0108] The set of normal point clouds includes ground point clouds P1, normal point clouds P3 excluded by the model, and normal point clouds P5 misdetected by the model;

[0109] Step seven): Input the set of filtered normal point clouds into the CenterPoint object detection network to achieve the perception of surrounding obstacles in complex weather.

[0110] As mentioned above, the embodiments given in the accompanying drawings are only the preferred solutions to achieve the object of the present invention. Those skilled in the art can obtain inspiration therefrom and directly derive other alternative structures that conform to the design concept of the present invention. The other structural features thus obtained should also fall within the scope of the solutions described in the present invention.

Claims

1. A deep learning-based laser radar target detection method in rainy and snowy weather, characterized by: It includes the following process steps: Step 1), data collection and training environment preparation; Step 2), read the data and extract the original characteristics of the environment; Extract the PCD of the point cloud in the rainy environment and obtain the point cloud features (x, y, z, i, diff) from it; where x is the X-axis coordinate value of the point cloud, y is the Y-axis coordinate value of the point cloud, z is the Z-axis coordinate value, i is the intensity value of the point cloud, and diff is the echo value of the point cloud; For each frame of input point cloud, the average echo value diff_average is calculated and compared with the preset threshold T to confirm whether it is light rain and snow weather or moderate to heavy rain and snow weather; T is the average echo value obtained by different laser radars based on the proportion of rain and snow noise points in the actual point cloud; The Euclidean distance, reflection intensity, and echo value of the input point cloud are converted into depth map, intensity map, and echo map respectively, and projected to the corresponding pixel position of the image; Step 3), image processing; Smoothly fill the blank values ​​in the depth map, intensity map and echo map; Step 4), feature extraction; The smoothed feature map is sent to the convolutional neural network for feature extraction, and the point cloud spatial sparse feature map is extracted through multi-layer downsampling modules and multi-layer upsampling modules; According to the judgment result of step 2), if it is light rain and snow, jump to step 6) for execution; if it is moderate to heavy rain and snow, continue to execute the following step 5); Step 5) Extraction of noise points in medium and heavy rain and snow point clouds; The sparse distribution feature map of the point cloud space of the previous frame is fused with the current frame through a multi-head temporal attention mechanism. That is, by passing the time difference between the spatial information feature map of the point cloud of the previous frame and the two frames, the deformable attention mechanism is used to extract the rain and snow noise features in the sparse distribution map of the point cloud space of the previous frame. Step 6) Remove rain and snow noise; The sparse distribution map of the current frame point cloud is obtained, and the rain and snow noise results are predicted by the detection head, and the three-channel rain and snow noise result map of depth, intensity and echo is output; Combine the point cloud height, intensity value, echo value and the model's prediction results to filter out noise points in rainy and snowy days; Step 7) Input the filtered normal point cloud set into the centerpoint target detection network to realize the perception of surrounding obstacles in complex weather conditions.

2. The deep learning-based laser radar target detection method in rainy and snowy weather according to claim 1, characterized in that: The step 1) comprises the following steps, Install the LiDAR on the vehicle to collect data in rainy and snowy environments; Convert the data collected by the LiDAR into point cloud PCD frames and save them with time correspondence information; Build a training environment and construct a training network architecture including pre-processing, single-frame network part, temporal network part, post-processing and loss function code; Prepare to extract point cloud data, perform training, and output the model.

3. The deep learning-based laser radar target detection method in rainy and snowy weather according to claim 1, characterized in that: In the step 2), each pixel position of the depth map, intensity map and echo map is derived from the angle conversion of the point cloud relative to the laser radar coordinate system; wherein the height H of the image resolution is determined by the laser radar beam, and the width W of the image resolution is determined by the maximum number of points of the horizontal scan of the laser radar; The maximum angle range of the up and down angles of the laser radar determines the α in the height direction. H The angle range of β in the width direction is determined according to the maximum angle range of the left and right angles of the laser radar. W Angular range; Set the image resolution based on the number of LiDAR beams and the maximum number of points in the horizontal scan; Among them, the angle value α of the current point cloud point in the height direction is calculated by the arc sine of the z value of the current point cloud and the distance d value of the current point from the center of the radar; α is the angle α in the height direction H The product of the ratio and the total height is the row coordinate r of the current point in the image. The calculation formula is as follows: α = arcsin (z, d), r = α / α H *H; The angle value β of the current point cloud point in the width direction is determined according to the angle of the point cloud in the (X, Y) plane from the X axis of the laser radar coordinate system. β accounts for the angle β in the width direction. W The product of the ratio and the total width is the column coordinate c of the current point in the image. The calculation formula is as follows: β = arctan (y, x), c = β / β W *W.

4. The method for detecting targets by laser radar in rainy and snowy weather based on deep learning according to claim 1, characterized in that: The step three) uses a Savitzky-Golay filter for smooth filling, sets the length m of the window M to 5, and retains more high-frequency information while focusing on the smoothing effect; at the same time, the left and right boundaries of the image are filled with reflect, and the blanks in the image input vector (3*H*W) of the three-channel feature map of depth, intensity and echo are filtered and smoothed along the width direction, and the original values ​​are still maintained where there are values, thereby improving data accuracy without changing the signal trend; finally, the three-channel feature map of depth, intensity and echo (3*H*W) after smooth filling is obtained; The smoothing function expression of the Savitzky-Golay filter is as follows: Among them, i is the index of the channel, j is the index in the height direction, k is the index in the width direction, and M n is the weight value within the filter window, and m is the total length of the window.

5. The method for detecting target by laser radar in rainy and snowy weather based on deep learning according to claim 1, characterized in that: In the step 4), the downsampling module adopts a combination of a wavelet transform layer, a convolution layer, a Batchnorm layer, a RELU nonlinear activation layer and a Dropout layer; each downsampling doubles the number of channels and reduces the length and width of the feature map by half; The upsampling module uses a residual network to accept features from downsampling. The specific processing layer uses a combination of convolutional layer, Batchnorm layer, RELU nonlinear activation layer, Dropout layer and inverse wavelet transform layer. Each upsampling reduces the number of channels by half and doubles the length and width of the feature map. Through convolution, discrete wavelet transform downsampling and upsampling, all sub-bands are used as input after each transformation, and it has a stronger ability to model spatial context and dependencies between sub-bands; finally, the point cloud spatial distribution feature map (3*H*W) of the current frame is obtained.

6. The deep learning-based laser radar target detection method in rainy and snowy weather according to claim 1, characterized in that: The step five) configures a learnable bias parameter (N*2*H*W) and a learnable weight coefficient parameter (N*H*W) for the point cloud spatial distribution feature map (3*H*W) obtained in step four) for the feature map of each channel; wherein the learnable bias parameter includes N relative offset results (x_pixel, y_pixel) of each pixel of the feature map of this frame to the same pixel position of the previous frame; 2 represents the x_pixel pixel values ​​offset in the pixel width direction and the y_pixel pixel values ​​offset in the height direction.

7. The deep learning-based laser radar target detection method in rainy and snowy weather according to claim 1, characterized in that: The step six) predicts the rain and snow noise results based on a detection head represented by a 1x1 three-channel two-dimensional convolutional layer.

8. The deep learning-based laser radar target detection method in rainy and snowy weather according to claim 1, characterized in that: The step six) comprises the following steps, 6.1) Set all the original point clouds of this frame as P0; All ground points whose point cloud z values ​​are lower than the ground plane z_ground are first excluded and set as non-rain and snow noise ground points P1, leaving the remaining unclassified point cloud P2; 6.2) Determine the point cloud category in the remaining unclassified point cloud P2. According to the prediction results of the three-channel rain and snow noise result map, process the results of the three channels in each pixel to predict the point cloud category in this pixel; among them, for each pixel's depth result d_result, intensity result i_result and echo result diff_result, determine whether the result of the point cloud in the remaining unclassified point cloud P2 is greater than the threshold δ, that is, the following expression: d_result*i_result 2 *diff_result>δ If the result of the point cloud in the remaining unclassified point cloud P2 is less than the threshold δ, these points are judged to be normal point clouds P3 excluded by the model; if the result of the point cloud in the remaining unclassified point cloud P2 is greater than the threshold δ, these points are judged to be candidate rain and snow noise points P4; 6.3) Further distinguish the candidate rain and snow noise points in the candidate rain and snow noise point P4; The points with intensity values ​​between 4 and 255 and echo values ​​between 0 and 3 cm are set as the normal point cloud P5 detected by mistake in the model, and the remaining part of P4 is set as the rain and snow noise point P6 after verification in this frame; 6.4) Summarize and obtain a set of normal point clouds; The normal point cloud set includes the ground point cloud P1, the normal point cloud P3 excluded by the model, and the normal point cloud P5 misdetected by the model.

Citation Information

Patent Citations

  • Laser point cloud noise reduction method for rainy and snowy weather

    CN112213735A

  • Vehicle, rain / snow weather identification method, vehicle control method and system

    CN116660934A

  • Short temporary rainfall prediction method based on space-time attention and data fusion

    CN114764602A

  • Mechanical laser radar point cloud data denoising method in snowy driving environment

    CN116645295A