LiDAR Data Processing Method and System
Through Gaussian filtering, bilateral filtering and multi-peak extraction technology combined with 3D CNN and Transformer networks, the problems of solid-state lidar data noise removal and feature retention in complex environments are solved, and the data quality is significantly improved, which is suitable for autonomous driving and robot navigation.
Patent Information
- Application Number
- CN202411022022.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-07-29
AI Technical Summary
The prior art is difficult to effectively remove noise in solid-state lidar data in complex environments, while retaining useful features, resulting in a decline in data quality and affecting the accuracy and real-time nature of autonomous driving and robot navigation.
Gaussian filtering and bilateral filtering, multi-peak extraction technology is used to combine 3D convolutional neural network (3D CNN) and Transformer network for feature extraction, multi-peak analysis and feature fusion, and finally optimize data quality through median filtering.
Effectively remove noise and retain useful features in complex environments, improving the quality and reliability of solid-state lidar data, and is suitable for fields such as autonomous driving and robot navigation.
Smart Images

Figure CN119087392B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a method and system for processing lidar data. Background Art
[0002] With the rapid development of technologies such as autonomous driving and robot navigation, solid-state lidar, as an important sensor in these fields, has been widely used due to its advantages of no mechanical components, fast response speed, small size, and low cost. The solid-state lidar emits laser pulses and receives their reflected signals to generate accurate three-dimensional point cloud data. However, during the data acquisition process, the solid-state lidar is susceptible to the influence of noise and ambient light interference, resulting in a decline in data quality, thereby affecting subsequent processing and application effects. Therefore, how to effectively filter and optimize solid-state lidar data has become the focus of research.
[0003] Traditional lidar data processing methods mainly rely on techniques such as Gaussian filtering and median filtering to reduce noise and optimize data. Although these methods can smooth data and reduce noise to a certain extent, they have obvious disadvantages in complex environments. First, these methods are sensitive to noise. When processing high-noise data, they often weaken useful features, especially at edges and details, resulting in the loss of important information. Second, traditional filtering methods are prone to over-smoothing of data, blurring sharp features and edges, making it difficult to retain fine structures and affecting the accuracy of subsequent object detection and classification. In addition, traditional methods are sensitive to parameter selection. Different environments and application scenarios require adjustment of different filtering parameters, increasing the complexity of practical applications. Finally, in complex scenarios, traditional filtering methods have a high computational complexity and are difficult to meet the requirements of real-time processing, which is a significant limitation for applications such as autonomous driving and robot navigation that require rapid response. Summary of the Invention
[0004] The present invention provides a method and system for processing lidar data, and the technical problem to be solved is: how to effectively remove noise in solid-state lidar data in a complex environment, while retaining useful features, improving the quality and reliability of solid-state lidar data, so as to meet the requirements of fields such as autonomous driving and robot navigation.
[0005] To solve the above technical problems, the present invention provides a method for processing lidar data, including the steps of:
[0006] S1. Use Gaussian filtering and bilateral filtering to perform noise reduction processing on the original histogram data of the lidar respectively;
[0007] S2. By using the multi-peak extraction technique, the histogram data after Gaussian filtering and bilateral filtering are respectively divided into multiple adjacent regions, and the local maximum value and its corresponding bin number are extracted in each region, obtaining two corresponding sets of feature group data;
[0008] S3. The two sets of feature group data are dimensionally elevated and respectively input into a 3D convolutional neural network, namely 3D CNN, and a Transformer network for local feature and global feature extraction;
[0009] S4. By using the multi-peak analysis technique, multi-peak analysis is performed on the feature vectors extracted by the 3D CNN and the Transformer network, and four types of features, namely the peak classification probability, time-of-flight value, distance, and probability density function, are extracted, obtaining corresponding multi-peak analysis feature vectors;
[0010] S5. The feature vectors extracted by the 3D CNN and the Transformer network are fused with the multi-peak analysis feature vectors to obtain fused features;
[0011] S6. Median filtering is performed on the fused features;
[0012] S7. Histogram data reconstruction is performed on the filtered fused features to generate optimized histogram data with the same dimension as the original histogram data.
[0013] Further, the step S5 specifically includes the steps:
[0014] S51. The feature vectors extracted by the 3D CNN and the Transformer network are feature-cascaded to obtain cascaded features;
[0015] S52. Weighted summation is performed on the two sets of multi-peak analysis feature vectors to obtain multi-peak analysis fused features;
[0016] S53. The cascaded features and the multi-peak analysis fused features are cascaded again to obtain the final fused features.
[0017] Further, the 3D CNN includes an extended dimension layer, a first convolutional layer, a first 3D max pooling layer, a second convolutional layer, a second 3D max pooling layer, a first flattening layer, and a first fully connected layer sequentially arranged from input to output; the extended dimension layer is used to elevate the dimension of the feature group data corresponding to Gaussian filtering to obtain a three-dimensional array adapted to the input of the 3D CNN; the first convolutional layer is used to extract local spatial features; the first 3D max pooling layer is used for max pooling to reduce the data dimension; the second convolutional layer is used to further extract local spatial features; the second 3D max pooling layer is used for max pooling to further reduce the data dimension; the first flattening layer is used to flatten the three-dimensional features into a one-dimensional feature vector using a flattening operation; the first fully connected layer is used to convert the flattened feature vector into an output vector.
[0018] Further, the Transformer network includes an encoder module, a decoder module, a second flattening layer, and a second fully connected layer connected in sequence. The encoder module includes a plurality of encoders connected in sequence, the decoder module includes a plurality of decoders connected in sequence, the output of the previous decoder is the input of the next decoder, and the output of the last encoder is input to each decoder.
[0019] Further, each encoder from input to output includes:
[0020] The encoder multi-head attention module is used to take the input data as the query matrix Q, the key matrix K, and the value matrix V, and perform weighted summation according to these weight matrices to obtain a new representation;
[0021] The encoder first residual connection and layer normalization module is used to add the output of the encoder multi-head attention module to the original input, and then perform layer normalization;
[0022] The encoder feed-forward neural network is used to perform non-linear mapping through two linear transformations and a ReLU activation function;
[0023] The encoder second residual connection and layer normalization module is used to add the output of the encoder feed-forward neural network to its input, and then perform layer normalization.
[0024] Further, each decoder from input to output includes:
[0025] The masked multi-head self-attention module is the same as the encoder multi-head attention module, but with a masking operation added;
[0026] The decoder first residual connection and layer normalization module is used to add the output of the masked multi-head attention module to the original input, and then perform layer normalization;
[0027] The decoder multi-head attention module: is used to take the output of the last encoder as the key and value, and the input of the decoder as the query, and generate a new representation through the multi-head attention mechanism;
[0028] The decoder second residual connection and layer normalization module is used to add the output of the decoder multi-head attention module to the original input, and then perform layer normalization;
[0029] The decoder feed-forward neural network is used to perform further non-linear transformation on the features of each position;
[0030] The decoder third residual connection and layer normalization module is used to add the output of the decoder feed-forward neural network to its input, and then perform layer normalization.
[0031] Further, in the step S1, the formula for performing a convolution operation on the three-dimensional data data(x, y, z) using Gaussian filtering is as follows:
[0032]
[0033] where k is the radius of the Gaussian filtering window, G(i, j, l) represents the Gaussian filtering function, (i, j, l) represents the coordinates within the Gaussian filter window, and G_data(x, y, z) represents the three-dimensional data obtained after performing Gaussian filtering on the three-dimensional data data(x, y, z);
[0034] The Gaussian filtering function G(i, j, l) is expressed as follows:
[0035]
[0036] where σ represents the standard deviation σ set in the Gaussian filtering, and e is the natural base;
[0037] The three-dimensional data data(x, y, z) is processed using bilateral filtering to obtain the smoothed data B_data(x, y, z):
[0038]
[0039] G r (data(x, y, z), data(x + o, y + p, z + q))
[0040] where (o, p, q) represents the coordinates within the spatial domain Gaussian filter window, and s is the radius of the spatial domain Gaussian filtering window; W p is the weight normalization term, which is used to ensure that the pixel values after filtering are within a reasonable range, and W p Specifically:
[0041]
[0042] G s (o, p, q) represents the spatial domain Gaussian filtering function, which is expressed as:
[0043]
[0044] where σ s is the standard deviation of the spatial domain Gaussian filtering function;
[0045] G r (data(x, y, z), data(x + o, y + p, z + q)) represents the gray domain Gaussian filtering function, which is expressed as:
[0046]
[0047] Among them, σ r is the standard deviation of the Gaussian filtering function in the grayscale domain;
[0048] The original histogram data is converted into an array with a dimension size of (H, W, D), denoted as three-dimensional data data(x, y, z). x, y, and z are the x-axis, y-axis, and z-axis coordinates of any point in the original histogram data, where x, y, z ≥ 0, and H, W, D are the x-axis range, y-axis range, and z-axis range corresponding to the original histogram data.
[0049] Furthermore, the step S2 includes the steps:
[0050] S21. Divide the Gaussian-filtered data G_data(x, y, z) and the bilateral-filtered data B_data(x, y, z) into N adjacent regions respectively, and the width of each region is W r , and the number N of regions is determined according to the size of the histogram data as R m is the measurement range of the histogram data. Then, the i-th region G_region i of the Gaussian-filtered data G_data(x, y, z) and the i-th region B_region i of the bilateral-filtered data B_data(x, y, z) can be expressed as:
[0051] G_region i ={G_data(x, y, z)|i·W r ≤ x < (i + 1)·W r}
[0052] B_region i ={B_data(x, y, z)|i·W r ≤ x < (i + 1)·Wr}
[0053] S22. In each region, extract the local maximum and its corresponding bin number bn i , and form a feature group denoted as The extraction formula is as follows:
[0054]
[0055] bn i = argmax(region i )
[0056] region i refers to G_region i or B_region i。
[0057] Further, the step S4 specifically includes the steps of:
[0058] S41. Use a fully connected feedforward neural network to calculate the activation value FNN(M) of the output of the 3D CNN and Transformer networks:
[0059] FNN(M) = ReLU(MW1 + b1)W2 + b2
[0060] where W1 and W2 are the weight matrices of the fully connected layers, b1 and b2 are the bias terms, M represents the output of the 3D CNN or Transformer network, FNN(M) is the output of the fully connected feedforward neural network, and ReLU() is the ReLU activation function;
[0061] S42. Use the softmax function to convert the activation value FNN(M) into the corresponding classification probability Class(M):
[0062] Class(M) = softmax(FNN(M)W c +b c )
[0063] where W c is the weight matrix of the classification layer; b c is the bias term of the classification layer;
[0064] S43. Calculate the corresponding time-of-flight value T TOF :
[0065] T TOF = Class(M)W t +b t
[0066] where W t is the weight matrix of the TOF recovery layer; b t is the bias term of the TOF recovery layer;
[0067] S44. Use the speed of light c and the time-of-flight TOF(M) to calculate the actual distance D TOF :
[0068]
[0069] S45. Use the generated time-of-flight T TOF to calculate the probability density function F PDF :
[0070]
[0071] where rB and r LB are the decay constants at different time periods, and r L is the photon arrival rate of the laser pulse, and T p is the pulse width, and t represents the current moment;
[0072] S46. The classification probabilities, time of flight, distance, and probability density function corresponding to the output features of the calculated 3D CNN and Transformer networks are respectively used to form corresponding multi-peak analysis feature vectors output_3d_mpe and output_t_mpe.
[0073] The present invention also provides a lidar data processing system, which is characterized in that it includes a data preprocessing unit, a multi-peak extraction unit, a neural network feature extraction unit, a multi-peak analysis unit, a feature fusion unit, a median filtering unit, and a data reconstruction unit, which are respectively used to execute steps S1 to S7 in the above-mentioned lidar data processing method.
[0074] The lidar data processing method and system provided by the present invention use bilateral filtering and Gaussian filtering to provide a more accurate and stable basis for multi-peak extraction and feature extraction; the use of multi-peak extraction effectively improves the robustness of the signal and reduces noise interference; the use of 3D CNN to extract local spatial features to capture details and local structure information in the histogram data; the use of the Transformer network to extract global features, which can capture long-range dependencies in the data; the use of multi-peak analysis to classify and analyze the peaks extracted by the neural network to further optimize the data quality; the features output by the neural network and the features of multi-peak analysis are fused to achieve the complementarity of local and global features and improve the robustness and accuracy of feature extraction; the result after feature fusion is processed by median filtering to further remove isolated noise points and improve the data quality; the feature data after median filtering is reconstructed to generate optimized histogram data with the same dimension as the original histogram data, so that it can be directly applied to subsequent analysis and processing. The method and system of the present invention can effectively remove noise in the original histogram data in a complex environment, retain useful features, significantly improve the quality and reliability of the solid-state lidar histogram data, and are applicable to fields such as autonomous driving and robot navigation that require high-precision data processing. Description of the Drawings
[0075] Figure 1 is a flowchart of the lidar data processing method provided by an embodiment of the present invention;
[0076] Figure 2 is a structural diagram of the 3D CNN provided by an embodiment of the present invention;
[0077] Figure 3It is the structural diagram of the Transformer network provided by the embodiments of the present invention;
[0078] Figure 4 It is the refined structural diagram of the Transformer network provided by the embodiments of the present invention. Specific embodiments
[0079] The following specifically illustrates the embodiments of the present invention in conjunction with the accompanying drawings. The given embodiments are only for illustrative purposes and should not be construed as a limitation of the present invention. The included drawings are for reference and illustration only and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0080] With the development of deep learning technology, methods for data processing and optimization using neural networks have gradually received attention. 3D convolutional neural networks (3D CNNs) perform excellently in local feature extraction, while Transformer networks have significant advantages in global feature capture. In addition, multi-peak extraction (MPE) and multi-peak analysis (MPA) techniques have shown significant advantages in solid-state lidar data processing. The multi-peak extraction technique improves the robustness of the signal by extracting local maxima from the histogram, while the multi-peak analysis technique further optimizes the data quality by classifying and analyzing the extracted peaks through a neural network.
[0081] Currently, this embodiment effectively combines multi-peak extraction and deep learning technology to efficiently filter and optimize solid-state lidar histogram data, overcome the shortcomings of traditional methods, effectively remove noise, retain useful features, improve the real-time performance and robustness of data processing, enhance data quality and reliability, and meet the requirements for high-precision lidar data in fields such as autonomous driving and robot navigation.
[0082] Specifically, the embodiments of the present invention first provide a method for lidar data processing, as Figure 1 shown in the flowchart of, including the steps:
[0083] S1. Data preprocessing: Use Gaussian filtering and bilateral filtering to perform noise reduction processing on the original histogram data of the lidar respectively;
[0084] S2. Multi-peak extraction: Through the multi-peak extraction technique, divide the histogram data after Gaussian filtering and bilateral filtering into multiple adjacent regions respectively, and extract the local maximum value and its corresponding bin number in each region to obtain two corresponding sets of feature group data;
[0085] S3. Neural Network Feature Extraction: The two groups of feature group data are dimensionally elevated and respectively input into a 3D convolutional neural network (3D CNN) and a Transformer network. The 3D convolutional neural network (3D CNN) extracts local features, and the Transformer network extracts global features;
[0086] S4. Multi-Peak Analysis: Through multi-peak analysis technology, multi-peak analysis is performed on the feature vectors extracted by the 3D CNN and the Transformer network to extract four types of features: their respective peak classification probabilities, time-of-flight (TOF) values, distances, and probability density functions (PDFs), obtaining corresponding multi-peak analysis feature vectors;
[0087] S5. Feature Fusion: The feature vectors extracted by the 3D CNN and the Transformer network are fused with the multi-peak analysis feature vectors to obtain fused features;
[0088] S6. Median Filtering: Median filtering is performed on the fused features;
[0089] S7. Data Reconstruction: Histogram data reconstruction is performed on the filtered fused features to generate optimized histogram data with the same dimension as the original histogram data.
[0090] In step S1, bilateral filtering combines information in the spatial domain and the gray scale domain, capable of retaining edge details while smoothing the data. The filtered data has better edge retention characteristics, further reducing noise interference and improving data quality. The data smoothed by Gaussian filtering and the data processed by bilateral filtering are used as the input data for the subsequent multi-peak extraction steps, providing a more accurate and stable basis for multi-peak extraction and feature extraction. In step S2, multi-peak extraction effectively improves the robustness of the signal, reduces noise interference, and ensures that the extracted peaks can reflect real object information. In step S3, neural network feature extraction is performed. The 3D CNN extracts local spatial features through multiple convolutional layers and pooling layers, suitable for capturing details and local structural information in the histogram data. The 3D CNN is good at extracting local features and can effectively identify detailed changes in the histogram data; the Transformer network performs excellently in extracting global features and can capture long-range dependencies in the data; in step S4, multi-peak analysis classifies and analyzes the peaks extracted by the neural network to further optimize data quality; in step S5, feature fusion is performed, and the combination of the two achieves the complementarity of local and global features, improving the robustness and accuracy of feature extraction; in step S6, median filtering is used to process the result after feature fusion to further remove isolated noise points and improve data quality; in step S7, the median-filtered feature data is reconstructed to generate optimized histogram data with the same dimension as the original histogram data, enabling it to be directly applied to subsequent analysis and processing.
[0091] The following elaborates on each step.
[0092] (1) Step S1: Data preprocessing
[0093] Step S1 includes a Gaussian filtering process and a bilateral filtering process for the original three-dimensional histogram data. Before performing Gaussian filtering and bilateral filtering, first read the original histogram data H(x, y, z) of the lidar, and then convert it into an array with a dimension size of (H, W, D) denoted as data(x, y, z). Here, x, y, z are the x-axis, y-axis, and z-axis coordinates of any point in the original histogram data, where x, y, z ≥ 0, and H, W, D are the x-axis range, y-axis range, and z-axis range corresponding to the original histogram data. As a preferred example, in this case, the dimension of data(x, y, z) is (150, 360, 288).
[0094] The specific Gaussian filtering process is as follows:
[0095] Perform a convolution operation on the three-dimensional data data(x, y, z) using Gaussian filtering:
[0096]
[0097] Among them, k is the radius of the Gaussian filtering window, G(i, j, l) represents the Gaussian filtering function, (i, j, l) represents the coordinates within the Gaussian filter window, and G_data(x, y, z) represents the three-dimensional data obtained after performing Gaussian filtering on the three-dimensional data data(x, y, z).
[0098] The Gaussian filtering function G(i, j, l) is expressed as follows:
[0099]
[0100] Among them, σ represents the standard deviation σ set in Gaussian filtering, and e is the natural base.
[0101] As a preferred example, in this case, σ is set to 1.2, and k is set to 3σ (i.e., 3.6).
[0102] The specific bilateral filtering process is as follows:
[0103] Process the three-dimensional data data(x, y, z) using bilateral filtering to obtain the smoothed data B_data(x, y, z):
[0104]
[0105] G r (data(x, y, z), data(x + o, y + p, z + q))
[0106] (o, p, q) represents the coordinates within the spatial-domain Gaussian filter window, s is the radius of the spatial-domain Gaussian filter window, G s (o, p, q) represents the spatial-domain Gaussian filter function, expressed as:
[0107]
[0108] where σ s is the standard deviation of the spatial-domain Gaussian filter function;
[0109] G r (data(x, y, z), data(x + o, y + p, z + q)) represents the gray-scale domain Gaussian filter function, expressed as:
[0110]
[0111] where σ r is the standard deviation of the gray-scale domain Gaussian filter function.
[0112] W p is the weight normalization term, used to ensure that the pixel values after filtering are within a reasonable range, W p Specifically:
[0113]
[0114] As a preferred example, σ s is set to 3.0, σ r is set to 0.5, and s is set to 3σ s (i.e., 9.0).
[0115] Bilateral filtering combines the information in the spatial domain and the gray-scale domain, and can retain edge details while smoothing the data. The filtered data B_data(x, y, z) has better edge retention characteristics, further reducing noise interference and improving data quality.
[0116] The data G_data(x, y, z) after Gaussian filtering and smoothing and the data B_data(x, y, z) after bilateral filtering will be used as the input data for the subsequent multi-peak extraction step, providing a more accurate and stable basis for multi-peak extraction and feature extraction.
[0117] (2) Step S2: Multi-peak extraction
[0118] Step S2 is specifically as follows:
[0119] S2. By using the Multi-Peak Extraction (MPE) technique, the data G_data(x, y, z) after Gaussian filtering and the data B_data(x, y, z) after bilateral filtering are respectively processed by multi-peak extraction.
[0120] The process of multi-peak extraction is as follows: G_data(x, y, z) and B_data(x, y, z) are respectively divided into multiple adjacent regions, and the local maximum value and its corresponding bin number bn are extracted in each region to form a feature group i where i ∈ [1, N], and N is the number of regions.
[0121] More specifically, step S2 includes the steps of:
[0122] S21. Partitioning: The data G_data(x, y, z) after Gaussian filtering and the data B_data(x, y, z) after bilateral filtering are respectively divided into N adjacent regions, and the width of each region is W r The number of regions N is determined according to the size of the histogram data as R m is the measurement range of the histogram data. Then, the i-th region G_region of the data G_data(x, y, z) after Gaussian filtering i and the i-th region B_region of the data B_data(x, y, z) after bilateral filtering i i can be expressed as:
[0123] G_region i = {G_data(x, y, z)|i·W r ≤ x < (i + 1)·W r}
[0124] B_region i = {B_data(x, y, z)|i·W r ≤ x < (i + 1)·Wr}
[0125] S22. Local maximum value and bin number extraction: In each region, the local maximum value and its corresponding bin number bn are extracted to form a feature group i denoted as The extraction formula is as follows:
[0126]
[0127] bn i = argmax(region i )
[0128] region i denotes G_region i or B_region i Finally, the feature group corresponding to G_data(x,y,z) is denoted as The feature group corresponding to B_data(x,y,z) is denoted as
[0129] Step S2 uses multi-peak extraction to effectively improve the robustness of the signal, reduce noise interference, and ensure that the extracted peaks can reflect the true object information.
[0130] (3) Step S3: Neural network feature extraction
[0131] Step S3 specifically includes using a 3D convolutional neural network (3D CNN) to extract features from the feature group data and using a Transformer network to extract features from the feature group data.
[0132] The structure of the 3D CNN is as Figure 2 shown, including an expansion dimension layer, a first convolutional layer, a first 3D max pooling layer, a second convolutional layer, a second 3D max pooling layer, a first flattening layer, and a first fully connected layer, which are sequentially arranged from input to output. The expansion dimension layer is used to increase the dimension of the feature group data to obtain a three-dimensional array input_3d(x,y,z) suitable for the input of the 3D CNN. The first convolutional layer is used to extract local spatial features using 32 3×3×3 convolutional kernels. The first 3D max pooling layer is used to perform max pooling using a 2×2×2 max pooling layer to reduce the data dimension and retain the main features. The second convolutional layer is used to further extract local spatial features using 32 3×3×3 convolutional kernels. The second 3D max pooling layer is used to further reduce the data dimension using a 2×2×2 max pooling layer and retain the main features. The first flattening layer is used to flatten the three-dimensional features into a one-dimensional feature vector using a flattening operation. The first fully connected layer (using the activation function ReLU) is used to convert the flattened feature vector into an output vector with a dimension of 512.
[0133] Among them, the expansion dimension layer inputs the feature group data into the expansion dimension layer for dimension increase to obtain a three-dimensional array input_3d(x,y,z) suitable for the input of the 3D CNN:
[0134]
[0135] The size of the original feature group is (N, 2), and after dimensionality increase processing, a three-dimensional array input_3d(x, y, z) with a size of (N, H, W, D) is formed.
[0136] As Figure 3 shown, the Transformer network includes an encoder module, a decoder module, a second flattening layer, and a second fully connected layer connected in sequence. Among them, the encoder module includes multiple encoders connected in sequence, the decoder module includes multiple decoders connected in sequence, the output of the previous decoder is the input of the next decoder, and the output of the last encoder is input to each decoder.
[0137] The feature group data after multi-peak extraction is already a feature vector of a fixed size and can be directly input into the Transformer network, denoted as input_t.
[0138] The refined structure of the Transformer network is as Figure 4 shown, and each encoder from input to output includes:
[0139] 1. Encoder multi-head attention module: Capture long-range dependencies in the data. Take the input data input_t as the query matrix Q, key matrix K, and value matrix V, and perform weighted summation according to these weight matrices to obtain a new representation;
[0140] 2. Encoder first residual connection and layer normalization module: Add the output of the encoder multi-head attention module to the original input, and then perform layer normalization to keep the gradients of the model stable;
[0141] 3. Encoder feed-forward neural network (FNN): Perform non-linear mapping through two linear transformations and a ReLU activation function;
[0142] 4. Encoder second residual connection and layer normalization module: Add the output of the encoder feed-forward neural network to its input, and then perform layer normalization.
[0143] Each decoder from input to output includes:
[0144] 1. Masked multi-head self-attention module: During the process of generating the output sequence, only consider the current time step and previous time steps to avoid leakage of future information. Similar to the multi-head attention module of the encoder, but with a masking operation added;
[0145] 2. Decoder first residual connection and layer normalization module: Add the output of the masked multi-head self-attention module to the original input, and then perform layer normalization to keep the gradients of the model stable;
[0146] 3. Decoder Multi-Head Attention Module: Taking the output of the last encoder as keys and values, and the input of the decoder as queries, generating new representations through the multi-head attention mechanism; combining the output of the last encoder with the current input of the decoder to capture the dependencies between the input and output.
[0147] 4. Decoder Second Residual Connection and Layer Normalization Module: Adding the output of the decoder multi-head attention module to the original input, and then performing layer normalization to maintain the stability of the model's gradients.
[0148] 5. Decoder Feed-Forward Neural Network: Performing further non-linear transformations on the features at each position.
[0149] 6. Decoder Third Residual Connection and Layer Normalization Module: Adding the output of the decoder feed-forward neural network to its input, and then performing layer normalization.
[0150] As a preferred example, 6 encoders and 6 decoders are each set.
[0151] Among them, the specific operations of the residual connection and layer normalization in the encoder and decoder are as follows:
[0152] Residual Connection: Preserving the information of the input X, and at the same time adding the output of the multi-head attention mechanism to it. The formula is:
[0153] X res = X + MultiHead(Q, K, V)
[0154] Among them, X res is the output of the residual connection, and MultiHead(Q, K, V) is the multi-head attention output by the previous module.
[0155] Layer Normalization: Normalizing the result of the residual connection to ensure the stability of the model during training. The formula is:
[0156] X′ = LayerNorm(X res )
[0157]
[0158] Among them, LayerNorm() represents the layer normalization operation, X′ represents the output of the layer normalization, μ is the mean of X res ; is the variance of X res ; ∈ is a small constant to prevent the denominator from being zero, set to 10 -6 ; γ and β are learnable parameters used to rescale and shift the normalized output.
[0159] Finally, the output of the Transformer network is converted into a one-dimensional feature vector through a second flattening layer, and then the flattened feature vector is converted into an output vector of 512 dimensions using a second fully connected layer.
[0160] The Transformer network uses multiple encoders, multiple decoders, and flattening layers to extract global features through steps such as the multi-head self-attention mechanism, feed-forward neural network, residual connection, and layer normalization.
[0161] In step S3, neural network feature extraction is performed. The 3D CNN extracts local spatial features through multiple convolutional layers and pooling layers, which is suitable for capturing details and local structural information in histogram data; the 3D CNN is good at extracting local features and can effectively identify detailed changes in histogram data; the Transformer network performs excellently in extracting global features and can capture long-range dependencies in the data.
[0162] (4) Step S4: Multi-peak analysis
[0163] Step S4 specifically includes the steps:
[0164] S41. Use a fully connected feed-forward neural network (FNN) to calculate the activation value FNN(M) of the output of the 3D CNN and the Transformer network:
[0165] FNN(M) = ReLU(MW1 + b1)W2 + b2
[0166] Where, W1 and W2 are the weight matrices of the fully connected layers, b1 and b2 are the bias terms, M represents the output of the 3D CNN or the Transformer network, FNN(M) is the output of the fully connected feed-forward neural network (FNN), and ReLU() is the ReLU activation function.
[0167] S42. Use the softmax function to convert the activation value FNN(M) into the corresponding classification probability Class(M):
[0168] Class(M) = softmax(FNN(M)W c + b c )
[0169] Where, W c is the weight matrix of the classification layer; b c is the bias term of the classification layer.
[0170] S43. Calculate the corresponding time of flight (TOF) value T TOF :
[0171] T TOF= Class(M)W t + b t
[0172] where W t is the weight matrix of the TOF recovery layer; b t is the bias term of the TOF recovery layer.
[0173] S44. Calculate the actual distance D using the speed of light c and the time of flight TOF(M) TOF :
[0174]
[0175] S45. Calculate the probability density function F using the generated time of flight T TOF : PDF where r
[0176]
[0177] and r B are attenuation constants at different time periods, r LB is the photon arrival rate of the laser pulse, T L is the pulse width, and t represents the current time. p
[0178] S46. Respectively form the corresponding multi-peak analysis feature vectors output_3d_mpe and output_t_mpe for the four types of features of the classification probability, time of flight, distance, and probability density function corresponding to the output features of the calculated 3D CNN and Transformer networks.
[0179] Step S4 performs multi-peak analysis to classify and analyze the peaks extracted by the neural network, further optimizing the data quality.
[0180] (5) Step S5: Feature fusion
[0181] Step S5 specifically includes the steps:
[0182] S51. Perform feature concatenation on the feature vectors output_3d and output_t extracted by the 3D CNN and Transformer networks to obtain the concatenated feature fused_feature. The dimensions of output_3d and output_t are both (N, 512), and the dimension of fused_feature is (N, 1024);
[0183] S52. Perform weighted summation on the multi-peak analysis feature vectors output_3d_mpe and output_t_mpe to obtain the multi-peak analysis fusion feature feature_mpe:
[0184] feature_mpe = ρ·output_3d_pe+(1 - ρ)·output_t_mpe
[0185] Among them, ρ is the weight coefficient, taking 0.4; the dimensions of output_3d_mpe and output_t_mpe are (N, 4), so the dimension of the obtained feature_mpe is (N, 4).
[0186] S53. Cascade the cascaded feature fused_feature and the multi-peak analysis fusion feature feature_mpe again. The dimensions of fused_feature and feature_mpe are (N, 1024) and (N, 4) respectively, and the result after cascading is denoted as feature, with a dimension of (N, 1028).
[0187] Step S5 performs feature fusion. The combination of the two realizes the complementarity of local and global features, improving the robustness and accuracy of feature extraction.
[0188] (6) Step S6: Median filtering
[0189] Perform median filtering on the result after feature fusion to further remove isolated noise points and improve data quality. Median filtering is also applicable when processing high-dimensional features, and it can effectively smooth feature data and retain important structural information.
[0190] For the feature vector feature, apply a median filter to smooth the data and remove noise points. The size of the filtering window is v = 3, and the median filtering operation is as follows:
[0191] median_filtered ij = median(feature i,j-v:j+v )
[0192] Among them, median() represents taking the median value of the data within the window; i is the index of the feature vector; j represents the jth dimension of the feature vector.
[0193] The feature vector after median filtering is median_filtered, and its dimension is still (N, 1028). This feature vector is used as the finally optimized feature vector for subsequent applications and analyses.
[0194] (6) Step S7: Reconstruct histogram data
[0195] The original dimension of the histogram data is (150, 360, 288), and the reconstruction process is as follows:
[0196] S71. First, map each value in the feature vector median_filtered to the corresponding coordinate in the histogram.
[0197] S72. Then, distribute the eigenvalues to the corresponding bins through the interpolation method interpolate():
[0198] histogramvalue(x, y, z) = interpolate(median_filtered, x, y, z)
[0199] Finally, output the reconstructed histogram data histogramvalue(x, y, z), whose dimension is (150, 360, 288). This histogram data can be directly used for subsequent analysis and processing.
[0200] To apply the above method, an embodiment of the present invention further includes a lidar data processing system, which includes a data preprocessing unit, a multi-peak extraction unit, a neural network feature extraction unit, a multi-peak analysis unit, a feature fusion unit, a median filtering unit, and a data reconstruction unit, which are respectively used to execute steps S1 to S7 in the above method.
[0201] Through the above steps, the method and system of the present invention can effectively remove the noise in the original histogram data in a complex environment, retain useful features, significantly improve the quality and reliability of the solid-state lidar histogram data, and are applicable to fields such as autonomous driving and robot navigation that require high-precision data processing.
[0202] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A method for processing lidar data, characterized in that, Including the steps: S1. Use Gaussian filtering and bilateral filtering respectively to perform noise reduction processing on the original histogram data of the lidar; S2. Through the multi-peak extraction technique, divide the histogram data after Gaussian filtering and bilateral filtering into multiple adjacent regions respectively, and extract the local maximum value and its corresponding bin number in each region to obtain two corresponding sets of feature group data; S3. Dimensionally expand the two sets of feature group data and input them into a 3D convolutional neural network, namely 3D CNN, and a Transformer network respectively to extract local features and global features; S4. Through the multi-peak analysis technique, perform multi-peak analysis on the feature vectors extracted by the 3D CNN and the Transformer network, and extract four types of features: peak classification probability, flight time value, distance, and probability density function of each, to obtain corresponding multi-peak analysis feature vectors; S5. Fuse the feature vectors extracted by the 3D CNN and the Transformer network with the multi-peak analysis feature vectors to obtain fused features; S6. Perform median filtering on the fused features; S7. Reconstruct the histogram data of the filtered fused features to generate optimized histogram data with the same dimension as the original histogram data.
2. The lidar data processing method according to claim 1, wherein The specific steps of step S5 include: S51. Cascade the feature vectors extracted by the 3D CNN and the Transformer network to obtain cascaded features; S52. Perform weighted summation on the two sets of multi-peak analysis feature vectors to obtain multi-peak analysis fused features; S53. Cascade the cascaded features and the multi-peak analysis fused features again to obtain the final fused features.
3. The lidar data processing method according to claim 1, characterized in that: The 3D CNN includes an expansion dimension layer, a first convolutional layer, a first 3D max pooling layer, a second convolutional layer, a second 3D max pooling layer, a first flattening layer, and a first fully connected layer sequentially arranged from input to output; the expansion dimension layer is used to dimensionally expand the feature group data corresponding to Gaussian filtering to obtain a three-dimensional array suitable for the input of the 3D CNN; the first convolutional layer is used to extract local spatial features; the first 3D max pooling layer is used to perform max pooling to reduce the data dimension; the second convolutional layer is used to further extract local spatial features; the second 3D max pooling layer is used to perform max pooling to further reduce the data dimension; the first flattening layer is used to flatten the three-dimensional features into a one-dimensional feature vector using a flattening operation; the first fully connected layer is used to convert the flattened feature vector into an output vector.
4. The lidar data processing method according to claim 1, wherein: The Transformer network includes an encoder module, a decoder module, a second flattening layer, and a second fully connected layer connected in sequence; the encoder module includes multiple encoders connected in sequence, the decoder module includes multiple decoders connected in sequence, the output of the previous decoder is the input of the next decoder, and the output of the last encoder is input to each decoder.
5. The lidar data processing method according to claim 4, wherein, Each encoder includes from input to output: An encoder multi-head attention module, which is used to use the input data as the query matrix Q, the key matrix K, and the value matrix V, and perform weighted summation according to these weight matrices to obtain a new representation; The encoder first residual connection and layer normalization module is used to add the output of the encoder multi-head attention module to the original input and then perform layer normalization; The encoder feed-forward neural network is used to perform non-linear mapping through two linear transformations and a ReLU activation function; The encoder second residual connection and layer normalization module is used to add the output of the encoder feed-forward neural network to its input and then perform layer normalization.
6. The lidar data processing method according to claim 5, wherein Each decoder from input to output includes: The masked multi-head self-attention module is the same as the encoder multi-head attention module, but with a masking operation added; The decoder first residual connection and layer normalization module is used to add the output of the masked multi-head attention module to the original input and then perform layer normalization; The decoder multi-head attention module: is used to take the output of the last encoder as the key and value, and the input of the decoder as the query, and generate a new representation through the multi-head attention mechanism; The decoder second residual connection and layer normalization module is used to add the output of the decoder multi-head attention module to the original input and then perform layer normalization; The decoder feed-forward neural network is used to perform further non-linear transformation on the features at each position; The decoder third residual connection and layer normalization module is used to add the output of the decoder feed-forward neural network to its input and then perform layer normalization.
7. The method for processing lidar data according to claim 1, wherein In the step S1, the formula for performing convolution operation on the three-dimensional data data(x, y, z) using Gaussian filtering is: where k is the radius of the Gaussian filter window, G(i, j, l) represents the Gaussian filter function, (i, j, l) represents the coordinates within the Gaussian filter window, and G_data(x, y, z) represents the three-dimensional data obtained after performing Gaussian filtering on the three-dimensional data data(x, y, z); The Gaussian filter function G(i, j, l) is expressed as follows: where σ represents the standard deviation σ set in the Gaussian filtering, and e is the natural base; The three-dimensional data data(x, y, z) is processed using bilateral filtering to obtain the smoothed data B_data(x, y, z): where (o, p, q) represents the coordinates within the spatial domain Gaussian filter window, and s is the radius of the spatial domain Gaussian filter window; Wp is the weight normalization term used to ensure that the pixel values after filtering are within a reasonable range, and Wp is specifically: Gs(o, p, q) represents the spatial domain Gaussian filter function, expressed as: where σs is the standard deviation of the spatial domain Gaussian filter function; Gr(data(x, y, z), data(x + o, y + p, z + q)) represents the gray domain Gaussian filter function, expressed as: where σr is the standard deviation of the gray domain Gaussian filter function; The original histogram data is converted into an array with a dimension size of (H, W, D) denoted as the three-dimensional data data(x, y, z), where x, y, z are the x-axis, y-axis, and z-axis coordinates of any point in the original histogram data, x, y, z ≥ 0, and H, W, D are the x-axis range, y-axis range, and z-axis range corresponding to the original histogram data.
8. The lidar data processing method according to claim 1, characterized in that The step S2 includes the steps: S21. Divide the Gaussian-filtered data G_data(x, y, z) and the bilateral-filtered data B_data(x, y, z) into N adjacent regions respectively, where the width of each region is Wr, and the number of regions N is determined according to the size of the histogram data as Rm is the measurement range of the histogram data. Then, the i-th region G_regioni of the Gaussian-filtered data G_data(x, y, z) and the i-th region B_regioni of the bilateral-filtered data B_data(x, y, z) can be expressed as: G_regions = {G_data(x, y, z) | i·Wr ≤ x < (i + 1)·Wr} B_regions = {B_data(x, y, z) | i·Wr ≤ x < (i + 1)·Wr} S22. In each region, extract the local maximum value and its corresponding bin number bni to form a feature group Denoted as The extraction formula is as follows: bni = argmax(regions) regions refers to G_regions or B_regions.
9. The lidar data processing method according to claim 1, wherein The specific steps of step S4 include the following steps: S41. Use a fully connected feedforward neural network to calculate the activation value FNN(M) of the output of the 3D CNN and Transformer networks: FNN(M) = ReLU(MW1 + b1)W2 + b2 where W1 and W2 are the weight matrices of the fully connected layers, b1 and b2 are the bias terms, M represents the output of the 3D CNN or Transformer network, FNN(M) is the output of the fully connected feedforward neural network, and ReLU() is the ReLU activation function; S42. Use the softmax function to convert the activation value FNN(M) into the corresponding classification probability Class(M): Class(M) = softmax(FNN(M)Wc + bc) where Wc is the weight matrix of the classification layer; bc is the bias term of the classification layer; S43. Calculate the corresponding flight time value T according to the classification probability Class(M) TOF :[[-END]] T TOF = Class(M)Wt + bt where Wt is the weight matrix of the TOF recovery layer; bt is the bias term of the TOF recovery layer; S44. Calculate the actual distance D using the speed of light c and the time of flight TOF(M). TOF : S45. Use the generated time of flight T TOF Calculate the probability density function F PDF : where r B and r LB are decay constants for different time periods, r L is the photon arrival rate of the laser pulse, Tp is the pulse width, and t represents the current time; S46. Respectively form the corresponding multi-peak analysis feature vectors output_3d_mpe and output_t_mpe from the four types of features of the classification probability, time of flight, distance, and probability density function corresponding to the output features of the 3D CNN and Transformer networks calculated.
10. A lidar data processing system, characterized in that: It includes a data preprocessing unit, a multi-peak extraction unit, a neural network feature extraction unit, a multi-peak analysis unit, a feature fusion unit, a median filtering unit, and a data reconstruction unit, which are respectively used to execute steps S1 to S7 in the lidar data processing method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Complex scene photon counting laser radar point cloud denoising algorithm based on multi-peak Gaussian fitting
CN111524084A
Real-time multi-mode sensing high-precision map construction method
CN116817891A