Method and System for Processing LiDAR Signals Based on Deep Learning

The deep learning-based laser radar signal processing method using the ADMN network architecture addresses weak echo and low signal-to-noise challenges by enhancing feature extraction and noise separation, improving three-dimensional reconstruction accuracy and efficiency.

CN119861382BActive Publication Date: 2025-07-15INST OF OPTICS & ELECTRONICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510355261.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-15
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract signal characteristics and separate signal noise in weak echo and low signal-to-noise ratio scenarios. The traditional method has poor effect when the luminous flux and integration time are limited, and deep learning-based methods lack specific designs for single-photon lidar.

Method used

The laser radar signal processing method based on deep learning is adopted, and the time window sliding module, self-attention mechanism module and soft threshold denoising module are introduced. Signal feature extraction and noise separation are performed through the ADMN network architecture, and the network convergence speed and solution accuracy are improved by using multiple loss functions.

Benefits of technology

The speed and accuracy of three-dimensional structure reconstruction of lidar is improved, the photon utilization efficiency is enhanced, and the robust processing capability is adapted to different signal-to-noise ratio scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119861382B_ABST
    Figure CN119861382B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for lidar signal processing based on deep learning, belonging to the technical field of lidar. The method obtains a training dataset through lidar detection principle simulation, introduces a time window sliding module, a self-attention mechanism module, and a soft threshold denoising module for the training dataset to construct a lidar signal processing and three-dimensional structure reconstruction network. Finally, based on the constructed lidar signal processing and three-dimensional structure reconstruction network, a corresponding denoised three-dimensional structure reconstruction result is output according to the input original lidar data to be processed. Compared with the traditional histogram technology method, the present invention can realize lidar signal processing and three-dimensional structure reconstruction in weak echo and low signal-to-noise ratio scenarios, improve the detection speed for the target scene, and meet the requirements for high-precision reconstruction of the detection scene in lidar.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of lidar, and in particular relates to a lidar signal processing method and system based on deep learning. Background Technique

[0002] As an active detection technology with high precision and high time resolution, lidar can obtain high-precision three-dimensional structure information of the detection scene and is often used in fields such as spaceborne remote sensing and autonomous driving. Traditional lidar needs to receive sufficient echo signals to suppress the inherent background noise in the detection process and is not applicable in scenarios with limited integration time and optical flux. For the improvement of hardware performance, single-photon lidar systems have been widely used. However, there is a problem of long accumulation time in the detection process, and it is necessary to optimize the algorithm design to improve the detection performance index.

[0003] Currently, there are two main challenges that limit the calculation of the distance information of the detection target from the photon sequence histogram. The first is the weak echo scenario. In scenarios with limited optical flux and integration time, compared with the ambient light intensity background noise, the signal photons corresponding to the lidar received pulse echo are weak; the second is the low signal-to-noise ratio scenario: in actual application scenarios, due to strong background light noise and detector dark count noise, the photon sequence histogram data recorded contains a large number of noise counts.

[0004] In order to achieve efficient three-dimensional reconstruction counting, it is necessary to solve the two major difficulties of difficult signal feature extraction and difficult signal noise separation in weak echo and low signal-to-noise ratio scenarios. The depth reconstruction method based on statistical learning realizes efficient photon imaging through constrained convex optimization heuristic iteration, but has the disadvantages of requiring manual setting of hyperparameters and poor effect in low signal-to-noise ratio scenarios. The method based on deep learning can achieve robust and efficient three-dimensional structure reconstruction in multi-signal-to-noise ratio scenarios, but has the disadvantages of lack of interpretability and lack of specific design for the unique properties of the original data of single-photon lidar, and there is room for further improvement. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a lidar signal processing method and system based on deep learning, which solves the technical problems that the statistical learning method in the prior art is not applicable to weak echo and low signal-to-noise ratio scenarios, and improves the processing efficiency and accuracy.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A lidar signal processing method based on deep learning, the method includes:

[0008] Obtain multiple depth maps of detection scenes marked with the true distance of pixels and multiple intensity maps of detection scenes marked with reflection intensity;

[0009] Perform simulation processing on the acquired depth map and intensity map to obtain a training dataset;

[0010] Based on the training dataset, train the ADMN network architecture to obtain a lidar signal processing and three-dimensional structure reconstruction network; wherein, a time window sliding module for initially extracting signal features, a self-attention mechanism module for fully extracting signal features, and a soft threshold denoising module for adaptively separating signal photons and noise photons are introduced into the ADMN network architecture;

[0011] Based on the lidar signal processing and three-dimensional structure reconstruction network, output a corresponding three-dimensional structure diagram of the denoised and reconstructed detection target according to the original lidar data to be processed.

[0012] On the other hand, the present invention provides a lidar signal processing system based on deep learning, and the system includes:

[0013] A data acquisition module for acquiring a plurality of depth maps of detection scenes marked with true pixel distances and a plurality of intensity maps of detection scenes marked with reflection intensities;

[0014] A simulation module for performing simulation processing on the acquired depth map and intensity map to obtain a training dataset;

[0015] A model construction module for training the ADMN network architecture based on the training dataset to obtain a lidar signal processing and three-dimensional structure reconstruction network; wherein, a time window sliding module for initially extracting signal features, a self-attention mechanism module for fully extracting signal features, and a soft threshold denoising module for adaptively separating signal photons and noise photons are introduced into the ADMN network architecture;

[0016] A data processing module for outputting a corresponding three-dimensional structure diagram of the denoised and reconstructed detection target according to the original lidar data to be processed based on the lidar signal processing and three-dimensional structure reconstruction network model.

[0017] In a third aspect, the present invention provides an electronic device, including: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing lidar signal processing method based on deep learning.

[0018] In a fourth aspect, the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is enabled to implement the foregoing lidar signal processing method based on deep learning.

[0019] The beneficial effects of the present invention are as follows:

[0020] The present invention introduces a time window sliding module, which facilitates the subsequent denoising module to more intuitively calculate the denoising threshold, and realizes feature extraction and enhancement without significantly increasing the number of parameters and computational cost; introduces a soft threshold denoising module based on residual connection, which does not require manual threshold setting, and adaptively sets the threshold through deep learning methods and better filters out noise photons; introduces a self-attention mechanism to learn the long-range correlation of the photon sequence histogram, which helps feature extraction and distance calculation; introduces joint constraints of multiple loss functions to improve the network convergence speed and calculation accuracy. Generally speaking, it improves the speed of lidar three-dimensional structure reconstruction, improves the accuracy of three-dimensional structure reconstruction of the detection scene, and meets the requirements of improving photon utilization efficiency in lidar detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the network structure of the lidar signal processing method based on deep learning of the present invention;

[0022] Figure 2 It is a schematic diagram of the algorithm principle of the soft threshold denoising module of the present invention;

[0023] Figure 3 It is one of the implementation effect diagrams based on the method of the present invention;

[0024] Figure 4 It is the second implementation effect diagram based on the method of the present invention;

[0025] Figure 5 It is a schematic diagram of the structure of the lidar signal processing system based on deep learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The present invention will be further described below with reference to the drawings and embodiments.

[0027] The present invention provides a lidar signal processing method based on deep learning, as Figure 1 shown, the method includes: obtaining a depth map and an intensity map of a detection scene of a PNG format image marked with depth information and intensity information of a detection target; respectively performing data augmentation and simulation on the obtained registered and aligned depth map and intensity map according to the lidar detection principle to obtain three-dimensional lidar raw data, and obtaining a training dataset; based on the training dataset, training an ADMN network architecture to obtain a lidar signal processing and three-dimensional structure reconstruction model; the ADMN network architecture introduces a time window sliding module for initially extracting signal features, a self-attention mechanism module for fully extracting signal features, and a soft threshold denoising module for adaptively separating signal photons and noise photons; based on the lidar signal processing and three-dimensional structure reconstruction model, output the corresponding denoised and calculated reconstructed three-dimensional structure according to the lidar raw data to be processed.

[0028] Furthermore, the simulation processing of the obtained depth map and intensity map to obtain the training dataset includes: randomly rotating the depth map and intensity map in different directions, randomly flipping the images, and adjusting the image size to increase the distribution diversity of the training data. Based on the depth map and intensity map after data augmentation by matching alignment, different signal-to-noise ratios are set, and the corresponding photon sequence histograms are simulated for each pixel to obtain the original lidar data as the training dataset.

[0029] Furthermore, the ADMN network architecture includes a time window sliding module for initially extracting signal features, a self-attention mechanism module for fully extracting signal features, and a soft threshold denoising module for adaptively separating signal photons and noise photons. Among them, the time window sliding module receives the original lidar data input by the user and outputs the initially processed feature data. Then, through a series of connections of convolutional layers, Relu activation function layers, and dilated convolutional layers, the initial features are extracted and input into the self-attention mechanism module. The self-attention mechanism module performs global correlation processing and global feature extraction on the feature data of the input initial features, and outputs the global correlation feature data considering global correlation to the soft threshold denoising module. The soft threshold denoising module performs adaptive soft threshold generation and noise photon removal on the global correlation feature data that has fully extracted feature information, and outputs the target feature data containing echo signal feature information after removing noise photons. This feature data undergoes a series of processes of transposed convolutional layers (i.e., the illustrated transposed convolution), Relu activation function layers, and fully connected layers, and finally, weighted summation calculation is performed according to the centroid method for the depth dimension to obtain the depth map distribution of the detection scene after denoising and calculation.

[0030] Furthermore, the time window sliding module in the ADMN network architecture for initially extracting signal features includes: the value of the convolution kernel used in the time window sliding module is related to the value of the system laser pulse width, and the convolution kernel value is set to a Gaussian convolution kernel of the same size as the laser pulse width and does not change during the network training process.

[0031] Furthermore, the self-attention mechanism module in the ADMN network architecture for fully extracting signal features includes first transforming and rearranging the input initial feature map, which is reflected in merging the feature dimension and depth dimension of the feature map, calculating the global correlation between the feature maps at different pixel positions, then calculating the corresponding self-attention features and performing residual connection, and finally deforming the self-attention features again to restore the same dimension as the input.

[0032] As Figure 2As shown, the ADMN network architecture contains a soft-threshold denoising module for adaptively separating signal photons and noise photons. The soft-threshold denoising module includes a first shaping operator, an absolute value operator, a normalization module, a three-dimensional feature extraction layer, an element-wise multiplication operator, a soft-threshold calculation operator, a second shaping operator, an element-wise summation operator, and an activation function output unit. Among them, the first shaping operator inputs the feature data considering global correlation output by the self-attention mechanism module processed by a series of convolutional layers and Relu activation function layers, and is used to rearrange the dimensions of the feature data considering global correlation to obtain a feature map corresponding to the three-dimensional spatial dimension. The absolute value operator, connected to the first shaping operator, is used to take the absolute value of the feature map corresponding to the three-dimensional spatial dimension to obtain an absolute value feature map. The normalization module, connected to the absolute value operator, is used to normalize the absolute value feature map along the depth dimension to obtain a normalized feature map. The three-dimensional feature extraction layer, connected to the normalization module, is used to extract and fuse features from the input normalized feature map to obtain signal distribution features. Among them, the three-dimensional feature extraction layer includes a series of alternately connected convolutional layers and Relu activation function layers, as well as a final sigmoid activation function layer. The element-wise multiplication operator, connected to the signal distribution feature output by the three-dimensional feature extraction layer and the absolute value feature output by the absolute value operator, is used to multiply the signal distribution feature and the absolute value feature map element-wise to obtain a soft-threshold feature. The soft-threshold calculation operator, connected to the element-wise multiplication operator and the first shaping operator, is used to fuse the feature map corresponding to the three-dimensional spatial dimension and the soft-threshold feature to separate the signal photon and noise photon features and obtain a feature map after soft-threshold denoising. The second shaping operator is connected to the soft-threshold calculation operator and is used to rearrange the dimensions of the feature map after soft-threshold denoising to obtain a final feature map. The element-wise summation operator is connected to the feature output by the self-attention mechanism module and the second shaping operator, and is used to add the feature data considering global correlation and the final feature map. The activation function output unit is used to receive the added feature data considering global correlation and the final feature map, process them using the Relu activation function, and output the target feature data containing echo signal feature information after removing noise photons to the subsequent processing network. Finally, it is connected to a series of transposed convolutional layers, Relu activation function layers, and fully connected layers for subsequent feature integration and depth resolution.

[0033] Further, the loss function for training the lidar signal processing and three-dimensional structure reconstruction network includes: cross-entropy loss, mean square error loss, and ordinal regression loss.

[0034] Further, the ordinal regression loss is:

[0035]

[0036] Where is the one-hot encoding of the time bins of the real data detection distance distribution, is the echo waveform distribution obtained through network calculation, is the -th time bin corresponding value of the one-hot encoding of the real data detection distance distribution, is the -th time bin corresponding value of the echo waveform distribution obtained through network calculation, is the entire detection period, is the value of the time bin corresponding to the real data detection distance, and cumsum is the cumulative summation operation.

[0037] The NYUv2 dataset is used, which contains 1449 pairs of intensity images and depth images of 27 different scenarios. The simulated average signal-to-noise ratios are set to 2:2 / 2:10 / 2:50 / 5:2 / 5:10 / 5:50 / 10:2 / 10:10 / 10:50 respectively. The training set, validation set, and test set are allocated according to 7:2:1. After completing the network training, the signal-to-noise ratio is set to 2:100, which does not appear in the training set, for testing on 8 scenarios of Middlebury. The quantization results are as Figure 3 shown. The network structures for comparison include U-Net++, PRSNet, Non-local, and the network proposed in the present invention. The optimal results are marked in bold, and the sub-optimal results are underlined. The quantization results in the figure show that compared with other deep learning methods, the method proposed in the present invention achieves better calculation results in terms of the two quantization indexes of root mean square error and threshold accuracy. In addition, the method proposed in the present invention is used to process the original lidar data of the real collected weak echo scenario, and the calculated depth map is as Figure 4 shown. For the low signal-to-noise ratio scenario (signal-to-noise ratio of 2:100) and the weak echo scenario (average number of echo photons per pixel is 3.48), compared with U-Net++, the method proposed in the present invention has relatively fewer noise points in the calculated depth map, and the edge calculation is clearer, which is more similar to the detection target distribution, and the overall calculation result is better, proving the effectiveness of the network structure proposed in the present invention.

[0038] As Figure 5 shown, to achieve the above and other related purposes, the present invention provides a lidar signal processing system based on deep learning. Each module included therein can implement each step of the foregoing method. The system includes: a data acquisition module for acquiring multiple detection scenario depth maps marked with the true distance of pixels and multiple detection scenario intensity maps marked with reflection intensities;

[0039] a simulation module for performing simulation processing on the acquired depth maps and intensity maps to obtain a training dataset;

[0040] A model construction module, configured to train an ADMN network architecture based on the training dataset to obtain a lidar signal processing and three-dimensional structure reconstruction network; wherein, a time window sliding module for preliminarily extracting signal features, a self-attention mechanism module for fully extracting signal features, and a soft threshold denoising module for adaptively separating signal photons and noise photons are introduced into the ADMN network architecture.

[0041] A data processing module, configured to output a corresponding three-dimensional structure diagram of the reconstructed detection target after denoising based on the lidar signal processing and three-dimensional structure reconstruction network model according to the raw lidar data to be processed.

[0042] In a third aspect, the present invention provides an electronic device, including: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing lidar signal processing method based on a deep learning method.

[0043] In a fourth aspect, the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is enabled to implement the foregoing lidar signal processing method based on a deep learning method.

[0044] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for processing lidar signals based on deep learning, characterized in that The method includes: Obtaining a plurality of depth maps of detection scenes marked with the true distance of pixels and a plurality of intensity maps of detection scenes marked with reflection intensity; Performing simulation processing on the obtained depth maps and intensity maps to obtain a training dataset; Based on the training dataset, training the ADMN network architecture to obtain a lidar signal processing and three-dimensional structure reconstruction network; in the ADMN network architecture, a time window sliding module for initially extracting signal features, a self-attention mechanism module for fully extracting signal features, and a soft threshold denoising module for adaptively separating signal photons and noise photons are introduced; wherein, the self-attention mechanism module is used to receive the initial features output by the time window sliding module, perform global correlation processing and secondary feature extraction on the initial features, output feature data considering global correlation and input it into the soft threshold denoising module; Based on the lidar signal processing and three-dimensional structure reconstruction network, output a corresponding three-dimensional structure diagram of the reconstructed detection target after denoising according to the raw lidar data to be processed.

2. The method for processing lidar signals based on deep learning according to claim 1, characterized in that, The performing simulation processing on the obtained depth maps and intensity maps to obtain a training dataset includes: Performing random direction rotation, random image flipping, and image size adjustment on the collected depth maps and intensity maps to obtain depth maps and intensity maps after data augmentation; Based on the depth maps and intensity maps after data augmentation that are matched and aligned, setting different signal-to-noise ratios, and performing simulation on each pixel to obtain a corresponding photon sequence histogram, obtaining raw lidar data as the training dataset.

3. The method for processing lidar signals based on deep learning according to claim 1, wherein, The value of the convolution kernel in the time window sliding module is set to a Gaussian convolution kernel of the same size as the laser pulse width and does not change during the network training process.

4. A method for processing lidar signals based on deep learning according to claim 1, characterized in that, The soft threshold denoising module includes a first shaping operator, an absolute value operator, a normalization module, a three-dimensional feature extraction layer, an element-wise multiplication operator, a soft threshold calculation operator, a second shaping operator, an element-wise summation operator, and an activation function output unit that are sequentially communicatively connected; wherein, The first shaping operator receives the feature data considering global correlation output by the self-attention mechanism module and is used to rearrange the dimensions of the feature data considering global correlation to obtain a feature map corresponding to the three-dimensional space dimension; The absolute value operator is used to take the absolute value of the feature map corresponding to the three-dimensional space dimension to obtain an absolute value feature map; The normalization module is used to normalize the absolute value feature map along the depth dimension to obtain a normalized feature map; The three-dimensional feature extraction layer is used to perform feature extraction and fusion on the normalized feature map to obtain signal distribution features; The element-wise multiplication operator is used to multiply the signal distribution features and the absolute value feature map element-wise to obtain soft threshold features; The soft threshold calculation operator is used to fuse the feature map corresponding to the three-dimensional space dimension and the soft threshold features to separate signal photon and noise photon features and obtain a feature map after soft threshold denoising; The second shaping operator is used to rearrange the dimensions of the feature map after soft threshold denoising to obtain a final feature map; The element-wise summation operator is used to add the feature data considering global correlation and the final feature map; The activation function output unit is configured to receive the feature data considering global correlation and the final feature map after addition, process them using the Relu activation function, and output target feature data containing the feature information of the echo signal with noise photons removed to the subsequent processing network.

5. A method for processing lidar signals based on deep learning according to claim 1, characterized in that The loss function for training the lidar signal processing and three-dimensional structure reconstruction network includes cross-entropy loss, mean squared error loss, and ordinal regression loss.

6. A method for processing lidar signals based on deep learning according to claim 5, characterized in that The ordinal regression loss is: Among them, is the one-hot encoding of the time bins of the real data detection distance distribution, is the echo waveform distribution obtained through network calculation, is the value corresponding to the th time bin of the one-hot encoding of the real data detection distance distribution, is the value corresponding to the th time bin of the echo waveform distribution obtained through network calculation, is the entire detection period, is the value of the time bin corresponding to the real data detection distance, is the cumulative summation operation.

7. A lidar signal processing system based on deep learning, characterized in that, The system includes: A data acquisition module, configured to acquire a plurality of depth maps of detection scenes marked with the true distance of pixels and a plurality of intensity maps of detection scenes marked with reflection intensity; A simulation module, configured to perform simulation processing on the acquired depth maps and intensity maps to obtain a training dataset; A model construction module, configured to train the ADMN network architecture based on the training dataset to obtain a lidar signal processing and three-dimensional structure reconstruction network; a time window sliding module for initially extracting signal features, a self-attention mechanism module for fully extracting signal features, and a soft threshold denoising module for adaptively separating signal photons and noise photons are introduced in the ADMN network architecture; wherein, the self-attention mechanism module is configured to receive the initial features output by the time window sliding module, perform global correlation processing and secondary feature extraction on the initial features, output feature data considering global correlation, and input it to the soft threshold denoising module; A data processing module, configured to output a corresponding three-dimensional structure diagram of the reconstructed detection target after denoising based on the lidar signal processing and three-dimensional structure reconstruction network according to the raw lidar data to be processed.

8. An electronic device, characterized in that, Includes: One or more processors; A memory, configured to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement a lidar signal processing method based on deep learning according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, An executable instruction is stored thereon, and when the instruction is executed by a processor, the processor can implement a lidar signal processing method based on deep learning according to any one of claims 1-6.

Citation Information

Patent Citations

  • Water depth remote sensing inversion method based on depth residual shrinkage network

    CN113639716A

  • Time-of-flight distance measurement method, time-of-flight distance measurement device, electronic equipment and computer readable storage medium

    CN115308718A