Optimization Method and System for Solid-State LiDAR Histogram

The noise interference problem is solved and the accuracy and efficiency of ranging and imaging are improved by establishing a noise estimation probability model and a cross-frame feature aggregation recurrent neural network.

CN118674631BActive Publication Date: 2025-07-11CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410699998.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-31
Publication Date
2025-07-11
Estimated Expiration
2044-05-31

AI Technical Summary

Technical Problem

The noise interference in the solid-state lidar histogram data is severe, affecting the ranging accuracy and object recognition effect. The efficiency and accuracy of existing noise reduction methods are not high.

Method used

Establish a noise estimation probability model, perform preliminary noise reduction on the two-dimensional histogram of solid-state lidar and map it to three-dimensional space, and use cross-frame feature aggregation recurrent neural network for feature enhancement and secondary noise reduction, including residual dense channel attention module and spatiotemporal feature enhancement module.

Benefits of technology

It significantly improves the signal-to-noise ratio, enhances the ability to capture fine structures, improves the accuracy and efficiency of ranging and imaging, and reduces noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118674631B_ABST
    Figure CN118674631B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology, and specifically discloses an optimization method and system for the histogram of a solid-state lidar. First, a noise estimation probability model is established for different types of noise suffered by the solid-state lidar, and preliminary noise reduction processing is performed on the two-dimensional histogram data of each pixel point collected by the solid-state lidar to suppress the influence of background noise and hardware accidental noise. Subsequently, the processed two-dimensional histogram data is mapped into a three-dimensional histogram, and a cross-frame feature aggregation recurrent neural network is used to perform feature enhancement and secondary noise reduction optimization on the mapped three-dimensional histogram, further improving the signal-to-noise ratio of the data, enhancing the ability of the solid-state lidar to capture fine structures, and significantly improving the accuracy and efficiency of ranging and imaging. Compared with the prior art, the present invention can effectively reduce noise interference, improve imaging quality and positioning accuracy, and has strong practicability and broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to an optimization method and system for a solid-state lidar histogram. Background Art

[0002] With the rapid development of autonomous driving technology, solid-state lidar (LiDAR), as one of the core perception components, is increasingly applied to commercial vehicles and advanced driver assistance systems (ADAS) due to its advantages such as no mechanical moving parts, fast response speed, small size, and low cost. Solid-state lidar detects and measures the distance and shape of objects in the surrounding environment by emitting laser pulses and receiving the reflected light pulses. The generated histogram data is crucial for subsequent object detection, classification, and tracking.

[0003] Histogram data optimization and peak extraction are key technologies in solid-state lidar data processing, which are directly related to the measurement accuracy and reliability of the lidar. The peaks in the histogram data reflect the intensity of the received reflected signals, and the positions of the peaks usually represent the actual distances of the objects. Therefore, accurate peak extraction can not only improve the accuracy of distance measurement but also contribute to improving the subsequent object recognition and positioning processes.

[0004] However, the solid-state lidar histogram data faces various challenges. First, it is interfered by noise and multipath effects. In a complex traffic environment, due to multipath effects such as reflection, refraction, and scattering, the signals received by the lidar may generate multiple peaks, making it difficult to identify the true reflected signals. In addition, environmental noise and the electronic noise of the device itself also affect the signal quality, making it easy to make mistakes in the peak extraction process, thus affecting the performance of the entire system.

[0005] Currently, the main methods for reducing noise in histogram data are digital filtering technology and waveform analysis technology. These methods reduce random noise by smoothing the data, but they often blur the sharp features in the data, such as edges and corners, thus affecting the accurate detection of peaks. At the same time, they are also sensitive to parameter selection and are not easy to achieve real-time processing. In terms of peak extraction methods, traditional threshold search technology and statistics-based methods are mainly used. The efficiency and accuracy of these two methods are not high in an environment with multiple peaks or close peak spacing and high noise. Summary of the Invention

[0006] The present invention provides an optimization method and system for a solid-state lidar histogram, and the technical problem to be solved is: how to effectively reduce the noise interference in the solid-state lidar histogram data, optimize the histogram data features, and improve the imaging quality and positioning accuracy of the solid-state lidar.

[0007] To solve the above technical problems, the present invention provides a method for data noise reduction and feature enhancement of a solid-state lidar histogram, including the steps:

[0008] S1. Measure the noise data of the solid-state lidar and establish a noise estimation probability model based on the noise data;

[0009] S2. Use the noise estimation probability model to perform preliminary noise reduction on each frame of the original histogram of the solid-state lidar pixel by pixel to obtain multiple frames of two-dimensional histograms after preliminary noise reduction;

[0010] S3. Map the multiple frames of two-dimensional histograms after preliminary noise reduction to a three-dimensional space to obtain multiple frames of three-dimensional histograms after preliminary noise reduction;

[0011] S4. Perform cross-frame feature aggregation and noise reduction optimization on the multiple frames of three-dimensional histograms after preliminary noise reduction to obtain multiple frames of three-dimensional histograms after secondary noise reduction and feature enhancement;

[0012] For each frame of the three-dimensional histogram after preliminary noise reduction, the step S4 specifically includes the steps:

[0013] S41. Extract the initial features of 2M + 1 layers of a sequence of three-dimensional histograms after preliminary noise reduction with a length of 2M + 1 frames centered on a frame of the three-dimensional histogram after preliminary noise reduction, where M is a natural number not less than 1;

[0014] S42. Perform spatio-temporal feature enhancement on the other 2M layers of initial features except the initial features corresponding to the three-dimensional histogram after preliminary noise reduction of this frame to obtain 2M layers of spatio-temporal enhanced features;

[0015] S43. Aggregate the 2M layers of spatio-temporal enhanced features with the initial features of the three-dimensional histogram after preliminary noise reduction of this frame respectively to obtain the corresponding 2M layers of aggregated features;

[0016] S44. Combine the 2M layers of aggregated features with the initial features of the three-dimensional histogram after preliminary noise reduction of this frame to obtain the three-dimensional histogram after secondary noise reduction and feature enhancement of this frame.

[0017] Further, the step S41 specifically includes the steps:

[0018] S411. Obtain the three-dimensional histograms after preliminary noise reduction of the current frame and the previous M frames and the next M frames of the current frame to form a sequence of three-dimensional histograms after preliminary noise reduction of the current frame;

[0019] S412. Input the first frame of the three-dimensional histogram sequence after preliminary noise reduction of the current frame and the initial hidden state into the first residual dense channel attention module for feature extraction and hidden state generation to obtain the corresponding first-layer hidden state and first-layer initial features;

[0020] S413. Input the second frame of the pre-denoised 3D histogram sequence and the first-layer hidden state into the second residual dense channel attention module for feature extraction and hidden state generation to obtain the corresponding second-layer initial feature and second-layer hidden state; repeat this process until the (2M + 1)-th frame of the pre-denoised 3D histogram sequence and the 2M-th layer hidden state are input into the (2M + 1)-th residual dense channel attention module for feature extraction to obtain the corresponding (2M + 1)-th layer initial feature.

[0021] Furthermore, the first residual dense channel attention module to the 2M-th residual dense channel attention module all adopt the structure of the residual dense channel attention module.

[0022] The operations performed by the residual dense channel attention module are as follows:

[0023] First, perform downsampling on the input pre-denoised 3D histogram, then perform channel concatenation on the downsampled result and the input hidden state, then perform N ≥ 2 times of residual dense operations and channel attention operations on the channel concatenated result to obtain N intermediate features, then perform channel concatenation on the N intermediate features and then perform 1×1 convolution to obtain the initial feature of the pre-denoised 3D histogram, and then obtain the corresponding hidden state through the hidden state generation function for the initial feature.

[0024] The structure adopted by the 2M-th residual dense channel attention module lacks the hidden state generation function compared with the residual dense channel attention module.

[0025] Furthermore, each residual dense operation and channel attention operation includes performing a channel attention operation after performing a residual dense operation.

[0026] The residual dense operation is as follows:

[0027] Perform 3×3 convolution and ReLU function activation on the first input to obtain the first output.

[0028] Perform 3×3 convolution and ReLU function activation on the first input and the first output as the second input to obtain the second output.

[0029] Perform 3×3 convolution and ReLU function activation on the second output, the first input, and the second input as the third input to obtain the third output.

[0030] Perform 1×1 convolution on the third output, the first input, the second input, and the third input, and then add the result element-wise to the first input to obtain the residual dense output.

[0031] The channel attention operation is as follows:

[0032] Perform global average pooling on the residual dense output, followed by a fully connected layer, activation with the ReLU function, another fully connected layer, and activation with the Sigmoid function, and then multiply the result element-wise with the residual dense output to obtain the channel attention output;

[0033] The hidden state generation function sequentially performs a 3×3 convolution, a residual dense operation, and another 3×3 convolution once.

[0034] Further, in step S42, performing spatio-temporal feature enhancement for each layer of the 2M-layer initial features specifically includes: performing spatio-temporal context enhancement on the input initial features to obtain spatio-temporal context enhanced features; performing spatio-temporal feature extraction on the spatio-temporal context enhanced features to obtain the spatio-temporal enhanced features corresponding to the initial features;

[0035] Performing spatio-temporal context enhancement specifically is:

[0036] Perform a 3×3 convolution on the input initial features and then shift by θ, perform deformable convolution on the shifted features and the initial features, and then perform a 3×3 convolution to obtain spatio-temporal context enhanced features;

[0037] Performing spatio-temporal feature extraction specifically is:

[0038] Perform a 7×7 convolution on the input spatio-temporal context enhanced features and activation with the Sigmoid function, and then multiply the result element-wise with the spatio-temporal context enhanced features to obtain spatio-temporal enhanced features.

[0039] Further, in step S43, aggregating any one of the 2M-layer spatio-temporal enhanced features with the initial features of the initially denoised three-dimensional histogram of this frame specifically is:

[0040] Perform feature concatenation on this layer of spatio-temporal enhanced features and its initial features to obtain concatenated features;

[0041] Perform global average pooling fusion on the concatenated features to obtain fused features; perform channel attention extraction on the concatenated features to obtain attention features;

[0042] Multiply the fused features and the attention features element-wise to obtain the aggregated features corresponding to the initial features of this layer;

[0043] Global average pooling fusion specifically is: perform global average pooling, a fully connected layer, activation with the ReLU function, and a fully connected layer, activation with the Sigmoid function on the concatenated features to obtain the first intermediate features, perform 1×1 convolutions on the concatenated features to obtain the second intermediate features, multiply the first intermediate features and the second intermediate features element-wise and then perform a 1×1 convolution to obtain fused features.

[0044] Further, the step S44 is specifically as follows: Feature connection is performed on the 2M-layer aggregated features and the initial features of the three-dimensional histogram after preliminary noise reduction of this frame, and then 1×1 convolution is performed to obtain the three-dimensional histogram of secondary noise reduction and feature enhancement of this frame.

[0045] Further, the step S1 specifically includes the following steps:

[0046] S11. Measure the dark count rate noise and record it as the total number of DCR noises Nn (DCR) , and calculate the average value of the dark count rate noise in each time bin according to the number of time bins Ntb

[0047] S12. Measure the background light noise and record it as the total number of background light noises Nn ( bg ) , and calculate the average value of the background light noise in each time bin according to the number of time bins Ntb

[0048] S13. Calculate Nn (DCR) and Nn ( bg ) The sum is recorded as the total number of noise-induced events Nn, and calculate μn (DCR) and μn ( bg ) The sum is recorded as the average noise μn of each time bin;

[0049] S14. Measure the full width at half maximum FWHM of the solid-state lidar, and calculate the standard deviation of the normal distribution of noise events through FWHM

[0050] S15. Calculate the number of photon-induced events at each pixel point (x, y) of the solid-state lidar and record it as the total number of events N tot(,x,y) , through the total number of events N tot(,x,y) and the total number of noise-induced events N n Calculate the total number of photon-induced events at each pixel point N ph,(x,y) = N tot(,x,y) - N n ; and calculate the average number of photon events in each time bin according to the number of time bins N tb

[0051] S16. Establish a noise estimation probability model p for the actual distribution of noise output i,(x,y) :

[0052]

[0053] where p i,(x,y) represents the probability of any event occurring in time bin i, i = 1, 2,..., N tb, where k is an adjustment factor; when calculating p i,x,y it should follow:

[0054]

[0055] Further, in the step S2, the Monte Carlo method and the noise estimation probability model obtained in the step S16 are used to perform preliminary noise reduction on each frame of the original histogram of the solid-state lidar to obtain multiple frames of two-dimensional histograms after preliminary noise reduction;

[0056] The step S3 is specifically as follows:

[0057] The signal receiving end of the solid-state lidar has an SPAD array with x×y pixel points. Each pixel point corresponds to z time bins in the two-dimensional histogram data, and the length of each time bin is t b ; create a three-dimensional array H′ with dimensions x, y, and z (x,y,z) , and map the x×y pixel point maps of the original histogram H′ to the three-dimensional array H′ (x,y,z) where the first dimension x and the second dimension y represent the coordinates of the pixel points, and the third dimension z represents the number of photons received by the time bins corresponding to the histogram data of each pixel point.

[0058] The present invention also provides a noise reduction system for the histogram of a solid-state lidar, which is characterized in that: the noise reduction system includes a feature extraction module, a spatio-temporal feature enhancement module, a feature aggregation module, and a sequence frame reconstruction module, and the feature extraction module, the spatio-temporal feature enhancement module, the feature aggregation module, and the sequence frame reconstruction module are respectively used to execute the steps S1, S2, S3, and S4 in the above method.

[0059] The optimization method and system for the histogram of a solid-state lidar provided by the present invention first establish a noise estimation probability model for different types of noise suffered by the solid-state lidar, and perform preliminary noise reduction processing on the two-dimensional histogram data of each pixel point collected by the solid-state lidar to suppress the influence of background noise and hardware accidental noise. Subsequently, the processed two-dimensional histogram data is mapped into a three-dimensional histogram, and a cross-frame feature aggregation recurrent neural network is used to perform feature enhancement and secondary noise reduction optimization on the mapped three-dimensional histogram, further improving the signal-to-noise ratio of the data, enhancing the ability of the solid-state lidar to capture fine structures, and significantly improving the accuracy and efficiency of ranging and imaging. Compared with the prior art, the present invention can effectively reduce noise interference, improve imaging quality and positioning accuracy, and has strong practicability and broad application prospects. Description of the Drawings

[0060] Figure 1 is a flowchart of the data noise reduction and feature enhancement method and system for the histogram of a solid-state lidar provided by an embodiment of the present invention;

[0061] Figure 2 is the flowchart of step S4 provided by an embodiment of the present invention;

[0062] Figure 3 is the structural diagram of the continuous residual dense channel attention module RDCAB Cell provided by an embodiment of the present invention;

[0063] Figure 4 is the structural diagram of the spatio-temporal feature enhancement module STFEM provided by an embodiment of the present invention;

[0064] Figure 5 is the structural diagram of the feature aggregation module FAM provided by an embodiment of the present invention. Detailed implementation manners

[0065] The following specifically illustrates the implementation manners of the present invention in conjunction with the accompanying drawings. The given embodiments are only for illustrative purposes and should not be construed as limiting the present invention. The included drawings are only for reference and illustration and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.

[0066] The data denoising and feature enhancement method for the solid-state lidar histogram provided by the embodiment of the present invention is based on the cross-frame feature aggregation recurrent neural network algorithm (step S4), as Figure 1 shown in the flowchart, and the method includes the steps:

[0067] S1. Measure the noise data of the solid-state lidar and establish a noise estimation probability model based on the noise data;

[0068] S2. Use the noise estimation probability model to perform preliminary denoising on each frame of the original histogram of the solid-state lidar pixel by pixel to obtain multiple frames of two-dimensional histograms after preliminary denoising;

[0069] S3. Map the multiple frames of two-dimensional histograms after preliminary denoising into a three-dimensional space to obtain multiple frames of three-dimensional histograms after preliminary denoising;

[0070] S4. Perform cross-frame feature aggregation and denoising optimization on the multiple frames of three-dimensional histograms after preliminary denoising to obtain multiple frames of three-dimensional histograms after secondary denoising and feature enhancement.

[0071] The following elaborates on each step one by one.

[0072] (1) Step S1: Establish a noise estimation probability model

[0073] Step S1 specifically includes steps S11 to S16.

[0074] S11. Measure the dark count rate noise and record the total number of DCR noises Nn (DCR), calculate the average value of the dark count rate noise in each time bin according to the number of time bins Ntb

[0075] Place the solid-state lidar in a completely light-shielded environment to ensure that no external light enters, and set the laboratory environment to maintain a stable temperature. Ensure that the solid-state lidar is in a normal working state, check and adjust the power supply and connections to ensure no interference, and set the detection time of the solid-state lidar to T. Start collecting DCR noise data, record the signals generated by the single-photon avalanche diode (SPAD) of the solid-state lidar under the condition of no light illumination and no actively emitting laser detection signals, and denote the collected signals as the total number of DCR noises Nn (DCR) . Let the number of time bins of the solid-state lidar be Ntb, and calculate the average value of DCR noise μn in each time bin according to the number of time bins (DCR) :

[0076]

[0077] S12. Measure the background light noise and denote it as the total number of background light noises Nn ( bg ) , calculate the average value of the background light noise in each time bin according to the number of time bins Ntb

[0078] Measure the background light noise. Place the solid-state lidar in the specific natural environment where it will work. Ensure that the solid-state lidar is in a normal working state as in step S11, and set the same detection time T. Start collecting background light noise data, record the signals generated by the SPAD of the solid-state lidar under the condition of being in the specific natural environment and not actively emitting laser detection signals, and denote the collected signals as the total number of background light noises N n(bg) . Similarly to formula (1), calculate the average value of background light noise μ in each time bin according to the number of time bins n(bg) .

[0079] S13. Calculate N n(DCR) and N n(bg) The sum is denoted as the total number of noise-induced events N n , calculate μ n(DCR) and μ n(bg) The sum is denoted as the average noise μ in each time bin n .

[0080] Calculate the sum of DCR noise and background light noise and denote it as the total number of noise-induced events N n (as shown in formula (2) below). Within a specific time, the average number of noise events follows a uniform distribution over time. Calculate the sum of the average value of DCR noise and the average value of background light noise in each time bin and denote it as the average noise μ in each time binn :

[0081] N n = N n(DCR) + N n(bg) (2)

[0082] μ n = μ n(DCR) + μ n(bg) (3)

[0083] S14. Measure the full width at half maximum (FWHM) of the solid-state lidar, and calculate the standard deviation σ of the normal distribution of noise events through the FWHM.

[0084] The time jitter of the solid-state lidar sensor represents the combination of all uncertainties in measuring the time of received photons. Therefore, the normal distribution of noise events can be calculated through the measurement of time jitter. Usually, the time jitter is measured using the full width at half maximum (FWHM), and the standard deviation σ of the normal distribution of noise events can be calculated through the value of FHWM:

[0085]

[0086] The specific process of measuring the full width at half maximum (FWHM) of the solid-state lidar is as follows:

[0087] Prepare a lidar system composed of a laser emitter and a single-photon avalanche diode (SPAD) and connect it to an oscilloscope to record the reflected signal. Place the solid-state lidar in a well-controlled environment to reduce external interference, and select a target object with good reflection characteristics and place it within the measurement range of the lidar. Set the detection time to T, start the lidar to emit pulses, and capture the pulses reflected from the target object through the SPAD. Organize the captured pulse signals, analyze the signal waveform using MATLAB, find the maximum amplitude point of the signal, and calculate half of its value. Then, locate two points in the waveform with amplitudes equal to half of the maximum amplitude, and measure the time difference between these two points. This time difference is the FWHM. Repeat the measurement multiple times and calculate the average value of the FWHM.

[0088] S15. Calculate the number of photon-induced events at each pixel point (x, y) of the solid-state lidar, denoted as the total number of events N tot(,x,y) , and calculate the total number of photon-induced events N tot(,x,y) at each pixel point through the total number of events N n and the total number of noise-induced events N ph,(x,y) ; and calculate the average number of photon events μ tb in each time bin according to the number of time bins N ph,(x,y) .

[0089] Calculate the number of photon-induced events for each pixel of the solid-state lidar. The signal receiving end of the solid-state lidar consists of an SPAD array with x×y pixels, denoted by H (x,y) represents the original histogram data of each pixel generated by the solid-state lidar, where (x, y) represents the coordinates of each pixel. Place the solid-state lidar in the specific natural environment where it will operate, and ensure that the solid-state lidar is in a normal working state as in step S11. Set the detection time to T, activate the lidar to emit pulses, and capture the pulses reflected from the target object through the SPAD. Denote the total number of pulse signals of each pixel collected as the total number of events N tot(,x,y) . Since DCR and background light noise will have the same noise effect on each pixel, the total number of photon-induced events N tot(,x,y) of each pixel is calculated through the total number of events N n and the total number of noise-induced events N ph,(x,y) :

[0090] N ph,(x,y) = N tot(,x,y) - N n (5)

[0091] Similarly to formula (1), calculate the average number of photon events μ ph,(x,y) in each time bin according to the number of time bins.

[0092] S16. Establish a noise estimation probability model p i,(x,y) for the actual distribution of the noise output.

[0093] Through steps S13 - S15, combine the obtained simulated noise uniform distribution and normal distribution to generate the actual distribution of the noise output of the solid-state lidar sensor. Establish a noise estimation probability model p i,(x,y) for the actual distribution of the noise output:

[0094]

[0095] where p i,(x,y) represents the probability of any event occurring in time bin i; k is an adjustment factor that compensates for the limitation of the finite time bins of the solid-state lidar.

[0096] To ensure that is, the sum of the probabilities of all time bins is equal to 1, the following formula (7) should be followed when calculating p i,(x,y) .

[0097]

[0098] The advantage of the present invention in establishing the noise estimation probability model using the above steps is that:

[0099] ① Accuracy: This model combines the uniform distribution and the normal distribution, enabling it to more accurately simulate the actual noise distribution and reflect the true performance of solid-state lidar under different noise conditions.

[0100] ② Robustness: Through the accurate measurement and modeling of DCR and background light noise, this method can effectively eliminate the interference of different environmental noises and improve the robustness of the system.

[0101] ③ Flexibility: The adjustment factor k in the model can be flexibly adjusted to adapt to different working environments and noise levels, improving the applicability and universality of the model.

[0102] ④ Improvement of signal-to-noise ratio: Through accurate noise estimation and noise reduction processing, this method can significantly improve the signal-to-noise ratio of the solid-state lidar system, enhancing the data quality and subsequent processing effects.

[0103] In summary, the adoption of this noise estimation probability model not only improves the accuracy and efficiency of data noise reduction but also enhances the adaptability of the system under various environmental conditions, significantly improving the feature enhancement effect of the solid-state lidar histogram.

[0104] (2) Step S2: Perform preliminary noise reduction on each frame of the original histogram of the solid-state lidar

[0105] Specifically, this step uses the Monte Carlo method and the noise estimation probability model obtained in step S16 to generate simulated two-dimensional histogram data for the original histogram.

[0106] This step S2 specifically includes the following steps:

[0107] S21. For each pixel point (x, y) of the solid-state lidar, perform the following operations. For each time bin i, from 1 to N tb Conduct simulations, and set the initial event count to 0 for each time bin.

[0108] S22. For this pixel point, obtain the noise estimation probability p calculated in step S16 i,(x,y) .

[0109] S23. For each simulation, generate a random number r uniformly distributed in [0, 1] i . Compare the noise estimation probability p i,(x,y) of each time bin with the random number r i . If r i is less than or equal to p i,(x,y) , then add an event in the corresponding time window.

[0110] S24. Repeat step S23 N tot,(x,y) times for each pixel point to generate the simulated histogram data H sim,(x,y) .

[0111] S25. Use the method of statistical matching to utilize the simulated histogram data H sim,(x,y) For the original histogram data H generated by the solid-state lidar (x,y) Perform noise reduction. For each pixel point, calculate the mean and variance of the original histogram data and the simulated histogram data. Perform noise reduction on each time bin of each pixel point (as shown in formula (8) below) to obtain the number of events H' in each time bin of the noise-reduced histogram data i,(x,y) :

[0112]

[0113] Among them, H' i,(x,y) Represents the number of events in each time bin of the noise-reduced histogram data; H i,(x,y) Represents the number of events in each time bin of the original histogram data; μ orig And σ orig Represent the mean and variance of the original histogram data; μ sim And σ sim Represent the mean and variance of the simulated histogram data.

[0114] Sort out the number of events in each time bin of the obtained noise-reduced histogram data, and the noise-reduced histogram data H' can be generated (x,y) .

[0115] S26. For each pixel point, repeating step S25 can obtain the two-dimensional histogram data H' after preliminary noise reduction for all pixel points of the solid-state lidar.

[0116] The advantages of the present invention in performing preliminary noise reduction on each frame of the original histogram of the solid-state lidar by the above steps are as follows: By using the Monte Carlo method for simulation, this method can generate high-quality noise-reduced histogram data in a short time, improving the processing efficiency. In addition, the application of the noise estimation probability model ensures the data accuracy during the noise reduction process, enabling the processed histogram to more accurately reflect the real environmental situation, thereby improving the accuracy of the system. The method of statistical matching effectively reduces the interference of different noise sources on the histogram data, enhances the noise resistance of the system, and reflects the robustness of the method. At the same time, this method is applicable to solid-state lidars of different types and different working environments, has strong adaptability and universality, and shows a high degree of flexibility. Finally, the quality of the noise-reduced histogram data is higher, providing a more reliable data basis for subsequent three-dimensional feature aggregation and noise reduction optimization, and significantly improving the overall performance of the system.

[0117] (3) Step S3: Histogram data mapping

[0118] This step is specifically as follows:

[0119] The SPAD array at the signal receiving end of the solid-state lidar has x×y pixel points, and the two-dimensional histogram data corresponding to each pixel point has z time bins (bins), and the length of each time bin is t b . Create a three-dimensional array H′ with dimensions x, y, and z (x,y,z) , and map the two-dimensional histogram data obtained from the x×y pixel points through step S19 to the three-dimensional array H′ (x,y,z) . Among them, the first dimension x and the second dimension y represent the coordinates of the pixel points, and the third dimension z represents the number of photons received in the time bins corresponding to the histogram data of each pixel point.

[0120] (4) Step S4

[0121] For each frame of the three-dimensional histogram after preliminary noise reduction, S4 specifically includes steps S41 to S44:

[0122] S41. Extract the 2M + 1 initial features of the 2M + 1-layer three-dimensional histogram sequence of the three-dimensional histogram after preliminary noise reduction with a frame as the center, where M is a natural number not less than 1.

[0123] Step S41 specifically includes steps S411 to S413:

[0124] S411. Obtain the three-dimensional histogram after preliminary noise reduction of the current frame and the previous M frames and the next M frames of the current frame to form the three-dimensional histogram sequence of the current frame after preliminary noise reduction.

[0125] In this example, M = 11 is used for illustration. In this step, a continuous sequence of 11 extremely short (0.00036 ms / frame) three-dimensional histogram data frames is used as the input, and the set {H′ t-5,(x,y,z) , H′ t-4,(x,y,z) , … H′ t,(x,y,z) , … H′ t+4,(x,y,z) , H′ t+5,(x,y,z)} is used to represent the three-dimensional histogram sequence of the current frame after preliminary noise reduction, where H′ t,(x,y,z) represents the three-dimensional histogram of the current frame after preliminary noise reduction. H′ t-5,(x,y,z) to H′ t-1,(x,y,z) represent the three-dimensional histograms of the first five frames of H′ t,(x,y,z) after preliminary noise reduction, and H′ t+1,(x,y,z) to H′ t+5,(x,y,z) represent the three-dimensional histograms of the last five frames of H′ t,(x,y,z) after preliminary noise reduction.

[0126] S412. Input the first frame of the pre-denoised 3D histogram sequence of the current frame and the initial hidden state into the first residual dense channel attention module for feature extraction and hidden state generation, obtaining the corresponding first-layer hidden state and first-layer initial features;

[0127] S413. Input the second frame of the pre-denoised 3D histogram sequence and the first-layer hidden state into the second residual dense channel attention module for feature extraction and hidden state generation, obtaining the corresponding second-layer initial features and second-layer hidden state; and so on in a loop until the (2M + 1)-th frame of the pre-denoised 3D histogram sequence and the 2M-th layer hidden state are input into the (2M + 1)-th residual dense channel attention module for feature extraction, obtaining the corresponding (2M + 1)-th layer initial features.

[0128] Steps S412 and S413, simply put, are to input 11 frames of extremely short 3D histogram data sequences into 11 consecutive residual dense channel attention modules RDCAB Cell (the structure of RDCAB Cell is as Figure 2 shown), obtaining the hidden states {h′ t-5,(x,y,z) ,h′ t-4,(x,y,z) ,…h′ t,(x,y,z) ,…h′ t+4,(x,y,z) ,h′ t+5,(x,y,z)} of the sequence frames and multi-layer initial features {X′ t-5,(x,y,z) ,X′ t-4,(x,y,z) ,…X′ t,(x,y,z) ,…X′ t+4,(x,y,z) ,X′ t+5,(x,y,z)}.

[0129] In this embodiment, the first residual dense channel attention module to the 2M-th residual dense channel attention module all adopt the structure of the residual dense channel attention module. The structure of the residual dense channel attention module is as Figure 3 shown, and the operations it performs are:

[0130] First, perform a downsampling operation on the input pre-denoised 3D histogram, then perform channel connection on the downsampling result and the input hidden state, then perform N ≥ 2 times of residual dense operations and channel attention operations on the channel connection result to obtain N intermediate features, then perform channel connection on the N intermediate features and then perform 1×1 convolution to obtain the initial features of the pre-denoised 3D histogram, and then obtain the corresponding hidden state through the hidden state generation function. The structure adopted by the 2M-th residual dense channel attention module lacks the hidden state generation function compared with the residual dense channel attention module.

[0131] Each residual dense operation and channel attention operation includes performing a channel attention operation after performing a residual dense operation once;

[0132] The residual dense operation is as follows:

[0133] Perform a 3×3 convolution and ReLU function activation on the first input to obtain a first output;

[0134] Perform a 3×3 convolution and ReLU function activation on the first input and the first output as the second input to obtain a second output;

[0135] Perform a 3×3 convolution and ReLU function activation on the second output, the first input, and the second input as the third input to obtain a third output;

[0136] Perform a 1×1 convolution on the third output, the first input, the second input, and the third input, and then add the result element-wise to the first input to obtain the residual dense output;

[0137] The channel attention operation is as follows:

[0138] Perform global average pooling, a fully connected layer, ReLU function activation, a fully connected layer, and Sigmoid function activation on the residual dense output, and then multiply the result element-wise with the residual dense output to obtain the channel attention output;

[0139] The hidden state generation function sequentially performs a 3×3 convolution, a residual dense operation, and a 3×3 convolution once.

[0140] Taking the three-dimensional histogram data H′ of the t-th frame t,(x,y,z) as an example, H′ t,(x,y,z) and the hidden layer state h′ at time t-1 t-1,(x,y,z) are input into the RDCAB Cell (at the beginning of the algorithm, the hidden state h′ at time t-1 1,(x,y,z) is initialized to 0).

[0141] Perform downsampling on the input three-dimensional histogram data H′ of the t-th frame t,(x,y,z) and connect it with the hidden state h′ at the previous time t-1,(x,y,z) to obtain the shallow feature

[0142]

[0143] where concat() represents the operation of channel connection; f DS () represents the downsampling operation.

[0144] The shallow feature Feed into a network structure composed of N RDCABs. Each RDCAB structure consists of a Residual Dense Block (RDB) and a Channel Attention Module (CAM). For each RDCAB, first extract multi-level features through the RDB, and then strengthen the channel features through the CAM to obtain features

[0145]

[0146] Among them, f RDB () represents the RDB operation. The RDB consists of 3 3×3 convolutional layers, 1 1×1 convolutional layer, dense connection and residual connection; f CAM () represents the CAM operation. The CAM consists of Global Average Pooling (GAP), 2 fully connected layers Linear, ReLU activation function and Sigmoid function; R n represents the nth RDCAB.

[0147] Connect the features extracted by N RDCABs channel-wise, and then obtain the multi-level features X′ of the three-dimensional histogram through a 1×1 convolution t,(x,y,z) :

[0148]

[0149] Among them, concat represents feature connection; conv represents a 1×1 convolution.

[0150] Input the obtained multi-level three-dimensional histogram feature X′ t,(x,y,z) into the hidden state generation function H() composed of two 3×3 convolutions and an RDB to obtain the hidden state h′ at the current time t t,(x,y,z) :

[0151] h′ t,(x,y,z) = H(X′ t,(x,y,z) ) (12)

[0152] The advantages of using the above continuous residual dense channel attention module for initial feature extraction in the present invention are as follows: By using the continuous residual dense channel attention module for initial feature extraction, this method can significantly improve the efficiency and accuracy of feature extraction. First of all, the residual dense module (RDCAB) can effectively integrate feature information of different scales, enhance the network's ability to capture detailed features, and thus improve the accuracy of feature extraction. In addition, the channel attention module (CAM) adaptively adjusts the weights of different channels, enabling the network to pay more attention to important feature channels, suppressing irrelevant noise information, and improving the robustness of feature extraction. The design of the continuous residual dense channel attention module enables the network to layer by layer extract and integrate the feature information of the multi-frame preliminary denoised three-dimensional histograms while maintaining efficient computation, enhancing the adaptability and generalization ability of the overall system. Finally, this method not only improves the denoising effect and feature extraction quality, but also provides a reliable data basis for subsequent cross-frame feature aggregation and denoising optimization, significantly enhancing the performance of the solid-state lidar system in various complex environments.

[0153] S42. Apply the spatio-temporal feature enhancement module (STFEM) as shown in Figure 4 to the other 2M layers of initial features except for the initial features corresponding to the preliminary denoised three-dimensional histogram of this frame to perform spatio-temporal feature enhancement, obtaining 2M layers of spatio-temporal enhanced features.

[0154] In step S42, performing spatio-temporal feature enhancement for each layer of the 2M layers of initial features specifically includes: performing spatio-temporal context enhancement on the input initial features to obtain spatio-temporal context enhanced features; performing spatio-temporal feature extraction on the spatio-temporal context enhanced features to obtain the spatio-temporal enhanced features corresponding to this initial feature;

[0155] Performing spatio-temporal context enhancement specifically is:

[0156] Performing 3×3 convolution on the input initial features and shifting by θ, performing deformable convolution on the shifted features and the initial features, and then performing 3×3 convolution to obtain spatio-temporal context enhanced features;

[0157] Performing spatio-temporal feature extraction specifically is:

[0158] Performing 7×7 convolution and Sigmoid function activation on the input spatio-temporal context enhanced features, and then performing element-wise multiplication with the spatio-temporal context enhanced features to obtain spatio-temporal enhanced features.

[0159] Taking the multi-level features X′ t-1,(x,y,z) of the three-dimensional histogram data H′ t-1,(x,y,z) of the (t - 1)-th frame as an example, first input the multi-level features X′ t-1,(x,y,z) into the first spatio-temporal context enhancement structure TCE of the STFEM to obtain enhanced features

[0160]

[0161] Among them, TCE() represents the spatio-temporal context enhancement structure, which is composed of deformable convolution and 3×3 convolution; θ = conv(X′ t-1,(x,y,z) ) represents the offset generated by convolution; D() represents the deformable convolution operation.

[0162] Input the enhanced feature into the second spatio-temporal feature extraction structure SFE of STFEM to obtain the enhanced extraction feature

[0163]

[0164] Among them, SFE() represents the spatio-temporal feature extraction structure, which is composed of 7×7 convolution and Sigmoid function. First, compress the input enhanced feature along the channel dimension to obtain a single-channel feature map of the same size; then use the 7×7 convolution with a large receptive field and the Sigmoid function to calculate the spatial attention map, that is, the spatio-temporal domain weight; finally, apply the corresponding weight to the input feature. S() represents the Sigmoid operation; represents element-wise multiplication of tensors.

[0165] Performing the above operations on the multi-level features of the three-dimensional histogram sequence frames (except for the multi-level features X′ t,(x,y,z) at the t-th moment) {X′ t-5,(x,y,z) ,…,X′ t-1,(x,y,z) ,X′ t+1,(x,y,z) ,…,X′ t+5,(x,y,z)} in sequence can obtain the enhanced extraction features of the three-dimensional histogram sequence frames

[0166] The advantages of the present invention in using the above spatio-temporal feature enhancement module to enhance the initial features are as follows: Through the spatio-temporal feature enhancement module (STFEM), the present invention can significantly improve the feature expression ability of the three-dimensional histogram data of solid-state lidar. First, the spatio-temporal context enhancement structure (TCE) enables the network to flexibly adjust the position and shape of the convolution kernel by introducing deformable convolution operations, thereby more effectively capturing complex spatio-temporal change features. This flexibility greatly improves the accuracy and robustness of feature extraction. Second, the spatio-temporal feature extraction structure (SFE) generates a spatial attention map by combining a 7×7 convolution with a large receptive field and a Sigmoid activation function, thereby enhancing the attention to important feature regions and suppressing irrelevant background noise. The design of this module makes feature extraction more comprehensive and accurate, effectively improving the performance of the system in complex environments. In addition, the spatio-temporal feature enhancement module can significantly improve the spatio-temporal consistency and discriminative ability of features without increasing significant computational overhead, providing a more reliable data basis for subsequent feature aggregation and noise reduction optimization, and ultimately improving the performance and stability of the entire solid-state lidar system.

[0167] S43. Aggregate the 2M-layer spatio-temporal enhanced features with the initial features of the three-dimensional histogram after preliminary noise reduction of this frame respectively to obtain the corresponding 2M-layer aggregated features.

[0168] This step uses the Figure 5 feature aggregation module FAM shown as to aggregate the enhanced extraction features of the three-dimensional histogram sequence frames after being enhanced and extracted by the spatio-temporal feature enhancement module ( x, y, z ) and the multi-level features Xt′ at the t-th moment

[0169]

[0170] wherein, FAM() represents the feature aggregation structure, which is composed of a global average pooling fusion GAP Fusion module and a 1×1 convolution. First, the multi-level features Xt′ of the current frame ( x, y, z ) and the enhanced extraction features of the remaining adjacent frames are input into GAP Fusion for fusion in turn; then a channel attention mechanism is introduced to enhance the feature expression in the channel dimension to obtain channel-enhanced features. GAPF() represents the global average pooling fusion GAP Fusion operation, which is composed of global average pooling, two fully connected layers Linear, ReLU activation function, Sigmoid function and three 1×1 convolutions; Indicates element-wise multiplication of tensors.

[0171] The advantages of the present invention in using the above-mentioned feature aggregation module to aggregate initial features are as follows: Through the Feature Aggregation Module (FAM), the present invention can effectively fuse the feature information of different time frames, thereby enhancing the integrity and accuracy of feature representation. First, the FAM module fuses the multi-level features of the current frame with the enhanced extraction features of adjacent frames through the Global Average Pooling Fusion (GAP Fusion) operation, realizing the effective complement of feature information of different time frames. This can not only retain the detailed features of each frame but also further enhance the feature expression ability by aggregating the information of adjacent frames. Second, the introduced channel attention mechanism can highlight important feature channels and suppress the interference of irrelevant information by adaptively weighting the features in the channel dimension, significantly enhancing the discrimination ability and robustness of the features. Finally, the features are further compressed and fused through 1×1 convolution operations, effectively reducing the computational complexity and ensuring efficient feature processing. Overall, using the feature aggregation module to aggregate the initial features not only improves the accuracy and integrity of feature representation but also enhances the anti-interference ability of the system in complex environments, significantly improving the performance and stability of the solid-state lidar system.

[0172] S44. Combine the aggregated features of the 2M layer with the initial features of the preliminarily denoised three-dimensional histogram of this frame to obtain the secondarily denoised and feature-enhanced three-dimensional histogram of this frame.

[0173] The aggregated enhanced sequence features obtained through step S43 and the multi-level features X′ of the current frame (at the t-th moment) t,(x,y,z) are input into a three-dimensional histogram sequence frame reconstruction module composed of concat and 1×1 convolution to convert the aggregated enhanced histogram feature map into the current three-dimensional histogram data frame with feature enhancement and secondary denoising.

[0174]

[0175] Based on the above method, an embodiment of the present invention also provides a noise reduction system for a solid-state lidar histogram, as Figures 1 to 5 shown. This noise reduction system includes a feature extraction module, a spatio-temporal feature enhancement module, a feature aggregation module, and a sequence frame reconstruction module. The feature extraction module, spatio-temporal feature enhancement module, feature aggregation module, and sequence frame reconstruction module are respectively used to execute steps S1, S2, S3, and S4 in the above method.

[0176] In summary, for the method and system for optimizing the histogram of a solid-state lidar provided by the embodiments of the present invention, a noise estimation probability model is first established for different types of noise suffered by the solid-state lidar, and preliminary noise reduction processing is performed on the two-dimensional histogram data of each pixel point collected by the solid-state lidar to suppress the influence of background noise and hardware accidental noise. Subsequently, the processed two-dimensional histogram data is mapped into a three-dimensional histogram, and a cross-frame feature aggregation recurrent neural network is used to perform feature enhancement and secondary noise reduction optimization on the mapped three-dimensional histogram, further improving the signal-to-noise ratio of the data, enhancing the ability of the solid-state lidar to capture fine structures, and significantly improving the accuracy and efficiency of ranging and imaging. Compared with the prior art, the present invention can effectively reduce noise interference, improve imaging quality and positioning accuracy, and has strong practicability and broad application prospects.

[0177] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent substitution methods and are all included in the protection scope of the present invention.

Claims

1. A method for data noise reduction and feature enhancement of the histogram of a solid-state lidar, characterized in that, Including the steps: S1. Measure the noise data of the solid-state lidar and establish a noise estimation probability model based on the noise data; S2. Use the noise estimation probability model to perform preliminary noise reduction on each frame of the original histogram of the solid-state lidar pixel by pixel to obtain multiple frames of two-dimensional histograms after preliminary noise reduction; S3. Map the multiple frames of the two-dimensional histograms after preliminary noise reduction into a three-dimensional space to obtain multiple frames of three-dimensional histograms after preliminary noise reduction; S4. Perform cross-frame feature aggregation and noise reduction optimization on the multiple frames of the three-dimensional histograms after preliminary noise reduction to obtain multiple frames of three-dimensional histograms after secondary noise reduction and feature enhancement; For each frame of the three-dimensional histogram after preliminary noise reduction, the S4 specifically includes the steps: S41. Extract the 2M + 1-layer initial features of the three-dimensional histogram sequence after preliminary noise reduction with a length of 2M + 1 frames centered on the three-dimensional histogram after preliminary noise reduction of one frame, where M is a natural number not less than 1; S42. Perform spatio-temporal feature enhancement on the other 2M layers of initial features except the initial features corresponding to the three-dimensional histogram after preliminary noise reduction of this frame to obtain 2M layers of spatio-temporal enhanced features; S43. Aggregate the 2M layers of spatio-temporal enhanced features with the initial features of the three-dimensional histogram after preliminary noise reduction of this frame respectively to obtain the corresponding 2M layers of aggregated features; S44. Combine the 2M layers of aggregated features with the initial features of the three-dimensional histogram after preliminary noise reduction of this frame to obtain the three-dimensional histogram after secondary noise reduction and feature enhancement of this frame.

2. The data noise reduction and feature enhancement method for the solid-state lidar histogram according to claim 1, wherein, The step S41 specifically includes the steps: S411. Obtain the three-dimensional histograms after preliminary noise reduction of the current frame and the first M frames and the last M frames before the current frame to form a sequence of three-dimensional histograms after preliminary noise reduction of the current frame; S412. Input the first three-dimensional histogram after preliminary noise reduction of the sequence of three-dimensional histograms after preliminary noise reduction of the current frame and the initial hidden state into the first residual dense channel attention module for feature extraction and hidden state generation to obtain the corresponding first-layer hidden state and first-layer initial features; S413. Input the second three-dimensional histogram after preliminary noise reduction of the sequence of three-dimensional histograms after preliminary noise reduction and the first-layer hidden state into the second residual dense channel attention module to perform feature extraction and hidden state generation to obtain the corresponding second-layer initial features and second-layer hidden states; and so on in a loop until the (2M + 1)-th three-dimensional histogram after preliminary noise reduction of the sequence of three-dimensional histograms after preliminary noise reduction and the 2M-th hidden state are input into the (2M + 1)-th residual dense channel attention module for feature extraction to obtain the corresponding (2M + 1)-th layer of initial features.

3. The data noise reduction and feature enhancement method for the histogram of the solid-state lidar according to claim 2, wherein The first residual dense channel attention module to the 2M-th residual dense channel attention module all adopt the structure of the residual dense channel attention module; The operations performed by the residual dense channel attention module are: First, perform downsampling on the initially denoised three-dimensional histogram of the input. Then, concatenate the result of downsampling with the hidden state of the input along the channel dimension. Next, perform N≥2 times of residual dense operations and channel attention operations on the result of channel concatenation to obtain N intermediate features. Then, concatenate the N intermediate features along the channel dimension and perform 1×1 convolution to obtain the initial feature of the initially denoised three-dimensional histogram. Finally, obtain the corresponding hidden state from the initial feature through the hidden state generation function; The structure adopted by the 2M-th residual dense channel attention module lacks the hidden state generation function compared to the residual dense channel attention module.

4. The data noise reduction and feature enhancement method for the histogram of the solid-state lidar according to claim 2, wherein: Each residual dense operation and channel attention operation includes performing a channel attention operation after executing a residual dense operation once; The residual dense operation is as follows: Perform 3×3 convolution and ReLU function activation on the first input to obtain the first output; Perform 3×3 convolution and ReLU function activation on the first input and the first output as the second input to obtain the second output; Perform 3×3 convolution and ReLU function activation on the second output, the first input, and the second input as the third input to obtain the third output; Perform 1×1 convolution on the third output, the first input, the second input, and the third input, and then add the result element-wise to the first input to obtain the residual dense output; The channel attention operation is as follows: Perform global average pooling, fully connected layer, ReLU function activation, fully connected layer, and Sigmoid function activation on the residual dense output, and then multiply the result element-wise with the residual dense output to obtain the channel attention output; The hidden state generation function sequentially performs 3×3 convolution, a residual dense operation, and 3×3 convolution once.

5. The data noise reduction and feature enhancement method for the histogram of the solid-state lidar according to claim 1, characterized in that, In step S42, for each initial feature among the 2M layers of initial features, performing spatio-temporal feature enhancement specifically includes: performing spatio-temporal context enhancement on the input initial feature to obtain spatio-temporal context enhanced features; performing spatio-temporal feature extraction on the spatio-temporal context enhanced features to obtain the spatio-temporal enhanced features corresponding to the initial feature; Performing spatio-temporal context enhancement specifically is: Perform 3×3 convolution on the input initial feature and shift by θ, perform deformable convolution on the shifted feature and the initial feature, and then perform 3×3 convolution to obtain spatio-temporal context enhanced features; Performing spatio-temporal feature extraction specifically is: Perform 7×7 convolution and Sigmoid function activation on the input spatio-temporal context enhanced features, and then multiply the result element-wise with the spatio-temporal context enhanced features to obtain spatio-temporal enhanced features.

6. The data noise reduction and feature enhancement method for the solid-state lidar histogram according to claim 1, characterized in that In step S43, aggregating any one of the 2M layers of spatio-temporal enhanced features with the initial feature of the initially denoised three-dimensional histogram of this frame specifically is: Perform feature connection on this layer of spatio-temporal enhanced features and its initial feature to obtain connection features; Perform global average pooling fusion on the connection features to obtain fusion features; Extract channel attention from the connection features to obtain attention features; Multiply the fusion features and the attention features element-wise to obtain the aggregated features corresponding to this layer of initial features; The global average pooling fusion is specifically as follows: perform global average pooling, a fully connected layer, ReLU function activation, a fully connected layer, and Sigmoid function activation on the concatenated features to obtain the first intermediate feature, perform 1×1 convolution and 1×1 convolution on the concatenated features to obtain the second intermediate feature, multiply the first intermediate feature and the second intermediate feature element-wise, and then perform 1×1 convolution to obtain the fused feature.

7. The method for data noise reduction and feature enhancement of the histogram of the solid-state lidar according to claim 6, wherein The specific steps of step S44 are as follows: perform feature concatenation on the 2M-layer aggregated feature and the initial feature of the initially denoised three-dimensional histogram of this frame, and then perform 1×1 convolution to obtain the secondarily denoised and feature-enhanced three-dimensional histogram of this frame.

8. The method for data noise reduction and feature enhancement of the histogram of the solid-state lidar according to claim 1, wherein, The specific steps of step S1 include the following steps: S11. Measure the dark count rate noise, denoted as the total number of DCR noises Nn (DCR) , and calculate the average value of the dark count rate noise in each time bin according to the number of time bins Ntb S12. Measure the background light noise and record it as the total background light noise Nn ( bg ) , and calculate the average background light noise in each time bin according to the number of time bins Ntb S13. Calculate Nn (DCR) and Nn ( bg ) The sum of them is denoted as the total number of noise-induced events Nn, and calculate μn (DCR) and μ n(bg) The sum of them is denoted as the average noise μ per time bin n ; S14. Measure the full width at half maximum (FWHM) of the solid-state lidar, and calculate the standard deviation of the normal distribution of noise events based on the FWHM S15. Calculate the number of photon-induced events at each pixel point (x, y) of the solid-state lidar, denoted as the total number of events N. tot(,x,y) , through the total number of events N tot(,x,y) and the total number of noise-induced events N n Calculate the total number of photon-induced events at each pixel point N ph,(x,y) = N tot(,x,y) - N n ; and according to the number of time bins N tb Calculate the average number of photon events in each time bin S16. Establish a noise estimation probability model p for the actual distribution of the noise output i,(x,y) : where p i,(x,y) represents the probability of any event occurring in time bin i, where i = 1, 2, …, N tb , and k is a regulation factor; when calculating p i,x,y , the following should be followed:

9. The data denoising and feature enhancement method for the solid-state lidar histogram according to claim 8, characterized in that: In step S2, use the Monte Carlo method and the noise estimation probability model obtained in step S16 to perform initial denoising on the original histograms of each frame of the solid-state lidar to obtain multiple frames of initially denoised two-dimensional histograms; The specific steps of step S3 are as follows: The SPAD array at the signal receiving end of the solid-state lidar has x×y pixel points, and the two-dimensional histogram data corresponding to each pixel point has z time bins, and the length of each time bin is t b ; Create a three-dimensional array H′ with dimensions x, y, and z (x,y,z) , and map the x×y pixel map of the original histogram H′ to the three-dimensional array H ( ′x,y,z ) . Among them, the first dimension x and the second dimension y represent the coordinates of the pixel points, and the third dimension z represents the number of photons received by the time bins corresponding to the histogram data of each pixel point.

10. Noise reduction system for solid-state lidar histogram, characterized in that: The denoising system includes a feature extraction module, a spatio-temporal feature enhancement module, a feature aggregation module, and a sequence frame reconstruction module. The feature extraction module, the spatio-temporal feature enhancement module, the feature aggregation module, and the sequence frame reconstruction module are respectively used to execute steps S1, S2, S3, and S4 described in any one of claims 1 to 9.