Radar Data Processing Method and System
By using improved wavelet denoising algorithm and 3D convolutional neural network in solid-state lidar, combined with dynamic time stereo technology and motion compensation algorithm, the peak of the three-dimensional histogram was extracted, solving the accuracy and robustness of peak extraction in dynamic environments, and improving point cloud imaging quality and positioning accuracy.
Patent Information
- Application Number
- CN202410870564.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-07-01
AI Technical Summary
In solid-state lidar, peak extraction faces many challenges such as data noise and error, balance of resolution and computing complexity, multi-peak recognition and separation, dynamic environment adaptation, occlusion and reflection intensity changes, computing resource limitations, environmental complexity, data diversity, real-time requirements, and algorithm robustness and reliability, which affect point cloud imaging quality and positioning accuracy.
The three-dimensional histogram is denoised by an improved wavelet denoising algorithm based on an adjustable threshold function, and the bird's-eye viewing angle feature is converted into perspective viewing angle feature using a 3D convolutional neural network, and the depth estimation and peak value are adjusted through dynamic time stereo technology and motion compensation algorithm.
It effectively solves the problem of mismatch in local features of long-distance or small-size objects, improves the imaging quality and positioning accuracy of solid-state lidar point clouds, and is suitable for real-time applications such as autonomous driving, drone and robot navigation.
Smart Images

Figure CN118735811B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a method and system for processing radar data. Background Art
[0002] Solid-state lidar is a radar technology that uses semiconductor materials and solid-state components for optical measurement. It has a wide range of applications in many fields (such as autonomous driving, drones, robot navigation, etc.). Its main advantages include small size, high reliability, good seismic resistance, etc. Solid-state lidar measures the distance and shape of objects by emitting laser pulses and receiving the reflected signals, generating three-dimensional point cloud data.
[0003] The amount of original point cloud data is huge. By calculating the peak value of the three-dimensional histogram, key features in the data can be extracted for effective dimensionality reduction processing. This not only reduces the burden of data processing and storage, but also improves the running efficiency of subsequent algorithms. In a dynamic environment, tracking the movement trajectory of an object is the key to realizing dynamic obstacle avoidance and path planning. Through peak extraction, the position and movement state of the object can be accurately identified, thereby realizing accurate motion estimation and trajectory tracking. Real-time calculation of the peak value of the three-dimensional histogram can provide timely and accurate environmental perception information for the system, enhancing the robustness and reliability of the system. This is particularly important in applications such as autonomous driving and industrial robots that require real-time decision-making. Through effective peak extraction and utilization, the system performance can be optimized, the computational complexity can be reduced, and the processing speed and response time can be improved. This is crucial for resource-constrained embedded systems and real-time applications.
[0004] Peak extraction faces many challenges in the three-dimensional histogram processing of solid-state lidar, including data noise and errors, the balance between resolution and computational complexity, multi-peak identification and separation, adaptation to dynamic environments, occlusion and reflection intensity changes, computational resource limitations, environmental complexity, data diversity, real-time requirements, and the robustness and reliability of the algorithm. Noise and errors can lead to false peaks, affecting accuracy; increasing the resolution improves accuracy while increasing computational complexity; multi-peaks in complex scenarios require a fine algorithm for separation; changes in objects and backgrounds in a dynamic environment require the algorithm to be adjusted in real time; occlusion and reflection changes result in uneven data, which requires compensation and correction; computational resources are limited in real-time applications, and the algorithm must be efficient; complex terrain and multi-path reflections increase the difficulty; different application scenarios require the algorithm to have good generality and adaptability; real-time applications require the algorithm to respond quickly; ultimately, the algorithm must remain efficient and stable in complex and uncertain environments.
[0005] There are various methods for peak extraction, including methods based on gradient, curvature, threshold, smoothing and filtering, clustering, and machine learning. These methods face many deficiencies, such as being sensitive to noise, having high computational complexity, strong parameter dependence, poor adaptability to dynamic scenes, insufficient robustness and generality, and poor performance in dealing with occlusion and sparse data. Especially machine learning methods, which rely on a large amount of training data, may have poor generalization ability and performance for new scenarios or scenarios with little data. For objects at a distance or small-sized objects, local features do not match, and the histogram data density changes with distance, making the data very sparse, making it difficult for the histogram to extract effective peaks. Summary of the Invention
[0006] The present invention provides a radar data processing method and system, and the technical problem to be solved is: how to accurately extract histogram peaks to improve the point cloud imaging quality and positioning accuracy of solid-state lidar.
[0007] To solve the above technical problems, the present invention provides a radar data processing method, including:
[0008] Using an improved wavelet denoising algorithm based on an adjustable threshold function to perform denoising processing on the three-dimensional histogram of the solid-state lidar;
[0009] Using a 3D convolutional neural network to convert the bird's-eye view features of the denoised three-dimensional histogram into multiple non-ego-centric perspective view features to obtain the perspective view features converted from the bird's-eye view;
[0010] Converting each point in the denoised three-dimensional histogram from the current coordinate system to a perspective view coordinate system centered on a new origin, and performing linear interpolation on the eigenvalue in the perspective view coordinate system to obtain the perspective view features after linear interpolation;
[0011] Performing feature fusion on the perspective view features converted from the bird's-eye view and the perspective view features after linear interpolation to obtain fusion features;
[0012] Using the dynamic time stereo method, adjusting the depth estimation by analyzing the depth change of the continuous frame voxels of the fusion features, and using the depth center and depth range to focus on the possible spatial contour boundaries and peak regions to obtain a three-dimensional histogram depth feature map;
[0013] Identifying the local maximum points of the three-dimensional histogram depth feature map by setting the threshold and the size and shape of the voxels, and extracting the peaks of each pixel point.
[0014] Further, the using an improved wavelet denoising algorithm based on an adjustable threshold function to perform denoising processing on the three-dimensional histogram of the solid-state lidar is specifically:
[0015] Perform one-dimensional wavelet decomposition on each dimension of the three-dimensional histogram of the solid-state lidar to obtain wavelet coefficients;
[0016] Estimate the standard deviation of the noise through the high-frequency coefficients in the wavelet coefficients;
[0017] Calculate the threshold according to the estimated standard deviation of the noise and the signal length;
[0018] Perform threshold processing on the wavelet coefficients according to the determined threshold to obtain new wavelet coefficients;
[0019] Perform an inverse wavelet transform on the new wavelet coefficients that is inverse to obtaining the wavelet coefficients to obtain the denoised three-dimensional histogram data.
[0020] Furthermore, the performing one-dimensional wavelet decomposition on each dimension of the three-dimensional histogram of the solid-state lidar to obtain wavelet coefficients specifically is as follows:
[0021] Perform one-dimensional wavelet transform on the original three-dimensional histogram data D(x, y, z) on each fixed (y, z) plane to obtain detail coefficients cD1(i, y, z) and scaling coefficients cA1(i, y, z), where i represents the decomposition level in the first one-dimensional wavelet transform, starting from 1 and increasing step by step;
[0022] Perform one-dimensional wavelet transform on the scaling coefficients cA1(i, y, z) in each fixed x direction to obtain new detail coefficients cD2(x, j, z) and new scaling coefficients cA2(x, j, z), where j represents the decomposition level in the second one-dimensional wavelet transform, starting from 1 and increasing step by step;
[0023] Perform one-dimensional wavelet transform on the scaling coefficients cA2(x, j, z) in each fixed (x, y) direction to obtain the final detail coefficients cD3(x, y, k) and the final scaling coefficients cA3(x, y, k) as wavelet coefficients, where k represents the decomposition level in the second one-dimensional wavelet transform, starting from 1 and increasing step by step.
[0024] Furthermore, the standard deviation of the noise is estimated by the following formula:
[0025]
[0026] where γ is the standard deviation of the noise, median(|w 1,k |) represents calculating the median of the absolute value of w 1,k A is an adjustment factor, and w 1,k represents the first-level high-frequency coefficient of the wavelet coefficient;
[0027] The calculation formula for calculating the threshold according to the estimated standard deviation of the noise and the signal length is:
[0028]
[0029] Among them, λ is the threshold, and N is the signal length;
[0030] The wavelet coefficients are processed according to the determined threshold, and the calculation formula for the new wavelet coefficients is:
[0031]
[0032] Among them, the adjustable parameter a ∈ [0, 1], the adjustable parameter b > 0, sign(w j,k ) represents the sign function of the wavelet coefficient w j,k .
[0033] Furthermore, the 3D convolutional neural network is used to convert the bird's-eye view features of the denoised three-dimensional histogram into multiple non-ego-centric perspective view features, and the perspective view features converted from the bird's-eye view are obtained, specifically including:
[0034] The denoised three-dimensional histogram is used as the input of the 3D convolutional neural network, and is convolved through the first 3D convolutional layer and the convolution result is passed through the first activation function; the convolution result passed through the first activation function is pooled by the first max pooling layer;
[0035] The result of the first max pooling layer is convolved through the second 3D convolutional layer and the convolution result is passed through the second activation function; the convolution result passed through the first activation function is pooled by the second max pooling layer;
[0036] The result of the second max pooling layer is convolved through the third 3D convolutional layer and the convolution result is passed through the third activation function.
[0037] Furthermore, the first 3D convolutional layer uses 32 filters, that is, convolution kernels, for convolution, and the size of each filter is 3×3×3; the pooling window size of the first max pooling layer is 2×2×2;
[0038] The second 3D convolutional layer uses 64 filters, that is, convolution kernels, for convolution, and the size of each filter is 3×3×3; the pooling window size of the second max pooling layer is 2×2×2;
[0039] The third 3D convolutional layer uses 64 filters, that is, convolution kernels, for convolution, and the size of each filter is 3×3×3;
[0040] The first activation function, the second activation function, and the third activation function all adopt the ReLU activation function.
[0041] Further, the conversion of each point in the denoised three-dimensional histogram from the current coordinate system to the perspective view coordinate system centered at the new origin is specifically as follows:
[0042] For a point (x i , y i , z i ) in the three-dimensional histogram, it is converted to the perspective view feature (r p , θ p , φ p ) with the new observation point (x′ i , y i ′, z i ′) as the origin. Here, r i , θ i , φ i represent the radial distance, horizontal angle, and vertical angle respectively;
[0043] The radial distance r i is calculated by the following formula:
[0044]
[0045] The horizontal angle θ i is calculated by the following formula:
[0046]
[0047] The vertical angle φ i is calculated by the following formula:
[0048]
[0049] The linear interpolation of the eigenvalue in the perspective view coordinate system to obtain the linearly interpolated perspective view feature is specifically as follows:
[0050] For a feature point c pv = (x, y, z) in the perspective view coordinate system, the eigenvalues of its eight neighboring voxels are f 000 , f 001 , f 010 , f 011 , f 100 , f 101 , f 110 , f1 111 , and the corresponding coordinates are from (x0, y0, z0) to (x1, y1, z1). Then the linear interpolation of c pv = (x, y, z) is specifically as follows:
[0051] Interpolate in the x direction to obtain the corresponding interpolation result:
[0052]
[0053] Interpolate in the y direction to obtain the corresponding interpolation result:
[0054]
[0055] Interpolate in the z direction to obtain the corresponding interpolation result:
[0056]
[0057] f(x, y, z) is the perspective view feature after linear interpolation.
[0058] Furthermore, using the dynamic time volume method, the depth estimation is adjusted by analyzing the depth changes of the consecutive frame voxels of the fusion feature, and the depth center and depth range are used to focus on the possible spatial contour boundaries and peak regions to obtain a three-dimensional histogram depth feature map, which specifically includes the steps of:
[0059] Take the fusion feature at the current time t as the reference frame and the fusion feature at the previous time t-1 as the source frame and send them into a 3D convolutional neural network to extract the corresponding features and
[0060] Take the features and and jointly input them into two depth modules with shared weights for depth information estimation, and the corresponding three-dimensional histogram depth feature maps and
[0061] In each depth module, independently calculate the single-frame depth MD, context information of each frame, introduce the depth center μ and depth range ρ, and calculate the stereo depth SD using the dynamic time volume strategy. At the same time, iterate μ and ρ through the expectation-maximization algorithm (EM algorithm); use the shared weight network to fuse the single-frame depth MD and stereo depth SD of each frame to obtain the three-dimensional histogram depth feature maps corresponding to the reference frame and the source frame and
[0062] Take the three-dimensional histogram depth feature maps and respectively perform channel connection after voxel pooling to obtain the final three-dimensional histogram depth feature map
[0063] Furthermore, identify the local maximum points of the three-dimensional histogram depth feature map by setting thresholds and the size and shape of the voxels, and extract the peaks of each pixel point, specifically:
[0064] Calculate the scores of each depth feature
[0065] Identify the three-dimensional histogram by setting a threshold The local maximum points in it, adjust the detection strategy according to the size and shape of the voxels, and extract the peak PV of each pixel point t,(x,y,z) :
[0066]
[0067] Among them, (dx, dy, dz) represents the coordinate offset of the neighboring voxels of the pixel point (x, y, z).
[0068] The present invention also provides a radar data processing system, which is characterized in that: the peak value of the three-dimensional histogram of the solid-state lidar is extracted by using the above-mentioned radar data processing method.
[0069] The radar data processing method and system provided by the present invention first use an improved wavelet denoising algorithm based on an adjustable threshold function to denoise the histogram data, improve the smoothness of the signal and retain useful information. Subsequently, the bird's-eye view features of the three-dimensional histogram are converted into multiple non-ego-centric perspective view features, and a 3D convolutional neural network is used to extract the features of each view. Multiple three-dimensional histogram features are fused to generate a richer feature map. Using dynamic time stereo technology, analyze the depth change of the three-dimensional histogram data frame, adjust the depth estimation, and correct the depth change through a motion compensation algorithm to ensure the accuracy of the histogram peak in a dynamic environment. This method and system can effectively solve the problem of local feature mismatch of distant or small-sized objects, improve the point cloud imaging quality and positioning accuracy of solid-state lidar, and are applicable to real-time applications such as autonomous driving, drones, and robot navigation. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is the schematic diagram of the radar data processing method and system provided by the embodiment of the present invention;
[0071] Figure 2 is the flowchart of the improved wavelet denoising algorithm based on an adjustable threshold function provided by the embodiment of the present invention;
[0072] Figure 3 is the structural diagram of the three-dimensional convolutional network provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The following specifically illustrates the embodiments of the present invention in conjunction with the drawings. The given embodiments are only for illustrative purposes and should not be construed as limiting the present invention. The drawings are only for reference and illustration, and do not constitute a limitation on the scope of patent protection of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0074] The radar data processing method provided by the embodiments of the present invention, as Figure 1 shown, includes the steps of:
[0075] S1. Use an improved wavelet denoising algorithm based on an adjustable threshold function to perform denoising processing on the three-dimensional histogram of the solid-state lidar;
[0076] S2. Use a 3D convolutional neural network to convert the bird's-eye view features of the denoised three-dimensional histogram into multiple non-ego-centric perspective view features, obtaining the perspective view features converted from the bird's-eye view;
[0077] S3. Convert each point in the denoised three-dimensional histogram from the current coordinate system to a perspective view coordinate system centered on a new origin, and then perform linear interpolation on the eigenvalue in the perspective view coordinate system to obtain the perspective view features after linear interpolation;
[0078] S4. Perform feature fusion on the perspective view features converted from the bird's-eye view and the perspective view features after linear interpolation to obtain fusion features;
[0079] S5. Use dynamic time stereo technology to adjust depth estimation by analyzing the depth changes of consecutive frame voxels of the fusion features, and use the depth center μ and depth range σ to focus on possible spatial contour boundaries and peak regions to obtain a three-dimensional histogram depth feature map;
[0080] S6. Identify the local maximum points of the three-dimensional histogram depth feature map by setting thresholds and the size and shape of the voxels, and extract the peaks of each pixel point.
[0081] In step S1, the improved wavelet denoising algorithm based on the adjustable threshold function improves the deficiencies of the traditional wavelet threshold denoising algorithm, which results in abnormal peaks in the denoised signal and the loss of useful information due to over-smoothing of the signal. It can better retain the peaks of the required useful signals in the original signal, filter out abnormal spikes, and avoid the signal after processing being blurred and losing detailed information, resulting in distortion. The denoised signal will be smoother and the overall shape will basically remain unchanged.
[0082] In step S2, a 3D convolutional neural network is used to convert the bird's-eye view (BEV) features of the denoised three-dimensional histogram into multiple non-ego-centric perspective view (PV) features, and these view features simulate the scenes observed from different positions. The 3D convolutional neural network is used to extract the features of each view. It breaks through the constraint that the origin of the traditional ego-centric three-dimensional histogram view is consistent with the bird's-eye view (BEV), and improves the ability of multi-view fusion.
[0083] In step S3, the number of three-dimensional histograms is extended by mimicking different perspectives through transforming the origin, and the features generated in step S4 are fused with those in step S2 to obtain a feature map with richer levels.
[0084] In step S5, using dynamic time stereo technology, the depth change of the three-dimensional histogram data frame is analyzed, the depth estimation is adjusted, and the depth change is corrected through a motion compensation algorithm to ensure the accuracy of the histogram peak in a dynamic environment.
[0085] In step S6, local maximum points in the three-dimensional histogram are identified by setting a threshold, and the detection strategy is adjusted according to the size and shape of the voxels to accurately extract the peak PV of each pixel point in the feature map generated in step S5. t,(x,y,z) 。
[0086] This method can effectively solve the problem of local feature mismatch of distant or small-sized objects, improve the point cloud imaging quality and positioning accuracy of solid-state lidar, and is applicable to real-time applications such as autonomous driving, drones, and robot navigation.
[0087] The following elaborates on each step.
[0088] (1) Step S1: Denoising of original data
[0089] S1. Use an improved wavelet denoising algorithm based on an adjustable threshold function to denoise the three-dimensional histogram of the solid-state lidar.
[0090] As Figure 2 shown in the flowchart of, this step S1 specifically includes the steps:
[0091] S11. Read the three-dimensional histogram data file and convert it into an array of size (150, 360, 288), denoted as data, and decompose data using three-dimensional wavelet transform to obtain wavelet coefficients w j,k 。
[0092] In this example, the Coiflet wavelet basis function is selected, and the decomposition level is set to 3. One-dimensional wavelet decomposition is performed on each dimension (x, y, z) of the three-dimensional data respectively.
[0093] Perform one-dimensional wavelet transform on the first dimension: Perform one-dimensional wavelet transform on the original three-dimensional histogram data on each fixed (y, z) plane to obtain detail coefficients cD1(i, y, z) and scale coefficients cA1(i, y, z), as shown in the following formula (1), where φ i (x), ψ i (x) are the scale function and wavelet basis function for performing one-dimensional wavelet transform on the original data on each fixed (y, z) plane respectively.
[0094]
[0095] Among them, in cD1, i represents the detail coefficient at the i-th level, representing the high-frequency component and the scale coefficient of the signal at the current level. In cA1, i represents the scale coefficient at the i-th level, representing the low-frequency component of the signal at the current level. i represents the decomposition level in the first one-dimensional wavelet transform, starting from 1 and increasing gradually. As i increases, the resolution of the signal decreases, and the lengths of the obtained detail coefficients and scale coefficients also decrease.
[0096] Perform one-dimensional wavelet transform on the second dimension: Perform one-dimensional wavelet transform on the scale coefficient cA1(i, y, z) in each fixed x direction to obtain a new detail coefficient cD2(x, j, z) and a new scale coefficient cA2(x, j, z), as shown in the following formula (2).
[0097]
[0098] Among them, in cD2, j represents the detail coefficient at the j-th level in the second one-dimensional wavelet transform, representing the high-frequency component of the signal at the current level. In cA2, j represents the scale coefficient at the j-th level in the second one-dimensional wavelet transform, representing the low-frequency component of the signal at the current level. j represents the decomposition level in the second one-dimensional wavelet transform, starting from 1 and increasing gradually. φ j (y) and ψ j (y) are respectively the scaling function and the wavelet basis function for performing one-dimensional wavelet transform on the scale coefficient cA1(i, y, z) in each fixed x direction.
[0099] Perform one-dimensional wavelet transform on the third dimension: Perform one-dimensional wavelet transform on the scale coefficient cA2(x, j, z) in each fixed (x, y) direction to obtain the final detail coefficient cD3(x, y, k) and the final scale coefficient cA3(x, y, k).
[0100]
[0101] Among them, in cD3, k represents the detail coefficient at the k-th level in the third one-dimensional wavelet transform, representing the high-frequency component of the signal at the current level. In cA3, k represents the scale coefficient at the k-th level in the third one-dimensional wavelet transform, representing the low-frequency component of the signal at the current level. k represents the decomposition level in the third one-dimensional wavelet transform, starting from 1 and increasing gradually. φ k (z) and ψ k (z) are respectively the scaling function and the wavelet basis function for performing one-dimensional wavelet transform on the scale coefficient cA2(x, j, z) in each fixed (x, y) direction.
[0102] After the above process, the original data D(x, y, z) is decomposed into wavelet coefficients w at multiple different scales and directions j,k , including detail coefficients cD3(x, y, k) and scale coefficients cA3(x, y, k). These coefficients contain the information of the original data at each scale and direction and can be further used for threshold processing and signal reconstruction.
[0103] S12. Estimate the standard deviation γ of the noise through the high-frequency coefficients in the wavelet coefficients w j,k .
[0104] Usually, the first-level high-frequency coefficients are selected (because the high-frequency coefficients mainly contain noise components). Let the first-level high-frequency coefficients of the wavelet coefficients w j,k be w 1,k , then the estimation formula for the standard deviation γ of the noise is:[[]]
[0105]
[0106] where the constant A = 0.6745 is an adjustment factor to make the estimated standard deviation more in line with the characteristics of Gaussian noise, median(|w 1,k |) represents calculating the median of the absolute value of w 1,k , A is the adjustment factor, and w 1,k represents the first-level high-frequency coefficients of the wavelet coefficients.
[0107] S13. Calculate the threshold λ according to the estimated noise standard deviation σ and the signal length N.
[0108] Calculate the threshold λ according to the estimated noise standard deviation γ and the signal length N. Select the VisuShrink method. The VisuShrink method uses a single common threshold for all wavelet detail coefficients. This threshold is used to remove high-probability additive Gaussian noise because additive Gaussian noise often makes the image too smooth. By specifying sigma smaller than the true noise standard deviation, more visually consistent results can be obtained. Its threshold calculation formula is:[[]]
[0109]
[0110] S14. Perform threshold processing on the wavelet coefficients w j,k according to the determined threshold λ to obtain new wavelet coefficients w' j,k .
[0111] Here, an improved adjustable threshold function is applied, and the formula is:[[]]
[0112]
[0113] Among them, a ∈ [0, 1], b > 0, and the flexible denoising of different data systems is achieved by adjusting the parameters a and b. sign(w j,k ) represents the sign function of the wavelet coefficient w j,k . The definition of the sign function sign(x) is as follows: when x > 0, sign(x) = 1; when x < 0, sign(x) = -1; when x = 0, sign(x) = 0.
[0114] S15. Perform the inverse wavelet transform on the new wavelet coefficient w′ j,k to obtain the denoised three-dimensional histogram data H t (x, y, z).
[0115] Perform the inverse wavelet transform on the third dimension through formula (3): perform the inverse wavelet transform on the processed wavelet coefficient w′ j,k on each fixed (x, y) plane to obtain the restored scale coefficient cA2′(x, j, z).
[0116] Perform the inverse wavelet transform on the second dimension through formula (2): perform the inverse wavelet transform on the restored scale coefficient cA2′(x, j, z) in each fixed x direction to obtain the new restored scale coefficient cA1′(i, y, z).
[0117] Perform the inverse wavelet transform on the first dimension through formula (1): perform the inverse wavelet transform on the restored scale coefficient cA1′(i, y, z) on each fixed (y, z) plane to obtain the finally restored three-dimensional data H t (x, y, z).
[0118] Through the above steps, the processed wavelet coefficient can be inversely wavelet transformed to reconstruct the denoised three-dimensional histogram H t (x, y, z). The size of the reconstructed data should be the same as that of the original data, and the noise is effectively removed, and the signal quality is significantly improved.
[0119] (2) Step S2: Multi-view feature extraction
[0120] S2. Use a 3D convolutional neural network (3D-CNN) to convert the bird's-eye view features of the denoised three-dimensional histogram H t (x, y, z) into multiple non-ego-centric perspective view features to obtain the perspective view features f bev of the bird's-eye view conversion.
[0121] 3D-CNN can capture local features in space. By performing convolutional operations in three dimensions, it can extract feature information more comprehensively. The convolutional operations are as follows:
[0122]
[0123] Among them, y i,j,k is the value at position (x i , y j , z k ) in the output feature map; is the activation function, which is used to introduce non-linearity; is the weight of the convolutional kernel (filter) at position (d x , d y , d z ); is the value at position (i + d x , j + d y , k + d z ) in the input feature map; B is the bias term.
[0124] Figure 3 The following shows the structure diagram of the 3D convolutional neural network (3D-CNN) used in this example. As Figure 3 shown, the denoised three-dimensional histogram H t (x, y, z) is used as the input of the network. After passing through the first 3D convolutional layer, convolution is performed using 32 filters, and the size of each filter is 3×3×3. The first max-pooling layer is applied, and the size of the pooling window is 2×2×2. The second 3D convolutional layer is applied, using 64 filters, and the size of each filter is 3×3×3. The second max-pooling layer is applied, and the size of the pooling window is 2×2×2. The third 3D convolutional layer is applied, using 64 filters, and the size of each filter is 3×3×3. Each step plays an important role in the feature extraction process, gradually capturing more advanced features through convolution and pooling operations.
[0125] The convolution operation generates the value of each position in the output feature map by sliding a small filter (convolutional kernel) on the input feature map and calculating the weighted sum of the area covered by the filter. The 3D convolution operation slides not only on the two-dimensional plane but also in the third dimension, enabling it to capture more complex spatial relationships. The ReLU activation function is used as the activation function to introduce non-linearity and pass the result of the convolution operation to the activation function to generate the final output features. The max-pooling operation is used for reducing the size of the feature map, improving the computational efficiency, and retaining the most important features at the same time.
[0126] (3) Step S3: Perspective view feature extraction
[0127] S3. Convert each point in the denoised three-dimensional histogram from the current coordinate system to the perspective view coordinate system centered on the new origin, and then perform linear interpolation on the feature values in the perspective view coordinate system to obtain the linearly interpolated perspective view feature f pv .
[0128] In this embodiment, step S3 specifically includes the following steps:
[0129] S31. Select a new observation point according to the position of the target object and the requirements of scene coverage, and transform each point in the denoised three-dimensional histogram H t (x, y, z) from the current coordinate system to a perspective view coordinate system centered at the new origin.
[0130] For the point (x i , y i , z i ) in the three-dimensional histogram, transform it into a perspective view feature (r p , θ p , φ p ) with (x′ i , y′ i , z′ i ) as the origin. Here, r i , θ i , and φ i represent the radial distance, horizontal angle, and vertical angle respectively.
[0131] Calculate the radial distance:
[0132]
[0133] Calculate the horizontal angle:
[0134]
[0135] Calculate the vertical angle:
[0136]
[0137] S32. Use the linear interpolation method to interpolate the perspective view feature (r i , θ i , φ i ).
[0138] Since the points after the transformation in S31 may not exactly correspond to the integer coordinates on the grid, it is necessary to interpolate the perspective view feature (r i , θ i , φ i ). Use the linear interpolation method to estimate the feature values in the perspective view coordinate system. Among them, f pv (c pv ) is the feature value in the target perspective view coordinate system, c pv represents a feature point in the perspective view coordinate system, w i is the interpolation weight, and f iis the eigenvalue of the nearest neighbor points, and M in the feature interpolation step represents the number of feature points. During feature interpolation, M nearest neighbor feature points are used to calculate the eigenvalue f of the target point pv .
[0139] In this example, trilinear interpolation is adopted. For a feature point c pv =(x, y, z) in the perspective view coordinate system, the eigenvalues of its eight neighboring voxels are f 000 , f 001 , f 010 , f 011 , f 100 , f 101 , f 110 , f 111 , and the corresponding coordinates are from (x0, y0, z0) to (x1, y1, z1).
[0140] The interpolation calculation steps are as follows:
[0141] Interpolate in the x direction:
[0142]
[0143] Interpolate in the y direction:
[0144]
[0145] Interpolate in the z direction:
[0146]
[0147] (4) Step S4: Feature fusion
[0148] S4: Perform feature fusion on the perspective view features converted from the bird's-eye view and the perspective view features after linear interpolation to obtain the fused features.
[0149] Connect the BEV feature f bev and the PV feature f pv in the channel dimension (concat) to form the fused feature f fusion = concat(f bev , f pv ). The connection here means performing a connection operation in the channel dimension.
[0150] (5) Step S5: Obtain the three-dimensional histogram depth feature map using dynamic time stereo technology
[0151] S5. Using the dynamic time stereo method, adjust the depth estimation by analyzing the depth changes of consecutive frame voxels of the fused features, and use the depth center μ and the depth range σ to focus on the possible spatial contour boundaries and peak regions to obtain a three-dimensional histogram depth feature map.
[0152] In this embodiment, step S5 specifically includes the steps:
[0153] S51. Use the fused feature at the current time t as the reference frame and the fused feature at the previous time t-1 as the source frame and send them into the 3D convolutional neural network in step S2 to extract the corresponding features and
[0154] S52. Input the features and together into two depth modules with shared weights for depth information estimation, and the corresponding three-dimensional histogram depth feature maps and
[0155] In each depth module, independently calculate the single-frame depth MD and context information of each frame, introduce the depth center μ and the depth range ρ, and calculate the stereo depth SD using the dynamic time stereo strategy. At the same time, iterate μ and ρ through the expectation-maximization algorithm (EM algorithm); use the shared weight network to fuse the single-frame depth MD and the stereo depth SD of each frame to obtain the three-dimensional histogram depth feature maps corresponding to the reference frame and the source frame and
[0156] S53. Pass the three-dimensional histogram depth feature maps and through voxel pooling respectively and then perform channel concatenation (concat) to obtain the final three-dimensional histogram depth feature map
[0157] In step S52, use the depth center μ and the depth range ρ to generate depth candidates D candidates ={μ - nρ, μ - (n - 1)ρ, …, μ, …, μ + (n - 1)ρ, μ + nρ}, where n is a positive integer representing the number of layers extended from the depth center to both sides, and D candidates is a set of possible depth values, generated by the depth center μ and the depth range ρ, and used for subsequent depth estimation calculations. And use the depth center μ and the depth range ρ to calculate the depth confidence where d represents any depth value, and P(d) represents the confidence of the given depth value d. Use the EM algorithm to update the depth center and the depth range and μ new and ρ new represent the updated depth center and depth range respectively, and d i represents the candidate depth value, and P(d i ) represents the confidence of the candidate depth value d i .
[0158] In step S52, a motion compensation mechanism F compensated (t, x, y, z) = F(t, x + Δx, y + Δy, z + Δz) is introduced to correct the feature maps and to cope with the perspective changes in the dynamic environment. Finally, the weight network is used to fuse the single-frame depth MD and the stereo depth SD to obtain the three-dimensional histogram depth feature map w MD and w SD represent the weights of the shared single-frame depth MD and the stereo depth SD respectively.
[0159] (6) Step S6: Extract peaks
[0160] S6. Identify the local maximum points of the three-dimensional histogram depth feature map by setting thresholds and the size and shape of the voxels, and extract the peaks of each pixel point.
[0161] In this embodiment, size-aware circular (NMS) is used to calculate the score of each depth feature Identify the local maximum points in the three-dimensional histogram by setting thresholds , adjust the detection strategy according to the size and shape of the voxels, and accurately extract the peak PV t,(x,y,z) of each pixel point, and the formula is as follows:
[0162]
[0163] where (dx, dy, dz) represents the coordinate offset of the adjacent voxels.
[0164] Based on the above method, an embodiment of the present invention also provides a radar data processing system, which uses the above method to process radar data.
[0165] This embodiment is described by taking a solid-state lidar as an example, and the radar data processing method and system provided by the present invention are also applicable to other radars or other devices that can generate three-dimensional point cloud data.
[0166] In summary, for the radar data processing method and system provided by the embodiments of the present invention, firstly, an improved wavelet denoising algorithm based on an adjustable threshold function is used to denoise the histogram data, improving the smoothness of the signal and retaining useful information. Subsequently, the bird's-eye view features of the three-dimensional histogram are converted into multiple non-ego-centric perspective view features, and a 3D convolutional neural network is used to extract the features of each view. Multiple three-dimensional histogram features are fused to generate a richer feature map. Using dynamic time stereo technology, the depth change of the three-dimensional histogram data frame is analyzed, the depth estimation is adjusted, and the depth change is corrected through a motion compensation algorithm to ensure the accuracy of the histogram peak in a dynamic environment. This method and system can effectively solve the problem of local feature mismatch of distant or small-sized objects, improve the point cloud imaging quality and positioning accuracy of solid-state lidar, and are applicable to real-time applications such as autonomous driving, drones, and robot navigation.
[0167] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A radar data processing method, characterized in that: include: The 3D histogram of solid-state laser radar is denoised using an improved wavelet denoising algorithm based on an adjustable threshold function. Using a 3D convolutional neural network, the bird's-eye view features of the denoised three-dimensional histogram are converted into multiple non-egocentric perspective features to obtain perspective features of bird's-eye view conversion; Convert each point in the denoised three-dimensional histogram from the current coordinate system to a perspective view coordinate system centered at the new origin, and perform linear interpolation on the feature values in the perspective view coordinate system to obtain the perspective view features after linear interpolation; Perform feature fusion on the perspective view features converted from the bird's-eye view and the perspective view features after linear interpolation to obtain fused features; Using a dynamic temporal stereo method, adjusting the depth estimation by analyzing the depth changes of the voxels of consecutive frames of the fused features, using the depth center and depth range to focus on possible spatial contour boundaries and peak areas, and obtaining a three-dimensional histogram depth feature map; The local maximum point of the 3D histogram depth feature map is identified by setting the threshold and the size and shape of the voxel, and the peak value of each pixel is extracted.
2. The radar data processing method according to claim 1, characterized in that: The improved wavelet denoising algorithm based on an adjustable threshold function is used to perform denoising on the three-dimensional histogram of the solid-state laser radar, specifically: Perform one-dimensional wavelet decomposition on each dimension of the three-dimensional histogram of the solid-state laser radar to obtain the wavelet coefficients; The standard deviation of the noise is estimated by the high-frequency coefficients in the wavelet coefficients; The threshold is calculated based on the estimated standard deviation of the noise and the signal length; Threshold processing is performed on the wavelet coefficients according to the determined threshold value to obtain new wavelet coefficients; The new wavelet coefficients are subjected to an inverse wavelet transform which is the inverse of the obtained wavelet coefficients to obtain the denoised three-dimensional histogram data.
3. The radar data processing method according to claim 2, characterized in that: The three-dimensional histogram of the solid-state laser radar is subjected to one-dimensional wavelet decomposition for each dimension to obtain wavelet coefficients, which are specifically: Perform one-dimensional wavelet transform on the original three-dimensional histogram data D(x,y,z) on each fixed (y,z) plane to obtain detail coefficients cD1(i,y,z) and scale coefficients cA1(i,y,z), where i represents the decomposition level in the first one-dimensional wavelet transform, starting from 1 and gradually increasing; Perform a one-dimensional wavelet transform on each fixed scale coefficient cA1(i,y,z) in the x direction to obtain a new detail coefficient cD2(x,j,z) and a new scale coefficient cA2(x,j,z), where j represents the decomposition level in the second one-dimensional wavelet transform, starting from 1 and gradually increasing; A one-dimensional wavelet transform is performed on each fixed scale coefficient cA2(x,j,z) in the (x,y) direction to obtain the final detail coefficient cD3(x,y,k) and the final scale coefficient cA3(x,y,k) as wavelet coefficients. w represents the decomposition level in the second one-dimensional wavelet transform, and k starts from 1 and gradually increases.
4. The radar data processing method according to claim 3, characterized in that: The standard deviation of the noise is estimated by: Among them, γ is the standard deviation of the noise, median(|w 1,k |) indicates the calculation of w 1,k The median of the absolute value of 1,k Represents the first-level high-frequency coefficients of the wavelet coefficients; The threshold is calculated based on the estimated standard deviation of the noise and the signal length: Among them, λ is the threshold, N is the signal length; The wavelet coefficients are threshold processed according to the determined threshold, and the calculation formula for the new wavelet coefficients is: Among them, the adjustable parameter a∈[0,1], the adjustable parameter b>0, sign(w j,k ) represents the wavelet coefficient w j,k The symbol function of .
5. The radar data processing method according to claim 1, characterized in that: The 3D convolutional neural network is used to convert the bird's-eye view features of the denoised three-dimensional histogram into multiple non-egocentric perspective features, and the perspective features of the bird's-eye view conversion are obtained, including: The denoised three-dimensional histogram is used as the input of the 3D convolutional neural network, convolved through the first 3D convolutional layer and the convolution result is transferred by applying the first activation function; the convolution result transferred by the first activation function is pooled by applying the first maximum pooling layer; Applying a second 3D convolution layer to the result of the first maximum pooling layer for convolution and applying a second activation function to transfer the convolution result; applying a second maximum pooling layer to the convolution result transferred by the first activation function for pooling; A third 3D convolution layer is applied to the result of the second maximum pooling layer for convolution and a third activation function is applied to transfer the convolution result.
6. The radar data processing method according to claim 5, characterized in that: The first 3D convolution layer uses 32 filters, i.e., convolution kernels, for convolution, and the size of each filter is 3×3×3; the pooling window size of the first maximum pooling layer is 2×2×2; The second 3D convolution layer uses 64 filters, i.e., convolution kernels, for convolution, and the size of each filter is 3×3×3; the pooling window size of the second maximum pooling layer is 2×2×2; The third 3D convolutional layer uses 64 filters, i.e., convolution kernels, for convolution, and the size of each filter is 3×3×3; The first activation function, the second activation function, and the third activation function all use the ReLU activation function.
7. The radar data processing method according to claim 6, characterized in that: The method of converting each point in the denoised three-dimensional histogram from the current coordinate system to the perspective view coordinate system centered on the new origin is specifically as follows: For a point (x i ,y i ,z i ), transforming it to a new observation point (x′ p ,y′ p ,z′ p ) is the perspective feature of the origin (r i ,θ i ,φ i ), r i ,θ i ,φ i They represent radial distance, horizontal angle and vertical angle respectively; Radial distance r i Calculated by the following formula: Horizontal angle θ i Calculated by the following formula: Vertical angle φ i Calculated by the following formula: The linear interpolation is performed on the feature values in the perspective view coordinate system to obtain the perspective view features after linear interpolation, specifically: For a feature point c in a perspective coordinate system pv =(x,y,z), the eigenvalues of its eight neighboring voxels are f 000 ,f 001 ,f 010 ,f 011 ,f 100 ,f 101 ,f 110 ,f 111 , the corresponding coordinates are (x0, y0, z0) to (x1, y1, z1), then for c pv =(x,y,z) for linear interpolation: Interpolate in the x direction and get the corresponding interpolation result: Interpolate in the y direction and get the corresponding interpolation result: Interpolate in the z direction and get the corresponding interpolation result: f(x,y,z) is the perspective feature after linear interpolation.
8. The radar data processing method according to claim 5, characterized in that: The dynamic time stereo method is used to adjust the depth estimation by analyzing the depth changes of the continuous frame voxels of the fusion feature, and the depth center and depth range are used to focus on the possible spatial contour boundary and peak area to obtain a three-dimensional histogram depth feature map, which specifically includes the steps of: The fusion features of the current time t As the fusion feature of the reference frame and the previous time t-1 The source frame is fed into the 3D convolutional neural network to extract the corresponding features and The characteristics and Two depth modules with shared input weights are used to estimate depth information, and the corresponding three-dimensional histogram depth feature map and In each depth module, the single-frame depth MD and context information of each frame are calculated independently, and the depth center μ and depth range ρ are introduced. The stereo depth SD is calculated using the dynamic time stereo strategy, and μ and ρ are iterated by the maximum expectation algorithm. The single-frame depth MD and stereo depth SD of each frame are fused using a shared weight network to obtain the three-dimensional histogram depth feature map corresponding to the reference frame and the source frame. and The 3D histogram depth feature map and After voxel pooling, channels are connected to obtain the final three-dimensional histogram depth feature map.
9. The radar data processing method according to claim 8, characterized in that: The local maximum point of the three-dimensional histogram depth feature map is identified by setting the threshold and the size and shape of the voxel, and the peak value of each pixel is extracted, specifically: Calculate the score for each deep feature Identify 3D histogram by setting threshold The local maximum point in the voxel is adjusted according to the size and shape of the voxel to extract the peak PV of each pixel. t,(x,y,z) : Among them, (dx, dy, dz) represents the coordinate offset of the neighboring voxels of the pixel point (x, y, z).
10. A radar data processing system, characterized in that: The radar data processing method according to any one of claims 1 to 9 is used to extract the peak value of the three-dimensional histogram of the solid-state laser radar.
Citation Information
Patent Citations
Laser radar data processing method and system
CN119087392A