A method for processing lidar data

Through the method of point cloud data splicing and interpolation, the problem of sparse data points collected by low-line number lidar data is solved, and high-wire harness lidar is realized, high-quality point cloud data processing is achieved, and the R&D cost of smart vehicles is reduced.

CN114814873BActive Publication Date: 2025-05-27NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210367870.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-08
Publication Date
2025-05-27
Estimated Expiration
2042-04-08

AI Technical Summary

Technical Problem

The data points collected by medium and low-line lidar data in the prior art are sparse and have no obvious characteristics, while high-wire lidar is expensive and it is difficult to find a balance between cost and data quality.

Method used

Through point cloud data splicing and interpolation, the extracted target object features are used to splice adjacent point clouds and fuse them with the interpolated point cloud data to improve the density of point cloud data.

Benefits of technology

The low-wire harness lidar data is expanded into high-wire numerical lidar data, which improves the robustness and accuracy of point cloud data processing and reduces the R&D cost of smart vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114814873B_ABST
    Figure CN114814873B_ABST
Patent Text Reader

Abstract

The present invention discloses a laser radar data processing method, including: collecting initial sparse point cloud data in front of a vehicle, and vehicle position and posture data; extracting target object features of the i+1th frame point cloud data and the ith frame point cloud data from the initial sparse point cloud data; obtaining the translation vector and rotation matrix from the i+1th frame point cloud data to the ith frame point cloud data according to the target object features; splicing the point cloud data together after coordinate transformation to obtain spliced ​​point cloud data; interpolating the ith frame point cloud data to obtain interpolated point cloud data; averaging the spliced ​​point cloud data and the interpolated point cloud data as the final point cloud data. The present invention uses the extracted target object features to splice adjacent point clouds, and fuses them with the interpolated point cloud data, so as to expand the data collected by the low-beam laser radar into the point cloud data of the high-line laser radar, thereby improving the robustness and accuracy of the processed point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of lidar, and particularly relates to a lidar data processing method. Background Art

[0002] As one of the popular vehicle-mounted sensors, lidar is divided into single-line lidar and multi-line lidar. Multi-line lidar is mainly used for radar imaging of automobiles. Compared with single-line lidar, it has a qualitative change in dimension improvement and scene restoration, and can identify the height information of objects. Currently, 16-line, 32-line, and 64-line lidars are mainly introduced in the international market.

[0003] Multi-line lidar has multiple transmitters and receivers in the vertical direction. By rotating the motor, multiple beam bundles are obtained. The more the number of lines, the more perfect the surface contour of the object. Of course, the larger the amount of data to be processed, the higher the hardware requirements; multi-line lidar is mainly used in unmanned driving technology, which can calculate the height information of objects and perform 3D modeling of the surrounding environment. However, as the number of lines increases, the price also increases.

[0004] At present, rich achievements have been made in the theoretical research on improving image resolution and clarity; in the Chinese invention patent application No. CN202110268835.5, titled "A Method for Improving Image Resolution and Clarity", a method for improving image resolution and clarity is disclosed. By extracting the features of a blurred image, reconstructing a high-definition image, and performing feature mapping from the blurred image to the high-definition image, a high-resolution and high-clarity image is obtained.

[0005] From the above research, the existing research on lidar data processing is relatively scarce. Low-line lidar has a low cost, but the obtained point cloud data has poor effects and unclear features; high-line lidar can obtain point cloud data with obvious features, but the cost is high. Therefore, it is necessary to expand the data of low-line lidar into the data of high-line lidar to save costs. Summary of the Invention

[0006] Aiming at the deficiencies of the above-mentioned existing technologies, the purpose of the present invention is to provide a lidar data processing method to solve the problems in the existing technologies that the data points collected by low-line lidar are sparse and the features are not obvious; and the high-line lidar is expensive; the present invention uses point cloud data stitching and interpolation, and through the fusion of the two, lidar data with a higher point cloud density is obtained.

[0007] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0008] A lidar data processing method of the present invention comprises the following steps:

[0009] 1) Collect the initial sparse point cloud data in front of the vehicle, as well as the vehicle position and attitude data;

[0010] 2) Extract the target object features of the (i + 1)-th frame point cloud data and the i-th frame point cloud data from the initial sparse point cloud data;

[0011] 3) According to the target object features obtained in step 2), obtain the translation vector T i from the (i + 1)-th frame point cloud data to the i-th frame point cloud data i and the rotation matrix R;

[0012] 4) After coordinate transformation, splice the (i + 1)-th frame point cloud data and the i-th frame point cloud data together to obtain the spliced point cloud data i * =[x * y * z * l * , where x * , y * , z * are the three-dimensional coordinates of the spliced point cloud data; l * is the reflection intensity of the spliced point cloud data;

[0013] 5) Interpolate the i-th frame point cloud data to obtain the interpolated point cloud data as

[0014] 6) Take the average of the spliced point cloud data and the interpolated point cloud data as the final point cloud data.

[0015] Furthermore, step 1) specifically includes: collecting the initial sparse point cloud data through a low-line lidar installed on the top or front end of the vehicle, and then collecting the vehicle position and attitude data through a global positioning system and an inertial sensor (GPS-IMU).

[0016] Furthermore, the feature extraction method in step 2) is specifically:

[0017] 21) Encode the voxels of different neural layers in the entire scene into a small number of key points;

[0018] 22) Aggregate the key point features encoded in step 21) into the RoI grid, and extract the key point to RoI grid features based on the set abstraction operation, where RoI is the region of interest.

[0019] Furthermore, the encoding process in step 21) is specifically:

[0020] 211) Key point sampling, extract n key points K = {p 1 , …, pn}, using the collected key points to represent the entire scene;

[0021] 212) Encode the multi-scale semantic features from the convolutional neural network feature volume to the key points, obtaining regular voxels with multi-scale semantic features around the key points. The encoding process includes a voxel set abstraction module;

[0022] 213) Expand the voxel set abstraction module through the key point features of the original point cloud data of the i-th frame and the i+1-th frame and the bird's-eye view feature map obtained by downsampling;

[0023] 214) Predict the key point weights. After encoding the entire scene with a small number of key points, use the small number of key points to refine the candidate boxes.

[0024] Further, the voxel set abstraction module in step 212) performs set abstraction on different voxel feature vector sets;

[0025] Voxel feature vector set of the k-th layer of the convolutional neural network In the formula, is the feature vector of the N k th non-empty voxel in the k-th layer; The three-dimensional coordinates calculated by the voxel index and the actual voxel size of the k-th layer In the formula, is the three-dimensional coordinate of the N k th non-empty voxel in the k-th layer; N k is the number of non-empty voxels in the k-th layer; For each key point p i , identify its adjacent non-empty voxels at the k-th level within a radius r k to retrieve the voxel feature vector set:

[0026]

[0027] In the formula, is the feature vector of the j-th non-empty voxel; is 's relative position; is the adjacent voxel set of the key point p i ; r k is the range radius for identifying adjacent voxels; is the three-dimensional coordinate of the j-th non-empty voxel; is N k the position set of voxels; is N k the feature vector set of voxels;

[0028] The feature of the generated key point p i is:

[0029]

[0030] In the formula, M(·) is to randomly sample at most T voxels from the adjacent voxel set k ; G(·) is used in the neural network to encode voxel features and relative positions; max(·) is the max pooling operation, which maps the feature vectors of different numbers of adjacent voxels to the feature vector of the key point p i .

[0031] Concatenate the aggregated features at different levels to generate the multi-scale semantic features of the key point p i :

[0032]

[0033] Furthermore, the original point cloud data in the step 213) compensates for the quantization loss of the initial point cloud voxelization.

[0034] Furthermore, in the step 214), the key point features are re-weighted through point cloud segmentation, and the predicted weight of the feature of each key point is expressed as:

[0035] f i (p) = A(f i (p) )·f i (p) (4)

[0036] In the formula, A(·) is used to predict the confidence between [0, 1].[[]END]]

[0037] Furthermore, the key point to RoI grid feature extraction based on the set abstraction operation in the step 22) is specifically as follows:

[0038] 221) Implement the RoI grid through set abstraction. For each 3D candidate box, aggregate the key point features into the RoI grid with multiple receptive fields; uniformly sample 6×6×6 grid points in each 3D candidate box, denoted as G = {g 1 , …, g 216}; Adopt the set abstraction operation to aggregate the features of the grid points from the key point features. The adjacent key points of the grid point g i within the radius r are:

[0039]

[0040] In the formula, p j - g i is the local relative position indicating the feature and the key point, and Ψ is the set of adjacent key point features;

[0041] Aggregate adjacent key-point feature sets Ψ to generate the feature f of grid point g i of feature f i (g) :

[0042] f i (g) = max{G(M(Ψ))} (6)

[0043] where M(·) performs random sampling on adjacent key-point feature sets Ψ; G(·) is used in the neural network to encode voxel features and relative positions;

[0044] 222) 3D candidate box refinement and confidence prediction; Use the candidate box refinement network to learn and predict the residuals of the size and position (i.e., center, size, and orientation) relative to the input 3D candidate box.

[0045] Furthermore, the candidate box refinement network in step 222) has two branches: confidence prediction and outer box refinement, and jointly uses the 3D candidate box and the corresponding ground truth box as the training target;

[0046] Normalize the confidence training target y of the k-th 3D candidate box k to the interval [0, 1]:

[0047] y k = min(1, max(0, 2[0, 2IoU k - 0.5)) (7)

[0048] where IoU k is the k-th 3D candidate box and the corresponding ground truth box.

[0049] Furthermore, the confidence prediction branch minimizes the cross-entropy loss when predicting the confidence target.

[0050] Furthermore, in step 211), the farthest point sampling method is used to sample the key points.

[0051] Furthermore, the translation vector T in step 3) i = (x c , y c , z c ), where x c , y c , z c are the three-dimensional position coordinates of the vehicle; the rotation matrix R i = R x (γ)R y (β)R z (α), where R x (γ) is the angle of rotation of the lidar around the x-axis; Ry (β) is the angle of rotation of the lidar around the y-axis; R z (α) is the angle of rotation of the lidar around the z-axis;

[0052]

[0053]

[0054]

[0055] Where α is the heading angle; β is the pitch angle; γ is the roll angle.

[0056] Furthermore, the splicing process in step 4) is as follows:

[0057] 41) Perform pose transformation on the i-th and (i + 1)-th frame point cloud data obtained in step 1) according to the rotation matrix and translation vector in step 3);

[0058] 42) Perform point cloud registration on the point cloud data after pose transformation in step 41) to generate new point cloud data i * .

[0059] Furthermore, the interpolation process in step 5) is specifically as follows:

[0060] 51) Denote the number of lines of the initial lidar as σ 1 , and the lidar beam of the target data as σ 2 , where σ 2 > σ 1 ;

[0061] 52) Insert lines of data below each line of the i-th frame point cloud data of the lidar. A total of lines of data are inserted for each frame of lidar data, where m is the number of lines of the i-th frame point cloud data. The generated image is denoted as i 2 , is the original data line, is the middle line between two adjacent original data, k ∈ [0, σ 1 );

[0062] 53) Initialize the expanded (non-original) data obtained in step 52), that is, all four newly inserted channels are set to 0;

[0063] 54) Denote the values of each channel of the point cloud data to be calculated as x 0 * , y 0 * , z 0 * , l 0* For the values of each channel of two adjacent point cloud data above and below, they are x 1 , y 1 , z 1 , l 1 and x 2 , y 2 , z 2 , l 2 ; For the values of x 0 , y 0 , z 0 in the three-dimensional rectangular coordinate system of the point cloud data, they are calculated using the average value of the corresponding channel values of the two adjacent pixels above and below; for the value of the reflection intensity l 0 of the point cloud data, it is the same as the reflection intensity of the adjacent point cloud, and the average value of the reflection intensities of the two adjacent point clouds above and below is used as the reflection intensity of this point cloud. The specific calculation formula is as follows:

[0064]

[0065]

[0066]

[0067]

[0068] 55) is the original data row, is the non-zero data row, is the intermediate row. According to the method in step 54), the average value of the reflection intensities of the adjacent point clouds is used as the reflection intensity of this point cloud, and all non-zero rows are continuously filled.

[0069] Furthermore, the four channels in step 53) are x, y, z, l, where x, y, z are the coordinates in the three-dimensional rectangular coordinate system; l is the reflection intensity of the lidar at a certain point.

[0070] Furthermore, for the values of each channel of the final point cloud data i' in step 6), they are x', y', z', l' respectively. For the values of x', y', z' in the three-dimensional rectangular coordinate system of the point cloud data, they are calculated using the average value of the corresponding channel values obtained by splicing and interpolation; for the value of the reflection intensity l 0 of the point cloud data, it is the same as the reflection intensity of the adjacent point cloud, and the average value of the reflection intensities of the two adjacent point clouds above and below is used as the reflection intensity of this point cloud; the specific calculation formula is as follows:

[0071]

[0072] In the formula, i' is the processed point cloud data; i *is the spliced point cloud data; is the interpolated point cloud data; x * , y * , z * are the coordinates in the three-dimensional rectangular coordinate system of the spliced point cloud; l * is the reflection intensity of the lidar to a certain point; are the coordinates in the three-dimensional rectangular coordinate system of the spliced point cloud; is the reflection intensity of the lidar to a certain point.

[0073] Advantages of the present invention:

[0074] The present invention uses the features of the extracted target object to splice adjacent point clouds, fuses them with the interpolated point cloud data, expands the data collected by the low-beam lidar into the point cloud data of the high-beam lidar, improves the robustness and accuracy of the processed point cloud data, and has much lower data cost compared with the data collected by the existing high-beam lidar, reducing the R & D cost of intelligent vehicles. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 is the schematic diagram of the method of the present invention;

[0076] Figure 2 is the point cloud splicing flow chart;

[0077] Figure 3 is the schematic diagram of the point cloud data collected by the lidar;

[0078] Figure 4 is the schematic diagram of the feature extraction process;

[0079] Figure 5 is the schematic diagram of cyclic interpolation;

[0080] Figure 3 In, 1 - pedestrian; 2 - vehicle; 3 - bicycle. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] For the convenience of understanding by those skilled in the art, the present invention will be further described below in conjunction with the embodiments and the drawings. The content mentioned in the embodiments does not limit the present invention.

[0082] Referring to Figure 1 as shown, a lidar data processing method of the present invention is as follows:

[0083] 1) Collect the initial sparse point cloud data in front of the vehicle, as well as the vehicle position and attitude data. The point cloud data is referred to Figure 3 as shown;

[0084] Step 1) specifically includes: collecting initial sparse point cloud data through a low-line lidar installed on the top or front end of the vehicle, and then collecting vehicle position and attitude data through a global positioning system and an inertial sensor (GPS-IMU).

[0085] 2) Extract the target object (including pedestrians, vehicles, bicycles) features of the (i + 1)-th frame point cloud data and the i-th frame point cloud data from the initial sparse point cloud data;

[0086] Reference Figure 4 As shown, the feature extraction method is specifically as follows:

[0087] 21) Encode the voxels of different neural layers in the entire scene into a small number of key points;

[0088] 22) Aggregate the key point features encoded in step 21) into a RoI grid, and extract the key point to RoI grid features based on set abstraction operations, where RoI is the region of interest.

[0089] Specifically, the encoding process in step 21) is specifically as follows:

[0090] 211) Key point sampling, extract n key points K = {p 1 , …, p n} from the i-th frame point cloud data and the (i + 1)-th frame point cloud data, and use the collected key points to represent the entire scene;

[0091] 212) Encode the multi-scale semantic features of the convolutional neural network feature volume to the key points, and obtain regular voxels with multi-scale semantic features around the key points. The encoding process includes a voxel set abstraction module;

[0092] 213) Expand the voxel set abstraction module through the key point features of the i-th frame and the (i + 1)-th frame original point cloud data and the bird's-eye view feature map obtained by downsampling;

[0093] 214) Predict the key point weights. After encoding the entire scene with a small number of key points, use the small number of key points for candidate box refinement.

[0094] Among them, the voxel set abstraction module in step 212) performs set abstraction on different voxel feature vector sets;

[0095] The voxel feature vector set of the k-th layer of the convolutional neural network In the formula, is the feature vector of the N k -th non-empty voxel in the k-th layer; the three-dimensional coordinates calculated by the voxel index and the actual voxel size of the k-th layer In the formula, is the Nk The three-dimensional coordinates of a non-empty voxel; N k is the number of non-empty voxels in the k-th layer; for each key point p i , within a radius r k , the k-th level identifies its adjacent non-empty voxels within the radius to retrieve the set of voxel feature vectors:

[0096]

[0097] wherein, is the feature vector of the j-th non-empty voxel; is 's relative position; is the key point p i 's adjacent voxel set; r k is the range radius for identifying adjacent voxels; is the three-dimensional coordinates of the j-th non-empty voxel; is N k the position set of voxels; is N k the feature vector set of voxels;

[0098] The generated key point p i 's feature is:

[0099]

[0100] wherein, M(·) randomly samples at most T voxels from the adjacent voxel set k ; G(·) is used in the neural network to encode voxel features and relative positions; max(·) is the max pooling operation, which maps the feature vectors of different numbers of adjacent voxels to the feature vector of the key point p i ;

[0101] Concatenate the aggregation features of different levels to generate the multi-scale semantic feature of the key point p i :

[0102]

[0103] Among them, the original point cloud data in step 213) compensates for the quantization loss of the initial point cloud voxelization.

[0104] Among them, in step 214), the key point features are reweighted through point cloud segmentation, and the predicted weight of the feature of each key point is expressed as:

[0105] f i (p) = A(f i (p) )·fi (p) (4)

[0106] Wherein, A(·) is used to predict the confidence level between [0, 1].

[0107] Specifically, the key point to RoI grid feature extraction based on the set abstraction operation in step 22) is specifically as follows:

[0108] 221) Implement the RoI grid through set abstraction. For each 3D candidate box, aggregate the key point features into the RoI grid with multiple receptive fields; uniformly sample 6×6×6 grid points in each 3D candidate box, denoted as G = {g 1 ,…,g 216}; Adopt the set abstraction operation to aggregate the features of the grid points from the key point features. The adjacent key points of the grid point g i within the radius r are:

[0109]

[0110] Wherein, p j -g i is the local relative position of the indication feature and the key point, and Ψ is the set of adjacent key point features;

[0111] Aggregate the adjacent key point feature set Ψ to generate the feature f i of the grid point g i (g) :

[0112] f i (g) = max{G(M(Ψ))} (6)

[0113] Wherein, M(·) is to randomly sample the adjacent key point feature set Ψ; G(·) is used in the neural network to encode the voxel features and the relative position;

[0114] 222) 3D candidate box refinement and confidence prediction; Use the candidate box refinement network to learn and predict the residuals of the size and position (i.e., center, size, and direction) relative to the input 3D candidate box.

[0115] Among them, the candidate box refinement network in step 222) has two branches: confidence prediction and outer box refinement, and jointly uses the 3D candidate box and the corresponding ground truth box as the training target;

[0116] Normalize the confidence training target y k of the kth 3D candidate box to the interval [0, 1]:

[0117] y k= min(1, max(0, 2[0, 2IoU k - 0.5)) (7)

[0118] Where IoU k is the k-th 3D candidate box and the corresponding ground truth box.

[0119] The confidence prediction branch minimizes the cross-entropy loss when predicting the confidence target.

[0120] In step 211), the farthest point sampling method is used to sample the key points.

[0121] 3) According to the target object features extracted in step 2), obtain the translation vector T i from the (i + 1)-th frame of point cloud data to the i-th frame of point cloud data and the rotation matrix R i ;

[0122] The translation vector T in step 3) i = (x c , y c , z c ), where x c , y c , z c are the three-dimensional position coordinates of the vehicle; the rotation matrix R i = R x (γ)R y (β)R z (α), where R x (γ) is the angle of rotation of the lidar around the x-axis; R y (β) is the angle of rotation of the lidar around the y-axis; R z (α) is the angle of rotation of the lidar around the z-axis;

[0123]

[0124]

[0125]

[0126] Where α is the heading angle; β is the pitch angle; γ is the roll angle.

[0127] 4) Referring to Figure 2 as shown, the (i + 1)-th frame of point cloud data and the i-th frame of point cloud data are spliced together after coordinate transformation to obtain the spliced point cloud data i * = [x * y * z * l * , where x * , y *, z * are the three-dimensional coordinates of the point cloud data after splicing; l * is the reflection intensity of the point cloud data after splicing;

[0128] The splicing process is as follows:

[0129] 41) Perform pose transformation on the i-th and (i + 1)-th frame point cloud data obtained in step 1) according to the rotation matrix and translation vector in step 3);

[0130] 42) Perform point cloud registration on the point cloud data after the pose transformation in step 41) to generate new point cloud data i * ;

[0131] 5) Interpolate the i-th frame point cloud data to obtain the interpolated point cloud data as

[0132] Reference Figure 5 As shown, the interpolation process is specifically as follows:

[0133] 51) Denote the number of lines of the initial lidar as σ 1 , and the lidar beam of the target data as σ 2 , where σ 2 > σ 1 ;

[0134] 52) Insert lines of data below each line of the i-th frame point cloud data of the lidar. A total of lines of data are inserted for each frame of lidar data, where m is the number of lines of the i-th frame point cloud data. The generated image is denoted as i 2 , is the original data line, is the middle line between two adjacent original data, k ∈ [0, σ 1 );

[0135] 53) Initialize the expanded (non-original) data obtained in step 52), that is, all four newly inserted channels are set to 0;

[0136] 54) Denote the values of each channel of the point cloud data to be calculated as x 0 * , y 0 * , z 0 * , l 0 * , and the values of each channel of the upper and lower adjacent point cloud data are x 1 , y 1 , z 1 , l 1 and x 2, y 2 , z 2 , l 2 ; For the values of the three-dimensional rectangular coordinate system x 0 , y 0 , z 0 of the point cloud data, they are calculated using the average value of the corresponding channels of the two adjacent pixels above and below; for the value of the reflection intensity l 0 of the point cloud data, it is the same as the reflection intensity of the adjacent point cloud, and the average value of the reflection intensities of the two adjacent point clouds above and below is used as the reflection intensity of this point cloud. The specific calculation formula is as follows:

[0137]

[0138]

[0139]

[0140]

[0141] 55) is the original data row, is the non-zero data row, is the middle row. According to the method in step 54), the average value of the reflection intensities of the adjacent point clouds is used as the reflection intensity of this point cloud, and all non-zero rows are continuously filled.

[0142] Among them, the four channels in step 53) are x, y, z, l, where x, y, z are the coordinates in the three-dimensional rectangular coordinate system; l is the reflection intensity of the laser radar at a certain point.

[0143] 6) Take the average value of the point cloud data obtained after splicing and the point cloud data obtained after interpolation as the final point cloud data;

[0144] The values of each channel of the final point cloud data i' in step 6) are x', y', z', l' respectively. For the values of the three-dimensional rectangular coordinate system x', y', z' of the point cloud data, they are calculated using the average value of the corresponding channels of the spliced and interpolated data; for the value of the reflection intensity l 0 of the point cloud data, it is the same as the reflection intensity of the adjacent point cloud, and the average value of the reflection intensities of the two adjacent point clouds above and below is used as the reflection intensity of this point cloud; the specific calculation formula is as follows:

[0145]

[0146] In the formula, i' is the point cloud data after processing; i * is the point cloud data after splicing; is the point cloud data after interpolation; x * , y* , z * is the coordinate in the three-dimensional rectangular coordinate system of the point cloud after splicing; l * is the reflection intensity of the lidar to a certain point; is the coordinate in the three-dimensional rectangular coordinate system of the point cloud after splicing; is the reflection intensity of the lidar to a certain point.

[0147] The specific application ways of the present invention are numerous. The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of this technology, several improvements can be made without departing from the principle of the present invention, and these improvements should also be regarded as the protection scope of the present invention.

Claims

1. A method for processing lidar data, characterized in that, the steps are as follows: 1) Collect the initial sparse point cloud data in front of the vehicle, as well as the vehicle position and attitude data; 2) Extract the target object features of the (i + 1)-th frame point cloud data and the i-th frame point cloud data from the initial sparse point cloud data; 3) According to the target object features obtained by extraction in step 2), obtain the translation vector and rotation matrix from the (i + 1)-th frame point cloud data to the i-th frame point cloud data; 4) After coordinate transformation, splice the point cloud data of the (i + 1)-th frame and the point cloud data of the i-th frame together to obtain the spliced point cloud data i * ; 5) Interpolate the point cloud data of the i-th frame to obtain the interpolated point cloud data as 6) Take the average of the stitched point cloud data and the interpolated point cloud data as the final point cloud data; The feature extraction method in step 2) is specifically: 21) Encode the voxels of different neural layers in the entire scene into a small number of key points; 22) Aggregate the key point features encoded in step 21) into the RoI grid, and extract the key point to RoI grid features based on the set abstraction operation, where RoI is the region of interest; The encoding process in step 21) is specifically: 211) Key point sampling, extracting n key points K = {p 1 , ···, p n} from the i-th frame point cloud data and the (i + 1)-th frame point cloud data, and using the collected key points to represent the entire scene; 212) Encode the multi-scale semantic features from the convolutional neural network feature volume to the key points, and obtain regular voxels with multi-scale semantic features around the key points. The encoding process includes a voxel set abstraction module; 213) Expand the voxel set abstraction module through the key point features of the i-th frame and (i + 1)-th frame original point cloud data and the bird's-eye view feature map obtained by downsampling; 214) Predict the key point weights. After encoding the entire scene with a small number of key points, use the small number of key points for candidate box refinement; The voxel set abstraction module in step 212) performs set abstraction on different voxel feature vector sets; Set of voxel feature vectors of the k-th layer of the convolutional neural network wherein is the feature vector of the N k -th non-empty voxel in the k-th layer; the three-dimensional coordinates calculated by the voxel index and the actual voxel size of the k-th layer wherein is the three-dimensional coordinates of the N k -th non-empty voxel in the k-th layer; N k is the number of non-empty voxels in the k-th layer; for each key point p i , identify its adjacent non-empty voxels at the k-th level within a radius r k to retrieve the set of voxel feature vectors: Wherein, is the feature vector of the j-th non-empty voxel; is the relative position; is the adjacent voxel set of the key point p i ; r k is the range radius for identifying adjacent voxels; is the three-dimensional coordinate of the j-th non-empty voxel; is N k the position set of voxels; is N k the feature vector set of voxels; The generated key point p i is characterized by: where M(·) randomly samples at most T voxels from the set of adjacent voxels k ; G(·) is used in the neural network to encode voxel features and relative positions; max(·) is the max pooling operation that maps feature vectors of different numbers of adjacent voxels to the feature vector of the key point p i ​ Concatenate the aggregation features at different levels to generate the key point p i with multi-scale semantic features: The key point to RoI grid feature extraction based on the set abstraction operation in step 22) is specifically: 221) Implement the RoI grid through set abstraction. For each 3D candidate box, aggregate the key-point features into an RoI grid with multiple receptive fields; uniformly sample 6×6×6 grid points in each 3D candidate box, denoted as G = {g 1 , …, g 216}; adopt the set abstraction operation to aggregate the features of the grid points from the key-point features. The adjacent key points of the grid point g i within the radius r are: where p j -g i indicates the local relative position of the feature to the key point, and Ψ is the adjacent key point feature set; Aggregate adjacent key-point feature sets Ψ to generate the feature f of grid point g i of feature f i (g) : f i (g) = max{G(M(Ψ))} (6) where M(·) randomly samples the adjacent key point feature set Ψ; G(·) is used in the neural network to encode the voxel features and relative positions; 222) 3D candidate box refinement and confidence prediction; Use the candidate box refinement network to learn and predict the size and position residuals relative to the input 3D candidate box; The interpolation process in step 5) is specifically: Let the number of lines of the initial lidar be denoted as σ 1 and the lidar beam of the target data be σ 2 where σ 2 > σ 1 ; 52) Insert the following number of lines below each line of the i-th frame of lidar point cloud data: The number of inserted lines for each frame of lidar data is where m is the number of lines in the i-th frame of point cloud data. The generated image is denoted as i 2 , is the original data line, is the middle line between two adjacent original data lines, and k ∈ [0, σ 1 ); 53) Initialize the augmented data obtained in step 52), that is, all four newly inserted channels take 0; 54) Denote the values of each channel of the point cloud data to be calculated as x 0 * , y 0 * , z 0 * , l 0 * . The values of each channel of the upper and lower adjacent point cloud data are x 1 , y 1 , z 1 , l 1 and x 2 , y 2 , z 2 , l 2 . For the values of the three-dimensional rectangular coordinate system x 0 , y 0 , z 0 of the point cloud data, they are calculated using the average value of the corresponding channels of the upper and lower adjacent pixels. For the value of the reflection intensity l 0 of the point cloud data, it is the same as the reflection intensity of the adjacent point cloud, and the average value of the reflection intensities of the upper and lower two adjacent point clouds is used as the reflection intensity of this point cloud. The specific calculation formula is as follows: 55) is the original data line, is the non-zero data line, is the intermediate line. According to the method in step 54) above, the average reflection intensity of adjacent point clouds is used as the reflection intensity of this point cloud, and all non-zero lines are continuously filled up; In step 6), the values of each channel of the final point cloud data i' are x', y', z', and l' respectively. For the values of x', y', and z' in the three-dimensional rectangular coordinate system of the point cloud data, they are calculated using the average value of the corresponding channel values obtained by splicing and interpolation; for the reflection intensity l 0 of the point cloud data, its reflection intensity is the same as that of the adjacent point cloud, and the average value of the reflection intensities of the two adjacent upper and lower point clouds is used as the reflection intensity of this point cloud; the specific calculation formula is as follows: Where, i' is the point cloud data after processing; i * is the point cloud data after splicing; is the point cloud data after interpolation; x * , y * , z * are the coordinates in the three-dimensional rectangular coordinate system of the point cloud after splicing; l * is the reflection intensity of the lidar to a certain point; are the coordinates in the three-dimensional rectangular coordinate system of the point cloud after splicing; is the reflection intensity of the lidar to a certain point.

2. The lidar data processing method according to claim 1, characterized in that, The translation vector T in step 3) i =(x c , y c , z c ), where x c , y c , z c are the three-dimensional position coordinates of the vehicle; the rotation matrix R i =R x (γ)R y (β)R z (α), where R x (γ) is the angle of rotation of the lidar around the x-axis; R y (β) is the angle of rotation of the lidar around the y-axis; R z (α) is the angle of rotation of the lidar around the z-axis; where α is the heading angle; β is the pitch angle; γ is the roll angle.

3. The lidar data processing method according to claim 1, characterized in that, The stitching process in step 4) is: 41) Perform pose transformation on the i-th frame and (i + 1)-th frame point cloud data obtained in step 1) according to the rotation matrix and translation vector in step 3); 42) Perform point cloud registration on the point cloud data after the pose transformation in step 41) to generate new point cloud data i * , i * = [x * y * z * l * , where x * , y * , z * are the three-dimensional coordinates of the point cloud data after splicing; l * is the reflection intensity of the point cloud data after splicing.

4. The lidar data processing method according to claim 1, characterized in that, Step 1) specifically includes: Collect the initial sparse point cloud data through a low-line lidar installed on the top or front end of the vehicle, and then collect the vehicle position and attitude data through the global positioning system and inertial sensors.

Citation Information

Patent Citations

  • Method for improving picture resolution and definition

    CN112950476A

  • Traffic scene sparse laser point cloud splicing method

    CN109633665A

  • Three-dimensional point cloud reconstruction device and method based on multi-fusion sensor

    CN110415342A