A processing method for target speed prediction based on point cloud data
By performing target recognition, association and pose registration of lidar point clouds, the problem that traditional point cloud target detection models cannot detect target speed is solved, and the speed detection of each target in the point cloud is realized, reducing the system's dependence on millimeter wave radar, and improving the working efficiency of the perception system.
Patent Information
- Application Number
- CN202210581422.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-05-26
AI Technical Summary
The traditional point cloud target detection model cannot confirm the speed of various targets, resulting in the autonomous driving perception system requiring additional millimeter-wave radar for speed detection, which increases the complexity and computing burden of the system.
By identifying and correlating two frames of lidar point clouds at adjacent moments, point cloud pose registration is performed using the iterative nearest point (ICP) algorithm, pose transformation matrix T is calculated, and the pose of the target at the next moment is predicted based on the matrix, thereby estimating the velocity of the target.
The speed detection of each target in the point cloud is achieved, the speed detection defects of traditional models are overcome, the dependence on millimeter wave radar is reduced, the calculation amount of data fusion module is reduced, and the working efficiency of the perception system is improved.
Smart Images

Figure CN114966736B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a processing method for predicting the speed of an object based on point cloud data. Background Art
[0002] The lidar of an autonomous driving perception system can obtain corresponding point cloud data by performing laser scanning on the surrounding environment of the vehicle. The perception system uses a point cloud object detection model to detect the original point cloud data and can identify multiple types of obstacle objects (such as people, vehicles, bicycles, plants, animals, traffic signs, etc.) and the corresponding three-dimensional object detection frames (bounding boxes) for each type of object. However, the traditional point cloud object detection model cannot confirm the speeds of various objects. For this reason, the perception system also needs to be additionally configured with a millimeter-wave radar for speed measurement to cooperate with it, and the speed characteristics in the millimeter-wave radar point cloud are added to the lidar point cloud through a data fusion module. Summary of the Invention
[0003] The purpose of the present invention is to provide a processing method, an electronic device, and a computer-readable storage medium for predicting the speed of an object based on point cloud data in view of the defects of the prior art. First, target recognition is performed on two frames of lidar point clouds at adjacent times respectively, and then the target recognition frames on the two frames of point clouds are associated to obtain multiple pairs of associated target recognition frames. Then, based on the Iterative Closest Point (ICP) algorithm, the point clouds in each pair of associated target recognition frames are registered to obtain the pose transformation matrix T between the previous and the next times, and based on the pose transformation matrix T, the pose of a specified position point on the target recognition frame at the previous time is predicted at the next time, and the speed of the corresponding object is calculated according to the pose of the point at the previous time and the predicted pose at the next time. Through the present invention, the speeds of each object in the point cloud can be obtained by continuously associating the output of the traditional point cloud object detection model and registering the point cloud poses, which can not only overcome the defect that the traditional point cloud object detection model cannot perform speed detection, but also reduce the dependence of the perception system on the millimeter-wave radar, reduce the point cloud fusion calculation amount of the data fusion module, and improve the overall working efficiency of the perception system.
[0004] To achieve the above object, in the first aspect of an embodiment of the present invention, a processing method for predicting the speed of an object based on point cloud data is provided, and the method includes:
[0005] Obtain the front and rear two frames of lidar point cloud data at adjacent times as the corresponding front frame point cloud and rear frame point cloud;
[0006] Perform point cloud object detection processing on the front-frame and back-frame point clouds respectively using a point cloud object detection model to obtain a plurality of first object detection boxes and a plurality of second object detection boxes; the first object detection boxes correspond to the front-frame point cloud, and the second object detection boxes correspond to the back-frame point cloud;
[0007] Perform object association processing on the first and second object detection boxes of the front and back-frame point clouds;
[0008] Predict the speed of the corresponding object according to each pair of associated first and second object detection boxes.
[0009] Preferably, the point cloud object detection model includes a VoxelNet model, a SECOND model, and a PointPillars model, and the PointPillars model is used by default;
[0010] Each of the first object detection boxes corresponds to a set of first detection box parameters; the first detection box parameters include the center point coordinates (x1, y1, z1) of the first object box, the depth l1 of the first object box, the width w1 of the first object box, the height h1 of the first object box, and the orientation angle yaw1 of the first object box;
[0011] Each of the second object detection boxes corresponds to a set of second detection box parameters; the second detection box parameters include the center point coordinates (x2, y2, z2) of the second object box, the depth l2 of the second object box, the width w2 of the second object box, the height h2 of the second object box, and the orientation angle yaw2 of the second object box.
[0012] Preferably, the performing object association processing on the first and second object detection boxes of the front and back-frame point clouds specifically includes:
[0013] Count the number of the first detection box parameters and record it as the first number m, and count the number of the second detection box parameters and record it as the second number n;
[0014] Perform sequential encoding on the center point coordinates of each of the first object boxes based on the object box index i to obtain the corresponding center point coordinates (x 1,i , y 1,i , z 1,i ), 1 ≤ i ≤ m; perform sequential encoding on the center point coordinates of each of the second object boxes based on the object box index j to obtain the corresponding center point coordinates (x 2,j , y 2,j , z 2,j ), 1 ≤ j ≤ n;
[0015] Based on the Kalman filter, for the center point coordinates (x 1,i , y 1,i , z 1,iPredict the coordinates of the center point of the corresponding predicted target box at the next moment
[0016] For each of the predicted target box center point coordinates Calculate the distance from each of the second target box center point coordinates (x 2,j , y 2,j , z 2,j ) to obtain m * n center point distances s i,j ,
[0017]
[0018] Construct a matrix vector with a shape of m * n from the m * n center point distances s i,j and denote it as the first matrix vector; and input the first matrix vector into the Deep Hungarian Network DHN to calculate the matching degree based on the improved Hungarian algorithm to obtain the corresponding association matrix vector A with a shape of m * n; the association matrix vector A includes m * n matching degrees a i,j , 0 < a i,j ≤1; each of the matching degrees a i,j corresponds to a pair of the first target detection box and the second target detection box based on the target box index i and the target box index j;
[0019] Group the matching degrees a i,j of the second quantity n corresponding to the same target box index i into the same matching degree set D i ; and extract the maximum matching degree a i in each of the matching degree sets D i,j as the corresponding maximum matching degree b i ;
[0020] Judge each of the maximum matching degrees b i ; if the current maximum matching degree b i is less than the preset minimum matching degree threshold a min , then confirm that there is no association relationship between the pair of the first target detection box and the second target detection box corresponding to the current maximum matching degree b i ; if the current maximum matching degree b i is greater than or equal to the minimum matching degree threshold a min , then create an association relationship between the pair of the first target detection box and the second target detection box corresponding to the current maximum matching degree b i .
[0021] Preferably, predicting the speed of the corresponding target according to each pair of associated first and second target detection boxes specifically includes:
[0022] Denote each pair of associated first and second target detection boxes as corresponding first associated box and second associated box;
[0023] Extract the point cloud within the first associated box from the previous frame point cloud as the first point cloud; and extract the point cloud within the second associated box from the subsequent frame point cloud as the second point cloud;
[0024] Perform point cloud pose registration on the first and second point clouds based on the Iterative Closest Point (ICP) algorithm to obtain the corresponding pose transformation matrix T; R is the rotation matrix and t is the translation vector;
[0025] Denote the point cloud coordinates of the specified position point on the first associated box corresponding to the first point cloud as the first point coordinates (x 1,c , y 1,c , z 1,c ); and perform rigid body pose transformation on the first point coordinates (x 1,c , y 1,c , z 1,c ) based on the pose transformation matrix T to obtain the corresponding second point coordinates (x 2,c , y 2,c , z 2,c );
[0026] Denote the corresponding targets of the first and second associated boxes as the current target; and estimate the velocity vector V of the current target based on the first and second point coordinates and the time difference between the previous and current moments.
[0027] Further, the performing point cloud pose registration on the first and second point clouds based on the Iterative Closest Point (ICP) algorithm to obtain the corresponding pose transformation matrix T specifically includes:
[0028] Step 51, take the first point cloud as the current source point cloud, take the second point cloud as the current target point cloud; initialize the value of the iteration counter to 1; and take the preset initial rotation matrix R0 and initial translation vector t0 as the corresponding previous rotation matrix and previous translation vector;
[0029] Step 52, determine the set of associated point pairs in the current source point cloud and the current target point cloud based on a preset minimum distance threshold d min ; the set of associated point pairs includes multiple associated point pairs G k , where k is the associated point pair index and 1 ≤ k; the associated point pair G k includes a point p k in the current source point cloud and a point q k in the current target point cloud, and the point p k and the point q kThe distance is less than or equal to the minimum distance threshold d min ;
[0030] Step 53, count the number of the association point pairs G k in the association point pair set, and record the number as the third quantity H;
[0031] Step 54, construct a corresponding objective function f(R’, t’) based on the Euclidean distance transformation:
[0032]
[0033] R’ is a rotation matrix variable, t’ is a translation vector variable, and c k is the normal vector of the point q k ;
[0034] Step 55, solve the two variables R’ and t’ that minimize the objective function f(R’, t’), and use the solution results as the corresponding current rotation matrix R * and the current translation vector t * ;
[0035] Step 56, calculate the error between the current rotation matrix R * and the previous rotation matrix, and the error between the current translation vector t * and the previous translation vector to obtain the corresponding rotation matrix error and translation vector error; if the rotation matrix error and the translation vector error do not both satisfy their respective preset reasonable error ranges, go to Step 57; if the rotation matrix error and the translation vector error both satisfy their respective preset reasonable error ranges, go to Step 59;
[0036] Step 57, increment the value of the iteration counter by 1; if the value of the iteration counter does not exceed the preset maximum number of iterations after incrementing, go to Step 58; if the value of the iteration counter exceeds the maximum number of iterations after incrementing, go to Step 59;
[0037] Step 58, form a corresponding current pose transformation matrix T * from the current rotation matrix R * and the current translation vector t * , and perform a pose transformation on each point in the current source point cloud based on the current pose transformation matrix T * , and use the transformed point cloud as the new current source point cloud; and use the current rotation matrix R * as the new previous rotation matrix, and use the current translation vector t *As the new previous translation vector; and go to step 52 to continue the iterative operation according to the new current source point cloud, the previous rotation matrix, and the previous rotation matrix;
[0038] Step 59, stop the iterative operation and use the latest current rotation matrix R * and the current translation vector t * as the corresponding rotation matrix R and translation vector t, and construct the corresponding pose transformation matrix T from the rotation matrix R and the translation vector t,
[0039] Furthermore, estimating the velocity vector V of the current target according to the first and second point coordinates and the time difference between the front and back moments specifically includes:
[0040] Denote the time difference between the front and back moments as the time difference δ; and according to the first point coordinate (x 1,c , y 1,c , z 1,c ), the second point coordinate (x 2,c , y 2,c , z 2,c ) and the time difference δ, estimate the three velocity components of the current target to obtain the corresponding v x , v y and v z ; and form the velocity vector V of the current target from the three velocity components v x , v y and v z ; v x =(x 2,c -x 1,c ) / δ, v y =(y 2,c -y 1,c ) / δ, v z =(z 2,c -z 1,c ) / δ.
[0041] The second aspect of the embodiments of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0042] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;
[0043] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.
[0044] In a third aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.
[0045] The embodiments of the present invention provide a processing method, an electronic device, and a computer-readable storage medium for predicting the target speed based on point cloud data. First, target recognition is performed on two frames of lidar point clouds at adjacent times respectively, and then the target recognition frames on the two frames of point clouds are associated to obtain multiple pairs of associated target recognition frames. Then, based on the ICP algorithm, the point clouds in each pair of associated target recognition frames are registered to obtain the pose transformation matrix T between the front and back times, and based on the pose transformation matrix T, the pose of a specified position point on the target recognition frame at the previous time is predicted at the next time, and the speed of the corresponding target is calculated according to the pose of the point at the previous time and the predicted pose at the next time. Through the present invention, the speed of each target in the point cloud can be obtained by continuously associating the output of the traditional point cloud target detection model and registering the point cloud poses, which not only overcomes the defect that the traditional point cloud target detection model cannot perform speed detection, but also reduces the dependence of the perception system on the millimeter-wave radar, reduces the point cloud fusion calculation amount of the data fusion module, and improves the overall working efficiency of the perception system. Description of the Drawings
[0046] Figure 1 It is a schematic diagram of a processing method for predicting the target speed based on point cloud data provided in Embodiment 1 of the present invention;
[0047] Figure 2 It is a schematic diagram of the structure of an electronic device provided in Embodiment 2 of the present invention. Detailed Embodiments
[0048] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] Embodiment 1 of the present invention provides a processing method for predicting the target speed based on point cloud data. As Figure 1 shown in the schematic diagram of a processing method for predicting the target speed based on point cloud data provided in Embodiment 1 of the present invention, the method mainly includes the following steps:
[0050] Step 1, obtain the front and back two frames of lidar point cloud data at adjacent times as the corresponding front frame point cloud and back frame point cloud.
[0051] Here, let the current time be \(t\) and the previous time be \(t - 1\). Then the previous-frame point cloud is the lidar point cloud data obtained by one or a group of lidars scanning the vehicle's environment within a set scanning range at time \(t - 1\), and the subsequent-frame point cloud is the lidar point cloud data obtained by the same one or the same group of lidars scanning the vehicle's environment within the same scanning range at time \(t\).
[0052] Step 2: Based on the point cloud object detection model, perform point cloud object detection processing on the previous-frame and subsequent-frame point clouds respectively to obtain multiple first object detection boxes and multiple second object detection boxes;
[0053] Among them, the point cloud object detection model includes the VoxelNet model, the SECOND model, and the PointPillars model. By default, the PointPillars model is used; the first object detection boxes correspond to the previous-frame point cloud, and the second object detection boxes correspond to the subsequent-frame point cloud; each first object detection box corresponds to a set of first detection box parameters; the first detection box parameters include the center point coordinates \((x1, y1, z1)\) of the first object box, the depth \(l1\) of the first object box, the width \(w1\) of the first object box, the height \(h1\) of the first object box, and the orientation angle \(yaw1\) of the first object box; each second object detection box corresponds to a set of second detection box parameters; the second detection box parameters include the center point coordinates \((x2, y2, z2)\) of the second object box, the depth \(l2\) of the second object box, the width \(w2\) of the second object box, the height \(h2\) of the second object box, and the orientation angle \(yaw2\) of the second object box.
[0054] Here, embodiments of the present invention can use a variety of point cloud object detection models to complete object detection of the front and rear frame point clouds, including the VoxelNet model, the SECOND model, and the PointPillars model. The computational efficiency of these three models increases one by one, so the PointPillars model with the highest computational efficiency is adopted by default. For the specific implementations of the VoxelNet model, the SECOND model, and the PointPillars model, reference can be made to the corresponding technical papers "VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection", "SECOND: Sparsely Embedded Convolutional Detection", and "PointPillars: Fast Encoders for Object Detection from Point Clouds" respectively, which will not be elaborated further here. It should be noted that based on each point cloud object detection model, a corresponding object detection box (bounding box), that is, the first and second object detection boxes, can be output for each type of recognized object. Each object detection box corresponds to a set of object detection box parameters consisting of the center point coordinates of the object box, the depth l of the object box, the width w of the object box, the height h of the object box, and the yaw angle yaw of the object box, that is, the first and second detection box parameters. Here, the first object detection box is the object detection box of various objects detected on the front frame point cloud corresponding to the previous moment t - 1, and the second object detection box is the object detection box of various objects detected on the rear frame point cloud corresponding to the current moment t.
[0055] Step 3, perform object association processing on the first and second object detection boxes of the front and rear frame point clouds;
[0056] Here, performing object association processing on the first and second object detection boxes of the front and rear frame point clouds actually means associating two object detection boxes belonging to the same object in the front and rear two frame point clouds;
[0057] Specifically, it includes: Step 31, count the number of the first detection box parameters and record it as the first number m, and count the number of the second detection box parameters and record it as the second number n;
[0058] Here, the first quantity m is actually the total number of targets detected in the previous frame of point cloud, and the second quantity n is actually the total number of targets detected in the subsequent frame of point cloud. In an ideal state, m = n. However, in reality, the situation where the two are not equal often occurs. For example, at the previous moment, there are only two vehicles around the host vehicle, corresponding to m = 2. At the current moment, a third vehicle suddenly merges quickly from another road, corresponding to n = 3. Another example is that at the previous moment, there are three vehicles around the host vehicle, corresponding to m = 3, but one of them has an emergency brake. Then, at the current moment, there may be only two vehicles around the host vehicle, corresponding to n = 2;
[0059] Step 32: Sequentially encode the center point coordinates of each first target box based on the target box index i to obtain the corresponding center point coordinates (x 1,i , y 1,i , z 1,i ), where 1 ≤ i ≤ m; Sequentially encode the center point coordinates of each second target box based on the target box index j to obtain the corresponding center point coordinates (x 2,j , y 2,j , z 2,j ), where 1 ≤ j ≤ n;
[0060] For example, if 4 targets are detected in the previous frame of point cloud, corresponding to the first target detection boxes 1 - 4 respectively, then m = 4, and the 4 corresponding center point coordinates of the first target boxes are (x 1,i=1 , y 1,i=1 , z 1,i=1 ), (x 1,i=2 , y 1,i=2 , z 1,i=2 ), (x 1,i=3 , y 1,i=3 , z 1,i=3 ), (x 1,i=4 , y 1,i=4 , z 1,i=4 ); If 3 targets are detected in the subsequent frame of point cloud, corresponding to the second target detection boxes 1 - 3 respectively, then n = 3, and the 3 corresponding center point coordinates of the second target boxes are (x 2,j=1 , y 2,j=1 , z 2,j=1 ), (x 2,j=2 , y 2,j=2 , z 2,j=2 ), (x 2,j=3 , y 2,j=3 , z 2,j=3 );
[0061] Step 33: Based on the Kalman filter, for each center point coordinate (x 1,i , y 1,i , z 1,i)Predict the position at the next moment to obtain the corresponding predicted center point coordinates of the target bounding box
[0062] Here, the principle of the Kalman filter can refer to the well-known Kalman filter algorithm, which will not be further elaborated here; in the embodiments of the present invention, a corresponding object motion model (such as a stationary model, a uniform motion model, a uniformly accelerated motion model, etc.) is pre-specified for various types of targets (such as people, vehicles, bicycles, plants, animals, traffic signs, etc.), and then the state equation of the Kalman filter is configured based on the specified object motion model. After that, the configured state equation can be used to predict the positions of various targets at the next moment; here, the embodiments of the present invention default that each target will not deform when stationary or moving, that is, it satisfies the rigid motion form. In the rigid motion form, the motion states of all points on a single target can be regarded as the same, that is to say, the change amounts of the positions of all points at the previous and subsequent moments are the same. Therefore, only one point on the detection bounding box of each target, that is, the center point of the target bounding box, is selected for prediction; select the center point coordinates (x 1,i ,y 1,i ,z 1,i ) of the first target bounding box at the previous moment t-1 for prediction, and the ideal value of the center point coordinates at the current moment t can be obtained;
[0063] For example, 4 center point coordinates of the first target bounding box are obtained from the previous frame point cloud: (x 1,i=1 ,y 1,i=1 ,z 1,i=1 ), (x 1,i=2 ,y 1,i=2 ,z 1,i=2 ), (x 1,i=3 ,y 1,i=3 ,z 1,i=3 ), (x 1,i=4 ,y 1,i=4 ,z 1,i=4 ); then, 4 ideal predicted center point coordinates of the target bounding box can be obtained through the current step, which are respectively:
[0064] Step 34, for the center point coordinates of each predicted target bounding box Calculate the distances from the center point coordinates of each predicted target bounding box to the center point coordinates (x 2,j ,y 2,j ,z 2,j ) of each second target bounding box to obtain m*n center point distances s i,j , s i,j =
[0065]
[0066] For example, it is known that there are 3 center point coordinates of the second target bounding box, which are respectively (x2,j=1 , y 2,j=1 , z 2,j=1 ), (x 2,j=2 , y 2,j=2 , z 2,j=2 ), (x 2,j=3 , y 2,j=3 , z 2,j=3 ), there are 4 predicted target box center point coordinates, which are respectively:
[0067]
[0068] Then, 4 * 3 = 12 center point distances s i,j will be obtained, and they are respectively:
[0069] s i=1,j=1 、s i=1,j=2 、s i=1,j=3 ,
[0070] s i=2,j=1 、s i=2,j=2 、s i=2,j=3 ,
[0071] s i=3,j=1 、s i=3,j=2 、s i=3,j=3 ,
[0072] s i=4,j=1 、s i=4,j=2 、s i=4,j=3 ;
[0073] Step 35, from m * n center point distances s i,j construct a matrix vector with a shape of m * n, denoted as the first matrix vector; and input the first matrix vector into the Deep Hungarian Net (DHN) to calculate the matching degree based on the improved Hungarian algorithm to obtain the corresponding association matrix vector A with a shape of m * n;
[0074] Among them, the association matrix vector A includes m * n matching degrees a i,j , 0 < a i,j ≤1; each matching degree a i,j corresponds to a pair of the first target detection box and the second target detection box based on the target box index i and the target box index j;
[0075] Here, the Hungarian algorithm is a well-known algorithm for calculating the matching degree. The Deep Hungarian Net (DHN) is a neural network based on the improved Hungarian algorithm for calculating the matching degree of an input matrix with the shape of m*n and equivalently outputting an associated matrix vector A with the same shape of m*n. For the specific model implementation, please refer to the paper "How To Train Your Deep Multi-Object Tracker", which will not be elaborated here one by one. It should be noted that from the paper implementation, it can be seen that the structure of the associated matrix vector A is the same as that of the input matrix, that is, the first matrix vector, and they are in a corresponding relationship. The associated matrix vector A consists of m*n matching degrees a i,j constituting, and each matching degree a i,j is used to denote the matching degree of the corresponding first target detection box i and the second target detection box j;
[0076] For example, the first matrix vector composed of the obtained 4*3 = 12 center point distances s i,j is: Then,
[0077] the obtained associated matrix vector A should be
[0078] Step 36, classify the matching degrees a i,j corresponding to the same target box index i and with the second quantity n into the same matching degree set D i ; and extract the maximum matching degree a i in each matching degree set D i,j as the corresponding maximum matching degree b i ;
[0079] For example, given that the associated matrix vector A is Then:
[0080] The matching degree set D i=1 corresponding to the target box index i = 1 should be {a 1,1 , a 1,2 , a 1,3}, and the maximum matching degree b i=1 is the maximum value in {a 1,1 , a 1,2 , a 1,3};
[0081] The matching degree set D i=2 corresponding to the target box index i = 2 should be {a 2,1 , a 2,2 , a 2,3}, and the maximum matching degree b i=2 is the maximum value in {a 2,1 , a2,2 , a 2,3 the maximum value in
[0082] the set of matching degrees D corresponding to the target box index i = 3 i=3 should be {a 3,1 , a 3,2 , a 3,3}, and the maximum matching degree b i=3 is also the maximum value in {a 3,1 , a 3,2 , a 3,3};
[0083] the set of matching degrees D corresponding to the target box index i = 4 i=4 should be {a 4,1 , a 4,2 , a 4,3}, and the maximum matching degree b i=4 is also the maximum value in {a 4,1 , a 4,2 , a 4,3};
[0084] Step 37, judge each maximum matching degree b i ; if the current maximum matching degree b i is less than the preset minimum matching degree threshold a min , then confirm that there is no association relationship between the corresponding pair of the first target detection box and the second target detection box; if the current maximum matching degree b i is greater than or equal to the minimum matching degree threshold a i , then create an association relationship between the corresponding pair of the first target detection box and the second target detection box. min Here, the minimum matching degree threshold a i is a preset reference threshold. In the embodiments of the present invention, it is stipulated that whenever any matching degree a
[0085] is lower than this reference threshold, it means that the corresponding two target detection boxes, that is, the first target detection box i and the second target detection box j, are not matched, that is, there is no association between them, and they respectively correspond to different targets; conversely, if any matching degree a min is higher than this threshold, it means that the corresponding first target detection box i and second target detection box j may be matched. However, for the same first target detection box i, there may be multiple matching degrees a i,j higher than this reference threshold. In the embodiments of the present invention, the maximum matching degree, that is, the maximum matching degree b i,j , is taken. For any matching degree a i,j higher than this threshold, it means that the corresponding first target detection box i and second target detection box j may be matched. For the same first target detection box i, there may be multiple matching degrees a i,j higher than this reference threshold. In the embodiments of the present invention, the maximum matching degree, that is, the maximum matching degree b i,j , is taken. That is, the maximum matching degree b iThe corresponding second target detection box is used as the most matching object for the current first target detection box i, and an association relationship is established between the two.
[0086] For example, the obtained 4 maximum matching degrees b i=1 , b i=2 , b i=3 , b i=4 , where b i=1 , b i=2 , b i=4 are all greater than the minimum matching degree threshold a min , but b i=3 is less than the minimum matching degree threshold a min ; then, there is no second target detection box associated with the first target detection box 3, that is, the corresponding target of the first target detection box 3 in the previous frame point cloud does not appear in the subsequent frame point cloud at the current moment; let b i=1 be a 1,1 , b i=2 be a 2,3 , b i=3 be a 3,2 , then, the second target detection box 1 is associated with the first target detection box 1, the second target detection box 3 is associated with the first target detection box 2, and the second target detection box 2 is associated with the first target detection box 3.
[0087] Step 4, predict the speed of the corresponding target according to each pair of associated first and second target detection boxes;
[0088] Here, in fact, the point cloud pose transformation relationship generated by the object movement, that is, the pose transformation matrix T, is obtained by registering the point clouds in a pair of associated first and second target detection boxes. Then, based on the pose transformation matrix T, the pose of the specified position point of the first target detection box at the next moment is predicted. Then, the speed can be estimated according to the pose change of the specified position point at the previous and subsequent moments;
[0089] Specifically, it includes: Step 41, record each pair of associated first and second target detection boxes as the corresponding first association box and second association box;
[0090] Step 42, extract the point cloud within the first association box from the previous frame point cloud as the first point cloud; and extract the point cloud within the second association box from the subsequent frame point cloud as the second point cloud;
[0091] Step 43, perform point cloud pose registration on the first and second point clouds based on the iterative closest point ICP algorithm to obtain the corresponding pose transformation matrix T;
[0092] Among them, R is the rotation matrix and t is the translation vector;
[0093] Specifically, it includes: Step 431, taking the first point cloud as the current source point cloud, taking the second point cloud as the current target point cloud; initializing the value of the iteration counter to 1; and taking the preset initial rotation matrix R0 and initial translation vector t0 as the corresponding previous rotation matrix and previous translation vector;
[0094] Step 432, based on the preset minimum distance threshold d min determine the set of corresponding point pairs in the current source point cloud and the current target point cloud;
[0095] Among them, the set of corresponding point pairs includes multiple corresponding point pairs G k , where k is the corresponding point pair index, 1 ≤ k; the corresponding point pair G k includes a point p in the current source point cloud k and a point q in the current target point cloud k , and the distance between the point p k and the point q k is less than or equal to the minimum distance threshold d min ;
[0096] Here, there are multiple ways to establish the corresponding point pairs; one specific way is: traverse each point in the current target point cloud, and based on the minimum distance threshold d min construct a spherical space with the currently traversed point as the center of the sphere and a radius of d min , and include the points in the current source point cloud whose point cloud coordinates fall into this spherical space into the same matching point set. If the number of points in the matching point set is not 0, then take the currently traversed point as a point q k and select the point closest to the currently traversed point from this matching point set as the corresponding point p k thus obtaining a set of corresponding point pairs G k ; after the traversal, all the obtained corresponding point pairs G k form the set of corresponding point pairs;
[0097] Step 433, count the number of corresponding point pairs G k in the set of corresponding point pairs and record it as the third quantity H;
[0098] Step 434, construct the corresponding objective function f(R’, t’) based on the Euclidean distance transformation:
[0099]
[0100] Among them, R’ is the rotation matrix variable, t’ is the translation vector variable, and c k is the normal vector of the point q k ;
[0101] Here, the objective function f(R’, t’) is constructed according to the objective function structure of the point-to-plane ICP algorithm; based on the point-to-plane ICP algorithm, each p k is regarded as a source point, and each q k is regarded as a target point. The point (R′p k +t′) obtained after the transformation of p k should be closer to the target point q k . The projection of the distance vector (R′p k +t′-q k ) on the normal vector c k of q k , that is, (R′p k +t′-q k )·c k should be smaller; the sum of (R′p k +t′-q k )·c k for all source point-target point pairs of the two point sets (source point set, target point set) that complete registration, that is, the objective function f(R’, t’) should be a minimum value;
[0102] Step 435: Solve the two variables R’ and t’ that minimize the objective function f(R’, t’), and use the solution results as the corresponding current rotation matrix R * and the current translation vector t * ;
[0103] Step 436: Calculate the error between the current rotation matrix R * and the previous rotation matrix, and the error between the current translation vector t * and the previous translation vector to obtain the corresponding rotation matrix error and translation vector error; if the rotation matrix error and the translation vector error do not both satisfy their respective preset reasonable error ranges, then go to Step 437; if the rotation matrix error and the translation vector error both satisfy their respective preset reasonable error ranges, then go to Step 439;
[0104] Here, if the iteration based on the point-to-plane ICP algorithm does not set an exit condition, it may calculate indefinitely; to save computing resources, an exit condition is set in the embodiments of the present invention: a reasonable error range; there are two reasonable error ranges, corresponding to the rotation matrix error and the translation vector error respectively. When both the rotation matrix error and the translation vector error enter their respective reasonable error ranges, it means that the transformed source points are very close to the target points. At this time, the iteration can be stopped and go to Step 439;
[0105] Step 437, increment the value of the iteration counter by 1; if the value of the iteration counter after incrementing does not exceed the preset maximum number of iterations, go to Step 438; if the value of the iteration counter after incrementing exceeds the maximum number of iterations, go to Step 439;
[0106] Here, during the iterative process of the point-to-plane ICP algorithm, sometimes it is very difficult for the rotation matrix error and the translation vector error to converge reasonably to their respective reasonable error ranges at the same time. At this time, if no additional exit condition is set, it may also cause a large consumption of computing resources; to save computing resources, the embodiment of the present invention also sets an exit condition: the maximum number of iterations; regardless of the convergence of the rotation matrix error and the translation vector error, once the number of iterations represented by the iteration counter exceeds the maximum number of iterations, it will go to Step 439 to forcibly stop the iteration;
[0107] Step 438, from the current rotation matrix R * and the current translation vector t * constitute the corresponding current pose transformation matrix T * , and based on the current pose transformation matrix T * perform pose transformation on each point in the current source point cloud, and use the transformed point cloud as the new current source point cloud; and use the current rotation matrix R * as the new previous rotation matrix, and use the current translation vector t * as the new previous translation vector; and go to Step 432 to continue the iterative operation according to the new current source point cloud, the previous rotation matrix, and the previous rotation matrix;
[0108] Here, when the rotation matrix error and the translation vector error have not completely converged and the number of iterations has not reached the maximum number of iterations, the embodiment of the present invention will perform pose transformation on the current source point cloud according to the current rotation matrix R * and the current translation vector t * to obtain a new source point cloud that is closer to the target point cloud and use it as the current source point cloud for the next iteration. Correspondingly, the current rotation matrix R * and the current translation vector t * will also be used as the previous rotation matrix and the previous rotation matrix for the next iteration; after preparing the data required for the next iteration, the embodiment of the present invention will continue to go to Step 432 to perform the next iteration process according to the data required for the next iteration (the new current source point cloud, the previous rotation matrix, and the previous rotation matrix);
[0109] Step 439, stop the iterative operation and use the latest current rotation matrix R * and the current translation vector t * as the corresponding rotation matrix R and translation vector t, and construct the corresponding pose transformation matrix T from the rotation matrix R and the translation vector t,
[0110] Here, when the rotation matrix error and the translation vector error are completely convergent or the number of iterations reaches the maximum number of iterations, the embodiment of the present invention stops this round of iteration and uses the finally obtained current rotation matrix R * and the current translation vector t * as the corresponding rotation matrix R and translation vector t for output, and thus obtains the corresponding pose transformation matrix T;
[0111] Step 44, record the point cloud coordinates corresponding to the specified position points on the first association box corresponding to the first point cloud as the first point coordinates (x 1,c , y 1,c , z 1,c ); and perform a rigid body pose transformation on the first point coordinates (x 1,c , y 1,c , z 1,c ) based on the pose transformation matrix T to obtain the corresponding second point coordinates (x 2,c , y 2,c , z 2,c );
[0112] Here, there are various setting methods for the specified position points. One of them is: regardless of the target type, it is uniformly set as the center point coordinates (x 1,i , y 1,i , z 1,i ) of the first target box corresponding to the first association box; another one is: regardless of the target type, it is uniformly set as the center point of the bottom edge closest to the vehicle of the first association box; another one is defined according to the target type. Specifically, if the target type corresponding to the first association box is a person, a bicycle, an animal, a plant, or a traffic marker, then the specified position point is set as the center point coordinates (x 1,i , y 1,i , z 1,i ) of the first target box corresponding to the first association box. If the target type corresponding to the first association box is a vehicle, then the specified position point is set as the center point of the bottom edge closest to the vehicle of the first association box or the point on the first association box that matches the center point of the rear axle of the vehicle;
[0113] Step 45, record the corresponding targets of the first and second association boxes as the current target; and estimate the velocity vector V of the current target according to the first and second point coordinates and the time difference between the previous and current moments;
[0114] Specifically, it includes: recording the time difference between the previous and current moments as the time difference δ; and according to the first point coordinates (x 1,c , y 1,c , z 1,c ), the second point coordinates (x 2,c , y 2,c , z 2,c) and the time difference δ to estimate the three velocity components of the current target to obtain the corresponding v x , v y and v z ; and from the three velocity components v x , v y and v z to form the velocity vector V of the current target; v x = (x 2,c - x 1,c ) / δ, v y = (y 2,c - y 1,c ) / δ, v z = (z 2,c - z 1,c ) / δ.
[0115] Figure 2 is a schematic structural diagram of an electronic device provided in the second embodiment of the present invention. The electronic device may be the aforementioned terminal device or server, or may be a terminal device or server connected to the aforementioned terminal device or server to implement the method of the embodiment of the present invention. As Figure 2 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver actions of the transceiver 303. Various instructions may be stored in the memory 302 to be used to complete various processing functions and implement the processing steps described in the foregoing method embodiments. Preferably, the electronic device related to the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The aforementioned communication port 306 is used for the electronic device to connect and communicate with other peripherals.
[0116] In Figure 2 the system bus 305 mentioned may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 2 only a thick line is used to represent it in , but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0117] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0118] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium, in which instructions are stored. When the instructions run on a computer, the computer is enabled to execute the methods and processing procedures provided in the above embodiments.
[0119] The embodiments of the present invention also provide a chip for running instructions, and the chip is used to execute the processing steps described in the foregoing method embodiments.
[0120] The embodiments of the present invention provide a processing method, an electronic device, and a computer-readable storage medium for predicting the target speed based on point cloud data. First, target recognition is performed on two frames of lidar point clouds at adjacent times, and then the target recognition frames on the two frames of point clouds are associated to obtain multiple pairs of associated target recognition frames. Then, based on the ICP algorithm, the point clouds in each pair of associated target recognition frames are registered to obtain the pose transformation matrix T between the front and rear times, and based on the pose transformation matrix T, the pose of a specified position point on the target recognition frame at the previous time is predicted at the next time, and the speed of the corresponding target is calculated according to the pose of the point at the previous time and the predicted pose at the next time. Through the present invention, by continuously associating the outputs of the traditional point cloud target detection model and registering the point cloud poses, the speeds of each target in the point cloud can be obtained, which not only overcomes the defect that the traditional point cloud target detection model cannot perform speed detection, but also reduces the dependence of the perception system on the millimeter-wave radar, reduces the point cloud fusion calculation amount of the data fusion module, and improves the overall working efficiency of the perception system.
[0121] Those skilled in the art should also be able to further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0122] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0123] The specific embodiments described above have further elaborated on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A processing method for target speed prediction based on point cloud data, characterized in that, The method includes: Obtaining the front and rear two-frame lidar point cloud data at adjacent moments as the corresponding front-frame point cloud and rear-frame point cloud; Performing point cloud object detection processing on the front-frame and rear-frame point clouds respectively based on a point cloud object detection model to obtain a plurality of first object detection frames and a plurality of second object detection frames; the first object detection frames correspond to the front-frame point cloud, and the second object detection frames correspond to the rear-frame point cloud; Performing object association processing on the first and second object detection frames of the front and rear frame point clouds; Predicting the speed of the corresponding object according to each pair of associated first and second object detection frames; Wherein, each of the first object detection frames corresponds to a set of first detection frame parameters; the first detection frame parameters include the center point coordinates (x1, y1, z1) of the first target frame, the depth l1 of the first target frame, the width w1 of the first target frame, the height h1 of the first target frame, and the yaw angle yaw1 of the first target frame; Each of the second object detection frames corresponds to a set of second detection frame parameters; the second detection frame parameters include the center point coordinates (x2, y2, z2) of the second target frame, the depth l2 of the second target frame, the width w2 of the second target frame, the height h2 of the second target frame, and the yaw angle yaw2 of the second target frame; The performing object association processing on the first and second object detection frames of the front and rear frame point clouds specifically includes: Counting the number of the first detection frame parameters and recording it as the first number m, and counting the number of the second detection frame parameters and recording it as the second number n; Sequentially encode the center point coordinates of each of the first target boxes based on the target box index i to obtain the corresponding center point coordinates (x 1,i , y 1,i , z 1,i ), where 1 ≤ i ≤ m; Sequentially encode the center point coordinates of each of the second target boxes based on the target box index j to obtain the corresponding center point coordinates (x 2,j , y 2,j , z 2,j ), where 1 ≤ j ≤ n; Based on the Kalman filter, the coordinates (x 1,i , y 1,i , z 1,i ) of the center point of each of the first target bounding boxes at the next moment are predicted to obtain the corresponding predicted target bounding box center point coordinates For the center point coordinates of each of the predicted target boxes and the center point coordinates (x 2,j , y 2,j , z 2,j ) of each of the second target boxes, calculate the distances to obtain m*n center point spacings s i,j , composed of m*n of the center point spacing s i,j Construct a matrix vector with the shape of m*n, denoted as the first matrix vector; and input the first matrix vector into the Deep Hungarian Network (DHN) to calculate the matching degree based on the improved Hungarian algorithm, obtaining the corresponding association matrix vector A with the shape of m*n; the association matrix vector A includes m*n matching degrees a i,j , 0 < a i,j ≤1; each of the matching degrees a i,j corresponds to a pair of the first target detection box and the second target detection box based on the target box index i and the target box index j Group the second quantity n of the matching degrees a corresponding to the same target box index i i,j into the same matching degree set D i ; and extract the maximum matching degree a among the values in each matching degree set D i as the corresponding maximum matching degree b i,j ; i ; For each of the maximum matching degrees b i make a judgment; if the current maximum matching degree b i is less than the preset minimum matching degree threshold a min , then confirm that there is no association relationship between the pair of the first target detection box and the second target detection box corresponding to the current maximum matching degree b i ; if the current maximum matching degree b i is greater than or equal to the minimum matching degree threshold a min , then create an association relationship between the pair of the first target detection box and the second target detection box corresponding to the current maximum matching degree b i .
2. The processing method for target speed prediction based on point cloud data according to claim 1, characterized in that, The point cloud object detection model includes a VoxelNet model, a SECOND model, and a PointPillars model, and the PointPillars model is used by default.
3. The processing method for target speed prediction based on point cloud data according to claim 1, characterized in that, The predicting the speed of the corresponding object according to each pair of associated first and second object detection frames specifically includes: Denoting each pair of associated first and second object detection frames as the corresponding first associated frame and second associated frame; Extracting the point cloud within the first associated frame from the front-frame point cloud as the first point cloud; and extracting the point cloud within the second associated frame from the rear-frame point cloud as the second point cloud; Based on the Iterative Closest Point (ICP) algorithm, perform point cloud pose registration on the first and second point clouds to obtain the corresponding pose transformation matrix T; R is the rotation matrix and t is the translation vector; Denote the point cloud coordinates corresponding to the specified position points on the first associated box corresponding to the first point cloud as the first point coordinates (x 1,c , y 1,c , z 1,c ); and perform a rigid body pose transformation on the first point coordinates (x 1,c , y 1,c , z 1,c ) based on the pose transformation matrix T to obtain the corresponding second point coordinates (x 2,c , y 2,c , z 2,c ); Denoting the corresponding objects of the first and second associated frames as the current object; and estimating the speed vector V of the current object according to the first and second point coordinates and the time difference between the front and rear moments.
4. The processing method for target speed prediction based on point cloud data according to claim 3, characterized in that, The performing point cloud pose registration on the first and second point clouds based on the iterative closest point ICP algorithm to obtain the corresponding pose transformation matrix T specifically includes: Step 51, taking the first point cloud as the current source point cloud, taking the second point cloud as the current target point cloud; initializing the value of the iteration counter to 1; and taking the preset initial rotation matrix R0 and initial translation vector t0 as the corresponding previous rotation matrix and previous translation vector; Step 52, based on a preset minimum distance threshold d min determine a set of corresponding point pairs in the current source point cloud and the current target point cloud; the set of corresponding point pairs includes a plurality of corresponding point pairs G k , where k is the corresponding point pair index, 1 ≤ k; the corresponding point pair G k includes a point p in the current source point cloud k and a point q in the current target point cloud k , and the distance between the point p k and the point q k is less than or equal to the minimum distance threshold d min ; Step 53, count the number of the association point pairs G in the set of the association point pairs, and denote it as the third quantity H; k Step 54, construct a corresponding objective function f(R ’ ,t ’ ): R ’ is the rotation matrix variable, t ’ is the translation vector variable, c k is the normal vector of the point q k ; Step 55, solve for the two variables R ’ , t ’ ) that minimize the objective function f(R ’ , t ’ ), and use the solution results as the corresponding current rotation matrix R * and the current translation vector t * ; Step 56, for the current rotation matrix R * calculate the error with the previous rotation matrix and the current translation vector t * and the error with the previous translation vector to obtain the corresponding rotation matrix error and translation vector error; if the rotation matrix error and the translation vector error do not both satisfy their respective preset reasonable error ranges, go to step 57; if the rotation matrix error and the translation vector error both satisfy their respective preset reasonable error ranges, go to step 59; Step 57, adding 1 to the value of the iteration counter; if the value of the iteration counter after adding 1 does not exceed the preset maximum number of iterations, then go to step 58; if the value of the iteration counter after adding 1 exceeds the maximum number of iterations, then go to step 59; Step 58, from the current rotation matrix R * and the current translation vector t * to form the corresponding current pose transformation matrix T * , and based on the current pose transformation matrix T * perform pose transformation on each point in the current source point cloud, and use the point cloud obtained after the transformation as the new current source point cloud; and use the current rotation matrix R * as the new previous rotation matrix, and use the current translation vector t * as the new previous translation vector; and go to Step 52 to continue the iterative calculation according to the new current source point cloud, the previous rotation matrix, and the previous rotation matrix; Step 59, stop the iterative operation and use the latest current rotation matrix R * and the current translation vector t * as the corresponding rotation matrix R and translation vector t, and construct the corresponding pose transformation matrix T from the rotation matrix R and the translation vector t 5. The processing method for target speed prediction based on point cloud data according to claim 3, wherein, Estimating the velocity vector V of the current target based on the first and second point coordinates and the time difference between the previous and current moments specifically includes: Denote the time difference between the previous and current moments as the time difference δ; and estimate the three velocity components of the current target based on the first point coordinates (x 1,c , y 1,c , z 1,c ), the second point coordinates (x 2,c , y 2,c , z 2,c ) and the time difference δ to obtain the corresponding v x , v y and v z ; and form the velocity vector V of the current target from the three velocity components v x , v y and v z ; v x = (x 2,c - x 1,c ) / δ, v y = (y 2,c - y 1,c ) / δ, v z = (z 2,c - z 1,c ) / δ.
6. An electronic device, wherein, Including: A memory, a processor, and a transceiver; The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method according to any one of claims 1-5; The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.
7. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1-5.
Citation Information
Patent Citations
Method and device for generating object detection box, equipment, storage medium and vehicle
CN109188457A
Vehicle speed detection and collision early warning method and electronic equipment
CN114332153A