3D object detection frame processing method and device, vehicle and storage medium

By smoothing the historical detection frame of the 3D object detection results of lidar in autonomous driving, the problem of inconsistency in the time dimension of the detection results is solved, and the stability and reliability of the detection results are improved.

CN119942523APending Publication Date: 2025-05-06AUTOMOTIVE INTELLIGENCE & CONTROL OF CHINA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411999458.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In autonomous driving technology, the 3D object detection technology based on lidar has inconsistency in the time dimension of the detection results, which leads to a jump in the detection results and affects subsequent path planning and control tasks.

Method used

By obtaining the historical detection box smoothing result and historical normalization factor of the target object, and combining the current detection box data, smoothing processing is performed to obtain stable detection box data, thereby determining the 3D object detection box of the target object.

Benefits of technology

It improves the stability and reliability of 3D object detection results, reduces the jump of detection results, and enhances the timing consistency recognition of target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942523A_ABST
    Figure CN119942523A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic driving, and discloses a 3D object detection frame processing method and device, a vehicle and a storage medium, and the method comprises the steps: obtaining a historical detection frame smoothing result of a target object in a last point cloud image frame, and a corresponding historical normalization factor, the historical normalization factor is determined based on a target smoothing weight corresponding to a historical detection frame smoothing result; acquiring current detection frame data of the target object in the current point cloud image frame; when the current detection frame data is non-singular data, smoothing the current detection frame data based on a historical detection frame smoothing result and a historical normalization factor to obtain target detection frame data; determining a 3D object detection frame of the target object based on the target detection frame data; according to the invention, the stability and reliability of the 3D object detection result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a 3D object detection frame processing method, device, vehicle and storage medium. Background Art

[0002] With the advancement of autonomous driving perception technology, 3D object detection technology based on LiDAR has also made significant progress. LiDAR point cloud data is highly favored due to its accurate 3D ranging capability, so many autonomous driving vehicles are equipped with LiDAR.

[0003] When using LiDAR to perceive the surrounding environment, artificial intelligence (AI) algorithms or traditional algorithms are usually used to identify and understand surrounding objects. For example, the pointpillars algorithm and the traditional PCL (point cloud library) algorithm commonly used in the industry are used to detect obstacles around the vehicle and determine the 3D detection frame of each obstacle.

[0004] However, due to the dynamic changes of LiDAR data and the accuracy of the algorithm, the detection results of consecutive frames are often different. This phenomenon manifests itself in the time dimension as the 3D detection box of the obstacle is inconsistent in consecutive frames, resulting in a jump in the detection results. If the jump is too drastic, it may have an adverse effect on subsequent path planning and control tasks. Summary of the invention

[0005] In view of this, the present invention provides a 3D object detection frame processing method, device, vehicle and storage medium to solve the problem of jumps in 3D object detection frame processing results in autonomous driving related technologies.

[0006] In the first aspect, the present invention provides a 3D object detection frame processing method, the method comprising: obtaining the historical detection frame smoothing result of the target object in the previous point cloud image frame, and the corresponding historical normalization factor, the historical normalization factor is determined based on the target smoothing weight corresponding to the historical detection frame smoothing result; obtaining the current detection frame data of the target object in the current point cloud image frame; based on the historical detection frame smoothing result and the historical normalization factor, smoothing the current detection frame data to obtain the target detection frame data; determining the 3D object detection frame of the target object based on the target detection frame data. Through the above process, the stability and reliability of the 3D object detection result can be improved.

[0007] In some optional implementations, based on the historical detection frame smoothing results and the historical normalization factor, the current detection frame data is smoothed to obtain the target detection frame data, including:

[0008] Input the historical detection frame smoothing result and the historical normalization factor into the detection frame smoothing model to obtain the smoothing result;

[0009] The target detection frame data is obtained based on the smoothing processing results.

[0010] In some optional implementations, the detection box smoothing model is:

[0011] Y i =(Y i-1 *smoothweight+X i )

[0012] Z i =Y i / norfactor i-1

[0013] normfactor i =1+normfactor i-1 *smoothweight

[0014] Among them, i is the frame number of the point cloud image frame, Y -1 =0, normfactor -1 =1,Y i is the historical detection frame smoothing data corresponding to the i-th point cloud image frame, X i is the detection frame data corresponding to the i-th point cloud image frame, Z i is the smoothing result of the detection frame data corresponding to the i-th point cloud image frame, normfactori i is the normalization coefficient corresponding to the i-th point cloud image frame, smoothweight is the target smoothing weight, Y -1 =0,norfactor -1 =1.

[0015] In some optional implementations, after obtaining the current detection box data of the target object in the current point cloud image frame, the method further includes:

[0016] Get the historical detection box data of the target object;

[0017] Determine whether the current detection frame data is singular data based on the historical detection frame data;

[0018] When the current detection frame data is not singular data, a step of performing smoothing processing on the current detection frame data based on the historical detection frame smoothing result and the historical normalization factor to obtain the target detection frame data;

[0019] When the current detection frame data is singular data, the current detection frame data is discarded, and the 3D object detection frame of the target object is determined based on the historical detection frame smoothing result.

[0020] In some optional implementations, judging whether current detection box data is singular data based on historical detection box data includes:

[0021] Determine the calibration detection frame data based on the current detection frame data and the historical detection frame data;

[0022] Calculate the data difference between the current detection frame data and the calibration detection frame data;

[0023] Based on the data difference, determine whether the current detection box data is singular data.

[0024] In some optional implementations, determining the calibration detection frame data based on the current detection frame data and the historical detection frame data includes:

[0025] Obtain the union of current contour data in the current detection frame data and historical contour data in the historical detection frame data to obtain a detection frame contour data set;

[0026] Get the contour statistical mean corresponding to each dimension of the contour points in the detection frame contour data set;

[0027] Calculate the standard deviation of each dimension of the contour points in the current contour data and the corresponding contour statistical mean to obtain the contour standard deviation;

[0028] The calibration detection frame data is determined based on the contour statistical mean and contour standard deviation of each dimension of contour points.

[0029] In some optional implementations, determining the calibration detection frame data based on the current detection frame data and the historical detection frame data further includes:

[0030] Calculate the variance of each dimension of the contour points in the current contour data and the corresponding contour statistical mean to obtain the contour variance;

[0031] Determine the contour covariance between contour points in each dimension based on the contour variance;

[0032] Construct a profile covariance matrix based on profile variance and profile covariance;

[0033] Based on the vector between the contour statistical means and the contour covariance matrix, the calibration detection box data is determined.

[0034] In the second aspect, the present invention provides a 3D object detection frame processing device, which mainly includes: a data acquisition module, which is used to obtain the historical detection frame smoothing result of the target object in the previous point cloud image frame, and the corresponding historical normalization factor, the historical normalization factor is determined based on the target smoothing weight corresponding to the historical detection frame smoothing result; a data processing module, which is used to obtain the current detection frame data of the target object in the current point cloud image frame; a smoothing processing module, which is used to smooth the current detection frame data based on the historical detection frame smoothing result and the historical normalization factor to obtain the target detection frame data; a detection frame determination module, which is used to determine the 3D object detection frame of the target object based on the target detection frame data. Through the above modules, the stability and reliability of the 3D object detection results can be improved.

[0035] In a third aspect, the present invention provides a vehicle, comprising: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the 3D object detection frame processing method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0036] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the 3D object detection frame processing method of the first aspect or any corresponding embodiment thereof.

[0037] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions, wherein the computer instructions are used to enable a computer to execute the 3D object detection frame processing method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0039] Figure 1 It is a schematic diagram of 3D object detection of laser point cloud in related technology;

[0040] Figure 2 is a flowchart of a 3D object detection frame processing method according to an embodiment of the present invention;

[0041] Figure 3 is another flowchart of the 3D object detection frame processing method according to an embodiment of the present invention;

[0042] Figure 4 is a rendering of a 3D object detection frame processing method according to an embodiment of the present invention;

[0043] Figure 5 is a structural block diagram of a 3D object detection frame processing device according to an embodiment of the present invention;

[0044] Figure 6 Schematic diagram of the hardware structure of a vehicle according to an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0046] With the advancement of autonomous driving perception technology, 3D object detection technology based on LiDAR has also made significant progress. LiDAR point cloud data is highly favored due to its accurate 3D ranging capability. Therefore, many autonomous driving vehicles are equipped with LiDAR. The schematic diagram of its laser point cloud 3D object detection is shown below. Figure 1 As shown in the figure, 3D object perception technology based on LiDAR plays a vital role in the perception module of autonomous driving. For advanced autonomous driving systems above L3 level, LiDAR has become an indispensable key sensor.

[0047] In the process of using LiDAR to perceive the surrounding environment, related technologies usually use AI algorithms or traditional algorithms to detect and perceive surrounding objects in 3D, such as pointpillars algorithms and traditional PCL point cloud computing algorithms. However, these types of methods can mostly detect the types of obstacles around the vehicle and the 3D-box and heading angle of each obstacle. However, from the perspective of timing, due to changes in laser point cloud data, algorithm accuracy and other reasons, there are differences in the detection results of the previous and next frames. Intuitively speaking, the 3D-box and heading angle of the detected obstacle will continue to change in the time dimension, resulting in jumps in the detection results. If the jump is too large, it may even affect the downstream planning and control tasks.

[0048] Based on this, an embodiment of the present invention provides a 3D object detection frame processing method, which includes: obtaining the historical detection frame smoothing result of the target object in the previous point cloud image frame, and the corresponding historical normalization factor, the historical normalization factor is determined based on the target smoothing weight corresponding to the historical detection frame smoothing result; obtaining the current detection frame data of the target object in the current point cloud image frame; based on the historical detection frame smoothing result and the historical normalization factor, smoothing the current detection frame data to obtain the target detection frame data; determining the 3D object detection frame of the target object based on the target detection frame data. Therefore, by deleting the abnormal detection results in the detection frame of the same object in time sequence and smoothing the detection frame at the same time, the 3D object detection result is kept consistent in the time dimension.

[0049] According to an embodiment of the present invention, an embodiment of a 3D object detection frame processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0050] In this embodiment, a 3D object detection frame processing method is provided, which can be used for the above-mentioned vehicle. Figure 2 is a flowchart of a 3D object detection frame processing method according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:

[0051] Step 210, obtaining the historical detection frame smoothing result of the target object in the previous point cloud image frame and the corresponding historical normalization factor.

[0052] Among them, the historical normalization factor is determined based on the target smoothing weight corresponding to the historical detection box smoothing result, and the value range of the target smoothing weight is 0.6 to 0.99.

[0053] In some optional implementations, when obtaining the historical detection frame smoothing result of the target object in the previous point cloud image frame, the historical detection frame data sequence corresponding to the target object in the historical point cloud image frame can be first obtained, and after filtering out the singular data in the historical detection frame data sequence, the initial normalization coefficient corresponding to the first detection frame data in the historical detection frame data sequence is obtained, and the initial normalization coefficient and the first detection frame data are input into the detection frame smoothing model to obtain the first detection frame smoothing result. Then, based on the initial normalization coefficient and the corresponding target smoothing weight, the first normalization coefficient is determined; the first detection frame smoothing result, the second detection frame data in the historical detection frame data sequence, and the first normalization coefficient are input into the detection frame smoothing model to obtain the second detection frame smoothing result...

[0054] Finally, the last detection frame data in the historical detection frame data sequence, the previous detection frame smoothing result, and the corresponding normalization coefficient are input into the detection frame smoothing model to obtain the historical detection frame smoothing result of the target object in the previous point cloud image frame.

[0055] In the actual operation process, for the target object that has completed 3D object detection and tracking, it is assumed that its historical detection frame data (perception data) is a 7-dimensional vector of x, y, z, l, w, h, θ, where x, y, z represent the coordinates of the center point of the historical detection frame in the historical detection frame data, l, w, h represent the historical contour data in the historical detection frame data (the length, width and height of the historical detection frame), and the angle θ represents the yaw angle of the historical detection frame (3D detection frame) in the bird's eye view. After the tracking algorithm, the tracking algorithm here can be performed by aligning the previous and next frames and performing IOU matching. In fact, other tracking algorithms can also be used here to add the detection results of the previous historical frame and the current frame to a queue with a fixed length (such as setting the queue length to 20), so that a historical detection frame data (3D state quantity data) of this target object [x i ,y i ,z i ,l i ,w i ,h i ,θ i ],i∈[0,20], in this queue, first divide these state quantities into two parts. The first part is the center point x,y,z of the historical detection frame. This part of data can be smoothed by the Kalman filter method. The main reason is that the state transition equation of this part of data in time series can be obtained, which will not be repeated here. The second part is the contour state data (state quantity data) l,w,h,θ of the detection frame. The stability of this part of data directly determines the consistency of the 3D perception result, so it is necessary to filter out singular data and smooth its detection frame.

[0056] Step 220, obtaining current detection frame data of the target object in the current point cloud image frame.

[0057] The current detection frame data of the target object in the current point cloud image frame may be l, w, h, θ. By obtaining the current detection frame data of the target object in the current point cloud image frame, the current detection frame data is smoothed using the historical detection frame smoothing result of the target object to avoid the detection frame of the target object jumping in time sequence, thereby affecting the accuracy and reliability of tracking the target object.

[0058] Step 230, based on the historical detection frame smoothing results and the historical normalization factor, the current detection frame data is smoothed to obtain the target detection frame data.

[0059] As described above, the current detection frame data is smoothed based on the historical detection frame smoothing result and the historical normalization factor to obtain the target detection frame data, so as to determine the 3D object detection frame of the target object based on the target detection frame.

[0060] In some optional implementations, the current detection frame data is smoothed based on the historical detection frame smoothing results and the historical normalization factor. When the target detection frame data is obtained, the historical detection frame smoothing results and the historical normalization factor can be input into the detection frame smoothing model to obtain a smoothing result; the target detection frame data is obtained based on the smoothing result.

[0061] Among them, the detection box smoothing model is:

[0062] Y i =(Y i-1 *smoothweight+X i )

[0063] Z i =Y i / norfactor i-1

[0064] normfactor i =1+normfactor i-1 *smoothweight

[0065] Among them, i is the frame number of the point cloud image frame, Y -1 =0, normfactor -1 =1,Y i is the historical detection frame smoothing data corresponding to the i-th point cloud image frame, X i is the detection frame data corresponding to the i-th point cloud image frame (used to characterize the corresponding dimension attribute data values ​​in l, w, h, θ before smoothing, that is, l, w, h represent the length, width and height of the detection frame, and the angle θ represents the yaw angle of the detection frame in the bird's-eye view), Z i is the smoothing result of the detection frame data corresponding to the i-th point cloud image frame (used to characterize the attribute data value of a certain dimension in l, w, h, θ after smoothing), normfactor i is the normalization coefficient corresponding to the i-th point cloud image frame, smoothweight is the target smoothing weight, Y -1 =0, normfactor -1 =1.

[0066] Specific:

[0067] Y0=Y -1 *smoothweight+X0=X0

[0068] Z0=Y0 / normfactor -1 =Y0=X0

[0069] normfactor0=1+normfactor -1 *smoothweight=1+smoothweight

[0070] Y1=(Y0*smoothweight+X1)=(X0*smoothweight+X1)

[0071] Z1=Y1 / normfactor0=(X0*smoothweight+X1) / (1+smoothweight)

[0072] normfactor1=1+normfactor0*smoothweight=1+smoothweight*(1+smoothweight)

[0073] =1+smoothweight+smoothweight 2

[0074] Y2=(Y1*smoothweight+X2)=(X0*smoothweight+X1)*smoothweight+X2

[0075] =X0*smoothweight 2 +X1*smoothweight+X2

[0076]

[0077] normfactor2=1+normfactor1*smoothweight

[0078] =1+smoothweight*(1+smoothweight+smoothweight 2 )

[0079] =1+smoothweight+smoothweight 2 +smoothweight 3

[0080] ……

[0081] It can be seen that the historical normalization factor of this embodiment changes as a whole with the change of time series, and is not a fixed weight; at the same time, in the time series dimension, the calculation of the historical normalization factor of the current point cloud image frame suffers from "historical forgetting", that is, in the time series dimension, the weighted weight of the measurement value that is farther away from the current frame will decrease through the increase of the exponential value of smoothweight, thereby slowly forgetting the historical measurement value; therefore, there is no need to store all historical measurement values, but only two variables, the historical detection frame smoothing result and the historical normalization factor of the previous point cloud image frame, need to be stored to achieve the same effect in engineering implementation. In addition, for the selection of smoothweight, we did some comparative experiments. Taking the length data smoothing of the 3D detection frame as an example, we set the values ​​of smoothweight to 0.6, 0.8, 0.95, and 0.99 respectively. We can see that the data before smoothing is as follows Figure 3 The dashed line in the figure and the smoothed data are shown in Figure 3 That is, after the detection amount of the 3D object is smoothed and processed, from the display of the data in the time dimension, the value of its state quantity is significantly and stably enhanced, thereby optimizing the result of 3D object detection.

[0082] Step 240 : determining a 3D object detection frame of the target object based on the target detection frame data.

[0083] Specifically, in the process of determining the 3D object detection frame, it is first necessary to analyze the target detection frame data to identify the characteristics of the target object. This step usually involves extracting information such as the size, shape, and position of the target object. Then, using this feature information, combined with the historical detection frame smoothing results and historical normalization factors of the current point cloud image frame, a specific algorithm model is used to predict the 3D position and size of the target object. Finally, one or more 3D object detection frames are output, which can accurately cover the target object and provide accurate target object location information for the autonomous driving vehicle, thereby assisting the vehicle in making safe decisions and path planning.

[0084] The 3D object detection frame processing method provided in this embodiment first obtains the historical detection frame smoothing result of the target object in the previous point cloud image frame and the corresponding historical normalization factor; then obtains the current detection frame data of the target object in the current point cloud image frame; then, based on the historical detection frame smoothing result and the historical normalization factor, the current detection frame data is smoothed to obtain the target detection frame data; finally, the 3D object detection frame of the target object is determined based on the target detection frame data. Through the above process, the stability and reliability of the 3D object detection result can be improved.

[0085] In this embodiment, a 3D object detection frame processing method is provided, which can be used for the above-mentioned vehicle. Figure 4is a flowchart of another embodiment of a 3D object detection frame processing method according to an embodiment of the present invention. Figure 4 As shown, the process includes the following steps:

[0086] Step 410 , obtaining the historical detection frame smoothing result of the target object in the previous point cloud image frame and the corresponding historical normalization factor.

[0087] For details, please see Figure 2 Step 210 of the illustrated embodiment will not be described in detail here.

[0088] Step 420, obtaining current detection frame data of the target object in the current point cloud image frame.

[0089] Specifically, the above step 420 includes:

[0090] Step 4201, obtaining historical detection frame data of the target object.

[0091] When obtaining the historical detection frame data of the target object, the approximate position of the target object in the historical point cloud image frame can be determined according to the motion state and direction of the target object; the point cloud data related to the target object can be extracted from the historical point cloud image frame using the predicted position information; the extracted point cloud data can be processed by the pointpillars algorithm or the PCL point cloud computing algorithm to generate the historical detection frame data of the target object.

[0092] Step 4202, determine whether the current detection frame data is singular data based on the historical detection frame data.

[0093] Specifically, the calibrated detection frame data can be first determined based on the current detection frame data and the historical detection frame data; then the data difference between the current detection frame data and the calibrated detection frame data is calculated; and then, based on the data difference, it is determined whether the current detection frame data is singular data.

[0094] In some optional embodiments, when determining the calibrated detection frame data based on the current detection frame data and the historical detection frame data, the union of the current contour data in the current detection frame data and the historical contour data in the historical detection frame data can be obtained to obtain the detection frame contour data set; the contour statistical mean corresponding to each dimensional contour point in the detection frame contour data set is obtained; the standard deviation of each dimensional contour point in the current contour data and the corresponding contour statistical mean is calculated to obtain the contour standard deviation; and the calibrated detection frame data is determined based on the contour statistical mean and contour standard deviation of each dimensional contour point.

[0095] Among them, the contour points of each dimension in the detection box contour dataset [l i ,w i ,h i ],i∈[0,20], the calculation model of the corresponding profile statistical mean is:

[0096]

[0097] Among them, μ l for l i The mean value of the contour statistics of the contour points, μ w w i The mean value of the contour statistics of the contour points, μ h h i The mean of the contour statistics of the contour points.

[0098] The standard deviation calculation model of each dimension of the contour points in the current contour data and the corresponding contour statistical mean is:

[0099]

[0100] Among them, σ l for l i The standard deviation of the contour points, σ w w i The standard deviation of the contour points is h i The standard deviation of the contour points.

[0101] When the calibration detection frame data is determined based on the contour statistical mean and contour standard deviation of each dimensional contour point, the calibration detection frame data can be determined based on the sum and difference of the contour statistical mean and contour standard deviation of each dimensional contour point.

[0102] When calculating the data difference between the current detection frame data and the calibrated detection frame data, and judging whether the current detection frame data is singular data based on the data difference, it can be determined based on the union of non-singular data of contour points in each dimension.

[0103] That is, l∈(μ l -δ l ,μ l +δ l )and w∈(μ w -δ w ,μ w +δ w )and h∈(μ h -δ h ,μ h +δh, then the above detection box contour data is non-singular data.

[0104] In some optional implementations, the variance of each dimensional contour point in the current contour data and the corresponding contour statistical mean can also be calculated to obtain the contour variance; the contour covariance between each dimensional contour points is determined based on the contour variance; a contour covariance matrix is ​​constructed based on the contour variance and the contour covariance; based on the vector between the contour statistical means and the contour covariance matrix, the calibration detection frame data is determined, and the data difference between the current detection frame data and the calibration detection frame data is calculated.

[0105] Specifically, considering that autonomous driving often only considers the detection data from the BEV perspective and often ignores the height h, the calculation model of the above profile variance is:

[0106]

[0107] Among them, var l for l i The contour variance of the contour points, var w w i dimensional contour variance of contour points.

[0108] The calculation model of the above profile covariance is:

[0109]

[0110] Among them, var lw for l i dimensional contour points and w i The contour covariance of the contour points, var wl w i dimensional contour points and l i dimensional contour covariance of contour points.

[0111] The above profile covariance matrix is

[0112] Based on the vector between the mean values ​​of the contour statistics and the contour covariance matrix, the model for calibrating the detection frame data is determined as follows:

[0113]

[0114] Assuming that the calibration detection box data is 0.5, you can input based on the current detection box data The calculation result obtained in determines whether the current detection frame data is singular data. That is, when the current detection frame data is input When the calculated result obtained in is greater than the calibration detection frame data, it is determined that the current detection frame data is non-singular data.

[0115] In some optional implementations, the height h may not be ignored, that is, the input vector of the above f is changed to a three-dimensional vector (l, w, h), and the overall calculation is the same as The calculation method of is similar and will not be described here.

[0116] It can be understood that only when the current detection box data meets the above conditions, the detection of this detection box is considered to be a successful detection. At this time, the tail data in the historical queue can be deleted and the detection vector of this time can be added to the fixed queue. During the whole process, the data and measurement values ​​in the queue change in real time. For another 3D state data of the target object [θ i ],i∈[0,20], because there is a vehicle turning. Therefore, we hope that its measurement value changes continuously and stably in time series, that is, it is necessary to smooth the current detection frame data based on the historical detection frame smoothing results and historical normalization factors.

[0117] Step 430, when the current detection frame data is not singular data, the current detection frame data is smoothed based on the historical detection frame smoothing result and the historical normalization factor to obtain the target detection frame data.

[0118] In some optional implementations, when the current detection frame data is singular data, the current detection frame data is discarded, and the 3D object detection frame of the target object is determined based on the historical detection frame smoothing result.

[0119] For details, please see Figure 2 Step 230 of the illustrated embodiment will not be described in detail here.

[0120] Step 440 : determining a 3D object detection frame of the target object based on the target detection frame data.

[0121] For details, please see Figure 2 Step 240 of the illustrated embodiment will not be described in detail here.

[0122] The 3D object detection frame processing method provided in this embodiment first obtains the historical detection frame smoothing result of the target object in the previous point cloud image frame and the corresponding historical normalization factor; then, obtains the current detection frame data of the target object in the current point cloud image frame, and determines whether the current detection frame data is singular data; then, when the current detection frame data is not singular data, the current detection frame data is smoothed based on the historical detection frame smoothing result and the historical normalization factor to obtain the target detection frame data; finally, the 3D object detection frame of the target object is determined based on the target detection frame data. Through the above process, the stability and reliability of the 3D object detection results can be improved.

[0123] This embodiment provides a 3D object detection frame processing device, such as Figure 5 As shown, including:

[0124] A data acquisition module 510 is used to acquire a historical detection frame smoothing result of a target object in a previous point cloud image frame and a corresponding historical normalization factor, wherein the historical normalization factor is determined based on a target smoothing weight corresponding to the historical detection frame smoothing result;

[0125] The data processing module 520 is used to obtain the current detection frame data of the target object in the current point cloud image frame;

[0126] A smoothing processing module 530 is used to perform smoothing processing on the current detection frame data based on the historical detection frame smoothing results and the historical normalization factor to obtain the target detection frame data;

[0127] The detection frame determination module 540 is used to determine the 3D object detection frame of the target object based on the target detection frame data.

[0128] In some optional implementations, the smoothing module 530 includes:

[0129] A parameter input unit, used to input the historical detection frame smoothing result and the historical normalization factor into the detection frame smoothing model to obtain a smoothing processing result;

[0130] The smoothing processing unit is used to obtain the target detection frame data based on the smoothing processing result.

[0131] In some optional implementations, the detection box smoothing model is:

[0132] Y i =(Y i-1 *smoothweight+X i )

[0133] Z i =Y i / norfactor i-1

[0134] normfactor i =1+normfactor i-1 *smoothweight

[0135] Among them, i is the frame number of the point cloud image frame, Y -1 =0, normfactor -1 =1,Y i is the historical detection frame smoothing data corresponding to the i-th point cloud image frame, X i is the detection frame data corresponding to the i-th point cloud image frame, Z i is the smoothing result of the detection frame data corresponding to the i-th point cloud image frame, normfactori iis the normalization coefficient corresponding to the i-th point cloud image frame, smoothweight is the target smoothing weight, Y -1 =0, normfactor -1 =1.

[0136] In some optional implementations, the data processing module 520 includes:

[0137] A data acquisition unit, used to acquire historical detection frame data of a target object;

[0138] A data judgment unit, used to judge whether the current detection frame data is singular data based on the historical detection frame data;

[0139] A data processing unit, configured to smooth the current detection frame data based on the historical detection frame smoothing result and the historical normalization factor to obtain the target detection frame data when the current detection frame data is not singular data;

[0140] The result determination unit is used to discard the current detection frame data when the current detection frame data is singular data, and determine the 3D object detection frame of the target object based on the historical detection frame smoothing result.

[0141] In some optional implementations, the control information generation module 530 includes:

[0142] A data determination unit, configured to determine calibration detection frame data based on current detection frame data and historical detection frame data;

[0143] A data calculation unit, used to calculate the data difference between the current detection frame data and the calibration detection frame data;

[0144] The singular data judgment unit is used to judge whether the current detection frame data is singular data based on the data difference.

[0145] In some optional embodiments, the data determination unit is specifically used to obtain the union of current contour data in the current detection frame data and historical contour data in the historical detection frame data to obtain a detection frame contour data set; obtain the contour statistical mean corresponding to each dimensional contour point in the detection frame contour data set; calculate the standard deviation of each dimensional contour point in the current contour data and the corresponding contour statistical mean to obtain the contour standard deviation; determine the calibrated detection frame data based on the contour statistical mean and contour standard deviation of each dimensional contour point.

[0146] In some optional embodiments, the data determination unit is also used to calculate the variance of each dimensional contour point in the current contour data and the corresponding contour statistical mean to obtain the contour variance; determine the contour covariance between each dimensional contour points based on the contour variance; construct a contour covariance matrix based on the contour variance and the contour covariance; and determine the calibration detection box data based on the vector between the contour statistical means and the contour covariance matrix.

[0147] The 3D object detection frame processing device provided by the present invention first obtains the historical detection frame smoothing result of the target object in the previous point cloud image frame and the corresponding historical normalization factor through the data acquisition module; then, obtains the current detection frame data of the target object in the current point cloud image frame through the data processing module; then, the current detection frame data is smoothed based on the historical detection frame smoothing result and the historical normalization factor through the smoothing processing module to obtain the target detection frame data; finally, the 3D object detection frame of the target object is determined based on the target detection frame data through the detection frame determination module. Through the above modules, the stability and reliability of the 3D object detection results can be improved.

[0148] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0149] The 3D object detection frame processing device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0150] The embodiment of the present invention further provides a vehicle having the above Figure 5 The 3D object detection frame processing device shown.

[0151] See also Figure 6 , Figure 6 is a schematic diagram of the structure of a vehicle provided by an optional embodiment of the present invention, such as Figure 6As shown, the vehicle includes: one or more processors 610, a memory 620, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed in the vehicle, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple vehicles can be connected, and each device provides some necessary operations (for example, as a storage server array, a group of blade storage servers, or a multi-processor system). Figure 6 A processor 610 is taken as an example.

[0152] The processor 610 may be a central processing unit, a network processor or a combination thereof. The processor 610 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable logic gate array, a general purpose array logic or any combination thereof.

[0153] The memory 620 stores instructions executable by at least one processor 610, so that the at least one processor 610 executes the method shown in the above embodiment.

[0154] The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created based on the use of a vehicle displayed on a small program landing page, etc. In addition, the memory 620 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 620 may optionally include a memory remotely arranged relative to the processor 610, and these remote memories may be connected to the vehicle via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a server cluster, a mobile communication network, and a combination thereof.

[0155] The memory 620 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 620 may also include a combination of the above types of memory.

[0156] The vehicle also includes a communication interface 630 for the vehicle to communicate with other devices or communication networks.

[0157] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.

[0158] The embodiment of the present invention further provides a computer program product, including computer instructions, where the computer instructions are used to enable a computer to execute the 3D object detection frame processing method of the first aspect or any corresponding embodiment thereof.

[0159] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A 3D object detection frame processing method, characterized in that: The method comprises: Obtaining a historical detection frame smoothing result of the target object in the previous point cloud image frame and a corresponding historical normalization factor, wherein the historical normalization factor is determined based on a target smoothing weight corresponding to the historical detection frame smoothing result; Get the current detection box data of the target object in the current point cloud image frame; Based on the historical detection frame smoothing result and the historical normalization factor, the current detection frame data is smoothed to obtain target detection frame data; A 3D object detection frame of the target object is determined based on the target detection frame data.

2. The method according to claim 1, characterized in that The step of smoothing the current detection frame data based on the historical detection frame smoothing result and the historical normalization factor to obtain the target detection frame data includes: Inputting the historical detection frame smoothing result and the historical normalization factor into a detection frame smoothing model to obtain a smoothing result; The target detection frame data is obtained based on the smoothing result.

3. The method according to claim 2, characterized in that The detection frame smoothing model is: Y i =(Y i-1 *smoothweight+X i ) Z i =Y i / norfactor i-1 normfactor i =1+normfactor i-1 *smoothweight Among them, i is the frame number of the point cloud image frame, Y -1 =0, normfactor -1 =1,Y i is the historical detection frame smoothing data corresponding to the i-th point cloud image frame, X i is the detection frame data corresponding to the i-th point cloud image frame, Z i is the smoothing result of the detection frame data corresponding to the i-th point cloud image frame, normfactor i is the normalization coefficient corresponding to the i-th point cloud image frame, smoothweight is the target smoothing weight, Y -1 =0, normfactor -1 =1.

4. The method according to claim 1, characterized in that: After obtaining the current detection frame data of the target object in the current point cloud image frame, the method further includes: Obtaining historical detection frame data of the target object; Determining whether the current detection frame data is singular data based on the historical detection frame data; When the current detection frame data is not singular data, performing a step of smoothing the current detection frame data based on the historical detection frame smoothing result and the historical normalization factor to obtain target detection frame data; When the current detection frame data is singular data, the current detection frame data is discarded, and the 3D object detection frame of the target object is determined based on the historical detection frame smoothing result.

5. The method according to claim 4, characterized in that The determining whether the current detection frame data is singular data based on the historical detection frame data includes: Determine calibration detection frame data based on the current detection frame data and the historical detection frame data; Calculating the data difference between the current detection frame data and the calibrated detection frame data; Based on the data difference, it is determined whether the current detection frame data is singular data.

6. The method according to claim 5, characterized in that The determining, based on the current detection frame data and the historical detection frame data, the calibration detection frame data comprises: Acquire a union of current contour data in the current detection frame data and historical contour data in the historical detection frame data to obtain a detection frame contour data set; Obtaining the contour statistical mean corresponding to each dimension of the contour points in the detection frame contour data set; Calculate the standard deviation of the contour points of each dimension in the current contour data and the corresponding contour statistical mean to obtain the contour standard deviation; The calibration detection frame data is determined based on the contour statistical mean and the contour standard deviation of the contour points in each dimension.

7. The method according to claim 6, characterized in that The determining of the calibration detection frame data based on the current detection frame data and the historical detection frame data further includes: Calculate the variance of each dimension of the contour points in the current contour data and the corresponding contour statistical mean to obtain the contour variance; Determine the contour covariance between the contour points of each dimension based on the contour variance; Constructing a profile covariance matrix based on the profile variance and the profile covariance; The calibration detection frame data is determined based on the vector between the contour statistical means and the contour covariance matrix.

8. A 3D object detection frame processing device, characterized in that: The device comprises: A data acquisition module, used to acquire a historical detection frame smoothing result of a target object in a previous point cloud image frame, and a corresponding historical normalization factor, wherein the historical normalization factor is determined based on a target smoothing weight corresponding to the historical detection frame smoothing result; A data processing module is used to obtain the current detection frame data of the target object in the current point cloud image frame; A smoothing processing module, configured to perform smoothing processing on the current detection frame data based on the historical detection frame smoothing result and the historical normalization factor to obtain target detection frame data; A detection frame determination module is used to determine a 3D object detection frame of the target object based on the target detection frame data.

9. A vehicle, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the 3D object detection frame processing method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the 3D object detection frame processing method according to any one of claims 1 to 7.