Task collaborative optimization method for monocular 3D object detection

By converting the directional rectangle into a Gaussian distribution, combined with technical means such as uncertainty estimator and Kalman filtering, the shortcomings of the existing three-dimensional object detection methods in boundary processing and uncertainty quantization are solved, and higher detection accuracy and robustness are achieved.

CN119942067APending Publication Date: 2025-05-06SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411958659.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing three-dimensional object detection method is difficult to effectively represent and deal with boundary problems when dealing with orientation rectangles, resulting in a decrease in detection accuracy and lack of quantification of uncertainty, affecting the robustness of the model.

Method used

A task collaborative optimization method is proposed. By converting the directional rectangle into a Gaussian distribution, the dimension of the 3D bounding box is reduced, and the uncertainty estimator and uncertainty weighted center loss function are introduced, combining Kalman filtering and the loss function in the form of IoU, multiple attributes are optimized and computational efficiency is improved.

Benefits of technology

This method effectively solves the boundary problem, improves the accuracy and robustness of three-dimensional object detection, significantly improves the 3D object detection performance on the KITTI benchmark, and shows high potential in real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942067A_ABST
    Figure CN119942067A_ABST
Patent Text Reader

Abstract

The invention provides a task collaborative optimization method (TCS) for monocular 3D object detection, which is used for improving the performance and the calculation efficiency of three-dimensional object detection. The method reduces the dimensionality of the 3D bounding box by selecting prediction branches related to bird's-eye view attributes and applies the same transformation to true values to generate corresponding tags, thereby simultaneously optimizing multiple attributes and ensuring computational efficiency. To deal with the boundary problem, the directional rectangle is converted into two-dimensional Gaussian distributions, and the distance between the two Gaussian distributions is reduced using an uncertainty weighted loss function. In addition, an uncertainty estimator is introduced to quantify the confidence of depth prediction, an overlapping region between two Gaussian distributions is determined through Kalman filtering, and further loss in an IoU form is calculated. The final total loss function combines uncertainty weighted center loss and loss in the form of IoU, the accuracy of three-dimensional target detection is improved, and the calculation efficiency is ensured through efficient loss function design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and digital image processing, and in particular to a task collaborative optimization method (TCS) for monocular 3D object detection, which is used to improve the performance and computational efficiency of 3D target detection. Background Art

[0002] In 3D object detection, existing methods usually adopt multiple loss functions for each detection head, but these methods often ignore the relationship between different prediction quantities. For example, attributes such as center point, size, and depth are usually optimized independently, which leads to the under-utilization of the correlation between the various attributes. In addition, it is still impractical to design a unified loss function for the overall bounding box in 3D space due to computational complexity, which may lead to inefficient back-propagation.

[0003] Existing methods have difficulty effectively representing and handling boundary issues when dealing with oriented rectangles, resulting in reduced detection accuracy. Specifically, traditional bounding box representation methods are prone to blurred boundaries and difficulty in accurately calculating overlapping areas when dealing with rotation and misalignment. These problems not only affect the accuracy of detection, but also increase the complexity of model training.

[0004] In addition, existing depth estimation methods usually lack the quantification of uncertainty, which makes the model less robust when dealing with complex scenes. Especially in applications such as autonomous driving and robotic navigation, accurate depth estimation is crucial for safety and reliability. However, existing methods often fail to provide reliable depth estimation confidence, which affects the performance of the overall system. Therefore, a new method is needed to overcome the above problems, improve the performance and computational efficiency of 3D object detection, and effectively handle boundary issues and uncertainties. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a task collaborative optimization strategy (TCS) for improving the performance and computational efficiency of 3D object detection. The method reduces the dimension of the 3D bounding box by selecting prediction branches related to bird's-eye view attributes (including center, size and depth), and applies the same transformation to the true value to generate the corresponding label, thereby optimizing multiple attributes simultaneously and ensuring computational efficiency.

[0006] The technical solution adopted by the present invention to achieve the above-mentioned purpose is:

[0007] A task collaborative optimization method for monocular 3D object detection comprises the following steps:

[0008] 1) Convert the 3D object image into a 2D oriented rectangle;

[0009] 2) Orient rectangle B2d (x,z,w,l,θ) is converted to a two-dimensional Gaussian distribution N(μ,Σ) to calculate the area of ​​the oriented rectangle;

[0010] 3) An uncertainty estimator is added to the deep branch of the depth estimation network to supervise the z-axis and x-axis of the oriented rectangle through the uncertainty estimator;

[0011] 4) Construct uncertainty weighted center loss function

[0012] 5) Construct a loss function in the form of IoU

[0013] 6) Using Lc and Calculate the total loss function and by minimizing Optimize the target detection task.

[0014] The step 2) is specifically as follows:

[0015] According to the oriented rectangle B 2d (x,z,w,l,θ), calculate the rotation matrix R and the eigenvalue diagonal matrix Λ:

[0016]

[0017] Where x represents the x-coordinate of the center of the rectangle, z represents the z-coordinate of the center of the rectangle, w represents the width of the rectangle, l represents the length of the rectangle, and θ represents the rotation angle of the rectangle relative to the reference direction;

[0018] Calculate the mean μ and covariance matrix Σ and:

[0019] μ=(x,z) T

[0020] Σ=RΛR T

[0021] Compute the area of ​​the oriented rectangle using the covariance matrix:

[0022]

[0023] in, represents the area corresponding to the Gaussian distribution, eig(Σ) represents the eigenvalue of the matrix Σ, and n represents the dimension.

[0024] The step 3) is specifically as follows:

[0025] Assuming that the depth prediction of each object follows the Laplace distribution La(η,λ), where η represents the predicted value of the depth regression, that is, the location parameter of the Laplace distribution, which defines the center of the distribution, and λ represents the scale parameter of the distribution, that is, the uncertainty of the prediction, which controls the width of the distribution, then the depth loss function of the x-axis and z-axis coordinates is and They are:

[0026]

[0027] Among them, σ x and σ z represents the standard deviation of the x-axis and z-axis depth predictions, p x and p z Represents the x-axis and z-axis depth values ​​predicted by the network, x gt and z gt They represent the true values ​​of the x-axis and z-axis, that is, the true depth coordinates.

[0028] The step 4) is specifically as follows:

[0029]

[0030] Among them, σ d and σ x are uncertainty weights, corresponding to the confidence of depth and x-axis coordinates, smooth A smooth loss function.

[0031] The step 5) is specifically as follows:

[0032] Using the Kalman filter to determine the covariance Σ between the overlapping regions of two Gaussian distributions overlap :

[0033] Σ overlap =Σ p -Σ p (Σ p +Σ t ) -1 Σ p

[0034] Among them, Σ p is the covariance matrix of the predicted box, Σ t is the covariance matrix of the true value; KFIoU is calculated based on the area of ​​the oriented rectangle:

[0035]

[0036] in, is the area of ​​the overlapping region, is the area of ​​the prediction box, The area of ​​the truth box.

[0037] Simulate the loss function of IoU form for:

[0038]

[0039] The total loss function for:

[0040]

[0041] The present invention has the following beneficial effects and advantages:

[0042] The present invention proposes a task collaborative optimization method (TCS) for monocular 3D object detection, which solves the boundary problem by converting the oriented rectangle into a Gaussian distribution, and improves the accuracy and robustness of 3D target detection through the uncertainty weighted center loss function and the Kalman filter and the loss function in the form of IoU. First, the method avoids the boundary blur problem in the traditional method by converting the oriented rectangle into a Gaussian distribution, and improves the detection accuracy. Secondly, an uncertainty estimator is introduced to quantify the confidence of the depth prediction, and the processing capability of low-confidence samples is enhanced through the uncertainty weighted center loss function, which improves the robustness of the model. In addition, the processing of overlapping areas is further optimized through the Kalman filter and the loss function in the form of IoU, ensuring the high performance of the model in complex scenes. Experiments on the KITTI benchmark show that the present invention can effectively improve the performance of 3D target detectors and provide a more reliable and efficient solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic diagram of the algorithm structure of the present invention. DETAILED DESCRIPTION

[0044] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] The proposed method mainly consists of two main parts: Gaussian distribution transformation and uncertainty weighted center loss. Through these methods, we aim to improve the accuracy and computational efficiency of 3D object detection. Specifically, we transform the oriented rectangle B 2d(x,z,w,l,θ) is converted into a two-dimensional Gaussian distribution N(μ,∑), and multiple attributes are optimized through uncertainty weighted center loss and Kalman filter overlap region calculation while ensuring computational efficiency.

[0046] like Figure 1 As shown, a task collaborative optimization strategy for 3D object detection includes the following steps:

[0047] Step 1: Oriented rectangle conversion: To avoid boundary problems, the oriented rectangle B 2d (x,z,w,l,θ) is transformed into a two-dimensional Gaussian distribution N(μ,∑), and the rotation matrix R and the eigenvalue diagonal matrix Λ are defined as follows:

[0048]

[0049] Where x: x-coordinate of the center of the rectangle. z: z-coordinate of the center of the rectangle. w: width of the rectangle. l: length of the rectangle. θ: rotation angle of the rectangle (relative to the reference direction).

[0050] The covariance matrix Σ can be expressed as:

[0051] Σ=RΛR T ,μ=(x,z) T

[0052] Where R represents the rotation matrix and Λ represents the diagonal matrix.

[0053] The area of ​​a directed rectangle can be easily calculated using the covariance matrix:

[0054]

[0055] in, represents the area corresponding to the Gaussian distribution, eig(∑) represents the eigenvalue of the matrix ∑, and n represents the dimension, in this example n=2.

[0056] Step 2: Depth prediction and uncertainty estimation: An uncertainty estimator is added to the depth branch to quantify the network's confidence in the depth prediction (i.e., the coordinates of the object on the z-axis). Assume that the depth prediction of each object follows the Laplace distribution La(μ,λ), where μ represents the depth regression and λ is the uncertainty predicted together with μ. To simplify the process, the standard deviation σ is used instead, which can be expressed as The depth loss function is defined as:

[0057]

[0058] Among them, σ z represents the standard deviation of depth prediction, p z Represents the depth value predicted by the network, zgt Represents the ground truth of the z-axis, i.e. the actual depth coordinate. log(σ z ) is a regularization term, which is used to penalize predictions with excessive uncertainty and ensure that the model does not rely too much on predictions with high uncertainty.

[0059] The x-axis coordinates are handled similarly:

[0060]

[0061] After predicting the depth and its corresponding uncertainty, it is recommended to calculate the uncertainty weighted center loss Lc to align the center points of the two Gaussian distributions. Compared with directly using the ordinary smoothL1 regression center point coordinates, the addition of soft weight terms allows more weights to be assigned to those branches with lower confidence, thereby improving the robustness to possible inaccuracies.

[0062] Step 3: Uncertainty weighted center loss function: In order to accurately represent the overlap of two Gaussian distributions and simulate the calculation of intersection over union (IoU), an uncertainty weighted center loss function Lc is proposed to minimize the distance between the center points. The specific formula is:

[0063]

[0064] Among them, σ d and σ x are uncertainty weights, corresponding to the confidence of the depth and x-axis coordinate respectively.

[0065] Step 4: Kalman filter and IoU loss: Inspired by KFIOU, the Kalman filter is used to determine the covariance Σ of the overlapping area between two Gaussian distributions overlap :

[0066] Σ overlap =Σ p -Σ p (Σ p +Σ t ) -1 Σ p

[0067] Among them, Σ p is the covariance matrix of the predicted box, Σ t is the covariance matrix of the true value. According to formula (7), the corresponding area can be obtained, and then KFIoU can be calculated:

[0068]

[0069] in, is the area of ​​the overlapping region, is the area of ​​the prediction box, The area of ​​the truth box.

[0070] Simulate the loss function of IoU form It can be expressed as:

[0071]

[0072] Step 5: Total loss function: The final total loss function of the collaborative optimization method based on Gaussian distribution and uncertainty weight is:

[0073]

[0074] Through the above technical solution, the method of the present invention can comprehensively consider the relationship between different prediction quantities, improve the accuracy of 3D object detection, and ensure the computational efficiency through efficient loss function design. Experimental results show that the method of the present invention significantly improves AP on the KITTI dataset. 3D and AP BEV Indicators have proven its effectiveness and superiority.

[0075] In terms of experimental settings, we used the standard ResNet-50 as the base network and adopted the AdamW optimizer with a weight decay of 10. -4 The batch size is 8 and the training epochs are over 150. The learning rate is reduced to 0.1 at epoch 65 and epoch 100 respectively. The number of learnable queries N is set to 50. With this configuration, we train on a single RTX 2080Ti GPU.

[0076] In terms of performance evaluation, our method has AP 3D The indicators were improved by 2.03 / 2.28 / 1.98 respectively. In addition, AP BEV The AOS and AOS metrics show gains of 2.43 / 0.93 / 1.38 and 0.5 / 1.36 / 1.36, respectively. These results show that our method not only improves the accuracy of 3D object detection, but also shows better robustness under different difficulty settings.

[0077] In terms of running time analysis, we process each image in 40 milliseconds on a single RTX 2080Ti GPU with a batch size of 1. Compared with CaDDN and DDMP3D, our method is 15 times and 4.5 times faster, respectively. This shows that our method has high real-time processing capabilities while maintaining high accuracy, and is suitable for real-time detection needs in practical applications.

[0078] To verify the statistical significance of our method, we performed Friedman test and Nemenyi follow-up test. The results show that our method achieved the highest average ranking among all available methods. This further proves the superiority of our method in the 3D object detection task.

[0079] In summary, by converting oriented rectangles into Gaussian distributions and combining uncertainty weighted center loss and Kalman filter-based loss, our task collaborative optimization strategy (TCS) significantly improves the performance of 3D object detection while maintaining computational efficiency. Experimental results show that this method outperforms existing image-only methods in multiple metrics and has high potential for real-time applications.

Claims

1. A task collaborative optimization method for monocular 3D object detection, characterized in that: The following steps are involved: 1) Convert the 3D object image into a 2D oriented rectangle; 2) Orient rectangle B 2d (x,z,w,l,θ) is converted to a two-dimensional Gaussian distribution N(μ,Σ) to calculate the area of ​​the oriented rectangle; 3) An uncertainty estimator is added to the deep branch of the depth estimation network to supervise the z-axis and x-axis of the oriented rectangle through the uncertainty estimator; 4) Construct uncertainty weighted center loss function 5) Construct a loss function in the form of IoU 6) Using Lc and Calculate the total loss function and by minimizing Optimize the target detection task.

2. The task collaborative optimization method for monocular 3D object detection according to claim 1, characterized in that: The step 2) is specifically as follows: According to the oriented rectangle B 2d (x,z,w,l,θ), calculate the rotation matrix R and the eigenvalue diagonal matrix Λ: Where x represents the x-coordinate of the center of the rectangle, z represents the z-coordinate of the center of the rectangle, w represents the width of the rectangle, l represents the length of the rectangle, and θ represents the rotation angle of the rectangle relative to the reference direction; Calculate the mean μ and covariance matrix Σ and: μ=(x,z) T S = RΛR T Compute the area of ​​the oriented rectangle using the covariance matrix: in, represents the area corresponding to the Gaussian distribution, eig(Σ) represents the eigenvalue of the matrix Σ, and n represents the dimension.

3. The task collaborative optimization method for monocular 3D object detection according to claim 1, characterized in that: The step 3) is specifically as follows: Assuming that the depth prediction of each object follows the Laplace distribution La(η,λ), where η represents the predicted value of the depth regression, that is, the location parameter of the Laplace distribution, which defines the center of the distribution, and λ represents the scale parameter of the distribution, that is, the uncertainty of the prediction, which controls the width of the distribution, then the depth loss function of the x-axis and z-axis coordinates is and They are: Among them, σ x and σ z represents the standard deviation of the x-axis and z-axis depth predictions, p x and p z Represents the x-axis and z-axis depth values ​​predicted by the network, x gt and z gt They represent the true values ​​of the x-axis and z-axis, that is, the true depth coordinates.

4. The task collaborative optimization method for monocular 3D object detection according to claim 1, characterized in that: The step 4) is specifically as follows: Among them, σ d and σ x are uncertainty weights, corresponding to the confidence of the depth and x-axis coordinates, respectively, A smooth loss function.

5. The task collaborative optimization method for monocular 3D object detection according to claim 1, characterized in that: The step 5) is specifically as follows: The Kalman filter is used to determine the covariance Σoverlap of the overlap region between two Gaussian distributions: Overlap = S p -S p (S p +S t ) -1 S p Among them, Σ p is the covariance matrix of the predicted box, Σ t is the covariance matrix of the true value; KFIoU is calculated based on the area of ​​the oriented rectangle: in, is the area of ​​the overlapping region, is the area of ​​the prediction box, The area of ​​the truth box. Simulate the loss function of IoU form for:

6. The task collaborative optimization method for monocular 3D object detection according to claim 1, characterized in that: The total loss function for:

Citation Information

Patent Citations

  • Loss function optimization method and system for improving tiny target detection performance

    CN119068171A