Three-dimensional target detection method based on multistage fine tuning

By building a multi-level fine-tuned three-dimensional target detection network and combining target enhancement technology, the problem of complex data acquisition and low multi-objective detection accuracy in autonomous driving is solved, and the accuracy of multi-objective detection in traffic scenarios is achieved, which improves the safety of autonomous driving.

CN120260010APending Publication Date: 2025-07-04江苏源驶科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510312179.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing three-dimensional object detection technology has complex and expensive data acquisition in autonomous driving. The three-dimensional object detection model cannot adapt to multi-target scenarios, resulting in a decrease in detection accuracy and the inability to achieve accurate detection of multi-targets in traffic scenarios, reducing the safety of autonomous driving.

Method used

By building a multi-level fine-tuning three-dimensional object detection network, including point cloud voxelization, feature extraction, fusion and multi-level fine-tuning modules, combined with target enhancement technology, the effectiveness and detection accuracy of the data set are improved, and multi-objective detection is achieved.

Benefits of technology

The safety of autonomous driving is improved, background noise is reduced through full-scene-level enhancement, and target detection on images of different quality, so as to achieve accurate detection of multiple targets in traffic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260010A_ABST
    Figure CN120260010A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional target detection method based on multistage fine tuning. The method comprises the steps of obtaining point cloud data from a public data set KITTI and performing target enhancement on the point cloud data; constructing a three-dimensional target detection network formed by connecting a point cloud voxelization module, a voxel feature extraction module, a voxel feature fusion module, a candidate box generation module and a multi-stage fine tuning module in series; target enhanced point cloud data is input into a point cloud voxelization module for voxel division, the point cloud data subjected to voxel division is input into a voxel feature extraction module for voxel feature extraction, all voxel features are input into a voxel feature fusion module for feature fusion, point cloud features are obtained, and the point cloud feature fusion module is used for performing voxel feature extraction. A preliminary target detection frame is generated in the input candidate frame generation module; and inputting the point cloud features and the initial target detection frame into a multi-stage fine tuning module together, and determining a final target detection frame, thereby realizing three-dimensional target detection, realizing accurate detection of multiple targets in a traffic scene, and improving the safety of automatic driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection, and specifically, to a three-dimensional target detection method based on multi-level fine-tuning. Background Art

[0002] With the development of society and the rapid economic growth, the quality of people's life has been continuously improved, and at the same time, people have entered the intelligent era. As a means of transportation for travel, cars have become an inalienable part of people's daily lives. With the continuous emergence of intelligent products and the various conveniences they bring, people's need for the automatic driving function of cars is also increasing.

[0003] The automatic driving perception task aims to obtain and understand various sensor data from the environment to perceive and understand the surrounding environment. The automatic driving perception task is crucial for the safety and reliability of the automatic driving system, and it is necessary to capture key information about roads, obstacles, traffic signs, pedestrians, etc. for the decision-making and planning of automatic driving. Three-dimensional target detection is an important part of realizing automatic driving. Although significant progress has been made in three-dimensional target detection in the field of computer vision, there are still some problems:

[0004] (1) Compared with two-dimensional images, obtaining and annotating three-dimensional target data is more complex and expensive. Professional sensors are required to obtain three-dimensional data, and accurate annotation is required to train the model, which is difficult to apply to automatic driving involving large-scale data sets. At the same time, current three-dimensional target detection technologies enhance the scene-level data set, which will enhance irrelevant backgrounds, reduce the effectiveness of the enhanced data, and cannot fully provide feature information such as the shape, size, and posture of each target in the traffic scene.

[0005] (2) Current three-dimensional target detection models are constructed with only a single detection head and can only regress the detection box once. However, in the traffic scene of automatic driving, there are multiple targets, including vehicles, pedestrians, roads, signs, and lane lines. For different targets, the sizes of their candidate boxes are different, and the positioning ability of a single detection head is very limited, resulting in a decrease in target capture accuracy and the inability to accurately detect multiple targets in the traffic scene, thus reducing the safety of automatic driving. Summary of the Invention

[0006] Aiming at the problems existing in the prior art, the present invention provides a three-dimensional target detection method based on multi-level fine-tuning, which realizes target enhancement through a target enhancer to improve the effectiveness of the data set. At the same time, by constructing a three-dimensional target detection network with multi-level fine-tuning, the secondary regression ability of target detection is improved, so as to realize the accurate detection of multiple targets in the traffic scene and improve the safety of automatic driving.

[0007] To achieve the above technical objectives, the present invention adopts the following technical solutions: A three-dimensional object detection method based on multi-level fine-tuning, specifically including the following steps:

[0008] Step S1: Obtain point cloud data and annotation information from the public dataset KITTI, and perform object enhancement on the point cloud data to obtain object-enhanced point cloud data;

[0009] Step S2: Construct a three-dimensional object detection network connected in series by a point cloud voxelization module, a voxel feature extraction module, a voxel feature fusion module, a candidate box generation module, and a multi-level fine-tuning module;

[0010] Step S3: Input the object-enhanced point cloud data into the point cloud voxelization module for voxel division, input the voxelized point cloud data into the voxel feature extraction module for voxel feature extraction, input all voxel features into the voxel feature fusion module for feature fusion to obtain point cloud features, and input them into the candidate box generation module to generate preliminary object detection boxes; Input the point cloud features and the preliminary object detection boxes into the multi-level fine-tuning module to determine the final object detection boxes;

[0011] Step S4: Repeat Step S3 with the object-enhanced point cloud data in sequence until the loss function of the three-dimensional object detection network converges, and complete the training of the three-dimensional object detection network;

[0012] Step S5: Obtain the point cloud data in the traffic scene, input it into the trained three-dimensional object detection network, and predict the three-dimensional object through the position where the final object detection box is located.

[0013] Furthermore, the objects in the point cloud data in Step S1 include: vehicles, pedestrians, roads, signs, and lane lines.

[0014] Furthermore, the specific process of performing object enhancement on the point cloud data in Step S1 is: Perform global scene-level enhancement on the point cloud data to obtain globally enhanced point cloud data. For each frame in the globally enhanced point cloud data, select the object for downsampling and filling, add noise sampled from the normal distribution, and then input it into the enhancer for enhancement, and output the point-by-point displacement not exceeding the corresponding object bounding box to obtain the object-enhanced point cloud data.

[0015] Furthermore, the point cloud feature extraction module includes four cascaded feature encoding modules. Among them, the first feature encoding module is composed of two cascaded sub-manifold sparse 3D convolutions, the second and the third feature encoding modules are both composed of one sparse 3D convolution and two cascaded sub-manifold sparse 3D convolutions in sequence, and the fourth feature encoding module is composed of one sparse 3D convolution and one cascaded sub-manifold sparse 3D convolution; the voxelized point cloud data is input into the first feature encoding module to extract the first voxel feature, the extracted first voxel feature is input into the second feature encoding module to extract the second voxel feature, the extracted second voxel feature is input into the third feature encoding module to extract the third voxel feature, and the third voxel feature is input into the fourth feature encoding module to extract the fourth voxel feature.

[0016] Furthermore, the specific process of the candidate box generation module for generating the preliminary object detection box is as follows:

[0017] i. For each object in the point cloud features, initialize the object detection box, generate several candidate anchors for each object, initialize the candidate anchors to the average size of the corresponding object, and set the directions of the candidate anchors to 0° or 90° respectively;

[0018] ii. If the intersection over union of the candidate anchor of the object and the object detection box is greater than the threshold, take the candidate anchor as the candidate box of the corresponding object, and adjust the position of the candidate box through the position offset between the candidate anchor and the object detection box, and take the candidate box with the adjusted position as the preliminary object detection box; otherwise, take the object detection box as the preliminary object detection box.

[0019] Furthermore, the process of adjusting the position of the candidate box is as follows:

[0020]

[0021] Among them, w g represents the width of the object detection box, w a represents the width of the corresponding object candidate box, w' a represents the width of the preliminary object detection box; l g represents the length of the object detection box, l a represents the length of the corresponding object candidate box, l' a represents the length of the preliminary object detection box; h g represents the height of the object detection box, h a represents the height of the corresponding object candidate box, h' a represents the height of the preliminary object detection box.

[0022] Further, the multi-stage fine-tuning module includes: a first-stage object detection box adjustment module, a second-stage object detection box adjustment module, a third-stage object detection box adjustment module, and an object detection box voting module. The first-stage object detection box adjustment module, the second-stage object detection box adjustment module, and the third-stage object detection box adjustment module are respectively used for the preliminary object detection box adjustment of the point cloud features with gradually increasing quality, and each includes: an integrity estimation sample reweighting module and a dual-branch detection head module. The integrity estimation sample reweighting module is used to calculate the weight of each preliminary object detection box, and the dual-branch detection head module is composed of a fully connected layer detection head and a convolutional layer detection head, and is used to calculate the confidence of the preliminary object detection box by combining the weight of the preliminary object detection box; the object detection box voting module is used to perform weighted averaging on the confidences of the preliminary object detection boxes output by the first-stage object detection box adjustment module, the second-stage object detection box adjustment module, and the third-stage object detection box adjustment module to obtain the final object detection box.

[0023] Further, the specific process of the integrity estimation sample reweighting module for calculating the weight of each preliminary object detection box is as follows: taking the intersection over union of the three-dimensional ground truth box of each object and the minimum bounding box of the corresponding object as the point integrity score of the object, and calculating the weight of the preliminary object detection box of the object:

[0024]

[0025] where represents the normalization factor in the b-th iteration, M represents the number of objects in the target, s represents the index of M, represents the point integrity score of the s-th object in the b-th iteration, represents the point integrity score of the m-th object in the b-th iteration.

[0026] Further, the loss function L total of the three-dimensional object detection network is specifically:

[0027] L total = L rpn + L pr

[0028] where L rpn represents the loss function of the candidate box generation module, N a represents the number of candidate anchors, i represents the index of N a , represents the classification loss of the i-th candidate anchor, p t,i represents the foreground score predicted by the i-th candidate anchor, and α tLet \(\alpha\) represent the first parameter with a value of 0.25; \(\gamma\) represent the second parameter with a value of 2; \(\tau(f\geq1)\) represents the regression loss only for foreground anchors. Represents the direction regression loss of the \(i\)-th candidate anchor. \(\theta_{i}^{r}\) p Represents the predicted direction of the \(i\)-th candidate anchor, \(\theta_{i}^{p}\). t Represents the yaw angle between the \(i\)-th candidate anchor and the corresponding target detection box, and SmoothL1 represents the SmoothL1 loss function. Represents the regression loss of the \(i\)-th candidate anchor. Represents the direction classification loss of the \(i\)-th candidate anchor; \(L_{i}^{d}\). pr Represents the loss function of the multi-level fine-tuning module. Represents the confidence loss of the \(j\)-th target detection box adjustment module, \(\omega_{j}^{c}\). fc Represents the weight of Represents the regression loss of the \(j\)-th target detection box adjustment module, \(\omega_{j}^{r}\). conv Represents the weight of Represents the target box orientation loss function of the \(j\)-th target detection box adjustment module.

[0029] Compared with the prior art, the present invention has the following beneficial effects: The three-dimensional object detection method based on multi-level fine-tuning of the present invention performs full-scene enhancement on point cloud data, which can reduce background noise, thereby making the enhancement effect of the objects in the image more accurate; at the same time, the present invention constructs a three-dimensional object detection network with multi-level fine-tuning, inputs the point cloud features and the preliminary object detection boxes into the multi-level fine-tuning module together to determine the final object detection box. The multi-level fine-tuning module can adapt to object detection on images of different qualities, thereby realizing accurate detection of multiple objects in the traffic scene and improving the safety of autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Is the flowchart of the three-dimensional object detection method based on multi-level fine-tuning of the present invention;

[0031] Figure 2 Is the flowchart of object enhancement in the present invention;

[0032] Figure 3 Is the structural schematic diagram of the three-dimensional object detection network in the present invention;

[0033] Figure 4 Is the structural schematic diagram of the multi-level fine-tuning detection module in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The technical solutions of the present invention will be further explained below with reference to the drawings.

[0035] As Figure 1 This is a flowchart of the 3D object detection method based on multi-level fine-tuning of the present invention. The 3D object detection method specifically includes the following steps:

[0036] Step S1: Obtain point cloud data and annotation information from the publicly available KITTI dataset, and perform target enhancement on the point cloud data to obtain target-enhanced point cloud data. As Figure 2 , the specific process is as follows: Perform global scene-level enhancement on the point cloud data to obtain globally enhanced point cloud data. For each frame in the globally enhanced point cloud data, select the target for downsampling and filling, add noise sampled from a normal distribution, and then input it into the enhancer for enhancement. Output the point-by-point displacement that does not exceed the corresponding target bounding box to obtain target-enhanced point cloud data. By performing full-scene-level enhancement on the point cloud data, while creating new data, not only is background noise reduced, but also the shape of the target does not change much, thus making the enhancement effect of the target in the image more accurate.

[0037] The KITTI dataset of the present invention is a publicly available dataset widely used in autonomous driving and computer vision research. It was jointly created by the Karlsruhe Institute of Technology in Germany and the Max Planck Society in Germany and is based on sensors on vehicles, including multiple cameras and lidar. Among them, the image data includes various types of images, including grayscale images, color images, and depth images, which provide scene views captured from different positions and angles of the vehicle; the point cloud data includes point cloud data collected by lidar, which provides accurate distance and shape information about the objects in the scene. The KITTI dataset also provides annotations for annotating the image and point cloud data, including bounding boxes and semantic labels for vehicles, pedestrians, roads, signs, lane lines, etc.; in addition, other annotations such as vehicle pose, optical flow, disparity, and depth estimation are also provided. The dataset is divided into 7481 RGB images and the corresponding point cloud data in the scene, and 7518 images and the corresponding point cloud data in the scene.

[0038] Step S2: Construct a 3D object detection network connected in series by a point cloud voxelization module, a voxel feature extraction module, a voxel feature fusion module, a candidate box generation module, and a multi-level fine-tuning module. As Figure 3 shown, this multi-level fine-tuning module can adapt to object detection on images of different qualities, thus adapting to the accurate detection of multiple objects in traffic scenarios and improving the safety of autonomous driving.

[0039] Step S3: Input the point cloud data to be enhanced into the point cloud voxelization module for voxel division. Input the voxelized point cloud data into the voxel feature extraction module for voxel feature extraction, including the spatial geometric features and semantic features of the voxels. Input all voxel features into the voxel feature fusion module for feature fusion to obtain point cloud features, and input them into the candidate box generation module to generate preliminary object detection boxes. Input the point cloud features and the preliminary object detection boxes into the multi-level fine-tuning module to determine the final object detection boxes.

[0040] The point cloud voxelization module clips and retains the range of [0, 70.4] m on the X-axis, [-40, 40] m on the Y-axis, and [-3, 1] m on the Z-axis for each frame of point cloud data, sets the size of each voxel to [0.05, 0.05, 0.05] m, and divides a number of voxels.

[0041] The point cloud feature extraction module contains four cascaded feature encoding modules. Among them, the first feature encoding module is composed of two cascaded sub-manifold sparse 3D convolutions, the second and third feature encoding modules are both composed of one sparse 3D convolution and two cascaded sub-manifold sparse 3D convolutions in sequence, and the fourth feature encoding module is composed of one sparse 3D convolution and one cascaded sub-manifold sparse 3D convolution. Input the voxelized point cloud data into the first feature encoding module to extract the first voxel feature, input the extracted first voxel feature into the second feature encoding module to extract the second voxel feature, input the extracted second voxel feature into the third feature encoding module to extract the third voxel feature, and input the third voxel feature into the fourth feature encoding module to extract the fourth voxel feature.

[0042] The candidate box generation module adjusts the position and scale of the candidate boxes by applying position offsets according to the predicted object scores, filters the anchors, and selects some anchor boxes with the highest scores as candidate boxes. At the same time, adjusts the position and scale of the candidate boxes by applying position offsets to better match the real object boxes. At the same time, after non-maximum suppression, redundant candidate boxes are eliminated to select the final detection boxes. Non-maximum suppression will retain the most relevant boxes and suppress other boxes according to the overlap degree and object scores between the candidate boxes. Specifically, the specific process of the candidate box generation module generating preliminary object detection boxes is as follows:

[0043] i. For each object in the point cloud features, initialize the object detection box, generate a number of candidate anchors for each object, initialize the candidate anchors to the average size of the corresponding object, and set the directions of the candidate anchors to 0° or 90° respectively;

[0044] ii. If the intersection over union (IoU) between the candidate anchor of the target and the target detection box is greater than the threshold, use the candidate anchor as the candidate box for the corresponding target, and adjust the position of the candidate box according to the position offset between the candidate anchor and the target detection box. Then, use the candidate box with the adjusted position as the preliminary target detection box; otherwise, use the target detection box as the preliminary target detection box. If the IoU between the anchor of the vehicle and the target box is greater than 0.6, assign the anchor point of the vehicle to the target box. If the IoU is less than 0.45, consider it as a background anchor point; for pedestrians and cyclists, the foreground target matching threshold is 0.5, and the background matching threshold is 0.35.

[0045] The process of adjusting the position of the candidate box is as follows:

[0046]

[0047] where w g represents the width of the target detection box, w a represents the width of the corresponding target candidate box, w' a represents the width of the preliminary target detection box; l g represents the length of the target detection box, l a represents the length of the corresponding target candidate box, l' a represents the length of the preliminary target detection box; h g represents the height of the target detection box, h a represents the height of the corresponding target candidate box, h' a represents the height of the preliminary target detection box.

[0048] For example Figure 4 , the multi-stage fine-tuning module alleviates the impact of different-quality samples on the performance of the 3D object detection network and improves the detection accuracy. It includes: the first-stage target detection box adjustment module, the second-stage target detection box adjustment module, the third-stage target detection box adjustment module, and the target detection box voting module. The first-stage target detection box adjustment module, the second-stage target detection box adjustment module, and the third-stage target detection box adjustment module are respectively used to adjust the preliminary target detection boxes of the point cloud features with gradually increasing quality. Each of them includes: the integrity estimation sample reweighting module and the dual-branch detection head module. The integrity estimation sample reweighting module is used to calculate the weight of each preliminary target detection box. The dual-branch detection head module consists of a fully connected layer detection head and a convolutional layer detection head, and is used to calculate the confidence of the preliminary target detection box by combining the weight of the preliminary target detection box; the target detection box voting module is used to perform weighted averaging on the confidences of the preliminary target detection boxes output by the first-stage target detection box adjustment module, the second-stage target detection box adjustment module, and the third-stage target detection box adjustment module to obtain the final target detection box.

[0049] If the detection head simply adds the confidence levels and regression losses of the two branches directly, it will ignore the sample quality problem, inevitably resulting in a larger loss for a large number of low-quality samples, while high-quality samples receive no attention. Therefore, sample reweighting is performed to reduce the weights of low-quality samples to reshape the overall training objective, thereby focusing the training on high-quality samples to make the loss distribution more balanced. For each candidate box at each stage, the corresponding weight of the positive sample is calculated using the score obtained from the integrity estimation sample reweighting module. The specific process of the integrity estimation sample reweighting module for calculating the weight of each preliminary object detection box is as follows: The intersection over union of the three-dimensional ground truth box of each object and the minimum bounding box of the corresponding object is used as the point integrity score of the object, and the weight of the preliminary object detection box of the object is calculated:

[0050]

[0051] where, represents the normalization factor in the b-th iteration, M represents the number of objects in the target, s represents the index of M, represents the point integrity score of the s-th object in the b-th iteration, represents the point integrity score of the m-th object in the b-th iteration.

[0052] Since the fully connected layer detection head is more suitable for classification tasks, while the convolutional layer detection head is more suitable for object localization, therefore, different from the previous method using a fully connected single-branch structure, the detection head of the multi-stage fine-tuning module of the present invention adopts a dual-branch detection head module, using a fully connected layer structure for classification tasks and a convolutional layer structure for bounding box regression tasks. Since the two branches have a complementary effect, the final classification result at each stage is obtained by fusing the prediction results of the fully connected branch and the convolutional branch S = S fc + S conv (1 - S fc ), and the result of bounding box regression is taken from the prediction result of the convolutional branch. Among them, S fc is the classification score output by the fully connected layer branch, and S conv is the classification score output by the convolutional branch.

[0053] Object detection box voting module. Three-dimensional object detection usually requires predicting the object height and rotation angle, resulting in weak and strong predictions being output at each stage. At the same time, some noise errors will propagate to the downstream multi-stage framework, causing the detection result to deteriorate. In order to obtain the best detection result, during inference, some low-confidence boxes containing additional information are retained, and a weighted bounding box voting strategy is adopted. The confidence levels S output by N r stages are averaged to obtain the average confidence level And use the confidence to perform weighted voting on the candidate boxes to obtain the result Through weighted voting, a higher-quality target box can be obtained.

[0054] Step S4: Repeatedly perform Step S3 on the target-enhanced point cloud data in sequence until the loss function of the 3D object detection network converges, and complete the training of the 3D object detection network;

[0055] The loss function L of the 3D object detection network in the present invention total Specifically:

[0056] L total = L rpn + L pr

[0057] Among them, L rpn represents the loss function of the candidate box generation module, N a represents the number of candidate anchors, i represents the index of N, represents the classification loss of the i-th candidate anchor, p t,i represents the foreground score predicted by the i-th candidate anchor, α t represents the first parameter, with a value of 0.25; γ represents the second parameter, with a value of 2; τ(f≥1) represents performing regression loss only on foreground anchor points; represents the direction regression loss of the i-th candidate anchor, θ p represents the predicted direction of the i-th candidate anchor, θ t represents the yaw angle between the i-th candidate anchor and the corresponding object detection box, SmoothL1 represents the loss function of Smooth L1, and is used to calculate the angle regression loss; represents the regression loss of the i-th candidate anchor, represents the direction classification loss of the i-th candidate anchor; L pr represents the loss function of the multi-level fine-tuning module, represents the confidence loss of the j-th object detection box adjustment module, ω fc represents 's weight, represents the regression loss of the j-th object detection box adjustment module, ω conv represents 's weight, represents the target box orientation loss function of the j-th object detection box adjustment module.

[0058] Step S5: Obtain the point cloud data in the traffic scene, input it into the trained 3D object detection network, and predict the 3D object through the position where the final object detection box is located.

[0059] The experiment of the 3D object detection method of the present invention was carried out on the ubuntu16.04 system. The training of the 3D object detection network used the pytorch1.6 deep learning framework and 2 NVIDIA 2080Ti graphics cards with a total video memory of 22GB. During the training and inference processes, two preprocessed virtual point cloud scenes were fed into the network each time for forward propagation. To better extract features for the backbone network, it was then input into the candidate box generation module to generate candidate boxes P = {x, y, z, l, w, h, 0}, where x, y, and z represent the center point coordinate values of the candidate boxes, l, w, and h represent the length, width, and height of each candidate box respectively, and 0 represents the rotation angle value of each candidate box. Then, the last two layers of the backbone network were used to correct the candidate boxes in the multi-level fine-tuning module, and finally, 3D object detection boxes were obtained, with more accurate detection results.

[0060] The above are only the preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.

Claims

1. A three-dimensional object detection method based on multi-level fine-tuning, characterized in that, Specifically, it includes the following steps: Step S1: Obtain point cloud data and annotation information from the publicly available KITTI dataset, and perform target enhancement on the point cloud data to obtain target-enhanced point cloud data; Step S2: Construct a three-dimensional object detection network composed of a point cloud voxelization module, a voxel feature extraction module, a voxel feature fusion module, a candidate box generation module, and a multi-level fine-tuning module connected in series; Step S3: Input the target-enhanced point cloud data into the point cloud voxelization module for voxel division, input the voxel-divided point cloud data into the voxel feature extraction module for voxel feature extraction, input all voxel features into the voxel feature fusion module for feature fusion to obtain point cloud features, and input them into the candidate box generation module to generate preliminary object detection boxes; Input the point cloud features and the preliminary object detection boxes into the multi-level fine-tuning module together to determine the final object detection boxes; Step S4: Repeat Step S3 for the target-enhanced point cloud data in sequence until the loss function of the three-dimensional object detection network converges, and complete the training of the three-dimensional object detection network; Step S5: Obtain the point cloud data in the traffic scene, input it into the trained three-dimensional object detection network, and predict the three-dimensional object through the position where the final object detection box is located.

2. The three-dimensional object detection method based on multi-level fine-tuning according to claim 1, wherein The targets in the point cloud data in Step S1 include: vehicles, pedestrians, roads, signs, and lane lines.

3. The three-dimensional object detection method based on multi-level fine-tuning according to claim 2, wherein, The specific process of performing target enhancement on the point cloud data in Step S1 is as follows: Perform global scene-level enhancement on the point cloud data to obtain globally enhanced point cloud data. For each frame in the globally enhanced point cloud data, select the target for downsampling and filling, add noise sampled from a normal distribution, and then input it into the enhancer for enhancement, and output the point-by-point displacement not exceeding the corresponding target bounding box to obtain the target-enhanced point cloud data.

4. The three-dimensional object detection method based on multi-level fine-tuning according to claim 3, wherein The point cloud feature extraction module contains four feature encoding modules connected in series. Among them, the first feature encoding module is composed of 2 sub-manifold sparse 3D convolutions connected in series, the second feature encoding module and the third feature encoding module are both composed of 1 sparse 3D convolution and 2 sub-manifold sparse 3D convolutions connected in series in sequence, and the fourth feature encoding module is composed of 1 sparse 3D convolution and 1 sub-manifold sparse 3D convolution connected in series; Input the voxel-divided point cloud data into the first feature encoding module to extract the first voxel feature, input the extracted first voxel feature into the second feature encoding module to extract the second voxel feature, input the extracted second voxel feature into the third feature encoding module to extract the third voxel feature, and input the third voxel feature into the fourth feature encoding module to extract the fourth voxel feature.

5. A three-dimensional object detection method based on multi-level fine-tuning according to claim 4, characterized in that The specific process of the candidate box generation module for generating preliminary object detection boxes is as follows: i. For each target in the point cloud features, initialize the object detection box, generate several candidate anchors for each target, initialize the candidate anchors to the average size of the corresponding target, and set the directions of the candidate anchors to 0° or 90° respectively; ii. If the intersection over union of the candidate anchor of the target and the object detection box is greater than the threshold, use the candidate anchor as the candidate box for the corresponding target, and adjust the position of the candidate box through the position offset between the candidate anchor and the object detection box, and use the candidate box with the adjusted position as the preliminary object detection box; Otherwise, the target detection box is used as the preliminary target detection box.

6. The 3D object detection method based on multi-level fine-tuning according to claim 5, characterized in that, The adjustment process of the position of the candidate box is as follows: Among them, w g represents the width of the target detection box, w a represents the width of the corresponding target candidate box, w‘ a represents the width of the preliminary target detection box; l g represents the length of the target detection box, l a represents the length of the corresponding target candidate box, l‘ a represents the length of the preliminary target detection box; h g represents the height of the target detection box, h a represents the height of the corresponding target candidate box, h‘ a represents the height of the preliminary target detection box.

7. A three-dimensional object detection method based on multi-level fine-tuning according to claim 6, characterized in that, The multi-stage fine-tuning module includes: a first-stage target detection box adjustment module, a second-stage target detection box adjustment module, a third-stage target detection box adjustment module, and a target detection box voting module. The first-stage target detection box adjustment module, the second-stage target detection box adjustment module, and the third-stage target detection box adjustment module are respectively used for adjusting the preliminary target detection box of the point cloud features with gradually increasing quality, and all include: an integrity estimation sample reweighting module and a dual-branch detection head module. The integrity estimation sample reweighting module is used to calculate the weight of each preliminary target detection box. The dual-branch detection head module is composed of a fully connected layer detection head and a convolutional layer detection head, and is used to calculate the confidence of the preliminary target detection box by combining the weight of the preliminary target detection box. The target detection box voting module is used to perform weighted averaging on the confidence of the preliminary target detection boxes output by the first-stage target detection box adjustment module, the second-stage target detection box adjustment module, and the third-stage target detection box adjustment module to obtain the final target detection box.

8. A three-dimensional object detection method based on multi-level fine-tuning according to claim 7, characterized in that The specific process of the integrity estimation sample reweighting module for calculating the weight of each preliminary target detection box is as follows: taking the intersection over union of the three-dimensional ground truth box of each target and the minimum bounding box of the corresponding target as the point integrity score of the target, and calculating the weight of the preliminary target detection box of the target: Among them, represents the normalization factor in the b-th iteration, M represents the number of objects in the target, s represents the index of M, represents the point integrity score of the s-th object in the b-th iteration, represents the point integrity score of the m-th object in the b-th iteration.

9. The three-dimensional object detection method based on multi-level fine-tuning according to claim 8, wherein The loss function L of the three-dimensional object detection network total Specifically: L total = L rpn + L pr Among them, L rpn represents the loss function of the candidate box generation module, N a represents the number of candidate anchors, and i represents the index of N a index, represents the classification loss of the i-th candidate anchor, p t,i represents the foreground score predicted by the i-th candidate anchor, and α t represents the first parameter, with a value of 0.25; γ represents the second parameter, with a value of 2; τ(f≥1) represents the regression loss only for foreground anchor points; represents the direction regression loss of the i-th candidate anchor, θ p represents the predicted direction of the i-th candidate anchor, and θ t represents the yaw angle between the i-th candidate anchor and the corresponding object detection box, and SmoothL1 represents the loss function of SmoothL1; represents the regression loss of the i-th candidate anchor, represents the direction classification loss of the i-th candidate anchor; L pr represents the loss function of the multi-level fine-tuning module, represents the confidence loss of the j-th object detection box adjustment module, and ω fc represents weight of, represents the regression loss of the j-th object detection box adjustment module, and ω conv represents weight of, represents the target box orientation loss function of the j-th object detection box adjustment module.