Multi-modal data level fusion three-dimensional target detection method and detector paying attention to long-distance target
Through the multimodal data-level fusion of three-dimensional object detection method, the combination of lidar point cloud and camera images is used to solve the problem of inaccurate position estimation in long-distance object detection, and achieve higher detection accuracy and three-dimensional parameter estimation accuracy.
Patent Information
- Application Number
- CN202510210359.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-22
AI Technical Summary
The existing three-dimensional object detection network based on camera images has low accuracy in estimating the position of long-distance targets, resulting in large errors in detection results, especially at 60 meters, which can reach more than 4 meters, making it impossible to effectively detect long-distance targets.
The multimodal data-level fusion method is adopted to predict the target direction angle through the data-level fusion of the lidar point cloud and the camera image, and the depth completion algorithm is used to densely process the point cloud. Combined with the DLA-34 and DLAUp feature extraction network, the detection head is designed to directly predict the key points and two-dimensional and three-dimensional parameters of the target, and the dense depth map is used as a reference. Multi-interval classification + regression prediction is used to predict the target direction angle, and Focal Loss and L1 Loss regression loss functions are used for supervision and training.
It improves the detection accuracy of long-distance targets, reduces position estimation errors, and improves the accuracy of 3D parameter estimation, especially in target detection beyond 60 meters, which performs significantly better than other networks.
Smart Images

Figure CN120355889A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and particularly to a multi-modal data-level fusion three-dimensional target detection method and detector focusing on long-distance targets. Background Art
[0002] Long-distance three-dimensional target detection is a highly challenging topic. In the multi-modal fusion solutions proposed in the industry, many solutions mainly rely on point clouds and supplement with images, so there are problems with limited detection distance. In recent years, many three-dimensional target detection networks based on camera images have been proposed, such as M3D-RPN, SMOKE, MonoPair, MonoDLE, etc. Through experimental verification, these networks are indeed better at identifying long-distance targets by leveraging the advantage of the data density of camera images. However, these networks also have fatal problems: the position estimation accuracy is not high. There are obvious errors in using only monocular images for position estimation, and this error increases with the increase of distance. For a target at 60 meters, this position estimation error may reach more than 4 meters. The excessive error makes the detection results of these networks for long-distance targets unusable. Summary of the Invention
[0003] The object of the present invention is to address the problems in the background art and propose a multi-modal data-level fusion three-dimensional target detection method and detector (Long-Distance-Focused 3D Detector, LDFMM) that focuses on long-distance targets. To achieve the data-level fusion of lidar point clouds and camera images, thereby improving the accuracy of three-dimensional parameter estimation.
[0004] The technical solution of the present invention, a multi-modal data-level fusion three-dimensional target detection method and detector that focuses on long-distance targets, includes a multi-modal data-level fusion three-dimensional target detection method that focuses on long-distance targets, and specifically includes the following steps:
[0005] S1. Use a depth completion algorithm to densify the depth map (non-empty pixels) of the lidar point cloud;
[0006] S11. Depth inversion: First, perform depth inversion on the valid depth in the depth map, specifically depth inverted = 100 - depth, where depth is the valid depth in the depth map, and depth inverted is the valid depth after linear inversion;
[0007] S12. Rhombus dilation: Perform the dilation operation (Dilation) of OpenCV image morphology in this step, and this operation uses a custom rhombus kernel with a size of 5×5;
[0008] S13. Small hole closure: In this step, the closing operation (Closure) of OpenCV image morphology is used, with a kernel size of 5×5. The closing operation first dilates the image and then erodes it (Erode).
[0009] S14. Background filling: To fill the invalid depths in the background, a dilation operation with a size of 7×7 is used. This algorithm only fills the invalid depth areas and does not change the valid depth areas.
[0010] S15. Denoising and smoothing: Denoising and smoothing. In the denoising process, a median filter with a size of 5×5 is used, and in the smoothing process, a bilateral filter with a size of 5×5 is used.
[0011] S16. Re - inversion: At the end of depth completion, all valid depths are inverted again to restore the normal depth distribution, that is, depth final = 100 - depth inverted where depth inverted is the valid depth after the first depth inversion, and depth final is the valid depth restored to normal after the second depth linear inversion.
[0012] S2. Use DLA - 34 and DLAUp as the feature extraction backbone and neck network.
[0013] In step S2, DLA - 34 is used as the feature extraction backbone to extract features from the input camera image; and the output feature maps of the last four stages are used as reference feature maps, denoted as {F3, F4, F5, F6}.
[0014] In step S2, DLAUp is used as the feature extraction neck network to generate feature maps of specific sizes required for subsequent detection processes; DLAUp adopts a tree - shaped aggregation structure, including IDA connections and double upsampling connections; it performs iterative depth aggregation on the input four feature maps; during the feature transfer process of DLAUp, the deep - level feature maps are gradually enlarged and sequentially concatenated with the shallow - level feature maps, and finally feature maps of specific sizes that are rich in both deep - and shallow - level information are output.
[0015] S3. Design a detection head with the output of target category, confidence, 2D bounding box, and 3D bounding box.
[0016] In step S3, in the key - point prediction step of the detection head, the category and confidence of the target are directly given by the key - point prediction information; the 2D bounding box and 3D bounding box obtain the predicted values of the corresponding positions, sizes, and orientation angles according to the pixel positions of the key points in the feature map, so as to output the 2D bounding box and 3D bounding box. The specific process is as follows:
[0017] a. Prediction of the target center point: Assume that the pixel coordinates of its key point in the feature map are p = (u p , v p ). The two-dimensional center offset o 2D = (Δu 2D , Δv 2D ) and the three-dimensional center offset o 3D = (Δu 3D , Δv 3D ). Then the coordinates of the centers of the target two-dimensional and three-dimensional bounding boxes are c 2D = p + o 2D and c 3D = p + o 3D respectively;
[0018] b. The prediction of the target depth is specifically to use the key point to find the depth reference value in the dense depth map, and on this basis, predict a depth residual value within the range of (-1, 1); add the reference value and the residual value to obtain the final target depth. Then the position of the center of the target three-dimensional bounding box in the camera coordinate system is:
[0019]
[0020] where the camera intrinsic matrix is K ∈ R 3×3 , the projection point of the center of the three-dimensional bounding box is c 3D = (u 3D , v 3D ), and the target depth is z;
[0021] c. To avoid the ambiguity of the target orientation angle and its morphology in the camera image, a multi-interval classification + regression form is used to predict the viewing angle of the target.
[0022] S4. Design the loss function and perform data augmentation.
[0023] The output of the key point prediction branch is a heat map with a size of H / 4 × W / 4 × nc;
[0024] To calculate the loss of this heat map, a corresponding heat map filled with ground truth needs to be generated in advance;
[0025] The method of filling the ground truth is as follows: First, fill 1 at the key point; second, near the key point, perform mapping filling in the form of Gaussian distribution attenuation, and the radius of the mapping is half of the length of the short side of the target two-dimensional bounding box; finally, fill 0 in other areas;
[0026] The Focal Loss classification loss function in a point-by-point form is used to supervise the training of the prediction of this heat map; let the prediction value and the ground truth at the position of the heat map be and respectively, and and are defined by the following formula;
[0027]
[0028] For a target, the heatmap classification loss is shown as follows:
[0029]
[0030] n k is the number of ground-truth targets, and two hyperparameters β and γ are set to 4 and 2 respectively; among them, the term can reduce the weight near the ground-truth of the key points, that is, it allows the network to have a small amount of error when predicting the position of the key points; and the term is responsible for balancing samples of different difficulties;
[0031] For all regression tasks of LDFMM, the L1 Loss regression loss function is used for supervised training; these regression tasks include: two-dimensional center offset, two-dimensional size, three-dimensional center offset, three-dimensional size, residual of target orientation, and target depth; the definition of the L1 Loss regression loss function is shown as follows:
[0032]
[0033] In the formula, n k is the number of ground-truth targets, x i and y i are the predicted value and the ground truth respectively; for the target orientation classification task of LDFMM, the Cross-Entropy Loss classification loss function is used for supervised training; adding up all the above losses can obtain the overall loss of LDFMM.
[0034] In step S4, two data augmentation methods of random flipping and random cropping are adopted; by flipping the input image in the horizontal direction or after random scaling and translation, and then randomly cropping the image according to the size requirements of the input network, the camera image and the dense depth map input to LDFMM are synchronously processed in each transformation process to ensure the consistency of these two data.
[0035] A multi-modal data-level fusion three-dimensional object detector that focuses on distant targets uses the above method for detection, including a detection head and a calculation unit;
[0036] The detection head includes at least one data acquisition module for acquiring the original lidar point cloud data
[0037] The calculation unit is integrated in the detection head, and it includes a data processing module and a calculation module;
[0038] The data processing module is used to perform densification processing on the acquired data;
[0039] The calculation module is used to construct the backbone network and the neck network for feature extraction. After using the backbone network and the neck network to extract features from the collected large data, it calculates and outputs the results according to the corresponding loss function and data augmentation steps.
[0040] Compared with the prior art, the present invention has the following beneficial technical effects:
[0041] By inputting the camera image into the feature extraction network and directly predicting the key points (Keypoints) of the target and the corresponding two-dimensional and three-dimensional parameters in the detection head by means of the anchor-free method, the ability of the detector to identify distant targets is enhanced; by projecting the lidar point cloud onto the camera image plane to generate a dense depth map and selecting the depth at the key points from the dense depth map as a reference, the accuracy of three-dimensional parameter estimation is improved; an improvement is made on the CenterNet architecture, and the dense depth map generated by the lidar point cloud is used as the input of another modality of the LDFMM, improving the accuracy of depth estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the DLA-34 network structure;
[0043] Figure 2 It is a schematic diagram of the DLAUp network structure;
[0044] Figure 3 It is a schematic diagram of the multi-modal data-level fusion three-dimensional object detection method according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] Embodiment 1
[0046] A multi-modal data-level fusion three-dimensional object detection method for focusing on distant targets proposed by the present invention includes the following steps:
[0047] STEP1. Densify the depth map of the lidar point cloud by using a depth completion algorithm:
[0048] The first step is depth inversion. First, invert the effective depth (non-empty pixels) in the depth map. Specifically, depth inverted = 100 - depth, where depth is the effective depth in the depth map, and depth inverted is the effective depth after linear inversion.
[0049] The second step is diamond dilation. In this step, perform the dilation operation (Dilation) of OpenCV image morphology. The operation uses a custom diamond kernel with a size of 5×5.
[0050] Step 3, small hole closure. In this step, the closing operation (Closure) of OpenCV image morphology is used, and the kernel size is 5×5. The closing operation will first perform a dilation operation on the image and then an erosion operation (Erode).
[0051] Step 4, background filling. To fill the invalid depth in the background, a dilation operation with a size of 7×7 is used. This algorithm only fills the invalid depth area and does not change the valid depth area.
[0052] Step 5, denoising and smoothing. The denoising process uses a 5×5 median filter, and the smoothing process uses a 5×5 bilateral filter.
[0053] Step 6, reverse again. At the end of depth completion, all valid depths are reversed again to restore the normal depth distribution, that is, depth final = 100 - depth inverted where depth inverted is the valid depth after the first depth reversal, and depth final is the valid depth restored to normal after the depth linear reversal is performed again.
[0054] STEP2. Use DLA-34 and DLAUp as the feature extraction backbone and neck network; the DLA-34 network structure is as Figure 1 shown.
[0055] HDA is a tree-like aggregation form for different convolutional blocks, as shown in the red box in Figure 1 . HDA contains many aggregation nodes (AggregationNode), and these nodes can very efficiently transfer features from different convolutional blocks. IDA is a serial aggregation form for different network stages, as shown by the yellow connection lines in Figure 1 . The role of IDA is to stack the output feature maps of each network stage from shallow to deep, thereby fusing shallow and deep features.
[0056] DLA-34 has a total of six network stages, namely {S1, S2, S3, S4, S5, S6}. The last four stages use the HDA aggregation form. The convolutional blocks of HDA adopt the residual structure in ResNet-34, that is, it includes two 3×3 convolutional layers connected in series and a 1×1 convolutional layer with a skip connection. The aggregation nodes of HDA will perform channel concatenation on each incoming feature map, and then use a 1×1 convolutional layer for channel fusion and modify the number of channels to the number of channels of the final output of this stage. The blue connection lines in the figure represent two-fold downsampling. In addition, in the first convolutional layer of each stage, the number of channels of the feature map is changed to the number of channels of the final output of this stage. Thus, the downsampling multiples of the output feature maps of the six stages are 1, 2, 4, 8, 16, 32 respectively, and the number of channels are 16, 32, 64, 128, 256, 512 respectively. The DLAUp network structure is as Figure 2 shown.
[0057] The sizes of the four feature maps {F3, F4, F5, F6} input to DLAUp are H / 4×W / 4×64, H / 8×W / 8×128, H / 16×W / 16×256, H / 32×W / 32×512 respectively. During the aggregation process of DLAUp, F3 will not be upsampled, and F4, F5, F6 will be upsampled once, twice, and three times respectively. The size of the finally output feature map F f is the same as that of F3, which is H / 4×W / 4×64. The upsampling connection is composed of a 3×3 convolutional layer and a deconvolutional layer with two-fold upsampling. Among them, the 3×3 convolutional layer will modify the number of channels of the feature map to half of the input. The aggregation nodes will perform channel concatenation on the two incoming feature maps, and then use a 3×3 convolutional layer for channel fusion and compress the number of channels by half.
[0058] STEP3. Design a detection head with the target category, confidence, 2D bounding box, and 3D bounding box as the output;
[0059] The detection head of LDFMM includes three parts: key points, 2D parameters, and 3D parameters. Among them, the target category, confidence, 2D bounding box, and 3D bounding box together constitute the output of the LDFMM detection head.
[0060] The prediction of key points in the detection head is the key. The target category and confidence are directly given by the key point prediction information. The 2D bounding box and 3D bounding box can obtain the predicted values of the corresponding position, size, and orientation angle according to the pixel positions of the key points in the feature map, so as to output the 2D bounding box and 3D bounding box. The specific process is as follows.
[0061] (1) Prediction of the target center point: Assume that the pixel coordinates of its key point in the feature map are p = (u p , vp ), the two-dimensional center offset o 2D =(Δu 2D , Δv 2D ) and the three-dimensional center offset o 3D =(Δu 3D , Δv 3D ), then the coordinates of the center points of the target two-dimensional and three-dimensional bounding boxes are c 2D =p + o 2D , c 3D =p + o 3D .
[0062] (2) The prediction of the target depth is specifically to use the key points to find the depth reference value in the dense depth map and predict a relatively small depth residual on this basis. Adding the reference value and the residual value gives the final target depth. Then the position of the center of the target three-dimensional bounding box in the camera coordinate system is:
[0063]
[0064] where the camera intrinsic matrix is K ∈ R 3×3 , the projection point of the center of the three-dimensional bounding box is c 3D =(u 3D , v 3D ), and the target depth is z.
[0065] (3) To avoid the ambiguity of the target orientation angle and its shape in the camera image, LDFMM uses the form of multi-bin classification + regression to predict the observation angle of the target. Assuming that the entire circumference is divided into 12 parts, the prediction result is that the k-th part has the highest probability (1 ≤ k ≤ 12), and the residual angle in the corresponding part is δ, then the target observation angle can be obtained from Equation (1), and the corresponding target orientation angle can be obtained from Equation (2).
[0066]
[0067] STEP4. Design the loss function and perform data augmentation.
[0068] The physical meaning of the key points predicted by LDFMM is the center point of the visible part of the target of interest in the camera image. In LDFMM, the key point prediction branch outputs a heatmap with a size of H / 4 × W / 4 × nc. To calculate the loss of this heatmap, a corresponding heatmap filled with ground truth needs to be generated in advance. The way of filling the ground truth is as follows: First, fill 1 at the key points. Second, near the key points, fill in according to the form of Gaussian distribution attenuation, and the radius of the mapping is half of the length of the short side of the target two-dimensional bounding box. Finally, fill 0 in other areas.
[0069] This algorithm uses the point - by - point form of the Focal Loss classification loss function to supervise the training of the prediction of this heatmap. Let the predicted value and the ground truth value at the position of the heatmap be $\hat{y}$ and $y$ respectively, which are defined by Equation (3) and Equation (4). For a target, the heatmap classification loss is shown in Equation (5).
[0070]
[0071] In Equation (5), $n$ k is the number of ground - truth targets, and the two hyperparameters $\beta$ and $\gamma$ are set to 4 and 2 respectively. Among them, the term can reduce the weight near the ground truth of the key point, that is, it allows the network to have a small amount of error when predicting the position of the key point. And the term is responsible for balancing samples of different difficulties.
[0072] For all regression tasks of LDFMM, this algorithm uses the L1 Loss regression loss function for supervised training. These regression tasks include: 2D center offset, 2D size, 3D center offset, 3D size, the residual of the target orientation, and the target depth. The definition of the L1 Loss regression loss function is shown in Equation (6).
[0073]
[0074] In Equation (12), $n$ k is the number of ground - truth targets, $\hat{x}$ i and $\hat{y}$ i are the predicted value and the ground truth respectively. For the target orientation classification task of LDFMM, this algorithm uses the Cross - Entropy Loss classification loss function for supervised training. Adding up all the above losses can obtain the overall loss of LDFMM.
[0075] In addition, this algorithm uses two data augmentation methods: random flipping and random cropping. By flipping the input image in the horizontal direction or after random scaling and translation, and then randomly cropping the image according to the size requirements of the input network, the camera image and the dense depth map input to LDFMM are synchronously transformed, so as to ensure the consistency of these two kinds of data.
[0076] In the overall performance evaluation experiment of the present invention, by comparing with the monocular 3D object detector MonoDLE and the multi - modal 3D object detector MVMM, the advantages of LDFMM in long - distance object detection are verified.
[0077] The evaluation results of LDFMM and two comparison networks on the KITTI validation set are shown in Table 1. Table 1 shows the detection accuracies of the three networks under different detection categories and different distance segments. The AP is approximately calculated based on 11 recall positions on the PR curve. The two comparison networks are the monocular 3D object detector MonoDLE and the multi-modal 3D object detector MVMM proposed in Chapter 2 of this paper. The following two parts of conclusions can be drawn from Table 1:
[0078] First, compared with MonoDLE, LDFMM shows an absolute accuracy advantage in all aspects. Regardless of which detection category or which distance segment, the detection accuracy of LDFMM is significantly higher than that of MonoDLE. Due to the introduction of the dense depth map as the reference depth in the detection head, the positioning accuracy of LDFMM for targets is greatly improved, and the quality of the predicted 3D bounding boxes is much higher than that of MonoDLE. Therefore, LDFMM greatly surpasses MonoDLE in the 3D AP of various categories and distance segments.
[0079] Second, compared with MVMM, the performance of LDFMM has both advantages and disadvantages. Regarding the detection of cars: For targets within 60 meters, the detection accuracy of MVMM is higher, and LDFMM has a disadvantage of about 2 to 10 points in 3D AP; but for targets beyond 60 meters, the detection accuracy of LDFMM is significantly higher than that of MVMM, with a 6.80 higher in 3D AP from 60 to 70 meters and an 18.86 higher in 3D AP from 70 to 80 meters. Regarding the detection of pedestrians and cyclists: LDFMM is inferior to MVMM in all distance segments and there is a large gap. For this, one possible reason is that the sample quantities of pedestrians and cyclists are not sufficient. Under the premise of insufficient training samples, it is difficult for the network to deeply learn the prediction rules of key points. Another possible reason is that in this chapter, when training LDFMM, the loss balance of different categories is not well done, resulting in the network paying more attention to the prediction of cars.
[0080] Table 1 Comparison of 3D AP of LDFMM and two comparison networks on the KITTI validation set
[0081]
[0082] Example 2
[0083] This example provides a multi-modal data-level fusion 3D object detector that focuses on distant targets, and uses the method in Example 1 for detection, including a detection head and a calculation unit;
[0084] The detection head includes at least one data acquisition module for acquiring the original lidar point cloud data
[0085] The computing unit is integrated in the detection head, which includes a data processing module and a computing module;
[0086] The data processing module is used to densify the collected data;
[0087] The computing module is used to construct the backbone network and the neck network for feature extraction. After using the backbone network and the neck network to extract features from the collected large data, the data is calculated and the result is output according to the corresponding loss function and data augmentation steps.
[0088] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those skilled in the art.
Claims
1. A multi-modal data-level fusion three-dimensional object detection method for focusing on distant targets, characterized in that, It includes the following specific steps: S1. Densify the depth map of the lidar point cloud using a depth completion algorithm; S2. Use DLA-34 and DLAUp as the feature extraction backbone and neck network; S3. Design a detection head that outputs the target category, confidence, 2D bounding box, and 3D bounding box; S4. Design the loss function and perform data augmentation.
2. The multi-modal data-level fusion three-dimensional object detection method for focusing on long-distance targets according to claim 1, wherein Step S1 includes: S11. Depth inversion: First, perform depth inversion on the valid depths in the depth map. Specifically, depth inverted = 100 - depth, where depth is the valid depth in the depth map, and depth inverted is the valid depth after linear inversion; S12. Rhombus dilation: Perform the dilation operation of OpenCV image morphology, which uses a rhombus kernel with a custom size for the operation; S13. Small hole closing; Use the closing operation of OpenCV image morphology; The closing operation will first perform a dilation operation on the image and then an erosion operation; S14. Background filling; To fill the invalid depth in the background, use the dilation operation; This operation method only fills the invalid depth area and does not change the valid depth area; S15. Denoising and smoothing; The median filter is used for the denoising process, and the bilateral filter is used for the smoothing process; S16. Reverse again; at the end of depth completion, reverse all valid depths again to restore the normal depth distribution, i.e., depth final = 100 - depth inverted , where depth inverted is the valid depth after the first depth reversal, and depth final is the valid depth restored to normal after depth linear reversal again.
3. The multi-modal data-level fusion three-dimensional object detection method for focusing on long-distance targets according to claim 2, wherein, In step S11, the valid depth in the depth map refers to non-empty pixels.
4. The multi-modal data-level fusion three-dimensional object detection method for focusing on long-distance targets according to claim 1, wherein, In step S2, DLA-34 is used as the feature extraction backbone to extract features from the input camera image; And the output feature maps of the last four stages are used as the reference feature maps, denoted as {F3, F4, F5, F6}.
5. The multimodal data-level fusion three-dimensional object detection method for focusing on long-distance targets according to claim 1, characterized in that, In step S2, DLAUp is used as the feature extraction neck network to generate feature maps of a specific size required for the subsequent detection process.
6. The multi-modal data-level fusion three-dimensional object detection method for focusing on long-distance targets according to claim 5, wherein, DLAUp adopts a tree-like aggregation structure, including IDA connections and two-fold upsampling connections; It performs iterative depth aggregation on the input four feature maps; During the feature transfer process of DLAUp, the deep feature maps are gradually enlarged and sequentially concatenated with the shallow feature maps, and finally feature maps of a specific size that are rich in both deep and shallow layer information are output.
7. The multi-modal data-level fusion three-dimensional object detection method and detector for focusing on distant targets according to claim 1, wherein In step S3, in the key point prediction step of the detection head, the category and confidence of the target are directly given by the key point prediction information; The 2D bounding box and 3D bounding box obtain the predicted values of the corresponding position, size, and orientation angle according to the pixel positions of the key points in the feature map, so as to output the 2D bounding box and 3D bounding box. The specific process is as follows: a. Prediction of the target center point: Assume that the pixel coordinates of its key point in the feature map are p = (u p , v p ), the 2D center offset o 2D = (Δu 2D , Δv 2D ), and the 3D center offset o 3D = (Δu 3D , Δv 3D ). Then the coordinates of the center points of the target 2D and 3D bounding boxes are c 2D = p + o 2D , c 3D = p + o 3D ; b. The prediction of the target depth is specifically to use the key points to find the depth reference value in the dense depth map, and on this basis, predict a depth residual within the range of (-1, 1); Add the reference value and the residual value to obtain the final target depth. Then, the position of the center of the target 3D bounding box in the camera coordinate system is: where the camera intrinsic matrix is K ∈ R 3×3 , the central projection point of the 3D bounding box is c 3D =(u 3D , v 3D ), and the target depth is z; c. To avoid the ambiguity between the target orientation angle and its morphology in the camera image, the observation angle of the target is predicted in the form of multi-interval classification + regression.
8. The multi-modal data-level fusion three-dimensional object detection method for focusing on distant targets according to claim 1, wherein, In step S4, the output of the key point prediction branch is a heat map with a size of H / 4×W / 4×nc; To calculate the loss of this heat map, a heat map filled with the corresponding ground truth needs to be generated in advance; The filling method of the ground truth is as follows: First, fill 1 at the key points; Second, near the key points, perform mapping filling in the form of Gaussian distribution attenuation, and the radius of the mapping is half of the length of the short side of the target 2D bounding box; Finally, fill 0 in other areas; The prediction of the heatmap is supervised and trained using the Focal Loss classification loss function in a point-by-point form; let the predicted value and the ground truth value at the position of the heatmap be $\hat{y}$ and $y$ respectively, and $\hat{y}$ and $y$ are defined by the following formulas respectively; For a target, the heatmap classification loss is shown in the following formula: n k is the number of ground truth targets, and two hyperparameters β and γ are set to 4 and 2 respectively; among them, the term can reduce the weight near the ground truth of the key points, that is, it allows a small amount of error to occur when the network predicts the position of the key points; while the term is responsible for balancing samples of different difficulties; For all regression tasks of LDFMM, the L1 Loss regression loss function is used for supervised training; these regression tasks include: two-dimensional center offset, two-dimensional size, three-dimensional center offset, three-dimensional size, residual of target orientation, and target depth; the definition of the L1 Loss regression loss function is shown in the following formula: where n k is the number of true value targets, x i and y i are the predicted value and the true value respectively; for the target orientation classification task of LDFMM, the Cross-Entropy Loss classification loss function is used for supervised training; adding up all the above losses can obtain the overall loss of LDFMM.
9. The multi-modal data-level fusion three-dimensional object detection method for focusing on distant targets according to claim 8, characterized in that, In step S4, two data augmentation methods of random flipping and random cropping are adopted; by flipping the input image in the horizontal direction or after random scaling and translation, and then randomly cropping the image according to the size requirements of the input network, the camera image and the dense depth map input to LDFMM are synchronously subjected to each transformation process to ensure the consistency of these two data.
10. A multi-modal data-level fusion 3D object detector for focusing on distant targets, which performs detection using the method according to any one of claims 1-9, characterized in that It includes a detection head and a calculation unit; The detection head includes at least one data acquisition module for acquiring original lidar point cloud data The calculation unit is integrated in the detection head and includes a data processing module and a calculation module; The data processing module is used for densifying the acquired data; The calculation module is used to construct the backbone network and the neck network for feature extraction. After using the backbone network and the neck network to extract features from the acquired large data, the data is calculated and the results are output according to the corresponding loss function and the data augmentation steps.