Three-dimensional target detection method and system

By adopting a multi-stage cross-modal feature fusion method in the unmanned driving environment of tunnel rail vehicles, the radar point cloud is fused with image data, and using multiple regression networks for target detection, the problem of performance attenuation of lidar in harsh environments and insufficient camera-mm wave radar fusion accuracy is solved, and the target detection effect with high accuracy and robustness is achieved.

CN120182584AActive Publication Date: 2025-06-20SHIJIAZHUANG TIEDAO UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510653992.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

In the perception scenario of unmanned driving environment of tunnel rail vehicles, the high purchase cost of lidar and the performance attenuation in strict environments seriously restrict their large-scale deployment. At the same time, the accuracy and robustness of camera-mm-wave radar cross-modal fusion are insufficient, resulting in poor target detection accuracy.

Method used

The multi-stage cross-modal feature fusion method is adopted to fuse the radar point cloud data and image data. Through the co-coding of Swin Transformer and sparse convolution, multi-scale fusion of millimeter-wave radar voxel features and image window features is realized. The cone depth range is dynamically optimized in combination with the mobile window self-attention mechanism, and the multi-stage fusion feature is input to the multiple regression network for target detection.

Benefits of technology

It effectively improves the accuracy and robustness of target detection in tunnel rail vehicle scenarios, reduces system cost and maintenance difficulty, and combines the stable ranging/speed measurement performance of millimeter wave radar and the camera's semantic recognition capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182584A_ABST
    Figure CN120182584A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a three-dimensional target detection method and system, and the method comprises the steps: obtaining point cloud data and image data collected by a sensing device at the same moment; performing voxelization operation on the point cloud data to obtain voxel features; extracting image multi-scale features from the image data; the voxel features and the image multi-scale features are fused in a channel splicing mode, and multi-level fusion features are obtained; the multi-level fusion features are input into a regression model to output a detection result of the target, the regression model comprises a multiple regression network, and the detection result comprises one or more of the center position, the size, the posture, the speed and the category of the target. According to the scheme of the invention, the accuracy and robustness of target detection in a tunnel rail vehicle scene are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing. More specifically, the present invention relates to a three-dimensional object detection method and system. Background Art

[0002] The unmanned transportation system mainly consists of three parts: environmental perception, planning and decision-making, and control execution. Among them, the environmental perception system is the key prerequisite for realizing unmanned driving, which determines whether the subsequent decision-making and control systems have complete and accurate data as the basis.

[0003] The autonomous driving environmental perception system usually relies on the collaborative work of multiple sensors. Among them, cameras, lidars, and millimeter-wave radars are the current mainstream sensor solutions. Each type of sensor has its own advantages, but also has inherent limitations. Cameras can provide high-resolution image information, capture details in the environment, and achieve precise object detection, recognition, and classification. However, its performance significantly decreases in harsh weather conditions such as low light, rain, snow, haze, and complex environments such as tunnels and mines, making it difficult to ensure clear image quality.

[0004] Lidars perform well in complex environments with their excellent distance and angle resolution, especially maintaining high performance in low light conditions. However, under extremely harsh weather conditions (such as heavy rain, heavy snow, thick fog, etc.), its detection accuracy and effective detection distance will drop significantly. In addition, the cost of lidars is relatively high, making it difficult to popularize in some low-cost applications.

[0005] Millimeter-wave radars, with their strong anti-interference ability and all-weather adaptability, can operate stably in complex lighting and harsh environmental conditions and accurately measure the position and speed information of targets. However, due to its low resolution, millimeter-wave radars have limitations in the detection accuracy and detail recognition ability of static objects.

[0006] In recent years, the multi-modal sensor fusion technology based on lidars and cameras has made remarkable progress in the field of autonomous driving. By fusing the high-precision three-dimensional point clouds of lidars and the rich texture information of cameras, it has been widely verified that high-precision 3D object detection can be achieved.

[0007] However, in the unmanned driving environmental perception scenario of tunnel rail vehicles, such fusion solutions face the following engineering bottlenecks: 1. The tunnel interior has characteristics such as narrow cross-sections and complex support structures, resulting in lidars being vulnerable to mechanical vibrations and equipment collisions, and significantly increasing the sensor layout and maintenance costs; 2. The lighting conditions in the tunnel are usually affected, and the intermittent lighting leads to uneven lighting conditions or sudden changes in lighting.

[0008] 3. Construction dust causes a large amount of noise in the lidar point cloud. At the same time, the metal reflection surface in the turnout area of the track interferes with the laser beam, causing multiple echoes, which further reduces the quality of the point cloud.

[0009] Under the above working conditions, the high purchase cost of lidar and its performance attenuation in harsh environments seriously restrict its large-scale deployment in tunnel scenarios. Moreover, the accuracy and robustness of camera-millimeter wave radar cross-modal fusion in tunnel rail vehicle scenarios are insufficient.

[0010] Therefore, how to achieve accurate detection of targets in the harsh environment of mining areas has become the focus of current research. Summary of the Invention

[0011] The present invention provides solutions in the following aspects for the problem of poor accuracy of target detection in the harsh environment of mining areas.

[0012] In a first aspect, the present invention provides a three-dimensional target detection method, including: acquiring point cloud data and image data collected by a sensing device at the same moment; performing voxelization operation on the point cloud data to obtain voxel features; extracting image multi-scale features from the image data; fusing the voxel features and the image multi-scale features by means of channel splicing to obtain multi-level fusion features; inputting the multi-level fusion features into a regression model to output the detection result of the target, where the regression model includes multiple regression networks, and the detection result includes one or more of the center position, size, attitude, speed, and category of the target.

[0013] Preferably, the regression model includes a primary regression network and a secondary regression network, and the method further includes: inputting the multi-level fusion features into the primary regression network to obtain a primary detection result, where the primary detection result includes 2D / 3D detection frames; performing frustum association and radar point cloud mapping according to the 2D / 3D detection frames to determine heatmap enhanced features; inputting the heatmap enhanced features into the secondary regression network to output a secondary detection result; performing confidence evaluation on the primary detection result and the secondary detection result to determine the final detection result.

[0014] Preferably, performing frustum association and radar point cloud mapping according to the 2D / 3D detection frames to determine heatmap enhanced features includes: generating a 3D frustum using the 3D detection frame and the camera intrinsic matrix, and projecting the radar point cloud onto the image 2D frame and the 3D frustum to screen valid point clouds; for each radar target associated with the 3D frustum, generating a three-channel heatmap respectively representing the center, boundary, and direction information of the target according to its 2D detection frame position; splicing the three-channel heatmap with the original image features to generate heatmap enhanced features.

[0015] Preferably, the calculation formula for the multi-level fusion features is:

[0016] In the formula, is the multi-level fusion feature, is the radar feature, is the image feature, is the spatial weight, is the channel weight.

[0017] Preferably, the initial regression network adopts a joint loss function during training, including one or more of classification loss, regression loss, and attribute regression loss.

[0018] Preferably, the loss functions of classification loss, regression loss, and attribute regression loss are respectively:

[0019]

[0020]

[0021] In the formula, is the classification loss, N is the total number of samples participating in the calculation, represents the true heatmap label, generated by a Gaussian kernel, with 1 at the center point and decaying according to the distance around it. α and β are the hyperparameters of the focal loss, is the probability value of the (x, y) position and c category in the predicted heatmap, is the regression loss, is the true value of the th dimensional regression quantity, is the attribute regression loss, is the true value of the data.

[0022] Preferably, the depth range of the 3D frustum is:

[0023]

[0024]

[0025] In the formula, is the frustum depth range, is the depth of the 3D detection box output after the initial regression, is the adjustment parameter, dep_max / dep_min are the upper and lower bounds of the target depth estimation obtained from the preliminary image detection, dist_thresh takes half of the depth range (dep_max - dep_min) as the reference threshold for defining the initial expansion radius of the frustum in the depth direction, and expansion_ratio is the expansion coefficient.

[0026] Preferably, confidence evaluation is performed on the primary detection result and the secondary detection result to determine the final detection result, including: comparing the confidences corresponding to the primary detection result and the secondary detection result, and selecting the output with a higher confidence as the final detection result.

[0027] Preferably, it further includes: determining whether the depth and the rotation angle in the detection result match. If they do not match, the secondary detection result is selected as the final detection result. Specifically, when determining whether they match, the secondary regression is the corrected result after fusing radar features, and its confidence combines physical information such as the depth and speed of the radar. If the confidence difference between the two regression results (both regressions generate Gaussian distributions of the key-point heatmaps, reflecting the probability distribution of the target center point) exceeds a preset threshold, it is considered that there is a conflict. Since dep and rot are both in the two regression heads, only the second regression heads with higher accuracy are used.

[0028] In a second aspect, the present invention further provides a three-dimensional target detection system, including: a processor; a memory storing computer program instructions, which implement a three-dimensional target detection method according to the first aspect described above when the computer program instructions are executed by the processor.

[0029] The beneficial effects of the present invention are as follows: By performing multi-level cross-modal feature fusion on radar point cloud data and image data and using a multi-level regression network for target detection, the present invention effectively improves the accuracy and robustness of target detection in the tunnel rail vehicle scenario. Further, by constructing a camera-millimeter wave radar cross-modal fusion framework, the stable ranging / speed measurement performance of the millimeter wave radar and the semantic recognition ability of the camera can be effectively combined, significantly reducing the system cost and maintenance difficulty while ensuring the detection accuracy. Description of the Drawings

[0030] Figure 1 is a flowchart of a three-dimensional target detection method according to an embodiment of the present invention; Figure 2 is a flowchart of a method for a regression network to process multi-level fusion features according to an embodiment of the present invention; Figure 3 is a network structure diagram of three-dimensional target detection according to an embodiment of the present invention; Figure 4It is a structural diagram of a 3D object detection system according to an embodiment of the present invention. Detailed implementation manners

[0031] The core improvement of the present invention lies in constructing a hierarchical cross-modal fusion architecture. Through the collaborative coding of Swin Transformer and sparse convolution, multi-scale fusion of millimeter-wave radar voxel features and image window features is achieved, and the cone depth range is dynamically optimized in combination with the moving window self-attention mechanism, effectively ensuring the detection accuracy and recognition ability, and improving the robustness of the system.

[0032] The following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings.

[0033] Figure 1 It is a flowchart of a 3D object detection method 100 according to an embodiment of the present invention.

[0034] As Figure 1 shown, at step S101, point cloud data and image data collected by a sensing device at the same moment are acquired.

[0035] Synchronous data acquisition: Synchronously read millimeter-wave radar point cloud data and camera image data, and align the sensor coordinate systems through spatio-temporal calibration to ensure the spatio-temporal consistency of the data, thereby ensuring that the point cloud data and the image data are aligned in time and space, that is, calibrating the external parameters and internal parameters.

[0036] At step S102, a voxelization operation is performed on the point cloud data to obtain voxel features.

[0037] Point cloud Pillar Expansion encoding: Convert sparse radar point clouds into dense 3D columnar voxels (similar to the PointPillars method), enhance the spatial representation ability of the target, and facilitate subsequent fusion with image features. Specifically, the original radar point cloud is divided into 3D columnar voxels of a fixed size, and the discrete point cloud is converted into a dense 3D columnar structure through voxelization encoding, increasing the volume representation of the target in three-dimensional space, thereby enhancing the subsequent cross-modal matching probability.

[0038] At step S103, image multi-scale features are extracted from the image data.

[0039] In some embodiments, the extraction of image multi-scale features can adopt the method of image feature extraction and high-resolution fusion: Use Swin Transformer as the backbone network to extract image multi-scale features, and generate high-resolution feature maps (original Figure 1 / 4 and 1 / 8 sizes) at Stage1-2.

[0040] At step S104, the voxel features and the image multi-scale features are fused by means of channel splicing to obtain multi-level fusion features. Specifically, the radar voxel features after Pillar encoding can be encoded into voxel features spatially aligned with the image feature map through a sparse convolutional network. The radar voxel features and the image window features at different stages are fused by channel splicing to obtain the multi-level fusion features of the camera and the millimeter-wave radar.

[0041] At step S105, the multi-level fusion features are input into a regression model to output the detection result of the target. The regression model includes multiple regression networks, and the detection result includes one or more of the center position, size, pose, speed, and category of the target. In some embodiments, it can be processed through multiple regression networks, and the result can be evaluated in combination with the confidence of the detection result.

[0042] In the above embodiments of the present invention, the feature alignment between the radar and the image is achieved through the cross-modal fusion process, and the parameter optimization is achieved through multi-stage regression processing, effectively improving the accuracy of the target detection result under harsh conditions.

[0043] Figure 2 It is a flowchart of a method 200 for a regression network to process multi-level fusion features according to an embodiment of the present invention. The above regression model includes a primary regression network and a secondary regression network.

[0044] As Figure 2 shown, at step S201, the multi-level fusion features are input into the primary regression network to obtain a primary detection result, and the primary detection result includes 2D / 3D detection boxes. In some embodiments, based on the fused features, a preliminary 2D detection box and a 3D detection box of the target are output through the regression network, including the center position, size, depth, and observation angle of the 3D box. In this step, rough 2D / 3D detection boxes can be predicted based on the fused features, so as to filter and denoise the radar point cloud by using the 2D / 3D detection boxes.

[0045] At step S202, frustum association and radar point cloud mapping are performed according to the 2D / 3D detection boxes to determine the heatmap enhanced features. In some embodiments, a 3D frustum can be generated by using the 3D detection box and the camera intrinsic matrix, and the radar point cloud is projected onto the image 2D box and the 3D frustum to screen the effective point cloud. Three heatmaps of the target center, boundary, and direction are generated through the 2D box, and the heatmaps are spliced with the original image features to generate the heatmap enhanced features.

[0046] The dynamic 3D frustum construction process can be carried out in the following way: Using the 3D Box parameters obtained from the preliminary regression and the camera intrinsic matrix, generate 3D frustums corresponding to each 3D Box. The frustums are used as candidate regions to quickly screen valid radar point clouds and suppress background noise and invalid matches. Specifically, frustum association depends on the depth, observation angle alpha, and 3D box size dim of each 3D Box obtained after the initial regression, plus the camera calibration matrix to generate 3D frustums. Similar to Roi extraction, it speeds up the matching speed while filtering out invalid point clouds.

[0047] In an application scenario: (1) Project the radar Pillar onto the pixel coordinate system and perform spatial overlap matching with the 2D detection box to screen out the preliminary associated targets; (2) Project the Pillar onto the camera coordinate system and match it with the depth range of the 3D frustum to further filter out invalid point clouds beyond the frustum depth range. (3) For each radar target associated with the frustum, generate three heatmaps based on the position of its 2D detection box, which respectively represent the target center, boundary, and direction information; (4) Stitch the three-channel heatmap with the original image features and input them as additional features into the subsequent network to enhance the target context information.

[0048] At step S203, input the enhanced heatmap features into the secondary regression network to output the secondary detection results. In some embodiments, based on the enhanced features fused with the heatmap, optimize the depth, rotation angle, speed, and category attributes of the target through the secondary regression network.

[0049] At step S204, perform confidence evaluation on the initial detection results and the secondary detection results to determine the final detection results. In some embodiments, comprehensively consider the results of the initial regression and the secondary regression, and select the output with higher confidence as the final detection result; for conflicting parameters (such as depth and rotation angle), preferentially adopt the secondary regression result and output the 3D position, size, pose, speed, and category of the target.

[0050] In summary of this embodiment, through dynamic frustum construction, the search space can be adjusted in real time according to the preliminary prediction, effectively improving the search efficiency. At the same time, through the stitching process of the three-channel heatmap, the geometric attributes of the target can be displayed, alleviating the positioning ambiguity caused by the sparse radar point cloud. Further, through the two-stage regression and confidence fusion operations, the accuracy of the detection process is effectively improved.

[0051] Next, the above scheme will be described in detail in combination with the calculation process in specific embodiments. Figure 3 It is a network structure diagram of three-dimensional target detection according to an embodiment of the present invention.

[0052] Step 1: The point cloud and image data are data from heterogeneous sensors. Therefore, it is necessary to perform time and space synchronization processing on the collected sparse point cloud data and image data. As Figure 3 shown, read the original image and point cloud data, normalize the image data, preprocess the point cloud data, and perform voxelization on the point cloud data to obtain training data.

[0053] Step 2: Multilevel cross-modal feature fusion. The backbone network feature extraction is achieved through the following steps, and multilevel fusion features are fused to generate multilevel cross-modal feature maps.

[0054] (1) In the high-resolution features of Stage1-2 of the Swin Transformer, encode the radar point cloud into voxel features through sparse convolution, and fuse it with the image window features through channel concatenation.

[0055] (2) In Stage3-4, adopt the hierarchical feature fusion block (HFF) of HiFuse, perform spatial-channel attention weighting on the radar and image features respectively and then fuse them. The spatial attention generates a spatial weight map through a 3×3 convolution, and the formula is:

[0056] In the formula, is the radar feature, is the image feature, is the spatial weight. This spatial weight is obtained by element-wise adding the radar feature and the image feature , then performing 3×3 convolution, and passing through the Sigmoid activation.

[0057] The channel attention adopts the SE module (compression ratio = 16), and the formula is:

[0058] The channel weight is obtained by adding the radar feature and the image feature , then performing global average pooling (GAP), MLP (multi-layer perceptron, whose structure includes two fully connected layers and a bottleneck design) processing, and finally passing through the Sigmoid activation, that is, using an SE module.

[0059] Finally, perform weighted fusion output to obtain the multilevel fusion features of the camera and millimeter-wave radar, is the multilevel fusion feature, and the formula is:

[0060] Step 3: Preliminary regression and 3D frustum generation.

[0061] For the fused features, the initial 2D and 3D detection boxes of the target are output through the initial regression network, including the center position, size, depth, and observation angle of the 3D box.

[0062] The frustum association depends on the depth depth, observation angle alpha, and 3D box size dim of each 3D Box obtained after the initial regression, plus the camera calibration matrix to generate a 3D frustum. Similar to Roi extraction, it speeds up the matching speed while filtering out invalid point clouds.

[0063] During model training, a joint loss function is adopted to ensure the balanced optimization of each subtask. It mainly includes the following parts: 1) Classification (heatmap) loss.

[0064] To accurately locate the target center, after the heatmap output by the network is normalized, FastFocalLoss based on the Focal Loss idea is used to calculate the loss. This loss function assigns different weights to positive and negative samples, focuses on the difficult-to-detect center points, and effectively alleviates the problem of sample imbalance. The main role of Focal Loss is to reduce the influence of the background area on the loss, making the network pay more attention to the detection of the target center. The specific formula is as follows:

[0065] In the formula: N is the total number of samples participating in the calculation, Y represents the true heatmap label, which is generated through a Gaussian kernel, with 1 at the center point and decaying according to the distance around it. α and β are the hyperparameters of the focal loss.

[0066] 2) Regression loss.

[0067] For most regression heads, the L1 loss (Mean Absolute Error, MAE) is adopted. The L1 loss calculates the difference between the predicted value and the true value, and has good robustness. Especially when there are outliers, the L1 loss can effectively reduce the influence of these outliers on model training.

[0068]

[0069] In the formula: is the true value of the data, and N is the number of samples in the regression task.

[0070] 3) Attribute regression loss.

[0071] For the attribute regression heads (such as the depth, speed, etc. of the target), Binary Cross Entropy (BCE) Loss is used. The formula of BCE Loss is:

[0072] In the formula: is the true value of the data, and N is the number of samples in the regression task.

[0073] Step Four: Secondary matching and feature enhancement.

[0074] Based on the camera calibration matrix (intrinsic and extrinsic parameters), the 2D box of the target, and depth estimation, a 3D frustum is generated for each detected target. This frustum is the projection area of the target in 3D space, covering the space from the camera optical center to the depth range of the target. It is used to screen the potentially associated radar point cloud. The depth threshold calculation formula is:

[0075] Adjustment parameter The formula is:

[0076] The depth range of the frustum is:

[0077] In the formula, dep_max / dep_min are the upper and lower bounds of the target depth estimation obtained from the preliminary image detection, dist_thresh takes half of the depth range (dep_max - dep_min) as the reference threshold, which is used to define the initial expansion radius of the frustum in the depth direction. expansion_ratio is the expansion coefficient, adjustment parameter By scaling dist_thresh proportionally, the search range in the depth direction is increased.

[0078] Step Five: Two-stage regression and result integration.

[0079] (1) Based on the enhanced features of the fused Heatmap, optimize the depth, rotation angle, speed, and class attributes of the target through a secondary regression network. The 3D frustum is similar to an roi (region of interest) extraction. Using the depth depth, observation angle alpha, and 3D box size dim of each 3D Box obtained by the Primary Regression Heads, combined with the camera calibration matrix to generate the 3D frustum. By filtering the point cloud, only the point cloud inside the frustum is left, and then the point cloud and the multi-level fused features are concatenated for secondary regression.

[0080] (2)Integrate the results of the first regression and the second regression, and select the output with a higher confidence level as the final detection result. For conflicting parameters (such as depth and rotation angle), give priority to the results of the second regression and output the 3D position, size, pose, speed, and category of the target. The second regression is the corrected result after fusing radar features, and its confidence level combines physical information such as the depth and speed of the radar. If the confidence level difference between the two regression results (both regressions generate Gaussian distributions of the key-point heatmaps, reflecting the probability distribution of the target center point) exceeds the preset threshold, it is considered that there is a conflict.

[0081] Through the fusion of radar and image features, the present invention effectively retains the geometric structure information of the radar point cloud and the texture features of the image. After they are fused after spatial alignment, it makes up for the inherent defects of a single sensor, effectively improves the accuracy of the detection process, and at the same time enhances the local representation through multi-scale feature fusion and corrects errors through multi-stage regression, strengthening the anti-interference ability of the target detection method and improving the robustness of the target detection process in harsh environments.

[0082] Figure 4 It is a structural diagram of a three-dimensional target detection system according to an embodiment of the present invention.

[0083] The present invention also provides a three-dimensional target detection system. As Figure 4 shown, the system includes a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a three-dimensional target detection method according to the foregoing embodiments is implemented.

[0084] In the description of this specification, the meanings of "a plurality" and "several" are at least two, such as two, three or more, etc., unless otherwise clearly and specifically defined.

[0085] Although this specification has shown and described multiple embodiments of the present invention, it is obvious to those skilled in the art that such embodiments are provided only by way of example. Those skilled in the art will think of many changes, alterations, and alternative ways without departing from the spirit and idea of the present invention. It should be understood that various alternative solutions to the embodiments of the present invention described herein can be adopted in the process of practicing the present invention.

Claims

1. A three-dimensional target detection method, characterized in that: include: Obtaining point cloud data and image data collected by the sensing device at the same time; Perform voxelization on the point cloud data to obtain voxel features; Extracting multi-scale image features from image data; The voxel features are fused with the multi-scale features of the image by channel splicing to obtain multi-level fusion features; The multi-level fusion features are input into a regression model to output a detection result of the target, wherein the regression model includes a multiple regression network, wherein the detection result includes one or more of the center position, size, posture, speed and category of the target.

2. The three-dimensional target detection method according to claim 1, characterized in that: The regression model includes a primary regression network and a secondary regression network, and the method further includes: Inputting the multi-level fusion features into the primary regression network to obtain a primary detection result, wherein the primary detection result includes a 2D / 3D detection frame; Perform frustum association and radar point cloud mapping based on 2D / 3D detection boxes to determine thermal map enhancement features; Input the heat map enhanced features into the secondary regression network to output the secondary detection results; Conduct confidence assessment on the initial and secondary test results to determine the final test results.

3. The three-dimensional target detection method according to claim 2, characterized in that: Perform frustum association and radar point cloud mapping based on 2D / 3D detection boxes to determine heatmap enhancement features, including: Generate a 3D viewing cone using the 3D detection frame and the camera intrinsic parameter matrix, and project the radar point cloud to the image 2D frame and 3D viewing cone to filter the valid point cloud; For each radar target associated with the 3D frustum, a three-channel heat map representing the target center, boundary and direction information is generated according to the position of its 2D detection box; The three-channel heat map is concatenated with the original image features to generate heat map enhanced features.

4. The three-dimensional target detection method according to claim 1, characterized in that: The calculation formula of the multi-level fusion feature is: In the formula, is a multi-level fusion feature. is the radar signature, is the image feature, is the spatial weight, is the channel weight.

5. The three-dimensional target detection method according to claim 2, characterized in that: The initial regression network adopts a joint loss function during training, including one or more of classification loss, regression loss and attribute regression loss.

6. The three-dimensional target detection method according to claim 5, characterized in that: The loss functions for classification loss, regression loss, and attribute regression loss are: In the formula, is the classification loss, N is the total number of samples involved in the calculation, Represents the real heat map label, generated by a Gaussian kernel, with a value of 1 at the center and decaying by distance around it. α and β are hyperparameters of focal loss. is the probability value of the (x, y) position and c category in the predicted heat map, is the regression loss, It is The true value of the dimensional regressor, It is The predicted value of the dimensional regressor, is the attribute regression loss, is the true value of the data.

7. The three-dimensional target detection method according to claim 3, characterized in that: The depth range of the 3D viewing frustum is: In the formula, is the depth range of the visual frustum, is the depth of the 3D detection frame output after the initial regression, It is an adjustment parameter. dep_max / dep_min are the upper and lower bounds of the target depth estimation based on the preliminary detection of the image. dist_thresh takes half of the depth range (dep_max - dep_min) as the reference threshold to define the initial expansion radius of the frustum in the depth direction. expansion_ratio is the expansion coefficient.

8. The three-dimensional target detection method according to claim 2, characterized in that: Conduct confidence assessment on the initial and secondary test results to determine the final test results, including: Compare the confidence levels of the initial test result and the secondary test result, and select the output with higher confidence level as the final test result.

9. The three-dimensional target detection method according to claim 2, characterized in that: Also includes: Determine whether the depth and rotation angle in the detection result match. If not, select the secondary detection result as the final detection result.

10. A three-dimensional target detection system, characterized in that: include: processor; A memory storing computer program instructions, which, when executed by the processor, implements a three-dimensional target detection method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Target detection method and system based on radar and image data fusion

    CN113267779A

  • Mechanical arm six-degree-of-freedom visual closed-loop grabbing method based on TSDF three-dimensional reconstruction

    CN114851201A

  • 3D point cloud target detection method based on attention mechanism and image feature fusion

    CN115115917A

  • Traffic target detection method and system based on cross-modal cross attention mechanism

    CN117173399A

  • Millimeter wave radar and camera fused 3D target detection method

    CN118362997A