A three-dimensional target detection method and system

Through multi-level cross-modal feature fusion and multi-level regression network, the problems of insufficient accuracy and robustness of target detection in tunnel environments are solved, high-precision target detection is achieved, and system cost and maintenance difficulty are reduced.

CN120182584BActive Publication Date: 2025-09-12SHIJIAZHUANG TIEDAO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510653992.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-12
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

In the tunnel rail vehicle unmanned driving environment perception scenario, the existing lidar and camera-millimeter wave radar cross-modal fusion solutions face problems such as high sensor deployment and maintenance costs, uneven lighting, and construction dust and noise interference, resulting in insufficient target detection accuracy and robustness.

Method used

A multi-level cross-modal feature fusion method is adopted to obtain point cloud data and image data, perform voxelization operations and multi-scale image feature extraction, use Swin Transformer and sparse convolution for feature fusion, and combine the moving window self-attention mechanism to optimize the cone depth range, construct a hierarchical cross-modal fusion architecture, and combine it with a multi-level regression network for target detection.

Benefits of technology

It effectively improves the accuracy and robustness of target detection in tunnel environments, reduces system costs and maintenance difficulties, and combines the stable ranging performance of millimeter-wave radar with the semantic recognition capabilities of the camera to improve detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182584B_ABST
    Figure CN120182584B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of image processing technology, and more specifically to a three-dimensional target detection method and system. The method comprises acquiring point cloud data and image data collected by a sensing device at the same time; performing a voxelization operation on the point cloud data to obtain voxel features; extracting multi-scale image features from the image data; fusing the voxel features with the multi-scale image features through channel splicing to obtain multi-level fusion features; and inputting the multi-level fusion features into a regression model to output target detection results. The regression model comprises a multiple regression network, wherein the detection results include one or more of the target's center position, size, posture, speed, and category. According to the solution of the present invention, the accuracy and robustness of target detection in tunnel rail vehicle scenarios are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and more particularly to a three-dimensional object detection method and system. Background Art

[0002] Unmanned transportation systems are primarily composed of three components: environmental perception, planning and decision-making, and control and execution. The environmental perception system is a key prerequisite for autonomous driving, determining whether subsequent decision-making and control systems have complete and accurate data as a basis.

[0003] Autonomous driving environmental perception systems typically rely on the coordinated operation of multiple sensors, with cameras, lidar, and millimeter-wave radar being the current mainstream sensor solutions. Each sensor type has its own advantages, but also inherent limitations. Cameras can provide high-resolution image information, capturing environmental details and enabling accurate object detection, recognition, and classification. However, their performance significantly degrades in low light, in adverse weather conditions such as rain, snow, fog, and haze, and in complex environments like tunnels and mines, making it difficult to ensure clear image quality.

[0004] LiDAR excels in complex environments with its superior range and angular resolution, maintaining high performance particularly in low-light conditions. However, its detection accuracy and effective range can drop significantly in extreme weather conditions (such as heavy rain, snow, and dense fog). Furthermore, LiDAR's high cost makes it less widely adopted in some low-cost applications.

[0005] Millimeter-wave radar, with its strong anti-interference capabilities and all-weather adaptability, can operate stably in complex lighting and harsh environmental conditions, accurately measuring the position and velocity of targets. However, due to its low resolution, millimeter-wave radar has limitations in its detection accuracy and ability to identify details of static objects.

[0006] In recent years, multimodal sensor fusion technology based on lidar and cameras has made significant progress in the field of autonomous driving. By fusing the high-precision three-dimensional point cloud of lidar with the rich texture information of the camera, it has been widely verified to achieve high-precision 3D target detection.

[0007] However, in tunnel rail vehicle driverless environment perception scenarios, such fusion solutions face the following engineering bottlenecks:

[0008] 1. Tunnel interiors have narrow cross-sections and complex support structures, making lidar susceptible to mechanical vibration and equipment collisions, significantly increasing sensor deployment and maintenance costs.

[0009] 2. The lighting conditions in the tunnel are usually affected, and intermittent lighting leads to uneven lighting conditions or sudden changes in lighting.

[0010] 3. Construction dust causes a lot of noise in the lidar point cloud. At the same time, the interference of metal reflective surfaces in the track switch area causes multiple echoes of the laser beam, further reducing the quality of the point cloud.

[0011] Under these conditions, the high acquisition cost of lidar and its performance degradation in harsh environments severely restrict its large-scale deployment in tunnels. Furthermore, the accuracy and robustness of camera-millimeter-wave radar cross-modal fusion in tunnel rail vehicle scenarios are insufficient.

[0012] Therefore, how to achieve accurate detection of targets in the harsh environment of mining areas has become the focus of current research. Summary of the Invention

[0013] The present invention addresses the problem of poor target detection accuracy in the harsh environment of mining areas and provides solutions in the following aspects.

[0014] In a first aspect, the present invention provides a three-dimensional target detection method, comprising: acquiring point cloud data and image data collected by a sensing device at the same time; performing a voxelization operation on the point cloud data to obtain voxel features; extracting image multi-scale features from the image data; fusing the voxel features with the image multi-scale features by channel splicing to obtain multi-level fusion features; inputting the multi-level fusion features into a regression model to output a detection result of the target, the regression model comprising a multiple regression network, wherein the detection result includes one or more of the center position, size, posture, speed and category of the target.

[0015] Preferably, the regression model includes a primary regression network and a secondary regression network, and the method further includes: inputting the multi-level fusion features into the primary regression network to obtain a primary detection result, wherein the primary detection result includes a 2D / 3D detection frame; performing visual cone association and radar point cloud mapping according to the 2D / 3D detection frame to determine the thermal map enhancement feature; inputting the thermal map enhancement feature into the secondary regression network to output the secondary detection result; and performing confidence evaluation on the primary detection result and the secondary detection result to determine the final detection result.

[0016] Preferably, the viewing cone association and radar point cloud mapping are performed according to the 2D / 3D detection frame to determine the heat map enhancement features, including: using the 3D detection frame and the camera intrinsic parameter matrix to generate a 3D viewing cone, and projecting the radar point cloud to the image 2D frame and 3D viewing cone to screen the valid point cloud; for each radar target associated with the 3D viewing cone, a three-channel heat map is generated according to its 2D detection frame position, which respectively represents the target center, boundary and direction information; the three-channel heat map is spliced ​​with the original image features to generate the heat map enhancement features.

[0017] Preferably, the calculation formula of the multi-level fusion feature is:

[0018]

[0019] Where, It is a multi-level fusion feature. is the radar signature, is the image feature, is the spatial weight, is the channel weight.

[0020] Preferably, the primary regression network adopts a joint loss function during training, including one or more of classification loss, regression loss and attribute regression loss.

[0021] Preferably, the loss functions for classification loss, regression loss, and attribute regression loss are:

[0022]

[0023]

[0024]

[0025] Where, is the classification loss, N is the total number of samples involved in the calculation, Represents the true heat map label, generated by a Gaussian kernel, with a value of 1 at the center and decaying by distance around it. α and β are hyperparameters of focal loss. It is the probability value of the (x, y) position and c category in the predicted heat map, is the regression loss, It is The true value of the dimensional regressor, It is The predicted value of the dimensional regressor, is the attribute regression loss, is the true value of the data.

[0026] Preferably, the depth range of the 3D viewing frustum is:

[0027]

[0028]

[0029]

[0030] Where, is the depth range of the visual cone, is the depth of the 3D detection frame output after the initial regression, It is an adjustment parameter. dep_max / dep_min are the upper and lower bounds of the target depth estimation based on the preliminary detection of the image. dist_thresh uses half of the depth range (dep_max - dep_min) as the reference threshold to define the initial expansion radius of the frustum in the depth direction. expansion_ratio is the expansion coefficient.

[0031] Preferably, a confidence evaluation is performed on the initial detection result and the secondary detection result to determine the final detection result, including: comparing the confidences corresponding to the initial detection result and the secondary detection result, and selecting the output with higher confidence as the final detection result.

[0032] Preferably, the method further includes determining whether the depth and rotation angle in the detection results match. If not, selecting the secondary detection result as the final detection result. Specifically, when determining whether a match exists, the secondary regression is a correction result after integrating radar features. Its confidence level incorporates physical information such as the radar's depth and velocity. If the difference in confidence between the two regression results (both regressions generate a Gaussian distribution of the keypoint heat map, reflecting the probability distribution of the target center point) exceeds a preset threshold, a conflict is considered to exist. If both dep and rot are present in both regression heads, only the second regression head with higher accuracy is used.

[0033] In a second aspect, the present invention further provides a three-dimensional target detection system, comprising: a processor; a memory storing computer program instructions, which, when executed by the processor, implements a three-dimensional target detection method according to the first aspect above.

[0034] The beneficial effects of this invention lie in: by performing multi-level cross-modal feature fusion of radar point cloud data and image data and utilizing a multi-level regression network for target detection, the accuracy and robustness of target detection in tunnel rail vehicle scenarios are effectively improved. Furthermore, by constructing a camera-millimeter-wave radar cross-modal fusion framework, the stable ranging and velocity measurement performance of millimeter-wave radar is effectively combined with the semantic recognition capabilities of the camera, significantly reducing system cost and maintenance while ensuring detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a flowchart of a three-dimensional object detection method according to an embodiment of the present invention;

[0036] Figure 2 is a flowchart of a method for processing multi-level fusion features using a regression network according to an embodiment of the present invention;

[0037] Figure 3 is a network structure diagram of three-dimensional object detection according to an embodiment of the present invention;

[0038] Figure 4 2 is a structural diagram of a three-dimensional target detection system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0039] The core improvement of this invention lies in the construction of a hierarchical cross-modal fusion architecture. Through the collaborative encoding of Swin Transformer and sparse convolution, multi-scale fusion of millimeter-wave radar voxel features and image window features is achieved, and the moving window self-attention mechanism is combined to dynamically optimize the cone depth range, effectively ensuring detection accuracy and recognition ability, and improving the robustness of the system.

[0040] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0041] Figure 1 is a flowchart of a three-dimensional object detection method 100 according to an embodiment of the present invention.

[0042] like Figure 1 As shown, in step S101, point cloud data and image data collected by the perception device at the same time are obtained.

[0043] Synchronous data acquisition: Synchronously read millimeter-wave radar point cloud data and camera image data, and align the sensor coordinate system through spatiotemporal calibration to ensure the spatiotemporal consistency of the data, thereby ensuring that the point cloud data and image data are aligned in time and space, that is, calibrating the external and internal parameters.

[0044] In step S102 , a voxelization operation is performed on the point cloud data to obtain voxel features.

[0045] Pillar Expansion encoding: This method converts sparse radar point clouds into dense 3D cylindrical voxels (similar to the PointPillars method), enhancing the spatial representation of targets and facilitating subsequent fusion with image features. Specifically, the original radar point cloud is divided into fixed-size 3D cylindrical voxels. Voxel encoding converts the discrete point cloud into a dense 3D cylindrical structure, increasing the volume representation of the target in three-dimensional space and improving the probability of subsequent cross-modal matching.

[0046] In step S103 , multi-scale image features are extracted from the image data.

[0047] In some embodiments, the multi-scale features of an image can be extracted by combining image feature extraction with high-resolution fusion: Swin Transformer is used as the backbone network to extract multi-scale features of an image, and a high-resolution feature map is generated in Stage 1-2 (the original Figure 1 / 4 and 1 / 8 sizes).

[0048] In step S104, voxel features are fused with multi-scale image features through channel stitching to obtain multi-level fused features. Specifically, the Pillar-encoded radar voxel features can be encoded using a sparse convolutional network into voxel features that are spatially aligned with the image feature map. Channel stitching fuses radar voxel features with image window features at different stages to obtain multi-level fused features from the camera and millimeter-wave radar.

[0049] At step S105, the multi-level fusion features are input into a regression model to output a detection result of the target. The regression model includes a multiple regression network, where the detection result includes one or more of the target's center position, size, posture, speed, and category. In some embodiments, the detection result can be processed by the multiple regression network and evaluated in combination with the confidence level of the detection result.

[0050] In the above-mentioned embodiments of the present invention, feature alignment between radar and image is achieved through a cross-modal fusion process, and parameter optimization is achieved through multi-stage regression processing, thereby effectively improving the accuracy of target detection results under harsh conditions.

[0051] Figure 2 FIG2 is a flow chart of a method 200 for processing multi-level fusion features using a regression network according to an embodiment of the present invention. The above-mentioned regression model includes a primary regression network and a secondary regression network.

[0052] like Figure 2 As shown, at step S201, the multi-level fused features are input into a primary regression network to obtain a primary detection result, which includes a 2D / 3D detection box. In some embodiments, based on the fused features, the regression network outputs a preliminary 2D detection box and a preliminary 3D detection box of the target, including the center position, size, depth, and observation angle of the 3D box. In this step, a rough 2D / 3D detection box can be predicted based on the fused features, thereby using the 2D / 3D detection box to filter and denoise the radar point cloud.

[0053] At step S202, the 2D / 3D detection frame is used to perform visual cone association and radar point cloud mapping to determine heatmap enhancement features. In some embodiments, a 3D visual cone can be generated using the 3D detection frame and the camera intrinsic parameter matrix, and the radar point cloud is projected onto the image 2D frame and 3D visual cone to screen for valid point clouds. Three heatmaps are generated for the target center, boundary, and orientation using the 2D frame, and these heatmaps are concatenated with the original image features to generate heatmap enhancement features.

[0054] The dynamic 3D frustum construction process can be implemented as follows: Using the 3D box parameters obtained from the initial regression and the camera intrinsic parameter matrix, a 3D frustum corresponding to each 3D box is generated. The frustum serves as a candidate region for quickly filtering valid radar point clouds and suppressing background noise and invalid matches. Specifically, the frustum association relies on the depth, observation angle alpha, and 3D box size dim of each 3D box obtained after the initial regression, combined with the camera calibration matrix to generate the 3D frustum. Similar to RoI extraction, this accelerates matching while filtering out invalid point clouds.

[0055] In one application scenario, (1) the radar Pillar is projected into the pixel coordinate system and spatially overlapped with the 2D detection box to screen out preliminary associated targets. (2) The Pillar is projected into the camera coordinate system and matched with the depth range of the 3D viewing cone to further filter out invalid point clouds that exceed the depth range of the viewing cone. (3) For each radar target associated with the viewing cone, three heatmaps are generated based on the position of its 2D detection box, representing the target center, boundary, and orientation information respectively. (4) The three-channel heatmap is concatenated with the original image features and input into the subsequent network as additional features to enhance the target context information.

[0056] At step S203, the heatmap enhancement features are input into a quadratic regression network to output a secondary detection result. In some embodiments, the depth, rotation angle, speed, and category attributes of the target are optimized by the quadratic regression network based on the enhanced features of the fused heatmap.

[0057] At step S204, the confidence level of the primary and secondary detection results is evaluated to determine the final detection result. In some embodiments, the primary and secondary regression results are combined to select the output with the higher confidence level as the final detection result. For conflicting parameters (such as depth and rotation angle), the secondary regression result is prioritized to output the target's 3D position, size, pose, velocity, and category.

[0058] In summary, this embodiment demonstrates that dynamic frustum construction allows for real-time adjustment of the search space based on preliminary predictions, effectively improving search efficiency. Furthermore, the stitching of three-channel heatmaps reveals target geometry, alleviating positioning ambiguity caused by sparse radar point clouds. Furthermore, dual-stage regression and confidence fusion operations effectively enhance detection accuracy.

[0059] Next, the above solution will be described in detail in conjunction with the calculation process in a specific embodiment. Figure 3 3D target detection network diagram according to an embodiment of the present invention.

[0060] Step 1: Point cloud and image data are heterogeneous sensor data, so the collected sparse point cloud data and image data need to be processed synchronously in time and space. Figure 3 As shown, the original image and point cloud data are read, the image data is normalized, the point cloud data is preprocessed, and the point cloud data is voxelized to obtain training data.

[0061] Step 2: Multi-level cross-modal feature fusion. The following steps are used to extract backbone network features and fuse them into multi-level fusion features to generate a multi-level cross-modal feature map.

[0062] (1) In the Stage 1-2 high-resolution features of Swin Transformer, the radar point cloud is encoded into voxel features through sparse convolution and fused with the image window features through channel splicing.

[0063] (2) In Stage 3-4, HiFuse’s hierarchical feature fusion block (HFF) is used to perform spatial-channel attention weighting on the radar and image features and then fuse them. The spatial attention generates a spatial weight map through 3×3 convolution. The formula is:

[0064]

[0065] Where, is the radar signature, is the image feature, is the spatial weight. The spatial weight is determined by the radar characteristics and image features The result is obtained by element-wise addition, 3×3 convolution, and Sigmoid activation.

[0066] Channel attention uses SE module (compression ratio = 16), formula:

[0067]

[0068] Channel weight By radar characteristics and image features The two features are added together, followed by global average pooling (GAP), MLP (multi-layer perceptron, whose structure includes two fully connected layers and a bottleneck design), and finally Sigmoid activation, that is, using an SE module.

[0069] Finally, the weighted fusion output is performed to obtain the multi-level fusion features of the camera and millimeter-wave radar. It is a multi-level fusion feature, and the formula is:

[0070]

[0071] Step 3: Preliminary regression and 3D frustum generation.

[0072] For the fused features, the initial regression network outputs the preliminary 2D detection frame and 3D detection frame of the target, including the center position, size, depth and observation angle of the 3D frame.

[0073] The frustum association relies on the depth depth, observation angle alpha, and 3D box size dim of each 3D Box obtained after the initial regression, plus the camera calibration matrix to generate a 3D frustum, similar to RoI extraction, which speeds up the matching speed while filtering out invalid point clouds.

[0074] A joint loss function is used in model training to ensure balanced optimization of each subtask. It mainly includes the following parts:

[0075] 1) Classification (heatmap) loss.

[0076] To accurately locate the center of the target, the network output heatmap is normalized and the loss is calculated using FastFocalLoss, which is based on the concept of Focal Loss. This loss function assigns different weights to positive and negative samples, focusing on the difficult-to-detect center point, effectively alleviating the sample imbalance problem. The main function of Focal Loss is to reduce the impact of the background area on the loss, allowing the network to focus more on detecting the center of the target. The specific formula is as follows:

[0077]

[0078] Where N is the total number of samples involved in the calculation, Y represents the true heatmap label, which is generated by a Gaussian kernel with a value of 1 at the center and decaying with distance around it, and α and β are hyperparameters of the focal loss.

[0079] 2) Regression loss.

[0080] For most regression heads, L1 loss (Mean Absolute Error, MAE) is used. L1 loss calculates the difference between the predicted value and the true value and has good robustness. In particular, L1 loss can effectively reduce the impact of outliers on model training.

[0081]

[0082] Where: is the true value of the data, and N is the number of samples in the regression task.

[0083] 3) Attribute regression loss.

[0084] For the attribute regression head (such as the depth and speed of the target), Binary Cross Entropy (BCE) Loss is used. The formula of BCE Loss is:

[0085]

[0086] Where: is the true value of the data, and N is the number of samples in the regression task.

[0087] Step 4: Secondary matching and feature enhancement.

[0088] Based on the camera calibration matrix (intrinsic and extrinsic parameters), the target's 2D bounding box, and depth estimation, a 3D frustum is generated for each detected target. This frustum is the target's projection in 3D space, covering the space from the camera's optical center to the target's depth range. It is used to filter possible related radar point clouds. The depth threshold calculation formula is:

[0089]

[0090] Adjustment parameters The formula is:

[0091]

[0092] The frustum depth range is:

[0093]

[0094] Where dep_max / dep_min are the upper and lower bounds of the target depth estimated based on the initial image detection, and dist_thresh uses half of the depth range (dep_max - dep_min) as the reference threshold to define the initial expansion radius of the frustum in the depth direction. expansion_ratio is the expansion coefficient, which is used to adjust the parameters. By scaling up dist_thresh, the search range in the depth direction is increased.

[0095] Step 5: Two-stage regression and results synthesis.

[0096] (1) Based on the enhanced features of the fused Heatmap, the depth, rotation angle, speed, and category attributes of the target are optimized through a quadratic regression network. The 3D frustum is similar to a ROI (region of interest) extraction. The depth depth, observation angle alpha, and 3D box size dim of each 3D box obtained by the Primary Regression Heads are combined with the camera calibration matrix to generate the 3D frustum. By filtering the point cloud, only the point cloud inside the frustum is left. The point cloud is then spliced ​​with the multi-level fusion features for quadratic regression.

[0097] (2) Combining the results of the primary and secondary regression, the output with higher confidence is selected as the final detection result. For conflicting parameters (such as depth and rotation angle), the secondary regression result is used first to output the target's 3D position, size, attitude, speed and category. The secondary regression is the correction result after integrating radar features. Its confidence combines the radar's depth, speed and other physical information. If the difference in confidence between the two regression results (both regressions generate the Gaussian distribution of the key point heat map, reflecting the probability distribution of the target center point) exceeds the preset threshold, it is considered that there is a conflict.

[0098] The present invention effectively preserves the geometric structure information of the radar point cloud and the texture features of the image through radar and image feature fusion. The two are fused after spatial alignment, which makes up for the inherent defects of a single sensor and effectively improves the accuracy of the detection process. At the same time, it enhances local representation through multi-scale feature fusion and performs error correction through multi-stage regression, which strengthens the anti-interference ability of the target detection method and improves the robustness of the target detection process in harsh environments.

[0099] Figure 4 2 is a structural diagram of a three-dimensional target detection system according to an embodiment of the present invention.

[0100] The present invention also provides a three-dimensional target detection system. Figure 4 As shown, the system includes a processor and a memory, wherein the memory stores computer program instructions. When the computer program instructions are executed by the processor, a three-dimensional target detection method according to the above embodiment is implemented.

[0101] In the description of this specification, "multiple" and "several" mean at least two, such as two, three or more, etc., unless otherwise clearly defined.

[0102] While several embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous modifications, variations, and alternatives will occur to those skilled in the art without departing from the concept and spirit of the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in practicing the present invention.

Claims

1. A three-dimensional target detection method, characterized in that: include: Obtain point cloud data and image data collected by the sensing device at the same time; Perform voxelization on the point cloud data to obtain voxel features; Extracting multi-scale image features from image data; The voxel features are fused with the multi-scale features of the image through channel splicing to obtain multi-level fusion features; Inputting the multi-level fusion features into a regression model to output a detection result of the target, wherein the regression model includes a multiple regression network, wherein the detection result includes one or more of the center position, size, posture, speed and category of the target; The regression model includes a primary regression network and a secondary regression network, and the method further includes: Inputting the multi-level fusion features into the primary regression network to obtain a primary detection result, wherein the primary detection result includes a 2D detection box and a 3D detection box; Perform frustum association and radar point cloud mapping based on 2D and 3D detection boxes to determine heatmap enhancement features; Input the heat map enhanced features into the quadratic regression network to output the secondary detection results; Conduct confidence assessment on the initial test results and the secondary test results to determine the final test results; The calculation formula of the multi-level fusion feature is: Where, It is a multi-level fusion feature. is the radar signature, is the image feature, is the spatial weight, is the channel weight.

2. The three-dimensional target detection method according to claim 1, characterized in that: Perform frustum association and radar point cloud mapping based on 2D and 3D detection boxes to determine heatmap enhancement features, including: Generate a 3D viewing cone using the 3D detection frame and the camera intrinsic parameter matrix, and project the radar point cloud onto the image 2D frame and 3D viewing cone to filter valid point clouds; For each radar target associated with the 3D frustum, a three-channel heat map is generated based on the position of its 2D detection box, representing the target center, boundary, and direction information respectively. The three-channel heatmap is concatenated with the original image features to generate heatmap enhanced features.

3. The three-dimensional target detection method according to claim 1, characterized in that: The initial regression network adopts a joint loss function during training, including one or more of classification loss, regression loss and attribute regression loss.

4. The three-dimensional target detection method according to claim 3, characterized in that: The loss functions for classification loss, regression loss, and attribute regression loss are: Where, is the classification loss, N is the total number of samples involved in the calculation, Represents the true heat map label, generated by a Gaussian kernel, with a value of 1 at the center and decaying by distance around it. α and β are hyperparameters of focal loss. It is the probability value of the (x, y) position and c category in the predicted heat map, is the regression loss, It is The true value of the dimensional regressor, It is The predicted value of the dimensional regressor, is the attribute regression loss, is the true value of the data.

5. The three-dimensional target detection method according to claim 2, characterized in that: The depth range of the 3D viewing frustum is: Where, is the depth range of the visual cone, is the depth of the 3D detection frame output after the initial regression, It is an adjustment parameter. dep_max / dep_min are the upper and lower bounds of the target depth estimation based on the preliminary detection of the image. dist_thresh uses half of the depth range (dep_max - dep_min) as the reference threshold to define the initial expansion radius of the frustum in the depth direction. expansion_ratio is the expansion coefficient.

6. The three-dimensional target detection method according to claim 1, characterized in that: Conduct confidence assessment on the initial and secondary test results to determine the final test results, including: Compare the confidence levels of the initial and secondary detection results, and select the output with higher confidence as the final detection result.

7. The three-dimensional target detection method according to claim 1, characterized in that: Also includes: Determine whether the depth and rotation angle in the detection result match. If not, select the secondary detection result as the final detection result.

8. A three-dimensional target detection system, characterized in that: include: processor; A memory storing computer program instructions, which, when executed by the processor, implements a three-dimensional target detection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • 3D point cloud target detection method based on attention mechanism and image feature fusion

    CN115115917A

  • Robustness target detection method and system based on feature level Raysight fusion

    CN118429924A