Monocular 3D object detection method with coordinate sensing position embedding

By introducing CAPE module and 3D coordinate-aware position coding in the monocular 3D detection framework, the problem of ignoring spatial information and local optimization in the prior art is solved, and better monocular 3D object detection performance is achieved.

CN119942515APending Publication Date: 2025-05-06SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311803779.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing monocular 3D detection algorithm ignores the spatial information in 3D coordinates, and ignores the overall importance when optimizing each attribute of 3D detection, resulting in poor performance.

Method used

A new 3D detection framework is proposed, which provides position encoding for coordinate-guided feature sets through the CAPE module, combines the optimization of various attributes of the bounding box to avoid local optimization, and solves coordinates in 3D world space through label prediction instance-level depth maps and camera parameters.

Benefits of technology

This method can effectively improve the performance of the monocular 3D object detector, and realize the improvement of AP3D and APBEV on the KITTI data set, proving its effectiveness without requiring additional auxiliary information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942515A_ABST
    Figure CN119942515A_ABST
Patent Text Reader

Abstract

The invention relates to a monocular 3D object detection method with a coordinate sensing position embedding function, and the method comprises the steps: coding the position information of 3D coordinates into a feature map in a depth discretization manner, enabling a Transform to have the 3D coordinate sensing capability, and maintaining the lower overhead. Specifically, firstly, an instance-level foreground depth map is predicted, then coordinates of a three-dimensional world are obtained through camera parameters and serve as position codes to be added to a decoder to interact with object query, and the capacity of sensing pixel information by the three-dimensional position is enhanced. And finally estimating the 3D attribute of each target through object query. Besides, a joint optimization strategy is provided in the training stage, prediction outputs of different attributes are considered as a whole through Gaussian modeling, and the training process of the network is guided. A large number of experiments on a KITTI data set prove the effectiveness of the method. Compared with the latest method, the method has the advantages that the measurement of AP3D at the hard level is improved, and the measurement of APBEV at the easy / medium / hard setting is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and digital image processing, and in particular to a monocular 3D object detection method with coordinate-aware position embedding. Background Art

[0002] Autonomous driving is a challenging task that has attracted extensive research recently. In the field of autonomous driving, monocular 3D detection is an important and promising task because monocular sensors can avoid expensive and complex corrections. However, due to the lack of depth cues, recovering the attributes of objects in 3D space from a single image is extremely challenging.

[0003] Most existing methods still have the following shortcomings: (1) they ignore the spatial information in the 3D coordinates, which is important for the model to understand the relative relationship of objects in the scene; (2) they optimize each attribute of the 3D bounding box individually, but ignore the importance of the whole, resulting in suboptimal final results. Although some methods try to balance the spatial cues by introducing depth compensation, the information contained in a single dimension is still not enough to achieve ideal performance. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention provides a monocular 3D object detection method with coordinate-aware position embedding, which overcomes the shortcomings of the existing monocular camera 3D detection algorithm that ignores the spatial information in the 3D coordinates and optimizes each attribute of 3D detection but ignores the importance of the whole.

[0005] This paper proposes a new 3D detection framework that fully explores the 3D spatial information of objects to achieve a better understanding of the scene and jointly optimizes the various properties of the bounding box to avoid local optimization. Position encoding is provided for the coordinate-guided feature set through the CAPE module. In order to save computing resources and avoid system latency, the instance-level depth map is first predicted using the label, and then the coordinates in the 3D world space are solved using the camera parameters. Then, the 3D coordinate-aware position encoding is generated in a simple MLP and interacts with the object query in the decoder. This approach simultaneously avoids overconfidence and inaccurate priors caused by the introduction of off-the-shelf depth estimators. Extensive experiments on the KITTI dataset verify the effectiveness of the model.

[0006] The technical solution adopted by the present invention to achieve the above-mentioned purpose is: a monocular 3D object detection method with coordinate-aware position embedding, comprising the following steps:

[0007] Pass a single image through the backbone network to extract image features;

[0008] The image features are extracted through the encoder; at the same time, the image features are input into the CAPE generator through convolution for coordinate-aware position embedding to obtain the coordinate-aware position;

[0009] The features extracted by the encoder and the coordinate-aware positions obtained by the CAPE generator are sent to the decoder for feature aggregation, and the target information is obtained through the detection head.

[0010] The backbone network is ResNet-50.

[0011] The image features are input into the CAPE generator through convolution for coordinate-aware position embedding, including the following steps:

[0012] The image features output by the backbone network are passed through the convolutional layer to obtain the distribution probability confidence D pro ;

[0013] According to the depth value d used to represent the i-th target in the current image i , combined with bin index d c The relationship between the bin index d and the depth value is obtained c ;

[0014] According to the sentence distribution probability confidence D pro and the corresponding bin index d c Weighted summation to obtain the estimated depth map D map ;

[0015] Depth Map D map Each point in is represented as p(u, v, d), where u and v represent pixel coordinates and d is the depth value along the axis orthogonal to the image plane; the normalized coordinates in the camera coordinate system are obtained from the intrinsic parameter matrix;

[0016] Calculate the 3D coordinate p from the normalized coordinates in the camera coordinate system and the depth value 3d (x w ,y w ,z w ):

[0017]

[0018] Among them, x c ,y c 、z c represents the normalized coordinates in the camera coordinate system; R represents the rotation matrix, and t represents the translation matrix;

[0019] The 3D coordinates Feed it into a multi-layer perceptron to obtain the 3D coordinate perception position; where H and W represent the height and width of the feature respectively.

[0020] The distribution probability confidence It is used to indicate the confidence that the depth value of each pixel belongs to a certain depth bin. H and W represent the height and width of the feature respectively.

[0021] The bin index d c Obtained by the following formula:

[0022]

[0023] d c Indicates the bin index, d max d min are the upper and lower limits of the depth value respectively; D is the set value, d i is the depth value of the i-th target in the current image.

[0024] The normalized coordinates in the camera coordinate system are obtained from the intrinsic parameter matrix, where z c =1;

[0025]

[0026] x c ,y c 、z c represents the normalized coordinates in the camera coordinate system; K represents the intrinsic parameter matrix; u and v represent pixel coordinates.

[0027] The target information includes category, size, direction and position.

[0028] A monocular 3D object detection device with coordinate-aware position embedding comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement the monocular 3D object detection method with coordinate-aware position embedding when executing the computer program.

[0029] A computer-readable storage medium, characterized in that a computer program is stored on the storage medium, and when the computer program is executed by a processor, the monocular 3D object detection method with coordinate-aware position embedding is implemented.

[0030] The present invention has the following beneficial effects and advantages:

[0031] This paper proposes a new monocular 3D object detection network using the CAPE generator. Comparative experiments on the KITTI dataset show that this method can effectively improve the performance of the monocular 3D object detector without the need for additional auxiliary information. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1This is a structural diagram of the coordinate-aware position-embedded monocular 3D object detection network model of the present invention. DETAILED DESCRIPTION

[0033] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0034] The present invention relates to a monocular 3D object detection method with coordinate-aware position embedding. By encoding the position information of 3D coordinates into a feature map in a deep discretized manner, the Transformer has the ability to perceive 3D coordinates while maintaining low overhead. Specifically, the instance-level foreground depth map is first predicted, and then the coordinates of the three-dimensional world are obtained through camera parameters as position encoding and added to the decoder to interact with object queries, enhancing the ability to perceive pixel information in three-dimensional positions. Finally, the 3D properties of each target are estimated through object queries. A large number of experiments on the KITTI dataset have demonstrated the effectiveness of this method. Compared with the latest methods, this method improves the AP3D metric at the hard level and the APBEV metric at the easy / medium / hard settings.

[0035] Specifically, the neural network is constructed using an attention mechanism and an encoder-decoder structure. Given an image W×H×3 as input, the detector uses Resnet-50 as the backbone network to extract image features. By predicting the instance-level depth map in the pipeline and solving the coordinate information from it, the coordinate-aware position embedding is obtained through the CAPE generator. The output of the encoder is then fed to the coordinate-guided decoder together with the position embedding, and valid information is extracted by interacting with the object query. The updated object query passes through the detection head to output the prediction information.

[0036] Figure 1 It is a structural diagram of the coordinate-aware position-embedded monocular 3D object detection network model of the present invention;

[0037] This paper proposes a new 3D detection framework that fully explores the 3D spatial information of objects to achieve a better understanding of the scene and jointly optimizes the various properties of the bounding box to avoid local optimization. Specifically, the model is based on a codec and mainly contains a coordinate-aware position embedding (CAPE) generator. The CAPE module provides position encoding for coordinate-guided feature sets. In order to save computing resources and avoid system latency, we first use labels to predict instance-level depth maps, and then use camera parameters to solve the coordinates in 3D world space. Then, 3D coordinate-aware position encodings are generated in a simple MLP and interact with object queries in the decoder. This approach simultaneously avoids overconfidence and inaccurate priors caused by the introduction of off-the-shelf depth estimators. The specific steps are as follows:

[0038] Step 1: Use the attention mechanism and encoder-decoder structure to build a neural network.

[0039] Input in Monocular 3D Object Detection System with Coordinate-Aware Position Embedding Experiments are conducted on the KITTI 3D Object Detection Dataset, which includes 7481 images for training and 7518 images for testing. The training samples are divided into training set (3712) and validation set (3769).

[0040] Step 2: A single image W×H×3 is taken as input, and the detector uses ResNet-50 as the backbone network to extract image features.

[0041] The global parameters are set as follows: epoch = 150, batch_size = 8. The AdamW optimizer is used, the weight decay is, and the learning rate decays by 0.1 times at 65 and 100 epochs.

[0042] Step 3: Predict the instance-level depth map within the pipeline and solve the coordinate information from it.

[0043] The instance-level depth map is predicted using a lightweight predictor and the 2D bounding box information within the label. The depth map is converted to 3D coordinates through the camera parameters, avoiding the dense transformation calculation of the grid and obtaining 3D features in a more sparse manner.

[0044] Step 4: Obtain coordinate-aware position embedding through CAPE generator;

[0045] The continuous depth range is discretized into k+1 bins, which contain K types of foreground and background depth. A simple convolutional layer is used after the feature map output by the backbone network to obtain the distribution probability of the depth bin. This represents the confidence that the depth value of each pixel belongs to a certain depth bin. The depth bins are formulated using Linear Increasing Discretization (LID). The relationship between the bin index and the depth value is as follows:

[0046]

[0047] Estimated Depth Map By distribution probability confidence D pro The depth map D is obtained by weighted summing of the depth of the corresponding bin. map Each point in can be represented as p(u, v, d), where (u, v) are the coordinates of the pixel in the image and d is the depth value along the axis orthogonal to the image plane. The normalized coordinates in the camera coordinate system are obtained from the intrinsic matrix, where z c =1:

[0048]

[0049] The encoder output is then fed into the coordinate-guided decoder position embedding to extract effective information by interacting with the object query.

[0050] Use the normalized camera coordinates and depth to calculate the 3D coordinate p 3d (x w ,y w ,z w ):

[0051]

[0052] The 3D coordinates Feed into a multi-layer perceptron to obtain 3D coordinate-aware position embedding (CAPE).

[0053] Step 6: The output of the encoder and CAPE is fed into the decoder.

[0054] The position embedding and the encoder output are fused through a cross-attention mechanism, so that object queries can perceive the 3D position of the object and interact and update it.

[0055] Step 7: Construct the loss function

[0056] Different loss functions are defined and calculated to optimize the four attributes of the detection head output. These attributes include category, size, orientation, and depth. The choice and design of the loss function is to make the predicted 3D bounding box closer to the real 3D bounding box. Specifically, the goal of the 3D center branch is to directly predict the projection coordinates of the object's 3D center point on the image, rather than the usual offset prediction. The center is estimated using the L1 loss, denoted as There are four branches to predict the category, size, direction and depth of the object. We apply Focal loss, dimension-aware 3D IoU loss, MultiBin loss and uncertainty loss as the loss functions of the above branches, and they are respectively denoted as and

[0057] Step 8: The features are passed through the detection head to output the predicted box.

[0058] The image feature information is input into the decoder for feature integration and then sent to the detection head for prediction information.

[0059] Step 9: 3D object detection.

[0060] After training the monocular 3D object detection neural network model based on coordinate-aware position embedding, the input is autonomous driving image data in different scenarios, and the output is the category, size, orientation and position after 3D detection.

Claims

1. A monocular 3D object detection method with coordinate-aware position embedding, characterized by: The following steps are involved: Pass a single image through the backbone network to extract image features; The image features are extracted through the encoder; at the same time, the image features are input into the CAPE generator through convolution for coordinate-aware position embedding to obtain the coordinate-aware position; The features extracted by the encoder and the coordinate-aware positions obtained by the CAPE generator are sent to the decoder for feature aggregation, and the target information is obtained through the detection head.

2. A monocular 3D object detection method with coordinate-aware position embedding according to claim 1, characterized in that: The backbone network is ResNet-50.

3. The monocular 3D object detection method with coordinate-aware position embedding according to claim 1, characterized in that: The image features are input into the CAPE generator through convolution for coordinate-aware position embedding, including the following steps: The image features output by the backbone network are passed through the convolutional layer to obtain the distribution probability confidence D pro ; According to the depth value d used to represent the i-th target in the current image i , combined with bin index d c The relationship between the bin index d and the depth value is obtained c ; According to the sentence distribution probability confidence D pro and the corresponding bin index d c Weighted summation to obtain the estimated depth map D map ; Depth Map D map Each point in is represented as p(u, v, d), where u and v represent pixel coordinates and d is the depth value along the axis orthogonal to the image plane; the normalized coordinates in the camera coordinate system are obtained from the intrinsic parameter matrix; Calculate the 3D coordinate p from the normalized coordinates in the camera coordinate system and the depth value 3d (x w ,y w ,z w ): Among them, x c ,y c 、z c represents the normalized coordinates in the camera coordinate system; R represents the rotation matrix, and t represents the translation matrix; The 3D coordinates Feed it into a multi-layer perceptron to obtain the 3D coordinate perception position; where H and W represent the height and width of the feature respectively.

4. The monocular 3D object detection method with coordinate-aware position embedding according to claim 1, characterized in that: The distribution probability confidence It is used to indicate the confidence that the depth value of each pixel belongs to a certain depth bin. H and W represent the height and width of the feature respectively.

5. The monocular 3D object detection method with coordinate-aware position embedding according to claim 1, characterized in that: The bin index d c Obtained by the following formula: d c Indicates the bin index, d max ,d min are the upper and lower limits of the depth value respectively; D is the set value, d i is the depth value of the i-th target in the current image.

6. The monocular 3D object detection method with coordinate-aware position embedding according to claim 1, characterized in that: The normalized coordinates in the camera coordinate system are obtained from the intrinsic parameter matrix, where z c =1; x c ,y c 、z c represents the normalized coordinates in the camera coordinate system; K represents the intrinsic parameter matrix; u and v represent pixel coordinates.

7. The monocular 3D object detection method with coordinate-aware position embedding according to claim 1, characterized in that: The target information includes category, size, direction and position.

8. A monocular 3D object detection device with coordinate-aware position embedding, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement a monocular 3D object detection method with coordinate-aware position embedding as described in any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, a monocular 3D object detection method with coordinate-aware position embedding as described in any one of claims 1 to 6 is implemented.