Monocular three-dimensional target detection method based on geometric appearance perception and feature fusion

By adopting a method based on geometric appearance perception and feature fusion in autonomous driving scenarios, combined with multiple estimation modules for feature extraction and loss calculation, the shortcomings of three-dimensional attributes and dimension estimation in monocular three-dimensional object detection are solved, and high-precision three-dimensional object detection is achieved.

CN119942514APending Publication Date: 2025-05-06SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311803778.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the detection of single-object three-dimensional targets in autonomous driving scenarios, it is difficult to effectively solve the estimation problems of three-dimensional attributes and three-dimensional dimensions, resulting in insufficient detection accuracy.

Method used

Using a monocular three-dimensional object detection method based on geometric appearance perception and feature fusion, depth features are extracted through the CNN skeleton, and combined with the central projection offset, three-dimensional dimension estimation, key point compensation, direction estimation and depth estimation modules, the calculation and object detection of multi-task losses are realized.

Benefits of technology

High-precision three-dimensional object detection of the input image is realized, the shortcomings of the existing methods in three-dimensional attributes and dimension estimation are overcome, and the performance performance of the monocular three-dimensional object detection algorithm is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942514A_ABST
    Figure CN119942514A_ABST
Patent Text Reader

Abstract

The invention discloses a monocular three-dimensional target detection method based on geometric appearance perception and feature fusion. Comprising a feature extraction module used for extracting deep features of a monocular RGB image; the two-dimensional target detection module assists the network in better learning three-dimensional related features through regression of two-dimensional related attributes; the central projection offset estimation module considers that a coordinate discretization error exists in down-sampling, and obtains a more accurate central point by regressing the offset; the three-dimensional dimension estimation module learns the size of a three-dimensional bounding box in the real world through sample perception feature fusion; the key point compensation estimation module obtains the offset of each angular point relative to the target center point through the depth of each angular point; the direction estimation module predicts a local angle of a yaw direction; the depth estimation module converts a network output depth into an absolute depth. According to the method, geometric constraints of the key points of the object are added into direction estimation of the object through a new appearance feature extraction strategy, and the feature fusion module is introduced to optimize three-dimensional size estimation, so that the performance of the monocular detector in a three-dimensional target detection task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine vision technology, and in particular to a monocular three-dimensional target detection method based on geometric appearance perception and feature fusion. Background Art

[0002] Object detection in autonomous driving scenarios is one of the important research directions of machine vision, which is completed by complex visual sensor systems (such as lidar sensors, stereo sensors, monocular sensors, etc.). The latest progress in autonomous driving is to use monocular sensors to achieve efficient 3D object detection tasks with geometric constraints. These detectors improve explicit geometric projections and build a bridge between two-dimensional images and the three-dimensional world.

[0003] Such research often tends to focus on optimizing depth estimation while ignoring the equally important directional 3D properties and 3D dimensionality issues. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention provides a monocular 3D target detection method based on geometric appearance perception and feature fusion, which can achieve high-precision 3D target detection of input images, overcome the shortcomings of current methods in the three-dimensional attributes of direction and three-dimensionality, and can be widely used in autonomous driving, smart transportation, smart parking systems and other fields.

[0005] The technical solution adopted by the present invention to achieve the above-mentioned purpose is:

[0006] A monocular 3D target detection method based on geometric appearance perception and feature fusion comprises the following steps:

[0007] The feature extraction module extracts the deep features of the monocular RGB image through the CNN skeleton;

[0008] The center projection offset module regresses and estimates the offset of the deep feature to obtain the coordinates of its center point;

[0009] The 3D dimension estimation module estimates the size of the 3D bounding box of the object in the image through the depth features of the image;

[0010] The key point compensation estimation module calculates the depth of the three-dimensional corner points based on the depth features of the image;

[0011] The direction estimation module estimates the direction of the 3D bounding box through the depth of the 3D corner points;

[0012] The depth estimation module performs depth estimation on the depth features;

[0013] The Loss module calculates the loss of each module, and calculates the multi-task loss based on each loss to complete target detection.

[0014] The feature extraction module performs the following steps:

[0015] Using a three-channel RGB image I∈R through a CNN skeleton W×H×3 Takes as input and outputs a global feature map Where d is the downsampling coefficient, W and H are the width and height of the image respectively, and R is a real number;

[0016] A 3×3 convolutional layer and a 1×1 convolutional layer are sequentially deployed in each detection head;

[0017] Add a 2D detection branch as an auxiliary network;

[0018] DLA-34 is used as the backbone network, and two sub-detection heads are set according to FCOS to perform two-dimensional attribute regression, where one sub-head outputs a heat map h and the other outputs a two-dimensional distance d 2d .

[0019] The center projection offset module performs the following steps:

[0020] In the downsampling process of the backbone network, the offset is regressed to obtain the center coordinate x of the object's three-dimensional bounding box o .

[0021] The three-dimensional dimension estimation module performs the following steps:

[0022] Based on the global feature map F, a high-dimensional feature is generated and a sample-aware filter

[0023]

[0024] Among them, G is a 3×3 convolutional layer, is the parameter of G;

[0025] Based on the global feature map F, an intermediate feature is generated through a 3×3 kernel.

[0026] C f Each pixel in is reconstructed into a 3×3 convolution kernel, and features are extracted through convolution operation. Right now:

[0027] F i2 =F i1 *C f ;

[0028] F i2 and F mid Through the attention module, the number of fusion feature channels is compressed from 512 to 256 to obtain the feature map

[0029] The final estimate D of the three-dimensional dimension is generated through the Conv-BN-ReLU module 3d .

[0030] The key point compensation estimation module performs the following steps:

[0031] For 8 predefined key points x k ={x k1 ,x k2 ,…,x k8} Projection from 8 3D angles;

[0032] Divide the 8 key points into 4 groups along the vertical edges of the 3D bounding box {(x k1 ,x k5 ),…,(x k4 ,x k8 )}, each group is composed of three-dimensional height h 3d The two-dimensional height h obtained by projection 2d ={h1,h2,h3,h4};

[0033] Solve the depth z of the 3D corner point k ={z k1 ,z k2 ,z k3 ,z k4}:

[0034] z k =h 3d f / h 2d

[0035] Where f is the focal length.

[0036] The direction estimation module performs the following steps:

[0037] The depth of the key points is divided into two groups according to the front and back relationship of the key points, and the difference between each group is calculated Δz = {z1-z4, z2-z3};

[0038] If the difference is less than 0, the front surface is visible, otherwise the back surface is visible;

[0039] Use the same method of judging the difference to judge the visibility of the upper and lower surfaces and the left and right surfaces;

[0040] After obtaining the three visible surfaces of the object, they are used to regress θ l1 Feature map Extract the corresponding region F s , and transform the region through perspective transformation to obtain a unified feature map F u :

[0041] F u =P*F s

[0042] Among them, P is composed of F u The perspective transformation matrix obtained by solving the four corners and their corresponding coordinates;

[0043] The unified feature maps of each target are concatenated and the number of channels is compressed from 768 to 256 using a 3×3 convolutional layer;

[0044] Output θ through the sequentially connected Conv-BN-ReLU module and AvgPool module l2 ;

[0045] By adding the target level l2 With θ l1 Add together to get the final estimated result θ l .

[0046] The depth estimation module performs the following steps:

[0047] Use the sigmoid function to reduce the network output depth Convert to absolute depth z;

[0048] Adding uncertainty to a branch prediction joint optimization with a depth term

[0049] The Loss module performs the following steps:

[0050] Calculate the loss of each module separately, including: heat map loss L h , two-dimensional distance loss L 2d , center shift loss L c , key point offset loss L k , 3D angular depth loss L kd , three-dimensional size loss L 3d , direction loss L o and the depth loss L d ;

[0051] According to the definition of each loss, calculate the multi-task loss L:

[0052] L=α1L h +α2L 2d +α3L c +α4L k +α5L kd +α6L 3d +α7L o +α8L d

[0053] Among them, α1, α2, α3, α4, α5, α6, α7 and α8 are loss coefficients.

[0054] A monocular 3D object detection method based on geometric appearance perception and feature fusion, comprising:

[0055] Feature extraction module, used to extract deep features of monocular RGB images through CNN skeleton;

[0056] The center projection offset module is used to regress and estimate the offset of the deep feature to obtain the coordinates of its center point;

[0057] A three-dimensional dimension estimation module is used to estimate the size of the three-dimensional bounding box of the object in the image through the depth features of the image;

[0058] The key point compensation estimation module is used to calculate the depth of the three-dimensional corner points according to the depth characteristics of the image;

[0059] A direction estimation module is used to estimate the direction of a 3D bounding box based on the depth of 3D corner points;

[0060] A depth estimation module, used for performing depth estimation on depth features;

[0061] The Loss module is used to calculate the losses of each module and calculate the multi-task loss based on each loss to complete target detection.

[0062] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, a monocular three-dimensional target detection method based on geometric appearance perception and feature fusion is implemented.

[0063] The present invention has the following beneficial effects and advantages:

[0064] The present invention realizes high-precision 3D target detection of input images through the method of geometric appearance perception and feature fusion, overcomes the shortcomings of existing methods in estimating the 3D attributes of direction and 3D dimension, and can improve the performance of monocular 3D target detection algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a schematic diagram of the algorithm structure of the present invention;

[0066] Figure 2 is a schematic diagram of a three-dimensional dimension estimation module;

[0067] Figure 3 It is a schematic diagram of direction estimation. DETAILED DESCRIPTION

[0068] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0069] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0070] A monocular 3D target detection method based on geometric appearance perception and feature fusion, comprising: using a three-channel RGB image I∈R W×H×3 Takes as input and outputs a global feature map Where d is the downsampling coefficient. After that, each detection head deploys a 3×3 convolution layer and a 1×1 convolution layer respectively. The 3×3 convolution kernel is used to expand the feature channels from 64 to 256, and the 1×1 convolution kernel is used to regress 2D and 3D target attributes. The 2D detection branch is then used as an auxiliary to help the network better learn 3D related features. DLA-34 is used as the backbone network, and two sub-detection heads are set according to FCOS to regress 2D related attributes. One sub-head outputs a heat map h, and the other outputs a 2D distance d 2d :{l,t,r,b}, the two-dimensional distance represented by the target point x o =(u o ,v o ) and the top left vertex x of the bounding box l =(u l ,v l ), the lower right vertex x r =(u r ,v r ) calculated:

[0071] d 2d ={(u o -u l ),(v o -v l ),(u r -u o ),(v r -v o )}

[0072] Then, three-dimensional target detection is performed, which can be decomposed into five sub-modules, including center projection offset estimation module, three-dimensional dimension estimation module, key point compensation estimation module, direction estimation module, and depth estimation module. The specific introduction is as follows:

[0073] (1) Center projection offset estimation module

[0074] The target is set to its center point xo The representation point is the center projection of the 3D bounding box of the object rather than the center of the 2D bounding box. Since there is a coordinate discretization error in the downsampling process of the backbone network, it is necessary to regress the offset to obtain a more accurate center point coordinate.

[0075] (2) 3D dimension estimation module

[0076] In neural networks, fixed convolution parameters are usually learned to extract target features. However, unlike the two-dimensional case, the three-dimensional dimension represents the size of the three-dimensional bounding box in the real world. The fixed model can only extract features of general knowledge learned from the training set, while ignoring the appearance uniqueness of occlusion and truncation between different samples caused by three-dimensional space projection. To solve this problem, the following method is introduced in the three-dimensional size estimation part: Figure 2 The sample-aware feature fusion module.

[0077] This module consists of two parts, including a sample-aware convolution module and a three-dimensional regression module. The input is the global feature map F. The sample-aware convolution module on the upper side first generates a high-dimensional feature and a sample-aware filter

[0078]

[0079] Where G is a 3×3 convolutional layer, is the parameter of G. The three-dimensional regression module on the lower side generates intermediate features through a 3×3 kernel C f Each pixel in can be reconstructed into a 3×3 convolution kernel, and features are extracted through convolution operation. Right now:

[0080] F i2 =F i1 *C f

[0081] Then F i2 and F mid Through the attention module, the number of fusion feature channels is compressed from 512 to 256, that is, the feature map is Finally, the Conv-BN-ReLU module generates the final estimate D of the three-dimensional dimension 3d .

[0082] (3) Key point compensation estimation module

[0083] The regression target of the 8 key points is defined as δ k ={δ k1 ,δ k2 ,…,δ k8}, represents the center point heat map x o With the key point coordinate x k The offset between Figure 3 As shown, for 8 predefined key points x k ={x k1 ,x k2 ,…,x k8} Project from 8 3D angles. Further, the 8 key points are divided into 4 groups along the vertical edges of the 3D bounding box {(x k1 ,x k5 ),…,(x k4 ,x k8 )}, each group is composed of three-dimensional height h 3d The two-dimensional height h obtained by projection 2d ={h1,h2,h3,h4}, use the following formula to solve the depth z of the three-dimensional corner point k ={z k1 ,z k2 ,z k3 ,z k4}, where f is the focal length. These corner points are then used to calculate the visible surface of the 3D bounding box in the direction estimation module.

[0084] z k =h 3d f / h 2d

[0085] (4) Direction Estimation Module

[0086] Based on the assumption that the object is located on the horizontal plane, for a given camera's intrinsic matrix, the light angle can be solved to obtain the global angle. Since the human eye can usually obtain the three-dimensional features of the target from the visible surface. Therefore, each object can be roughly represented as a three-dimensional bounding box consisting of six faces. This appearance representation architecture has more advantages in direction regression than ordinary phase plane information. Therefore, we propose a geometry-aware strategy to construct a more reliable appearance feature representation for predicting direction.

[0087] The eight key points are predefined in a specific order and can be used to determine which surfaces are visible in the camera system. Specifically, it is assumed that the top surface of the object is always visible and the bottom surface is always invisible. Based on this assumption, only the visibility of the four other planes needs to be determined. First, the depth of the key points is divided into two groups according to the front-to-back relationship, and the difference Δz = {z1-z4, z2-z3} between each group is calculated. If the difference is less than 0, the front surface is visible, otherwise the back surface is visible. Since the coordinates of the three-dimensional angle projected along the u-axis are not as ambiguous as the depth projection, we directly use the distance in the phase plane to determine the visibility of the left and right surfaces. Similar to the above, the right side is visible when the distance is less than 0, otherwise the left side is visible.

[0088] according to Figure 1 After obtaining the three visible surfaces of the object, this paper uses it to regress θ l1 Feature map Extract the corresponding region F s , and transform these regions into a unified form F by the following perspective transformation u .

[0089] F u =P*F s

[0090] Where P is composed of F u The perspective transformation matrix is ​​solved by the four corners and their corresponding coordinates. Then the unified feature maps of each target are connected, and the number of channels is compressed from 768 to 256 using a 3×3 convolution layer. Finally, the Conv-BN-ReLU module and the AvgPool module are used to output θ l2 , by adding θ to the target level l2 With θ l1 Add together to get the final estimated result θ l .

[0091] (5) Depth Estimation Module

[0092] Since monocular imaging systems lack depth information, predicting instance depth is a challenging task. Considering that direct regression of depth is limited by a large search space, this paper uses the sigmoid function Output the network depth is converted into an absolute depth z. In addition, deep neural networks inevitably have epistemic uncertainty and unpredictable uncertainty. On this basis, we add a branch to predict the uncertainty of joint optimization with depth terms Loss Design

[0093] In order to clearly describe the losses of multiple tasks, loss terms are defined for each subtask branch, including the heat map loss L h , two-dimensional distance loss L 2d , center shift loss L c , key point offset loss L k , 3D angular depth loss L kd , three-dimensional size loss L 3d , direction loss L o and the depth loss L d In order to solve the problem of sample imbalance, L is optimized by penalty attenuation. h Considering the two-dimensional distance d 2d Used to get the size of the two-dimensional bounding box, L 2d GIoU loss is used. Center offset loss Lc , key point offset loss L k and the three-dimensional size loss L 3d The ordinary L1 loss form is used. Since the depth accuracy of the three-dimensional angle is very important in direction estimation, L is designed according to the L1 loss form. kd The items are:

[0094]

[0095] in is the true value of the 3D angular depth. In order to reduce the multimodal estimation problem, L o Multi-Bin loss is used. Depth loss L d Set to be the uncertainty u c The guided L1 loss takes the form:

[0096]

[0097] where z * is the true depth value.

[0098] According to the above loss definitions, the multi-task loss L is:

[0099] L=α1L h +α2L 2d +α3L c +α4L k +α5L kd +α6L 3d +α7L o +α8L d

[0100] The loss coefficients α1, α2, α3, α4, α5, α6, α7 and α8 are set to 1, 1, 0.5, 1, 0.2, 1, 1, 1 respectively.

Claims

1. A monocular 3D object detection method based on geometric appearance perception and feature fusion, characterized in that: The following steps are involved: The feature extraction module extracts the deep features of the monocular RGB image through the CNN skeleton; The center projection offset module regresses and estimates the offset of the deep feature to obtain the coordinates of its center point; The 3D dimension estimation module estimates the size of the 3D bounding box of the object in the image through the depth features of the image; The key point compensation estimation module calculates the depth of the three-dimensional corner points based on the depth features of the image; The direction estimation module estimates the direction of the 3D bounding box through the depth of the 3D corner points; The depth estimation module performs depth estimation on the depth features; The Loss module calculates the loss of each module, and calculates the multi-task loss based on each loss to complete target detection.

2. The monocular 3D target detection method based on geometric appearance perception and feature fusion according to claim 1, characterized in that: The feature extraction module performs the following steps: Using a three-channel RGB image I∈R through a CNN skeleton W×H×3 Takes as input and outputs a global feature map Where d is the downsampling coefficient, W and H are the width and height of the image respectively, and R is a real number; A 3×3 convolutional layer and a 1×1 convolutional layer are sequentially deployed in each detection head; Add a 2D detection branch as an auxiliary network; DLA-34 is used as the backbone network, and two sub-detection heads are set according to FCOS to perform two-dimensional attribute regression, where one sub-head outputs a heat map h and the other outputs a two-dimensional distance d 2d .

3. The monocular 3D target detection method based on geometric appearance perception and feature fusion according to claim 1, characterized in that: The center projection offset module performs the following steps: In the downsampling process of the backbone network, the offset is regressed to obtain the center coordinate x of the object's three-dimensional bounding box o .

4. The monocular 3D target detection method based on geometric appearance perception and feature fusion according to claim 1, characterized in that: The three-dimensional dimension estimation module performs the following steps: Based on the global feature map F, a high-dimensional feature is generated and a sample-aware filter Among them, G is a 3×3 convolutional layer, is the parameter of G; Based on the global feature map F, an intermediate feature is generated through a 3×3 kernel. C f Each pixel in is reconstructed into a 3×3 convolution kernel, and features are extracted through convolution operation. Right now: F i2 =F i1 *C f ; F i2 and F mid Through the attention module, the number of fusion feature channels is compressed from 512 to 256 to obtain the feature map The final estimate D of the three-dimensional dimension is generated through the Conv-BN-ReLU module 3d .

5. The monocular 3D target detection method based on geometric appearance perception and feature fusion according to claim 1, characterized in that: The key point compensation estimation module performs the following steps: For 8 predefined key points x k ={x k1 , x k2 , …, x k8 } Projection from 8 3D angles; Divide the 8 key points into 4 groups along the vertical edges of the 3D bounding box {(x k1 , x k5 ),…,(x k4 , x k8 )}, each group is composed of three-dimensional height h 3d The two-dimensional height h obtained by projection 2b ={h1,h2,h3,h4}; Solve the depth z of the 3D corner point k ={z k1 , z k2 , z k3 , z k4 }: z k =h 3d ·f / h 2d Where f is the focal length.

6. The monocular 3D target detection method based on geometric appearance perception and feature fusion according to claim 1, characterized in that: The direction estimation module performs the following steps: The depth of the key points is divided into two groups according to the front-back relationship of the key points, and the difference Δz between each group is calculated = {z1-z4, z2-z3}; If the difference is less than 0, the front surface is visible, otherwise the back surface is visible; Use the same method of judging the difference to judge the visibility of the upper and lower surfaces and the left and right surfaces; After obtaining the three visible surfaces of the object, they are used to regress θ l1 Feature map Extract the corresponding region F s , and transform the region by perspective transformation to obtain a unified feature map F u : F u =P*F s Among them, P is composed of F u The perspective transformation matrix obtained by solving the four corners and their corresponding coordinates; The unified feature maps of each target are concatenated and the number of channels is compressed from 768 to 256 using a 3×3 convolutional layer; Output θ through the sequentially connected Conv-BN-ReLU module and AvgPool module l2 ; By adding the target level l2 With θ l1 Add together to get the final estimated result θ l .

7. The monocular 3D target detection method based on geometric appearance perception and feature fusion according to claim 1, characterized in that: The depth estimation module performs the following steps: Use the sigmoid function to reduce the network output depth Convert to absolute depth z; Adding uncertainty to a branch prediction joint optimization with a depth term 8. The monocular 3D target detection method based on geometric appearance perception and feature fusion according to claim 1, characterized in that: The Loss module performs the following steps: Calculate the loss of each module separately, including: heat map loss L h , two-dimensional distance loss L 2d , center shift loss L c , key point offset loss L k , 3D angular depth loss L kd , three-dimensional size loss L 3d , direction loss L o and the depth loss L d ; According to the definition of each loss, calculate the multi-task loss L: L=α1L h +α2L 2d +α3L c +α4L k +α5L kd +α6L 3d +α7L o +α8L d Among them, α1, α2, α3, α4, α5, α6, α7 and α8 are loss coefficients.

9. A monocular three-dimensional object detection method based on geometric appearance perception and feature fusion, characterized in that: include: Feature extraction module, used to extract deep features of monocular RGB images through CNN skeleton; The center projection offset module is used to regress and estimate the offset of the deep feature to obtain the coordinates of its center point; A three-dimensional dimension estimation module is used to estimate the size of the three-dimensional bounding box of the object in the image through the depth features of the image; The key point compensation estimation module is used to calculate the depth of the three-dimensional corner points according to the depth characteristics of the image; A direction estimation module is used to estimate the direction of a 3D bounding box based on the depth of 3D corner points; A depth estimation module, used for performing depth estimation on depth features; The Loss module is used to calculate the losses of each module and calculate the multi-task loss based on each loss to complete target detection.

10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, a monocular three-dimensional target detection method based on geometric appearance perception and feature fusion as described in any one of claims 1 to 8 is implemented.