Monocular three-dimensional target detection method based on Transform auxiliary depth information fusion
By introducing the Transformer deep information fusion method and global adaptive attention module, the problem of lack of depth information in monocular three-dimensional object detection is solved, and the accuracy and detection accuracy of depth information estimation are improved.
Patent Information
- Application Number
- CN202311327238.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-12
- Publication Date
- 2025-07-18
AI Technical Summary
The existing monocular three-dimensional object detection methods lack direct depth information, resulting in the inability to fully understand the global depth context information, which affects the accurate identification of three-dimensional information.
The Transformer-based deep information fusion method is adopted, combined with the global adaptive attention module and the multi-objective border module, and the accuracy of depth information estimation is improved through adaptive feature fusion and depth information aggregation.
A more accurate estimation of the depth information of the target object is achieved, improving the accuracy and accuracy of monocular three-dimensional object detection.
Smart Images

Figure CN120340018A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to a monocular three-dimensional object detection method based on Transformer-assisted depth information fusion. Background Art
[0002] Monocular three-dimensional object detection has attracted increasing attention due to the use of a monocular RGB camera as a simpler and cheaper deployment setting; most existing monocular 3D detection methods follow the benchmarks of traditional 2D object detectors and first locate the objects in the image by detecting their 2D or projected 3D centers from heatmaps. Then, the visual features around each object center are aggregated to predict the 3D attributes of the object, such as depth, 3D size, and orientation; however, since the depth information is not directly reflected in the input, the monocular 3D detection task is inherently ambiguous, and existing monocular 3D object detection schemes mainly rely on local features, which results in their inability to fully understand the global depth context information, thus affecting the accurate recognition of three-dimensional information.
[0003] Therefore, in order to solve the above technical problems, it is urgent to propose a new technical means. Summary of the Invention
[0004] In view of this, in order to make full use of the global depth information, better capture the remote spatial depth information of the image, and improve the accuracy of depth information estimation, the present invention proposes a monocular three-dimensional object detection method based on Transformer-assisted depth information fusion, including the following steps:
[0005] S1. Obtain RGB pictures as a sample data set;
[0006] S2. Build a monocular three-dimensional object detection model and input the sample data set into the monocular three-dimensional object detection model for training; the monocular three-dimensional object detection model includes a backbone feature extraction network, a Neck network, a Transformer depth information aggregation module, and a Head network;
[0007] Among them, the output end of the backbone feature extraction network is connected to the input end of the Neck network, the output end of the Neck network is connected to the input end of the Transformer depth information aggregation module and the input end of the Head network, and the output end of the Transformer depth information aggregation module is connected to the input end of the Head network;
[0008] S3. Determine whether the monocular three-dimensional object detection model is trained. If so, go to step S4; if not, update the parameters in the monocular three-dimensional object detection model and continue training until the training is completed;
[0009] S4. Input the real-time collected images into the trained monocular 3D object detection model, and output the category information, depth information, 2D bounding box information, and 3D bounding box information of the objects in the images.
[0010] Further, in step S2, the backbone feature extraction network includes a ResNet101 network and an FPN network. The output end of the ResNet101 network is connected to the input end of the FPN network; the FPN network outputs feature maps P0, P1, P2, and P3.
[0011] The Neck network includes four parallel global adaptive attention modules. The global adaptive attention module includes a parallel convolutional attention module and a coordinate attention module. Among them, the convolutional attention module includes a channel attention module and a spatial attention module. The input feature of the channel attention module is multiplied by the output weight of the channel attention module and then input into the spatial attention module. The input feature of the spatial attention module and the output weight of the spatial attention module are multiplied and then adaptively aggregated with the output feature of the coordinate attention module and output; the feature maps P0, P1, P2, and P3 are respectively input into the four global adaptive attention modules, and the four global adaptive attention modules respectively output feature maps P4, feature map P5, feature map P6, and feature map P7; among them, feature map P0 corresponds to feature map P4, feature map P1 corresponds to feature map P5, feature map P2 corresponds to feature map P6, and feature map P3 corresponds to feature map P7.
[0012] The Transformer depth information aggregation module includes a depth position encoder and a depth decoder.
[0013] Among them, the depth position encoder includes an upsampling module, a downsampling module, a feature embedding module, a masked multi-head attention module, a position encoder, two residual modules, two normalization modules, and a feed-forward neural network; the downsampling module downsamples the feature map P5 to obtain the feature map P8, and the upsampling module upsamples the feature map P7 to obtain the feature map P9. The feature maps P8, P6, and P9 are feature-fused and then input into the feature embedding module. The feature output by the feature embedding module and the feature output by the position encoder are added and then input into the masked multi-head attention module. The input feature of the masked multi-head attention module and the output feature of the masked multi-head attention module are simultaneously input into the first residual module. The output end of the first residual module is connected to the input end of the first normalization module, and the output end of the first normalization module is connected to the input end of the feed-forward neural network. The input feature of the feed-forward neural network and the output feature of the feed-forward neural network are simultaneously input into the second residual module. The output end of the second residual module is connected to the input end of the second normalization module.
[0014] The depth decoder includes a feature cross-attention layer and a depth cross-attention layer; the output end of the feature cross-attention layer is connected to the input end of the depth cross-attention layer;
[0015] The feature cross-attention layer includes a masked multi-head attention module, a multi-head attention module, two residual modules, and two normalization modules; among them, the input feature and the output feature of the masked multi-head attention module are simultaneously input into the first residual module, the output end of the first residual module is connected to the input end of the first normalization module, the output end of the first normalization module is connected to the input end of the multi-head attention module, the output end of the multi-head attention module is connected to the input end of the second residual module, and the output end of the second residual module is connected to the input end of the second normalization module;
[0016] The depth cross-attention layer includes a masked multi-head attention module, two residual modules, two normalization modules, and a feed-forward neural network; among them, the input feature and the output feature of the masked multi-head attention module are simultaneously input into the first residual module, the output end of the first residual module is connected to the input end of the first normalization module, the output end of the first normalization module is connected to the input end of the feed-forward neural network, the input feature and the output feature of the feed-forward neural network are simultaneously input into the second residual module, and the output end of the second residual module is connected to the input end of the second normalization module;
[0017] The Head network includes a classification network, a regression network, and a multi-object bounding box module; the classification network and the regression network are parallel, and the features input to the classification network and the regression network are the same; the output end of the regression network is connected to the input end of the multi-object bounding box module; in the regression network, the depth information output by the regression network and the depth information output by the Transformer depth information aggregation module are adaptively aggregated to obtain new depth information;
[0018] The two-dimensional bounding box information and the new depth information output by the regression network are input into the multi-object bounding box module, and the multi-object bounding box module determines the best label by setting pseudo-labels, and obtains the three-dimensional bounding box information according to the best label.
[0019] Furthermore, the adaptive aggregation formula is as follows:
[0020] F Z =λ t F1+(1-λ t )F2
[0021] where, F ZRepresents the adaptable aggregated information. When the adaptable aggregation is the output features of the convolutional attention module and the output features of the coordinate attention module, F1 and F2 respectively represent the feature information output by the convolutional attention module and the feature information output by the coordinate attention module. When the adaptable aggregation is the depth information output by the deep network and the depth information output by the Transformer depth information aggregation module, F1 and F2 respectively represent the depth information output by the deep network and the depth information output by the Transformer depth information aggregation module, λ t Represents the learnable gating weight parameter.
[0022] Furthermore, the multi-object bounding box module finds the best label through the following steps:
[0023] First, set a group of size error tolerances Δz to perturb the depth Z of each label P=(class, X, Y, Z, H, W, L, yaw). Obtain the depth offset Δz·Z according to the size error tolerance Δz and the object depth Z. Perturb the object depth according to the depth offset to obtain the perturbed object depth as Z + Δz·Z, and further obtain the three-dimensional pseudo-label P'=(class, X, Y, Z + Δz·Z, H, W, L, yaw);
[0024] where class represents the category, (X, Y, Z) represents the coordinates of the object in the camera coordinate system, Z also represents the depth, H represents the height of the three-dimensional object, W represents the width of the three-dimensional object, L represents the length of the three-dimensional object, and yaw represents the yaw angle;
[0025] Second, obtain the two-dimensional bounding box coordinates of the label and the three-dimensional bounding box coordinates of the corresponding pseudo-label. Project the three-dimensional bounding box coordinates of the corresponding pseudo-label onto the two-dimensional bounding box plane of the label, and determine the projected bounding box according to the maximum and minimum values of the projected coordinates;
[0026] Then, measure the similarity between the two-dimensional bounding box of the label and the projected bounding box through the IoU label score, and determine the IoU label score with the highest similarity;
[0027] MAX IoU Label = IoU(B gt , B false )
[0028] where IoU Label represents the label score, B gt represents the two-dimensional bounding box of the label, and B false represents the projected bounding box;
[0029] Finally, determine the best label P”=(class, X t , Y t,Z+Δz·Z,H,W,L,yaw t );
[0030] Adjust the yaw angle according to the IoU label score with the highest similarity. The calculation formula is as follows:
[0031]
[0032] where yaw represents the yaw angle, and yaw t represents the adjusted yaw angle, represents the IoU label score with the highest similarity, H represents the height of the three-dimensional object, W represents the width of the three-dimensional object, L represents the length of the three-dimensional object, Y represents the Y-axis coordinate of the three-dimensional object in the camera coordinate system, and exp represents the exponential function with the natural constant e as the base;
[0033] Adjust the X-axis and Y-axis coordinates in the camera coordinate system according to the IoU label score with the highest similarity;
[0034] Calculate the adjusted X-axis and Y-axis coordinates based on the object depth Z+Δz·Z in the IoU label score with the highest similarity. The calculation formula is as follows:
[0035]
[0036]
[0037] where X t represents the adjusted X-axis coordinate, Y t represents the adjusted Y-axis coordinate, (u, v) represents the two-dimensional coordinates of the points on the image, and f x , f y , c x and c y represent the intrinsic parameters of the camera, and Z+Δz·Z represents the object depth corresponding to the IoU label score with the highest similarity.
[0038] Furthermore, in step S2, an SGD optimizer with a weight decay of 10 -4 is used for training with the following loss function L:
[0039] L = L cls + L offset + L size + L rotsin + L velo + L depth
[0040] where L cls represents the classification loss, L offset represents the geometric center point offset loss, L size represents the size loss, L rotsinDenote the rotational loss as L velo Denote the velocity loss as L size Denote the depth loss.
[0041] Furthermore, in step S3, when 48 epochs are trained, the monocular three-dimensional object detection model is trained and completed.
[0042] Advantages of the present invention: The present invention proposes a global adaptive attention module that adaptively fuses and enhances feature information through the global adaptive attention module, while extracting spatial long-range dependencies and retaining local information; it also proposes a Transformer depth information aggregation module, which is used to capture the long-range spatial depth information of the image and improve the accuracy of depth information estimation; a multi-object bounding box module is introduced, and the multi-object bounding box module uses soft labels to eliminate the strict limitations of the original hard labels, improving the accuracy of monocular 3D object detection. Brief Description of the Drawings
[0043] The present invention will be further described below in conjunction with the drawings and embodiments:
[0044] Figure 1 is the flow chart of the present invention;
[0045] Figure 2 is the schematic diagram of the monocular three-dimensional object detection model of the present invention. Detailed Embodiments
[0046] The present invention will be further described below in conjunction with the accompanying drawings of the specification:
[0047] A monocular three-dimensional object detection method based on Transformer-assisted depth information fusion provided by the present invention includes the following steps:
[0048] S1. Obtain RGB pictures as the sample data set;
[0049] S2. Build a monocular three-dimensional object detection model and input the sample data set into the monocular three-dimensional object detection model for training; the monocular three-dimensional object detection model includes a backbone feature extraction network, a Neck network, a Transformer depth information aggregation module, and a Head network;
[0050] Among them, the output end of the backbone feature extraction network is connected to the input end of the Neck network, the output end of the Neck network is connected to the input end of the Transformer depth information aggregation module and the input end of the Head network, and the output end of the Transformer depth information aggregation module is connected to the input end of the Head network;
[0051] S3. Determine whether the monocular 3D object detection model is trained. If so, proceed to step S4. If not, update the parameters in the monocular 3D object detection model and continue training until the training is completed;
[0052] S4. Input the real-time collected images into the trained monocular 3D object detection model, and output the category information, depth information, 2D bounding box information, and 3D bounding box information of the objects in the images. Through the above method, the depth information of the target object can be estimated more accurately, improving the accuracy of monocular 3D object detection.
[0053] In this embodiment, in step S1, the KITTI object detection dataset is used as the sample dataset. The KITTI object detection dataset is a commonly used dataset for monocular 3D object detection and is an existing dataset, which will not be elaborated here.
[0054] In this embodiment, in step S2, a monocular 3D object detection model is constructed and trained; monocular refers to a monocular camera. The monocular 3D object detection model constructed in this application is used to predict the information of RGB images captured by the monocular camera; the monocular 3D object detection model constructed in this application is improved on the basis of the PGD monocular 3D object detection model, and a Transformer depth information aggregation module, a global adaptive attention module, and a multi-object bounding box module are added. The specific structure is as follows:
[0055] The monocular 3D object detection model includes a backbone feature extraction network, a Neck network, a Transformer depth information aggregation module, and a Head network, as Figure 2 shown;
[0056] Among them, the output end of the backbone feature extraction network is connected to the input end of the Neck network, the output end of the Neck network is connected to the input end of the Transformer depth information aggregation module and the input end of the Head network, and the output end of the Transformer depth information aggregation module is connected to the input end of the Head network;
[0057] The backbone feature extraction network includes a ResNet101 network and an FPN network. The output end of the ResNet101 network is connected to the input end of the FPN network; the FPN network outputs feature maps P0, P1, P2, and P3; among them, the ResNet101 network and the FPN network are existing technologies, which will not be elaborated here;
[0058] The Neck network represents the neck network, including four global adaptive attention modules. The global adaptive attention module is the Adaptive Channel Spatial Location Attention MOBule (hereinafter referred to as ACSL). The global adaptive attention module includes a parallel Convolutional Block Attention MOBule (hereinafter referred to as CBAM) and a Coordinate Attention MOBule (hereinafter referred to as CA). Among them, the convolutional attention module includes a Channel Attention MOBule (hereinafter referred to as CAM) and a Spatial Attention MOBule (hereinafter referred to as SAM). The input feature of the channel attention module is multiplied by the output weight of the channel attention module and then input into the spatial attention module. The input feature of the spatial attention module and the output weight of the spatial attention module are multiplied and then adaptively aggregated with the output feature of the coordinate attention module and output; The feature maps P0, P1, P2, and P3 are respectively input into the four global adaptive attention modules, and the four global adaptive attention modules respectively output the feature map P4, the feature map P5, the feature map P6, and the feature map P7; Among them, the feature map P0 corresponds to the feature map P4, the feature map P1 corresponds to the feature map P5, the feature map P2 corresponds to the feature map P6, and the feature map P3 corresponds to the feature map P7; The correspondence between the feature map P0 and the feature map P4 means that after the feature map P0 is input into the global adaptive attention module, the feature map P4 is output, and the same applies to the other feature maps. Among them, the convolutional attention module, the channel attention module, the spatial attention module, and the coordinate attention module are all existing technologies and will not be elaborated here;
[0059] The global adaptive attention module specifically performs feature fusion and enhancement on the feature map through the following steps:
[0060] First, the channel attention module encodes each channel of the input feature map using a max pooling layer and an average pooling layer respectively; And share the features after passing through the max pooling layer and the average pooling layer to the MLP transformation function, and output an intermediate feature map Z1 that encodes the significant features and global features of the spatial information;
[0061] Z1 = δ(MLP(AvgPool(F)) + MLP(MaxPool(F)))
[0062] Among them, δ represents the sigmiod function, AvgPool represents average pooling, MaxPool represents max pooling, and F represents the feature input into the global adaptive attention module;
[0063] The intermediate feature map Z1 is transformed through a 1×1 convolution to obtain attention weights with the same number of channels as the intermediate feature map. Then, the attention weights and the features of the feature map input to the channel attention module are multiplied element-wise to obtain the output feature A1 of the channel attention module;
[0064] A1 = F × δ(T1(Z1))
[0065] where F represents the feature input to the global adaptive attention module, δ represents the sigmoid function, T1 represents the 1×1 convolution transformation function, and Z1 represents the intermediate feature map;
[0066] Secondly, the output feature of the channel attention module is input to the spatial attention module. The max pooling layer and the average pooling layer are used to encode each channel separately and then concatenated; the concatenated features are input to a shared 7×7 convolution transformation function to obtain attention weights with the same spatial dimensions; the output feature of the channel attention block and the attention weights are multiplied element-wise to obtain the output feature of the spatial attention module, which is also the output feature F1 of the convolutional attention module;
[0067] F1 = A1 × δ(T2([AvgPool(A1), MaxPool(a1)]))
[0068] where a1 represents the output feature of the channel attention module, δ represents the sigmoid function, AvgPool represents average pooling, MaxPool represents max pooling, T2 represents the 7×7 convolution transformation function, and [AvgPool(A1), MaxPool(A1)] represents the concatenation operation of A1 through the average pooling layer and A1 through the max pooling layer along the spatial dimension;
[0069] Then, in order to encourage the attention block to capture long-range interactions spatially by leveraging position information, the coordinate attention module decomposes the global pooling into a pair of 1D feature encoding operations. Specifically:
[0070] Pooling kernels of two spatial dimensions (H, 1) and (1, W) are used to encode each channel separately, and the features generated in the two spatial directions are concatenated and shared to a 1×1 convolution transformation function to obtain an intermediate feature map Z2 that encodes the spatial information in the horizontal and vertical directions;
[0071]
[0072] where δ represents the sigmoid function, T1 represents the 1×1 convolution transformation function, represents the concatenation operation along the spatial dimension, represents the feature of the c-th channel with height h, represents the feature of the c-th channel with width w;
[0073] Decompose the intermediate feature map Z2 into two independent tensors along the spatial dimension and Use a 1×1 convolution transformation to and become tensors with the same number of channels, expand the tensors with the same number of channels to be used as attention weights, multiply the input features of the coordinate attention by the attention weights and output the feature F2;
[0074]
[0075] Among them, F represents the feature input to the global adaptive attention module, represents the attention weight in the height direction, represents the attention weight in the width direction;
[0076] Finally, adaptively aggregate the features output by the convolutional attention module and the features output by the coordinate attention module to obtain the features output by the global adaptive attention module. The adaptive aggregation is performed through the following formula:
[0077] F Z = λ t F1+(1 - λ t )F2
[0078] Among them, F Z represents the information after adaptive aggregation, which represents the features output by the global adaptive attention module here. F1 and F2 respectively represent the features output by the convolutional attention module and the features output by the coordinate attention module. λ t represents the learnable gating weight parameter. λ is a learnable training parameter. Use the sigmoid function to project λ into the range of [0,1]. Therefore, λ t ∈[0,1];
[0079] When λ t approaches 0, the ACSL module pays more attention to the coordinate attention module and focuses on the long-range dependencies in the spatial direction; while when it approaches 1, the ACSL module pays more attention to the interrelationships between channels and spaces; through the above method, it is possible to adaptively find the best feature representation to make the model perform best on the given task;
[0080] The global adaptive attention module can adaptively fuse and strengthen the feature information, retain the local information while extracting the spatial long-range dependencies, so as to make the best use of the advantages of the convolutional attention module and the coordinate attention module;
[0081] The Transformer depth information aggregation module includes a depth position encoder and a depth decoder, where Transformer represents the attention mechanism;
[0082] Among them, the depth position encoder includes an upsampling module, a downsampling module, a feature embedding module, a masked multi-head attention module, a position encoder, two residual modules, two normalization modules, and a feed-forward neural network; the downsampling module downsamples the feature map P5 to obtain the feature map P8, the upsampling module upsamples the feature map P7 to obtain the feature map P9, the feature maps P8, P6, and P9 are feature-fused and then input into the feature embedding module, the features output by the feature embedding module and the features output by the position encoder are added and then input into the masked multi-head attention module, the input features and output features of the masked multi-head attention module are simultaneously input into the first residual module, the output end of the first residual module is connected to the input end of the first normalization module, the output end of the first normalization module is connected to the input end of the feed-forward neural network, the input features and output features of the feed-forward neural network are simultaneously input into the second residual module, and the output end of the second residual module is connected to the input end of the second normalization module; the feed-forward neural network uses an existing FFN network; among them, the part composed of the masked multi-head attention module, the residual module, the normalization module, and the feed-forward neural network is the depth encoder;
[0083] The depth decoder includes a feature cross-attention layer and a depth cross-attention layer; the output end of the feature cross-attention layer is connected to the input end of the depth cross-attention layer;
[0084] The feature cross-attention layer includes a masked multi-head attention module, a multi-head attention module, two residual modules, and two normalization modules; among them, the input features and output features of the masked multi-head attention module are simultaneously input into the first residual module, the output end of the first residual module is connected to the input end of the first normalization module, the output end of the first normalization module is connected to the input end of the multi-head attention module, the output end of the multi-head attention module is connected to the input end of the second residual module, and the output end of the second residual module is connected to the input end of the second normalization module;
[0085] The depth cross-attention layer includes a masked multi-head attention module, two residual modules, two normalization modules, and a feed-forward neural network. Among them, the input features of the masked multi-head attention module and the output features of the masked multi-head attention module are simultaneously input into the first residual module. The output end of the first residual module is connected to the input end of the first normalization module, and the output end of the first normalization module is connected to the input end of the feed-forward neural network. The input features of the feed-forward neural network and the output features of the feed-forward neural network are simultaneously input into the second residual module, and the output end of the second residual module is connected to the input end of the second normalization module. The feed-forward neural network uses an existing FFN network.
[0086] Specifically, the Transformer depth information aggregation module captures the far and near spatial depth information of the image through the following steps:
[0087] First, in the feature encoder, the downsampling module and the upsampling module respectively unify the sizes of the feature map P5 and the feature map P7 to the downsampling ratio, and output the feature maps P8 and P9. The feature maps P8, P6, and P9 are fused through element-wise addition and convolution operations to obtain the fused features. The fused features are input into the feature embedding module and combined with the position encoder to obtain the feature F s , and then the feature F s is input into the depth encoder for encoding, and the depth embedding feature T S is output. The calculation formula is as follows:
[0088] T S = LN(FFN(LN(F S + SelfAttention(F S )))+ LN(F S + SelfAttention(F S )))
[0089] Among them, LN represents the LayerNorm function, FFN in the formula represents the function in the FFN neural network, and SelfAttention represents the self-attention function.
[0090] The SelfAttention self-attention function linearly transforms the feature F s into the query Q S , the key K S , and the value V S .
[0091] Compared with the existing detection algorithms, using the Transformer depth position encoder to construct the global depth guidance region can capture the remote spatial depth information of the image and provide the network with global clues for the entire three-dimensional space.
[0092] Then, in the feature cross-attention layer, the query q aggregates the depth-embedded feature T through the cross-attention function s , and the cross-attention function linearly transforms the query q and the depth-embedded feature into Q D , key K D , and value V D ; and calculate the query q feature cross-attention feature T D according to Q D , key K D , and value V D , and the calculation formula is as follows:
[0093]
[0094] where C represents the number of channels, which is equal to 256, represents the transposed matrix of the key K D ;
[0095] Further calculate the depth-aware query q′ according to T D , and the calculation formula is as follows:
[0096] q′ = LN(SelfAttention(LN(T s + T D )) + LN(T s + T D )
[0097] where LN represents the LayerNorm function and SelfAttention represents the self-attention function;
[0098] The introduction of the feature cross-attention layer enables each object query to adaptively explore and obtain spatial cues according to the depth-guided region of the image, which is of great significance for understanding the global spatial structure, modeling the geometric relationship between objects, and assisting depth estimation;
[0099] Finally, in the depth cross-attention layer, the depth information of the feature image is guided by the depth-aware query; the depth-aware query q′, as the output of the feature cross-attention layer, contains depth information and can therefore guide the depth information of the feature image, and the specific formula is as follows;
[0100] D T = LN(FFN(LN(q′ + CrossAttention(q′, F E ))) + LN(q′ + CrossAttention(q′, F E )))
[0101] where D Trepresents the depth output by the Transformer depth information aggregation module, LN represents the LayerNorm function, SelfAttention represents the self-attention function, FFN in the formula represents the function in the FFN neural network, q′ represents the depth-aware query, and F E represents the feature of the feature map P6 output by the global adaptive attention module;
[0102] The Head network includes a classification network, a regression network, and a Multiple Objects Bounding Box Moudle (abbreviated as MOB); the classification network and the regression network are parallel, and the features input to the classification network and the regression network are the same. The classification network outputs the target category information of the image; the regression network outputs two-dimensional bounding box information and depth information; the output end of the regression network is connected to the input end of the multiple objects bounding box module; the regression network contains a depth network and a depth estimation module. The depth network outputs depth information, and the depth estimation module adaptively aggregates the depth information output by the depth network and the depth information output by the Transformer depth information aggregation module to obtain new depth information;
[0103] Input the two-dimensional bounding box information and the new depth information output by the regression network into the multiple objects bounding box module. The multiple objects bounding box module determines the best label by setting pseudo-labels and obtains the three-dimensional bounding box information according to the best label. Among them, the classification network and the regression network adopt the classification network and the regression network in the PGD monocular three-dimensional object detection model;
[0104] The depth information output by the depth network and the depth information output by the Transformer depth information aggregation module are adaptively aggregated to obtain new depth information. The adaptive aggregation formula is as follows:
[0105] D = λ b D L +(1 - λ b )D T
[0106] where D represents the new depth obtained by adaptive aggregation, D L represents the depth output by the depth network, D T represents the depth output by the Transformer depth information aggregation module, and λ b represents the learnable gating parameter. The λ is projected into the range of [0,1] by using the sigmoid function. Therefore, λ b ∈[0,1];
[0107] The multiple objects bounding box module MOB finds the best label through the following steps:
[0108] First, set a set of dimensional error tolerances Δz = [-4%, -2%, -1%, +1%, +2%, +4%] to perturb the depth Z of each label P = (class, X, Y, Z, H, W, L, yaw). Obtain the depth offset Δz·Z based on the dimensional error tolerance Δz and the object depth Z. Perturb the object depth according to the depth offset to get the perturbed object depth as Z + Δz·Z, and further obtain the 3D pseudo-label P' = (class, X, Y, Z + Δz·Z, H, W, L, yaw).
[0109] Among them, class represents the category, (X, Y, Z) represents the coordinates of the object in the camera coordinate system, Z also represents the depth, H represents the height of the 3D object, W represents the width of the 3D object, L represents the length of the 3D object, and yaw represents the yaw angle;
[0110] Secondly, obtain the 2D bounding box coordinates of the label and the 3D bounding box coordinates of the corresponding pseudo-label. Project the 3D bounding box coordinates of the corresponding pseudo-label onto the 2D bounding box plane of the label, and determine the projected bounding box according to the maximum and minimum values of the projected coordinates;
[0111] Then, measure the similarity between the 2D bounding box of the label and the projected bounding box through the IoU label score, and determine the IoU label score with the highest similarity;
[0112] MAX IoU Label = IoU(B gt ,B false )
[0113] Among them, IoU Label represents the label score, B gt represents the 2D bounding box of the label, and B false represents the projected bounding box;
[0114] IoU represents the intersection over union, which is used to evaluate the overlap degree between the actual box and the predicted box in the object detection network;
[0115] Finally, determine the best label P” = (class, X t , Y t , Z + Δz·Z, H, W, L, yaw t );
[0116] Adjust the yaw angle according to the IoU label score with the highest similarity. The calculation formula is as follows:
[0117]
[0118] Among them, yaw represents the yaw angle, and yaw t represents the adjusted yaw angle, Represents the IoU label score with the highest similarity, H represents the height of the three-dimensional object, W represents the width of the three-dimensional object, L represents the length of the three-dimensional object, Y represents the Y-axis coordinate of the three-dimensional object in the camera coordinate system, and exp represents the exponential function with the natural constant e as the base;
[0119] Adjust the X-axis coordinate and Y-axis coordinate in the camera coordinate system according to the IoU label score with the highest similarity;
[0120] Obtain the object depth Z + Δz·Z according to the IoU label score with the highest similarity, and then calculate the adjusted X-axis coordinate and Y-axis coordinate according to the object depth. The calculation formula is as follows:
[0121]
[0122]
[0123] Among them, X t Represents the adjusted X-axis coordinate, Y t Represents the adjusted Y-axis coordinate, (u, v) represents the two-dimensional coordinates of the points on the image, f x 、f y 、c x And c y Represents the intrinsic parameters of the camera, and Z + Δz·Z represents the object depth corresponding to the IoU label score with the highest similarity. The multi-object bounding box module eliminates the strict limitations of the original hard labels using pseudo-labels; after viewing multiple reasonable pseudo-labels, the most suitable pseudo-label is selected, which can improve the generalization ability of the network.
[0124] In this embodiment, in step S2, all attention modules use 8 heads, set the query times N to 50, set both the channel C and the latent feature dimensions of the detection heads based on FFN and based on mlp to 256, set the learning rate to 1×10 -3 , set the momentum factor to 0.9, set the weight decay coefficient to 0.0001, set the warm-up iteration times to 500, the warm-up ratio to 0.33, and adjust the learning rates at the 32nd and 44th times to the set learning rate values, and adopt the SGD optimizer with weight decay of 10 -4 And the following loss function L for training:
[0125] L = L cls + L offset + L size + L rotsin + L velo + L depth
[0126] Among them, L cls Represents the classification loss, and the FocalLoss loss function is adopted. L offsetRepresents the geometric center point offset loss, using the SmoothL1Loss loss function, L size Represents the size loss, using the SmoothL1Loss loss function, L rotsin Represents the rotation loss, using the SmoothL1Loss loss function, L velo Represents the speed loss, using the SmoothL1Loss loss function, L size Represents the depth loss, using the SmoothL1Loss loss function. Among them, the FocalLoss loss function and the SmoothL1Loss loss function are existing loss functions, which will not be elaborated here. By the above method, the convergence speed of the monocular three-dimensional object detection model is accelerated.
[0127] In this embodiment, in step S3, it is judged whether the monocular three-dimensional object detection model is trained. When 48 epochs are trained, the monocular three-dimensional object detection model is trained. Among them, epoch represents the number of rounds, which means that the complete data set has passed through the monocular three-dimensional object detection network once and returned once.
[0128] In this embodiment, in step S4, the real-time collected pictures are input into the trained monocular three-dimensional object detection model, and the category information, depth information, two-dimensional bounding box information, and three-dimensional bounding box information of the objects in the pictures are output. The output depth information is the depth information in the best label; the output three-dimensional bounding box information includes the coordinates of the center of the three-dimensional bounding box and the size of the three-dimensional bounding box in the three-dimensional space.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A monocular three-dimensional object detection method based on Transformer-assisted depth information fusion, characterized in that: It includes the following steps: S1. Obtain RGB pictures as the sample data set; S2. Build a monocular three-dimensional object detection model, and input the sample data set into the monocular three-dimensional object detection model for training; the monocular three-dimensional object detection model includes a backbone feature extraction network, a Neck network, a Transformer depth information aggregation module and a Head network; Among them, the output end of the backbone feature extraction network is connected to the input end of the Neck network, the output end of the Neck network is connected to the input end of the Transformer depth information aggregation module and the input end of the Head network, and the output end of the Transformer depth information aggregation module is connected to the input end of the Head network; S3. Judge whether the monocular three-dimensional object detection model is trained. If so, enter step S4. If not, update the parameters in the monocular three-dimensional object detection model and continue training until the training is completed; S4. Input the pictures collected in real time into the trained monocular three-dimensional object detection model, and output the category information, depth information, two-dimensional bounding box information and three-dimensional bounding box information of the objects in the pictures.
2. The monocular three-dimensional object detection method based on Transformer-assisted depth information fusion according to claim 1, wherein: In step S2, the backbone feature extraction network includes a ResNet101 network and an FPN network. The output end of the ResNet101 network is connected to the input end of the FPN network; the FPN network outputs feature maps P0, P1, P2 and P3; The Neck network includes four parallel global adaptive attention modules. The global adaptive attention module includes a parallel convolutional attention module and a coordinate attention module. Among them, the convolutional attention module includes a channel attention module and a spatial attention module. The input feature of the channel attention module is multiplied by the output weight of the channel attention module and then input into the spatial attention module. The input feature of the spatial attention module and the output weight of the spatial attention module are multiplied and then adaptively aggregated with the output feature of the coordinate attention module and output; the feature maps P0, P1, P2 and P3 are respectively input into the four global adaptive attention modules, and the four global adaptive attention modules respectively output feature map P4, feature map P5, feature map P6 and feature map P7; among them, feature map P0 corresponds to feature map P4, feature map P1 corresponds to feature map P5, feature map P2 corresponds to feature map P6, and feature map P3 corresponds to feature map P7; The Transformer depth information aggregation module includes a depth position encoder and a depth decoder; Among them, the depth position encoder includes an upsampling module, a downsampling module, a feature embedding module, a masked multi-head attention module, a position encoder, two residual modules, two normalization modules, and a feed-forward neural network; the downsampling module downsamples the feature map P5 to obtain the feature map P8, the upsampling module upsamples the feature map P7 to obtain the feature map P9, the feature maps P8, P6, and P9 are feature-fused and then input into the feature embedding module, the features output by the feature embedding module and the features output by the position encoder are added and then input into the masked multi-head attention module, the input features of the masked multi-head attention module and the output features of the masked multi-head attention module are simultaneously input into the first residual module, the output end of the first residual module is connected to the input end of the first normalization module, the output end of the first normalization module is connected to the input end of the feed-forward neural network, the input features of the feed-forward neural network and the output features of the feed-forward neural network are simultaneously input into the second residual module, and the output end of the second residual module is connected to the input end of the second normalization module; The depth decoder includes a feature cross-attention layer and a depth cross-attention layer; the output end of the feature cross-attention layer is connected to the input end of the depth cross-attention layer; The feature cross-attention layer includes a masked multi-head attention module, a multi-head attention module, two residual modules, and two normalization modules; among them, the input features of the masked multi-head attention module and the output features of the masked multi-head attention module are simultaneously input into the first residual module, the output end of the first residual module is connected to the input end of the first normalization module, the output end of the first normalization module is connected to the input end of the multi-head attention module, the output end of the multi-head attention module is connected to the input end of the second residual module, and the output end of the second residual module is connected to the input end of the second normalization module; The depth cross-attention layer includes a masked multi-head attention module, two residual modules, two normalization modules, and a feed-forward neural network; among them, the input features of the masked multi-head attention module and the output features of the masked multi-head attention module are simultaneously input into the first residual module, the output end of the first residual module is connected to the input end of the first normalization module, the output end of the first normalization module is connected to the input end of the feed-forward neural network, the input features of the feed-forward neural network and the output features of the feed-forward neural network are simultaneously input into the second residual module, and the output end of the second residual module is connected to the input end of the second normalization module; The Head network includes a classification network, a regression network, and a multi-object bounding box module; the classification network and the regression network are parallel, and the input features of the classification network and the regression network are the same; the output end of the regression network is connected to the input end of the multi-object bounding box module; in the regression network, the depth information output by the regression network and the depth information output by the Transformer depth information aggregation module are adaptively aggregated to obtain new depth information; The two-dimensional bounding box information output by the regression network and the new depth information are input into the multi-object bounding box module. The multi-object bounding box module determines the optimal label by setting pseudo-labels and obtains the three-dimensional bounding box information based on the optimal label.
3. The monocular three-dimensional object detection method based on Transformer-assisted depth information fusion according to claim 2, wherein: The adaptable aggregation formula is as follows: F Z = λ t F1 + (1 - λ t )F2 Among them, F Z represents the information adaptable to aggregation. When the information adaptable to aggregation is the output features of the convolutional attention module and the output features of the coordinate attention module, F1 and F2 respectively represent the feature information output by the convolutional attention module and the feature information output by the coordinate attention module. When the information adaptable to aggregation is the depth information output by the deep network and the depth information output by the Transformer depth information aggregation module, F1 and F2 respectively represent the depth information output by the deep network and the depth information output by the Transformer depth information aggregation module, λ t represents the learnable gating weight parameter.
4. The monocular three-dimensional object detection method based on Transformer-assisted depth information fusion according to claim 3, wherein: The multi-object bounding box module finds the optimal label through the following steps: First, set a group of size error tolerances Δz to interfere with the depth Z of each label P=(class, X, Y, Z, H, W, L, yaw). Obtain the depth offset Δz·Z according to the size error tolerance Δz and the object depth Z. Interfere with the object depth according to the depth offset to obtain the interfered object depth as Z + Δz·Z, and further obtain the three-dimensional pseudo-label P'=(class, X, Y, Z + Δz·Z, H, W, L, yaw); Among them, class represents the category, (X, Y, Z) represents the coordinates of the object in the camera coordinate system, Z also represents the depth, H represents the height of the three-dimensional object, W represents the width of the three-dimensional object, L represents the length of the three-dimensional object, and yaw represents the yaw angle; Second, obtain the two-dimensional bounding box coordinates of the label and the three-dimensional bounding box coordinates of the corresponding pseudo-label. Project the three-dimensional bounding box coordinates of the corresponding pseudo-label onto the two-dimensional bounding box plane of the label, and determine the projected bounding box according to the maximum and minimum values of the projected coordinates; Then, measure the similarity between the two-dimensional bounding box of the label and the projected bounding box through the IoU label score, and determine the IoU label score with the highest similarity; MAX IoU Label = IoU(B gt , B false ) Among them, IoU Label represents the label score, and B gt represents the two-dimensional bounding box of the label, and B false represents the projected bounding box; Finally, determine the best label P″ = (class, X t , Y t , Z + Δz·Z, H, W, L, yaw t ) according to the IoU label score with the highest similarity; Adjust the yaw angle according to the IoU label score with the highest similarity. The calculation formula is as follows: Among them, yaw represents the yaw angle, and yaw t represents the adjusted yaw angle, represents the IoU label score with the highest similarity, H represents the height of the three-dimensional object, W represents the width of the three-dimensional object, L represents the length of the three-dimensional object, Y represents the Y-axis coordinate of the three-dimensional object in the camera coordinate system, and exp represents the exponential function with the natural constant e as the base; Adjust the X-axis coordinate and Y-axis coordinate in the camera coordinate system according to the IoU label score with the highest similarity; Calculate the adjusted X-axis coordinate and Y-axis coordinate according to the object depth Z + Δz·Z in the IoU label score with the highest similarity. The calculation formula is as follows: Among them, X t represents the adjusted X-axis coordinate, Y t represents the adjusted Y-axis coordinate, (u, v) represents the two-dimensional coordinates of a point on the image, f x 、f y 、c x and c y represent the intrinsic parameters of the camera, and Z + Δz·Z represents the object depth corresponding to the IoU label score with the highest similarity.
5. The monocular three-dimensional object detection method based on Transformer-assisted depth information fusion according to claim 1, characterized in that: In step S2, Stochastic Gradient Descent (SGD) optimizer with weight decay of 10 -4 is adopted and trained with the following loss function L: L = L cls +L offset +L size +L rotsin +L velo +L depth Among them, L cls represents the classification loss, L offset represents the geometric center point offset loss, L size represents the size loss, L rotsin represents the rotation loss, L velo represents the speed loss, L size represents the depth loss.
6. The monocular three-dimensional object detection method based on Transformer-assisted depth information fusion according to claim 1, wherein: In step S3, when 48 epochs are trained, the monocular three-dimensional object detection model training is completed.
Citation Information
Cited By
Three-dimensional target detection method and system based on two-dimensional detection result prompt
CN120707994A