Monocular 3D target detection method and device
By using a trained monocular 3D object detection model, multi-scale in-depth features are generated using a backbone network and feature processing module. These features are then encoded using a depth and visual encoder, which solves the problem of limited feature extraction capability in complex scenes for monocular 3D object detection methods and achieves higher detection reliability and stability.
Patent Information
- Application Number
- CN202511167795.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-14
AI Technical Summary
Existing monocular 3D target detection methods have limited feature extraction capabilities in complex scenes, resulting in low detection reliability.
A pre-trained monocular 3D target detection model is used to extract multi-scale hierarchical features through a backbone network. The feature processing module captures features along multiple directions to generate multi-scale in-depth features, which are then encoded by a depth encoder and a visual encoder. Visually guided decoding is performed by a decoder, and finally, attribute detection is performed by predicting the head.
It enhances the model's generalization and detection reliability, maintains stable monocular 3D target detection performance in unseen environments, improves the limitations of traditional methods in feature extraction in complex scenes, and enhances detection accuracy.
Smart Images

Figure CN120953707A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target detection technology, and in particular to a monocular 3D target detection method and apparatus. Background Technology
[0002] Monocular 3D object detection aims to detect the category of an object and its position, size, and orientation in three-dimensional space from a two-dimensional image captured by a single camera. Traditional monocular 3D object detection methods typically rely on predefined anchor boxes with fixed orientations for object localization, which is difficult to adapt to targets of varying sizes and orientations in complex scenes, resulting in limited feature extraction capabilities and low detection reliability. Summary of the Invention
[0003] This invention provides a monocular 3D target detection method and apparatus, which solves the technical problem that existing monocular 3D target detection methods have limited feature extraction capabilities in complex scenes, resulting in low detection reliability.
[0004] The first aspect of this invention provides a monocular 3D target detection method, comprising: Obtain a trained monocular 3D object detection model; the trained monocular 3D object detection model includes a backbone network, a feature processing module, a depth encoder, a visual encoder, a decoder, and a prediction head; The backbone network is used to extract features from the monocular image under test to determine multi-scale hierarchical features. The feature processing module is used to capture features along multiple directions for the multi-scale hierarchical features, generating multi-scale in-depth features; Based on the multi-scale deepening features, depth input features and visual input features are determined, and corresponding input depth encoders and visual encoders are encoded to output depth encoded features and visual global features. After performing foreground depth prediction based on the depth input features, pixel depth estimation is performed to determine the depth position encoding, which is then added to the depth encoding features to generate global depth features. After the decoder performs visually guided decoding based on the visual global features and the depth global features, the input prediction head is used for attribute detection, and the target detection result is output.
[0005] Furthermore, the feature processing module includes a 1×1 convolution branch, a horizontal convolution branch, a vertical convolution branch, a dilated convolution branch, a global pooling branch, a convolutional residual layer, and a convolutional projection layer; the processing procedure of the feature processing module includes: The input hierarchical features are extracted in parallel using 1×1 convolutional branches, horizontal convolutional branches, vertical convolutional branches, dilated convolutional branches, and global pooling branches. The channels are then concatenated to generate intermediate processing features. The global pooling branch includes a spatial pyramid pooling module. After feature mapping of intermediate processed features through convolutional projection layers, the features are added element-wise with the hierarchical features processed by convolutional residual layers to output the corresponding deepened features.
[0006] Furthermore, in the feature processing modules corresponding to hierarchical features at different scales, the dilation rate of the horizontal convolution branch, the vertical convolution branch, and the dilated convolution branch decreases scale by scale.
[0007] Furthermore, determining the depth input features and visual input features based on the multi-scale deepening features includes: The smallest scale deepening feature among the multi-scale deepening features is used as the visual input feature; The deepening features of each scale in the multi-scale deepening features are aligned and then added element by element. A convolutional layer is then used for feature processing to determine the visual input features.
[0008] Furthermore, the decoder includes a cascaded visual cross-attention module, a self-attention module, a deep cross-attention module, and a feedforward network.
[0009] Furthermore, the acquisition of the trained monocular 3D object detection model includes: Based on the monocular image training set, a segmented collaborative adjustment mechanism for the learning rate is established using linear preheating, cooling-off period locking, plateau period adaptive decay, active boosting, and dynamic cosine annealing, combined with a preset total loss function. This mechanism trains the monocular 3D object detection model to be trained, thus determining the trained monocular 3D object detection model.
[0010] A second aspect of the present invention provides a monocular 3D target detection device, comprising: The model training module is used to obtain a trained monocular 3D object detection model; the trained monocular 3D object detection model includes a backbone network, a feature processing module, a depth encoder, a visual encoder, a decoder, and a prediction head; The feature extraction module is used to extract features from the monocular image to be tested through the backbone network and determine multi-scale hierarchical features. A multi-directional extraction module is used to capture features along multiple directions using the feature processing module to generate multi-scale in-depth features; The feature encoding module is used to determine the depth input features and visual input features based on the multi-scale deepening features, perform encoding processing on the corresponding input depth encoder and visual encoder, and output the depth encoded features and visual global features. The deep embedding module is used to perform foreground depth prediction based on the depth input features, then perform pixel depth estimation, determine the depth position code, and add it to the depth encoding features to generate a global depth feature. The decoding and detection module is used to perform visually guided decoding based on the visual global features and the depth global features by the decoder, input the prediction head for attribute detection, and output the target detection result.
[0011] A computer device provided in a third aspect of the present invention includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the monocular 3D target detection method as described in any of the preceding claims.
[0012] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the monocular 3D target detection method as described in any of the preceding claims.
[0013] The fifth aspect of the present invention provides a computer program product comprising a computer program / instructions, wherein the computer program / instructions, when executed by a processor, implement the monocular 3D target detection method as described in any of the preceding claims.
[0014] As can be seen from the above technical solutions, the present invention has the following advantages: The above-described solution of the present invention provides a monocular 3D target detection method, comprising: acquiring a trained monocular 3D target detection model; the trained monocular 3D target detection model includes a backbone network, a feature processing module, a depth encoder, a visual encoder, a decoder, and a prediction head; extracting features from the monocular image to be tested through the backbone network to determine multi-scale hierarchical features; using the feature processing module to capture features along multiple directions for the multi-scale hierarchical features to generate multi-scale deepening features; determining depth input features and visual input features based on the multi-scale deepening features, encoding them corresponding to the depth encoder and visual encoder, and outputting depth encoded features and visual global features; performing foreground depth prediction based on the depth input features and then pixel depth estimation to determine the depth position encoding, and adding it to the depth encoded features to generate depth global features; performing visually guided decoding based on the visual global features and depth global features through the decoder, inputting it to the prediction head for attribute detection, and outputting the target detection result. Based on the above scheme, the feature processing module provides the ability to capture targets of different sizes and shapes, while enhancing the generalization of the model, enabling it to maintain stable monocular 3D target detection performance in unseen environments. Visually guided feature decoding enables faster focusing on salient targets, providing clearer semantic anchors for subsequent integration of depth information, improving the limitations of traditional methods in feature extraction in complex scenes, and helping to improve detection reliability. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart of the steps of a monocular 3D target detection method provided in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the architecture of the monocular 3D target detection model provided in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the feature processing module provided in Embodiment 1 of the present invention; Figure 4 This is a flowchart of the monocular 3D target detection method using the Kitti dataset as an example, provided in Embodiment 1 of the present invention. Figure 5 This is a schematic diagram of the target detection results of the Kitti dataset provided in Embodiment 1 of the present invention. Figure 1 ; Figure 6 This is a schematic diagram of the target detection results of the Kitti dataset provided in Embodiment 1 of the present invention. Figure 2 ; Figure 7 This is a structural block diagram of a monocular 3D target detection device provided in Embodiment 2 of the present invention. Detailed Implementation
[0017] This invention provides a monocular 3D target detection method and apparatus to address the technical problem that existing monocular 3D target detection methods have limited feature extraction capabilities in complex scenes, resulting in low detection reliability.
[0018] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0019] Please see Figure 1 The present invention provides a monocular 3D target detection method, comprising: Step 101: Obtain the trained monocular 3D object detection model.
[0020] It should be noted that this embodiment is for a monocular 3D object detection task, and the model design is based on the MonoDETR network architecture, such as... Figure 2 As shown, the model includes a backbone network, a feature processing module, a depth encoder, a visual encoder, a decoder, and a prediction head. After the model is built, it is trained. When the iteration stopping condition is met, the trained monocular 3D object detection model is obtained.
[0021] Step 102: Extract features from the monocular image to be tested using a backbone network to determine multi-scale hierarchical features.
[0022] The monocular image to be tested refers to the image for which monocular 3D object detection is to be performed.
[0023] Multi-scale hierarchical features refer to hierarchical features that include multiple scales; hierarchical features refer to the features output at the corresponding level of feature extraction.
[0024] It should be noted that the RGB image obtained by the monocular camera, i.e. the monocular image to be tested, is used as the input of the trained monocular 3D object detection model. Multi-scale feature extraction is performed through the backbone network to obtain output features at different scales, thus providing a basis for subsequent depth estimation and feature processing. In specific implementation, the backbone network may include a ResNet50 network.
[0025] Step 103: Use the feature processing module to capture features along multiple directions for multi-scale hierarchical features and generate multi-scale in-depth features.
[0026] Multi-scale deepening features refer to deepening features that include multiple scales, with each level of the deepening feature corresponding to a specific level feature.
[0027] In one specific implementation of this embodiment, such as Figure 3 As shown, the feature processing module includes a 1×1 convolution branch, a horizontal convolution branch, a vertical convolution branch, a dilated convolution branch, a global pooling branch, a convolutional residual layer, and a convolutional projection layer; the processing procedure of the feature processing module includes: The input hierarchical features are extracted in parallel using 1×1 convolutional branches, horizontal convolutional branches, vertical convolutional branches, dilated convolutional branches, and global pooling branches. The channels are then concatenated to generate intermediate processing features. The global pooling branch includes a spatial pyramid pooling module. After feature mapping of intermediate processed features through convolutional projection layers, the features are added element-wise with the hierarchical features processed by convolutional residual layers to output the corresponding deepened features.
[0028] In a more specific implementation of this embodiment, in the feature processing modules corresponding to hierarchical features at different scales, the dilation rate of the horizontal convolution branch, the vertical convolution branch, and the dilated convolution branch decreases scale by scale.
[0029] It should be noted that this embodiment achieves efficient multi-scale feature modeling through a five-branch parallel architecture based on orientation-sensitive convolutional kernels, multi-level dilation rate configuration, and hierarchical parameter adaptive mechanism, resulting in a feature processing module. This module dynamically adjusts parameters such as convolutional kernel size, dilation rate, and padding parameters according to the size of features at different levels. Taking three-scale hierarchical features as an example, including high-level semantic Layer0, mid-level context Layer1, and low-level detail Layer2, the specific implementation is as follows: Table 1. Schematic diagram of the branch-parallel architecture of the feature processing module.
[0030] From Table 1 and Figure 3 It can be seen that the feature processing module includes a 1×1 convolution branch, a horizontal convolution branch, a vertical convolution branch, a dilated convolution branch, a global pooling branch, a convolutional residual layer, and a convolutional projection layer. The convolutional residual layer includes cascaded 1×1 convolutions and BN units, and the convolutional projection layer (Project layer) includes cascaded 1×1 convolutions, BN units, and ReLU activation functions. For the 1×1 convolution branch, it is used to preserve the original channel dimensions as a benchmark for multimodal fusion. For the horizontal convolution, an asymmetric kernel is used to expand the receptive field along the image width direction. In the high-level features (Layer 0), a dilation rate of 6 and padding of 12 are configured to capture long-distance horizontal patterns. In the low-level features (Layer 2), a (1,3) kernel with a dilation rate of 2 is used to enhance fine-grained edge detection. For the vertical convolution, the convolution kernel is designed through positive interpolation. In high-level features, a dilation rate of 3 and padding of 6 are configured to identify large vertical objects, while in low-level features, a dilation rate of 1 and padding of 1 are configured to focus on local vertical details. For dilated convolution, a standard 3×3 convolution kernel is used in conjunction with hierarchical dilation rates to expand the receptive field, enabling cross-level context modeling and avoiding the loss of downsampling information. For the global pooling branch, spatial pyramid pooling is used to extract image-level semantics and suppress local noise interference. The specific structure of the spatial pyramid pooling module can be found in existing technologies. It should be noted that since the model in this embodiment is divided into visual and depth branches, the feature processing module can be further refined according to the features required by the subsequent visual encoder and depth encoder, including both visual feature processing and depth feature processing modules, both based on the structure of the aforementioned feature processing module. Understandably, taking natural scenes as an example, targets in natural scenes often have significant directional structures, such as vehicles with a width-to-height ratio of approximately 3:1, pedestrians with a width-to-height ratio of 1:2, and streetlights with a width-to-height ratio of 1:8, etc. Traditional square convolutional kernels (such as 3×3 / 5×5) have isotropic receptive fields and respond uniformly in all directions, making it impossible to distinguish horizontal / vertical directional structures. This can easily lead to feature confusion and make them unsuitable for such geometric characteristics. In scenes with strong directional structures (such as lane lines and vehicle edges in road scenes), these kernels will introduce noise from irrelevant directions. Introducing noise from irrelevant directions, this embodiment improves upon the use of asymmetric convolution kernels, which helps the model learn objects of various real shapes in the scene. Horizontal convolution slides only along the width direction, making it more responsive to horizontal edges. For example, it can better adapt to horizontal targets such as long buses and the top / bottom edges of vehicles in road scenes. Vertical convolution slides only along the height direction, making it more sensitive to vertical edges and adapting to the detection of vertical structural objects, such as utility poles and pedestrian outlines. By decoupling the kernel shape from the physical size of the target, the feature activation region can more accurately cover the target subject without redundant information. At the same time, it uses dilated convolution to capture multi-scale context and global pooling to integrate semantic information. It is understandable that in this embodiment, the padding values corresponding to the dilation rate of each layer strictly adhere to the principle that the input and output feature map sizes remain unchanged.
[0031] Step 104: Determine the depth input features and visual input features based on multi-scale deepening features, encode the corresponding input depth encoder and visual encoder, and output the depth encoded features and visual global features.
[0032] In a more specific implementation of this embodiment, determining depth input features and visual input features based on multi-scale deepening features includes: The smallest scale deepening feature among the multi-scale deepening features is used as the visual input feature; After feature alignment of the deepening features at each scale in the multi-scale deepening feature set, the features are added element by element, and then a convolutional layer is used for feature processing to determine the visual input features.
[0033] It should be noted that multi-scale deepening features include deepening features at multiple scales. In this embodiment, in the vision part, the deepening features corresponding to the lowest scale layer features of the last layer of the backbone network can be used as visual input features to be input into the visual encoder for encoding processing. In the depth part, the deepening features at each scale are aligned (e.g., the deepening features at three scales are unified to the middle scale) and then added element by element. Then, convolution processing is performed through convolutional layers (e.g., two 3×3 convolutions) to determine the depth input features as the input of the depth encoder, thereby obtaining the outputs of the two encoders as depth encoded features and visual global features, respectively. It can be understood that, correspondingly, the visual feature processing module includes only one feature processing module structure, while the depth feature processing module includes multiple feature processing modules structure. The parameter design of different feature processing modules corresponds one-to-one with the size of the layer features at the corresponding level. Both the visual encoder and the depth encoder are based on the Transformer encoder module. Each Transformer encoder module includes a cascaded multi-head self-attention layer and a feed-forward network layer. Each layer is equipped with residual connections and layer normalization mechanisms to form a lightweight encoding structure, which helps to reduce computational complexity and the number of parameters. In specific implementations, the visual encoder may include one Transformer encoder module, and the depth encoder may include one Transformer encoder module.
[0034] Step 105: After predicting the foreground depth based on the depth input features, perform pixel depth estimation, determine the depth position encoding, and add it to the depth encoding features to generate global depth features.
[0035] It's worth noting that the depth branch, unlike the vision branch, has an additional feature: dividing the continuous depth space into discrete intervals, each called a bin. For example, a depth range of 0-80 meters can be divided into 40 bins, each covering 2 meters. Since direct regression of depth values (continuous regression) is susceptible to noise and gradient instability, discretization transforms it into a probabilistic classification problem, making it easier to optimize. Specifically, foreground depth features are obtained by predicting the foreground depth using 1×1 convolutions on the depth input features. , The key feature is that it can represent the depth value of the detected target, ignore the depth value of the background, and all pixels within a bounding box have the same depth value. The softmax activation function is used to predict the probability value of each pixel belonging to each bin, i.e., the depth class confidence of each pixel. Then, the probability values are weighted and summed with the corresponding bin depth to obtain the estimated depth of each pixel. Based on the estimated depth, the depth position code of each pixel is determined. For details, please refer to existing technologies. Finally, the depth position code is added element by element to the depth coding feature to generate the global depth feature.
[0036] Step 106: After visually guided decoding based on visual global features and depth global features by the decoder, the target detection result is output by inputting the prediction head for attribute detection.
[0037] It should be noted that this embodiment takes into account that visual features usually contain richer semantic information and can clearly represent the target. Therefore, it prioritizes using global visual features to focus on the thermal regions of objects in the feature map, and then uses depth features to reconstruct the depth relationships of objects in the entire scene. Figure 2 As shown, the decoder includes a cascaded visual cross-attention module, a self-attention module, a deep cross-attention module, and a feedforward network. It initializes a series of object queues, which, through a series of interactions, ultimately become decoded queues. The entire decoding process includes: 1. Visual Cross-Attention Module Input: Object Queries (initial) + visual global features (from the visual encoder); Process: The query vector acts as the "questioner," and the global visual features act as the "knowledge base." Each query calculates the association weight with all positions of the features, focusing on relevant areas (e.g., the headlight area has a higher weight for vehicle queries). Output: Updated query vector (with initial perception of visual context); 2. Self-Attention (self-attention module) Input: The query vector updated in the previous step; Process: Query attention weights are calculated between each other to establish relationships between targets (such as the exclusion relationship between "vehicle query" and "pedestrian query") and to eliminate redundant assumptions (such as suppressing one of the two queries when they repeatedly detect the same target). Output: Relation-enhanced query vector; 3. Deep Cross-Attention Module Input: Query vector from the previous relationship enhancement step + deep global features (from the deep encoder); Process: Match the query with the depth map as a whole, and adjust the target's position on the Z-axis using 3D spatial coordinates; Output: The final query vector carrying 3D spatial information (top Decoded Object Queries); Finally, the decoder output is fed into the prediction head for attribute detection. The prediction head can include a series of MLP-based head structures, which in turn output the target detection results.
[0038] In one specific embodiment of this example, step 101 includes the following sub-steps: Based on the monocular image training set, a segmented collaborative adjustment mechanism for the learning rate is established using linear preheating, cooling-off period locking, plateau period adaptive decay, active boosting, and dynamic cosine annealing, combined with a preset total loss function. This mechanism trains the monocular 3D object detection model to be trained, thus determining the trained monocular 3D object detection model.
[0039] It should be noted that the monocular image training set refers to the collection of monocular image samples used for model training. This can be specifically set according to the target detection task, and this embodiment does not impose any restrictions. To train the monocular 3D target detection model using the monocular image training set, the hyperparameters such as the number of training epochs, batch size, and learning rate can be initialized first. For example, the number of training epochs can be 350 and the batch size can be 16. This can be set according to the training needs. Secondly, the total loss function can be set to perceive the model training status. In this embodiment, to optimize model performance, a five-stage dynamic learning rate adjuster is designed. This adjuster achieves adaptive optimization through a collaborative mechanism of linear warm-up, cool-down period locking, adaptive decay during plateau periods, active boosting, and dynamic cosine annealing. Specifically, it includes: 1. Linear preheating stage (Warmup) Triggering condition: The first 20 training cycles; Adjustment mechanism: The learning rate increases linearly from an initial value (e.g., 0.1base_lr) to the base learning rate base_lr; Purpose: To avoid drastic gradient fluctuations in the early stages of training and improve stability; 2. Cooldown lock Triggering condition: After each learning rate adjustment, a cooldown period (8 epochs) begins. Adjustment mechanism: The learning rate is locked during the cooldown period, and further adjustments are prohibited; Function: To prevent oscillations caused by frequent adjustments and ensure stable parameter updates; 3. Adaptive decay during plateau period Triggering condition: The validation loss does not decrease for 6 consecutive epochs and is within a specified epoch window; the epoch window is understood as a range of multiple epochs divided within a set number of training rounds, such as the 20th to 300th epoch. Adjustment mechanism: Reduce the learning rate by a decay factor (e.g., plateau_factor=0.5), with the lower limit of the learning rate being min_lr; Function: Adaptive decay during plateau phase; 4. Active Boost Mechanism Triggering condition: If the current loss is higher than the reference value (e.g., 0.85*ref_loss) within a specified period window. Adjustment mechanism: Increase the learning rate by a boost factor (e.g., boost_factor=1.05), with the upper limit of the learning rate being max_lr; Function: Actively escapes local optima, enhancing the model's exploratory capabilities; 5. Dynamic Cosine Annealing Triggering condition: Starts in the later stages of training (epoch ≥ cosine annealing start period cosine_start = 300); Adjustment mechanism: The learning rate is adjusted periodically and smoothly based on the cosine function; Function: To balance exploration and development, and to approach the global optimum; It is understandable that the two adjustment strategies, adaptive decay during the plateau period and active boost mechanism, will not be triggered simultaneously. After one of the strategies is triggered, there is an 8-epoch cooling-off period during which the program is not allowed to adjust the learning rate so that the model can adapt to this learning rate and train effectively. Finally, a dynamic cosine annealing strategy is used in 300-350 epochs to gradually reduce the learning rate. For the loss function, the query features decoded by the decoder are fed into the prediction head, and the output includes object category, 2D size, projected 3D center, depth, 3D size, and orientation. To correctly match each query with the real object, the loss for each query-label pair is calculated, and the Hungarian algorithm is used to find the globally optimal match. For each match, the six attribute losses are divided into two groups: the first group is object category, 2D size, and projected 3D center, and the second group is depth, 3D size, and orientation. The second group belongs to the object's 3D spatial attributes. Based on the aforementioned grouping, the two groups of losses are summed separately, denoted as 2D loss and 3D loss. It is understandable that since the network's prediction of 3D attributes is usually less accurate than that of 2D attributes (especially in the early stages of training), the instability of 3D loss may interfere with the matching process. Therefore, only 2D loss is used as the matching cost for query-label pairing. After matching is completed, the results are selected from N queries. There are valid matching pairs, where N is the maximum number of detected targets; the total loss function includes:
[0040] In the formula, For the total loss function, To effectively match the number of pairs, For an index to be used for a valid match, For 2D attribute loss, For 3D attribute loss, Foreground depth features Focal loss.
[0041] To better illustrate the technical effects of this embodiment, refer to... Figure 4 Taking the KITTI dataset as an example, this embodiment uses the monocular 3D object detection model for object detection. Figure 5 and Figure 6 The corresponding target detection results are shown, which can be well adapted to the detection of vehicles of different sizes. It should be noted that... Figure 4 This document provides a brief overview of the general process of a monocular 3D target detection method. For details on the specific implementation of each step, please refer to the relevant content in the aforementioned embodiment.
[0042] In this embodiment of the invention, the feature processing module provides the ability to capture targets of different sizes and shapes, while enhancing the generalization of the model, enabling it to maintain stable monocular 3D target detection performance in unseen environments. Visually guided feature decoding enables faster focusing on salient targets, providing clearer semantic anchors for subsequent integration of depth information. This improves the limitations of traditional methods in feature extraction in complex scenes, helps to enhance detection reliability, and is applicable to fields such as autonomous driving and industrial quality inspection, possessing good technical advantages and commercial potential.
[0043] Please see Figure 7 The second embodiment of the present invention provides a monocular 3D target detection device, comprising: The model training module 701 is used to obtain a trained monocular 3D object detection model; the trained monocular 3D object detection model includes a backbone network, a feature processing module, a depth encoder, a visual encoder, a decoder, and a prediction head. Feature extraction module 702 is used to extract features from the monocular image to be tested through the backbone network and determine multi-scale hierarchical features; The multi-directional extraction module 703 is used to capture features along multiple directions for multi-scale hierarchical features using the feature processing module, and generate multi-scale in-depth features. The feature encoding module 704 is used to determine the depth input features and visual input features based on multi-scale deepening features, perform encoding processing on the corresponding input depth encoder and visual encoder, and output depth encoded features and visual global features. The deep embedding module 705 is used to perform pixel depth estimation after predicting the foreground depth based on the depth input features, determine the depth position encoding, and add it to the depth encoding features to generate global depth features. The decoding and detection module 706 is used to perform attribute detection by inputting the prediction head after visually guided decoding based on visual global features and depth global features by the decoder, and output the target detection result.
[0044] In one specific embodiment of this example, the feature processing module includes a 1×1 convolution branch, a horizontal convolution branch, a vertical convolution branch, a dilated convolution branch, a global pooling branch, a convolutional residual layer, and a convolutional projection layer; the processing procedure of the feature processing module includes: The input hierarchical features are extracted in parallel using 1×1 convolutional branches, horizontal convolutional branches, vertical convolutional branches, dilated convolutional branches, and global pooling branches. The channels are then concatenated to generate intermediate processing features. The global pooling branch includes a spatial pyramid pooling module. After feature mapping of intermediate processed features through convolutional projection layers, the features are added element-wise with the hierarchical features processed by convolutional residual layers to output the corresponding deepened features.
[0045] In a more specific implementation of this embodiment, in the feature processing modules corresponding to hierarchical features at different scales, the dilation rate of the horizontal convolution branch, the vertical convolution branch, and the dilated convolution branch decreases scale by scale.
[0046] In one specific implementation of this embodiment, determining depth input features and visual input features based on multi-scale deepening features includes: The smallest scale deepening feature among the multi-scale deepening features is used as the visual input feature; After feature alignment of the deepening features at each scale in the multi-scale deepening feature set, the features are added element by element, and then a convolutional layer is used for feature processing to determine the visual input features.
[0047] In one specific embodiment of this example, the decoder includes a cascaded visual cross-attention module, a self-attention module, a deep cross-attention module, and a feedforward network.
[0048] In one specific implementation of this embodiment, obtaining a trained monocular 3D object detection model includes: Based on the monocular image training set, a segmented collaborative adjustment mechanism for the learning rate is established using linear preheating, cooling-off period locking, plateau period adaptive decay, active boosting, and dynamic cosine annealing, combined with a preset total loss function. This mechanism trains the monocular 3D object detection model to be trained, thus determining the trained monocular 3D object detection model.
[0049] This invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program; when the computer program is executed by the processor, the processor performs the steps of the monocular 3D target detection method as described in any of the above embodiments.
[0050] This invention also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implement the steps of the monocular 3D target detection method as described in any of the above embodiments.
[0051] This invention also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the monocular 3D target detection method as described in any of the above embodiments.
[0052] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0053] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0054] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0055] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0056] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0057] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A monocular 3D target detection method, characterized in that, include: Obtain a trained monocular 3D object detection model; the trained monocular 3D object detection model includes a backbone network, a feature processing module, a depth encoder, a visual encoder, a decoder, and a prediction head; The backbone network is used to extract features from the monocular image under test to determine multi-scale hierarchical features. The feature processing module is used to capture features along multiple directions for the multi-scale hierarchical features, generating multi-scale in-depth features; Based on the multi-scale deepening features, depth input features and visual input features are determined, and corresponding input depth encoders and visual encoders are encoded to output depth encoded features and visual global features. After performing foreground depth prediction based on the depth input features, pixel depth estimation is performed to determine the depth position encoding, which is then added to the depth encoding features to generate global depth features. After the decoder performs visually guided decoding based on the visual global features and the depth global features, the input prediction head is used for attribute detection, and the target detection result is output.
2. The monocular 3D target detection method according to claim 1, characterized in that, The feature processing module includes a 1×1 convolution branch, a horizontal convolution branch, a vertical convolution branch, a dilated convolution branch, a global pooling branch, a convolutional residual layer, and a convolutional projection layer; the processing procedure of the feature processing module includes: The input hierarchical features are extracted in parallel using 1×1 convolutional branches, horizontal convolutional branches, vertical convolutional branches, dilated convolutional branches, and global pooling branches. The channels are then concatenated to generate intermediate processing features. The global pooling branch includes a spatial pyramid pooling module. After feature mapping of intermediate processed features through convolutional projection layers, the features are added element-wise with the hierarchical features processed by convolutional residual layers to output the corresponding deepened features.
3. The monocular 3D target detection method according to claim 2, characterized in that, In the feature processing modules corresponding to hierarchical features at different scales, the dilation rate of the horizontal convolution branch, the vertical convolution branch, and the dilated convolution branch decreases scale by scale.
4. The monocular 3D target detection method according to claim 1, characterized in that, The determination of depth input features and visual input features based on the multi-scale deepening features includes: The smallest scale deepening feature among the multi-scale deepening features is used as the visual input feature; The deepening features of each scale in the multi-scale deepening features are aligned and then added element by element. A convolutional layer is then used for feature processing to determine the visual input features.
5. The monocular 3D target detection method according to claim 1, characterized in that, The decoder includes a cascaded visual cross-attention module, a self-attention module, a deep cross-attention module, and a feedforward network.
6. The monocular 3D target detection method according to claim 1, characterized in that, The process of obtaining the trained monocular 3D object detection model includes: Based on the monocular image training set, a segmented collaborative adjustment mechanism for the learning rate is established using linear preheating, cooling-off period locking, plateau period adaptive decay, active boosting, and dynamic cosine annealing, combined with a preset total loss function. This mechanism trains the monocular 3D object detection model to be trained, thus determining the trained monocular 3D object detection model.
7. A monocular 3D target detection device, characterized in that, include: The model training module is used to obtain a trained monocular 3D object detection model; the trained monocular 3D object detection model includes a backbone network, a feature processing module, a depth encoder, a visual encoder, a decoder, and a prediction head; The feature extraction module is used to extract features from the monocular image to be tested through the backbone network and determine multi-scale hierarchical features. A multi-directional extraction module is used to capture features along multiple directions using the feature processing module to generate multi-scale in-depth features; The feature encoding module is used to determine the depth input features and visual input features based on the multi-scale deepening features, perform encoding processing on the corresponding input depth encoder and visual encoder, and output the depth encoded features and visual global features. The deep embedding module is used to perform foreground depth prediction based on the depth input features, then perform pixel depth estimation, determine the depth position code, and add it to the depth encoding features to generate a global depth feature. The decoding and detection module is used to perform visually guided decoding based on the visual global features and the depth global features by the decoder, input the prediction head for attribute detection, and output the target detection result.
8. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the steps of the monocular 3D target detection method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the monocular 3D target detection method as described in any one of claims 1-6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the monocular 3D target detection method as described in any one of claims 1-6.