A monocular 3D target detection method and device
By introducing multimodal feature fusion module and Contextual transformer module into the ControlNet model, high-quality 3D object detection data is generated, which solves the problem of low accuracy of monocular 3D object detection, and realizes more efficient data set expansion and model training.
Patent Information
- Application Number
- CN202411719713.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-28
AI Technical Summary
In the prior art, the accuracy of monocular 3D object detection is mainly due to the difficulty in producing data sets, insufficient diversity and high cost.
The Fusion-ControlNet model architecture with the multimodal feature fusion module added to the ControlNet model architecture is adopted, and it is introduced into the Stable Diffusion model to generate and label a large number of high-quality 3D object detection data, and the model's characterization ability is enhanced through the Contextual transformer module.
It reduces the cost and difficulty of data set production, improves the accuracy of 3D object detection, and solves the problem of low accuracy of monocular 3D object detection.
Smart Images

Figure CN119206196B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a monocular 3D target detection method and device. Background Art
[0002] Accurate detection of 3D objects in various scenarios has wide applications in autonomous driving, virtual reality, and robotics. To obtain accurate perception of 3D information, many 3D object detection methods rely on depth sensors (such as LiDAR), which, while effective, often face the obstacles of high cost and relatively sparse data provided. In contrast, monocular 3D object detection has attracted increasing attention due to the use of monocular RGB cameras as a simpler and cheaper deployment setup.
[0003] In 3D object detection, the amount of data in a dataset plays a very important role. More data often means better final results. However, the production of 3D object detection datasets often faces some limitations, such as: (1) difficulty in labeling, especially in complex scenes; (2) insufficient data diversity, which may not cover all practical situations; (3) large amount of data, but may not be enough for certain object categories or scenes; (4) high data collection cost. Due to the above reasons, it is very difficult to produce or expand 3D object detection datasets by yourself. This leads to the problem of low accuracy of monocular 3D object detection. Summary of the invention
[0004] In view of this, an object of the present invention is to provide a monocular 3D target detection method and device, aiming to solve the problem of low accuracy in identifying monocular 3D targets in the prior art.
[0005] An object of the present invention is to provide a monocular 3D target detection method, the method comprising:
[0006] Obtain a target detection image to be detected;
[0007] Inputting the target detection image into a pre-trained target detection model to obtain target detection information of the target detection image;
[0008] Among them, the training process of the target detection model is:
[0009] Add a multimodal feature fusion module to the ControlNet model architecture to establish the Fusion-ControlNet model architecture;
[0010] The Fusion-ControlNet model architecture is introduced as a guidance module for conditional generation into the StableDiffusion model architecture to obtain a data set expansion model.
[0011] Obtain a KITTI dataset, and use the dataset expansion model to expand the KITTI dataset to obtain a KITTI expanded dataset;
[0012] Use the Contextual transformer module to replace the self-attention mechanism in the MonoDETR model architecture to establish the CoT-MonoDETR model architecture;
[0013] The KITTI augmented data set is input into the CoT-MonoDETR model framework for training to obtain the target detection model.
[0014] Furthermore, in the above-mentioned monocular 3D target detection method, the multimodal feature fusion module is a multimodal feature fusion encoder, which is used to input the semantic image and the depth image into two ResNet50 pre-trained models respectively, and through the adaptive feature fusion module connected between the two ResNet50 models, so that the two images of different modalities can adaptively fuse different modal features.
[0015] Furthermore, in the above monocular 3D target detection method, the step of introducing the Fusion-ControlNet model architecture as a conditional generation guidance module into the Stable Diffusion model architecture to obtain a data set expansion model includes:
[0016] The weights of the trained Stable Diffusion model architecture are frozen, and the Fusion-ControlNet model architecture is trained to obtain the data set expansion model.
[0017] Furthermore, in the above monocular 3D target detection method, the step of obtaining a KITTI dataset and using the dataset expansion model to expand the KITTI dataset to obtain a KITTI expanded dataset includes:
[0018] The dataset expansion model is used to generate a KITTI diffusion dataset using the KITTI dataset, and the annotations of the KITTI dataset are directly projected into the corresponding newly generated KITTI diffusion dataset through a projection method, thereby obtaining pseudo labels of new images and obtaining the final KITTI expansion dataset.
[0019] Furthermore, in the above monocular 3D target detection method, the training process of the CoT-MonoDETR model architecture includes:
[0020] Using a pre-trained convolutional neural network as the backbone feature extractor of the CoT-MonoDETR model architecture;
[0021] Extract multi-scale feature maps from the RGB images of the KITTI augmented dataset and input them into the CoT-Transformer encoder of the CoT-MonoDETR model architecture, where the CoT-Transformer encoder consists of a multi-layer CoT attention mechanism and a feed-forward neural network, and the long-distance dependencies between objects in the image are captured through the multi-layer CoT attention mechanism;
[0022] In the multi-layer CoT attention mechanism, given an input X, the key vector, query vector, and value vector are defined as K=X, Q=X, V=XWV, respectively, where W is the weight matrix;
[0023] Perform context encoding on all adjacent key vectors in the K*V grid in space to obtain the context key vector representation ;
[0024] The context key vector is represented as It is concatenated with the query vector on the channel, and then the attention matrix A is obtained through two consecutive 1*1 convolutions;
[0025] According to the attention matrix A, the weighted feature map is obtained by aggregating all V. , the final output and Fusion of
[0026] 3D object detection is performed through the Transformer decoder. The decoder uses a learnable object query vector, performs a cross-attention operation on the query vector and the features output by the CoT-Transformer encoder, and finally generates the detection result of the 3D object.
[0027] Furthermore, in the above monocular 3D target detection method, the loss function of the data set expansion model is:
[0028] ;
[0029] in, represents the original image, represents the actual noise added, represents the image after adding noise, t represents the diffusion time step, Indicates text conditions, Indicates potential conditions, Represents a noisy estimate of the model output.
[0030] Furthermore, in the above monocular 3D target detection method, the loss function of the CoT-MonoDETR model architecture is:
[0031] ;
[0032] in, is the classification loss, is the geometric center point offset loss, is the size loss, is the rotation loss, is the speed loss, is the depth loss.
[0033] Another object of the present invention is to provide a monocular 3D object detection device, the device comprising:
[0034] An acquisition module is used to acquire a detection image of a target to be detected;
[0035] A detection module, used for inputting the target detection image into a pre-trained target detection model to obtain target detection information of the target detection image;
[0036] Among them, the training process of the target detection model is:
[0037] Add a multimodal feature fusion module to the ControlNet model architecture to establish the Fusion-ControlNet model architecture;
[0038] The Fusion-ControlNet model architecture is introduced as a guidance module for conditional generation into the StableDiffusion model architecture to obtain a data set expansion model.
[0039] Obtain a KITTI dataset, and use the dataset expansion model to expand the KITTI dataset to obtain a KITTI expanded dataset;
[0040] Use the Contextual transformer module to replace the self-attention mechanism in the MonoDETR model architecture to establish the CoT-MonoDETR model architecture;
[0041] The KITTI augmented data set is input into the CoT-MonoDETR model framework for training to obtain the target detection model.
[0042] Another object of the present invention is to provide a readable storage medium having a computer program stored thereon, wherein the program implements the steps of the above method when executed by a processor.
[0043] Another object of the present invention is to provide an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.
[0044] The present invention establishes a Fusion-ControlNet model architecture by adding a multimodal feature fusion module to the ControlNet model architecture; introduces the Fusion-ControlNet model architecture as a conditional generation guidance module into the Stable Diffusion model architecture to obtain a data set expansion model, automatically generates and annotates a large amount of high-quality 3D target detection data, reduces the cost and difficulty of data set production, thereby ensuring the accuracy of model training and improving the accuracy of 3D target detection. The problem of low accuracy in identifying monocular 3D targets in the prior art is solved.
[0045] In addition, the present invention introduces the Contextual transformer module, which enhances the ability to characterize target features by combining static and dynamic context information, and significantly improves the performance of the model in 3D detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A flowchart of monocular 3D target detection provided by an embodiment of the present invention;
[0047] Figure 2 4 is a structural block diagram of a monocular 3D target detection device in a third embodiment of the present invention.
[0048] The following specific implementation manner will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0049] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0050] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0052] How to improve the accuracy of 3D object detection will be described in detail below with reference to specific embodiments and drawings.
[0053] Embodiment 1
[0054] See also Figure 1 , which shows a monocular 3D target detection method in a first embodiment of the present invention, and the method includes steps S10 to S11.
[0055] Step S10, obtaining a detection image of a target to be detected.
[0056] The target detection image is a monocular image containing a 3D target. Specifically, the target detection image is generally an RGB image.
[0057] Step S11, inputting the target detection image into a pre-trained target detection model to obtain target detection information of the target detection image.
[0058] Among them, the target detection model masters the detection logic of the information contained in the target detection image, so that the target detection information of the target detection image can be detected by the target detection model. Specifically, the target detection information includes the category of the target, the parameters of the 2D bounding box, the parameters of the 3D bounding box, the depth information and the orientation information;
[0059] More specifically, the training and generation of the target detection model of the embodiment of the present invention includes two parts, one part is to obtain high-quality detection data of 3D target detection images, and the other part is to use the detection data of high-quality 3D target detection images to perform specific training of the target detection model, wherein the existing ControlNet model architecture is adopted, and a multimodal feature fusion module is added to the ControlNet model architecture to establish a Fusion-ControlNet model architecture, so that it can effectively combine semantic information with depth information, and introduce the Fusion-ControlNet model architecture into the Stable Diffusion model as a guidance module for conditional generation, so that the generated image can accurately reflect semantics and depth information while retaining the sense of reality, and finally, the Fusion-ControlNet model architecture is introduced into the Stable Diffusion model architecture as a guidance module for conditional generation to obtain a data set expansion model for training so that it can automatically generate and annotate a large amount of high-quality 3D target detection data, thereby reducing the cost and difficulty of data set production for training the target detection model.
[0060] Furthermore, the KITTI dataset is used as the basic dataset for training the target detection model. The dataset consists of 7481 images for training and validation sets, and 7518 images for the test set, including a total of 80256 labeled objects. The main categories are cars, bicycles, pedestrians, and trucks. According to the height, degree of occlusion, and degree of truncation of the object, the detection results are divided into three levels: easy, medium, and difficult. The training images are divided into 5985 training sets and 1496 validation sets. When training the dataset expansion model, the autoencoder of the Stable Diffusion model architecture is first trained to encode the input image into the Latent Space. After the DDIM process, the information in the latent space is decoded back to the image. After training the autoencoder, import the weight file of the trained autoencoder model, use Latent Diffusion Model (LDM) as the basis, combine the UNet architecture and multi-scale attention mechanism, and use the loss function to train the StableDiffusion model. For example, the input and output of the model are both 3-channel 64x64 images. The linear learning rate scheduler is used during training, and the learning rate is gradually increased after 10,000 steps of warm-up. The number of channels of UNet is set to 192, the attention mechanism is applied at multiple resolutions, and upsampling and downsampling are achieved through residual blocks. The ultimate goal is to generate high-quality images. On a single GeForce RTX 4090 GPU, 300 epoches are trained with a learning rate of , using weight decay The SGD optimizer is used with a momentum factor of 0.9, a weight decay coefficient of 0.0001, 10,000 warm-up iterations, and a warm-up ratio of 0.33.
[0061] Furthermore, the Stable Diffusion model with Fusion-ControlNet model architecture is post-trained. First, data preprocessing is performed. In order to improve the interaction efficiency between the modalities, the original modal information needs to be processed. The potential deep features are recorded as , the semantic feature sequence is defined as , then the algorithm alignment input can be recorded as ,in is the number of targets in the scene, and the network input is The multimodal feature fusion module in the Fusion-ControlNet model architecture is a multimodal feature fusion encoder, which includes two Resnet50 models and an adaptive feature fusion module connected to the Resnet50 model. Specifically, for the acquired 3D target detection image, the semantic image and depth image in the 3D target detection image are input into two Resnet50 pre-trained models respectively, and the adaptive feature fusion block is used to transfer feature cues from one modality to another. This allows the two Resnet50 models of different modalities to adaptively fuse features of different modalities.
[0062] Specifically, the processing flow of the entire multimodal feature fusion encoder is shown in the following formula:
[0063] ;
[0064] in, is the depth map input, is the semantic graph input, is a multimodal feature fusion encoder, To output multimodal fusion features. It should be noted that since the UNet of the Stable Diffusion model accepts latent features instead of original images, the image-based conditions need to be converted to a 64×64 feature space in the multimodal feature fusion encoder to match the convolution size.
[0065] Specifically, in practice, the loss function of the dataset expansion model is:
[0066] ;
[0067] in, represents the original image, represents the actual noise added, represents the image after adding noise, t represents the diffusion time step, Indicates text conditions, Indicates potential conditions, Represents a noisy estimate of the model output.
[0068] During the training process, 50% of the text conditions are randomly replaced with empty strings. . This helps to better understand the meaning of the input condition graph.
[0069] Finally, the dataset expansion model is used to generate the KITTI diffusion dataset using the KITTI dataset. The annotations of the KITTI dataset are directly projected into the corresponding newly generated KITTI diffusion dataset through the projection method, thereby obtaining the pseudo labels of the new images and obtaining the final KITTI expansion dataset. Specifically, the pseudo labels of the new images are obtained by directly mapping the original image labels of the corresponding KITTI dataset. Since the generated image has the same objects and similar 3D properties as the original image, the pseudo labels can correspond one-to-one with the targets in the new image, and then constitute new data samples with the newly generated image.
[0070] In addition, the KITTI dataset is expanded from the original 7481 images to 10438 images using the expansion method described in the embodiment of the present invention, including a total of 130256 labeled objects, the main categories of which are cars, bicycles, pedestrians and trucks. According to the height, occlusion and truncation of the object, the detection results are divided into three levels: easy, medium and difficult. The training images are divided into 8350 training sets and 2088 validation sets.
[0071] In summary, the monocular 3D target detection method in the above embodiment of the present invention establishes a Fusion-ControlNet model architecture by adding a multimodal feature fusion module to the ControlNet model architecture; the Fusion-ControlNet model architecture is introduced as a conditional generation guidance module into the Stable Diffusion model architecture to obtain a data set expansion model, which automatically generates and annotates a large amount of high-quality 3D target detection data, reduces the cost and difficulty of data set production, and thus ensures the accuracy of model training and improves the accuracy of 3D target detection. The problem of low accuracy in identifying monocular 3D targets in the prior art is solved.
[0072] Embodiment 2
[0073] This embodiment also proposes a monocular 3D target detection method. The monocular 3D target detection method in this embodiment is different from the monocular 3D target detection method in the first embodiment in that:
[0074] The training process of the CoT-MonoDETR model architecture includes:
[0075] Using a pre-trained convolutional neural network as the backbone feature extractor of the CoT-MonoDETR model architecture;
[0076] Extract multi-scale feature maps from the RGB images of the KITTI augmented dataset and input them into the CoT-Transformer encoder of the CoT-MonoDETR model architecture, where the CoT-Transformer encoder consists of a multi-layer CoT attention mechanism and a feed-forward neural network, and the long-distance dependencies between objects in the image are captured through the multi-layer CoT attention mechanism;
[0077] In the multi-layer CoT attention mechanism, given an input X, the key vector, query vector, and value vector are defined as K=X, Q=X, V=XWV, respectively, where W is the weight matrix;
[0078] Perform context encoding on all adjacent key vectors in the K*V grid in space to obtain the context key vector representation ;
[0079] The context key vector is represented as It is concatenated with the query vector on the channel, and then the attention matrix A is obtained through two consecutive 1*1 convolutions;
[0080] According to the attention matrix A, the weighted feature map is obtained by aggregating all V. , the final output and Fusion of
[0081] 3D object detection is performed through the Transformer decoder. The decoder uses a learnable object query vector, performs a cross-attention operation on the query vector and the features output by the CoT-Transformer encoder, and finally generates the detection result of the 3D object.
[0082] Among them, the existing Transformer-based methods usually only use the self-attention mechanism and the cross-attention mechanism to capture the attention matrix of isolated query-key pairs (query vector-key vector), and fail to fully utilize the contextual information around the key vector, which limits the performance of the model. To address this limitation, this embodiment introduces the Contextual Transformer module, which enhances the representation ability of target features by combining static and dynamic context information, and significantly improves the performance of the model in 3D detection tasks.
[0083] Specifically, to train the CoT-MonoDETR model, first, the KITTI augmented dataset expanded in the previous step is used in the data preparation stage. In order to adapt to the input requirements of the model, the images are normalized and the image size is adjusted. In order to increase the diversity of training data, data augmentation techniques such as random cropping, rotation, and scaling are applied;
[0084] In the model construction stage, the CoT-MonoDETR model uses a pre-trained convolutional neural network (CNN) such as ResNet101 as the backbone feature extractor to extract multi-scale feature maps from RGB images. These feature maps are then input into the CoT-Transformer encoder to capture the long-distance dependencies between objects in the image through a multi-layer CoT attention mechanism. The CoT-Transformer encoder consists of a multi-layer CoT attention mechanism and a feed-forward neural network (FFN) to enhance the representation capability of features;
[0085] In the multi-layer CoT attention mechanism, given an input X, the key vector (key), query vector (query), and value vector (value) are defined as K=X, Q=X, V=XW respectively. V, Perform context encoding on all adjacent key vectors in the K*V grid in space to obtain the context key vector representation , which reflects the static context information of the local neighboring positions. Next, the context key vector is represented as The attention matrix A is obtained by concatenating the query vector with the query vector on the channel and performing two consecutive 1*1 convolutions. For each attention head, the local attention matrix (i.e., K*K grid size) of each spatial position of A is learned based on the query features and the key features with context, rather than isolated query-key pairs. This approach is effective in static contexts. Self-attention learning is enhanced with additional guidance from
[0086] Next, the model performs 3D object detection through the decoder. The decoder uses learnable object query vectors, which perform cross-attention operations with the features output by the encoder to finally generate 3D object detection results. The features output by the decoder are respectively predicted by the classification head, bounding box regression head, and depth estimation head to predict the target category, 3D bounding box parameters, and depth information of the target. During the training process, the model uses a variety of loss functions for optimization, specifically;
[0087] ;
[0088] in, is the classification loss, is the geometric center point offset loss, is the size loss, is the rotation loss, is the speed loss, is the depth loss.
[0089] In terms of training strategy, some layers of the backbone network are frozen at the beginning to accelerate model convergence, and then all layers are gradually unfrozen. For example, on a single GeForce RTX 4090 GPU, FCOSD3Dformer is trained for 48 epochs with a learning rate of ,The optimizer selects SGD, and combines the learning rate scheduler to dynamically adjust the learning rate, the momentum factor is set to 0.9, and the weight decay coefficient is set to 0.0001.
[0090] In summary, the monocular 3D target detection method in the above embodiment of the present invention establishes a Fusion-ControlNet model architecture by adding a multimodal feature fusion module to the ControlNet model architecture; the Fusion-ControlNet model architecture is introduced as a conditional generation guidance module into the Stable Diffusion model architecture to obtain a data set expansion model, which automatically generates and annotates a large amount of high-quality 3D target detection data, reduces the cost and difficulty of data set production, and thus ensures the accuracy of model training and improves the accuracy of 3D target detection. The problem of low accuracy in identifying monocular 3D targets in the prior art is solved.
[0091] Embodiment 3
[0092] See also Figure 2 , shown is a monocular 3D target detection device proposed in the third embodiment of the present invention, the device comprising:
[0093] The acquisition module 100 is used to acquire a detection image of a target to be detected;
[0094] A detection module 200 is used to input the target detection image into a pre-trained target detection model to obtain target detection information of the target detection image;
[0095] Among them, the training process of the target detection model is:
[0096] Add a multimodal feature fusion module to the ControlNet model architecture to establish the Fusion-ControlNet model architecture;
[0097] The Fusion-ControlNet model architecture is introduced as a guidance module for conditional generation into the StableDiffusion model architecture to obtain a data set expansion model.
[0098] Obtain a KITTI dataset, and use the dataset expansion model to expand the KITTI dataset to obtain a KITTI expanded dataset;
[0099] Use the Contextual transformer module to replace the self-attention mechanism in the MonoDETR model architecture to establish the CoT-MonoDETR model architecture;
[0100] The KITTI augmented data set is input into the CoT-MonoDETR model framework for training to obtain the target detection model.
[0101] Furthermore, in the above-mentioned monocular 3D target detection device, the multimodal feature fusion module is a multimodal feature fusion encoder, which is used to input the semantic image and the depth image into two ResNet50 pre-trained models respectively, and through the adaptive feature fusion module connected between the two ResNet50 models, so that the two images of different modalities can adaptively fuse different modal features.
[0102] Furthermore, in the above-mentioned monocular 3D target detection device, the step of introducing the Fusion-ControlNet model architecture as a guidance module for conditional generation into the Stable Diffusion model architecture to obtain a data set expansion model includes:
[0103] The weights of the trained Stable Diffusion model architecture are frozen, and the Fusion-ControlNet model architecture is trained to obtain the data set expansion model.
[0104] Furthermore, in the above-mentioned monocular 3D target detection device, the step of obtaining a KITTI data set and using the data set expansion model to expand the KITTI data set to obtain a KITTI expanded data set includes:
[0105] The dataset expansion model is used to generate a KITTI diffusion dataset using the KITTI dataset, and the annotations of the KITTI dataset are directly projected into the corresponding newly generated KITTI diffusion dataset through a projection method, thereby obtaining pseudo labels of new images and obtaining the final KITTI expansion dataset.
[0106] Furthermore, in the above-mentioned monocular 3D object detection device, the training process of the CoT-MonoDETR model architecture includes:
[0107] Using a pre-trained convolutional neural network as the backbone feature extractor of the CoT-MonoDETR model architecture;
[0108] Extract multi-scale feature maps from the RGB images of the KITTI augmented dataset and input them into the CoT-Transformer encoder of the CoT-MonoDETR model architecture, where the CoT-Transformer encoder consists of a multi-layer CoT attention mechanism and a feed-forward neural network, and the long-distance dependencies between objects in the image are captured through the multi-layer CoT attention mechanism;
[0109] In the multi-layer CoT attention mechanism, given an input X, the key vector, query vector, and value vector are defined as K=X, Q=X, V=XWV, respectively, where W is the weight matrix;
[0110] Perform context encoding on all adjacent key vectors in the K*V grid in space to obtain the context key vector representation ;
[0111] The context key vector is represented as It is concatenated with the query vector on the channel, and then the attention matrix A is obtained through two consecutive 1*1 convolutions;
[0112] According to the attention matrix A, the weighted feature map is obtained by aggregating all V. , the final output and Fusion of;
[0113] 3D object detection is performed through the Transformer decoder. The decoder uses a learnable object query vector, performs a cross-attention operation on the query vector and the features output by the CoT-Transformer encoder, and finally generates the detection result of the 3D object.
[0114] Furthermore, in the above monocular 3D object detection device, the loss function of the data set expansion model is:
[0115] ;
[0116] in, represents the original image, represents the actual noise added, represents the image after adding noise, t represents the diffusion time step, Indicates text conditions, Indicates potential conditions, Represents a noisy estimate of the model output.
[0117] Furthermore, in the above-mentioned monocular 3D target detection device, the loss function of the CoT-MonoDETR model architecture is:
[0118] ;
[0119] in, is the classification loss, is the geometric center point offset loss, is the size loss, is the rotation loss, is the speed loss, is the depth loss.
[0120] The functions or operation steps implemented when the above modules are executed are substantially the same as those in the above method embodiments, and will not be repeated here.
[0121] Embodiment 4
[0122] Another aspect of the present invention further provides a readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method described in any one of the above-mentioned embodiments 1 to 2 are implemented.
[0123] Embodiment 5
[0124] Another aspect of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in any one of the above-mentioned embodiments 1 to 2 when executing the program.
[0125] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0126] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be specifically implemented in any computer-readable storage medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For purposes of this specification, a "computer-readable storage medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0127] More specific examples (a non-exhaustive list) of computer-readable storage media include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable storage medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.
[0128] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0129] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0130] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A monocular 3D target detection method, characterized in that: The method comprises: Obtain a target detection image to be detected; Inputting the target detection image into a pre-trained target detection model to obtain target detection information of the target detection image; Among them, the training process of the target detection model is: Add a multimodal feature fusion module to the ControlNet model architecture to establish the Fusion-ControlNet model architecture; The Fusion-ControlNet model architecture is introduced as a guidance module for conditional generation into the Stable Diffusion model architecture to obtain a data set expansion model. Obtain a KITTI dataset, and use the dataset expansion model to expand the KITTI dataset to obtain a KITTI expanded dataset; Use the Contextual transformer module to replace the self-attention mechanism in the MonoDETR model architecture to establish the CoT-MonoDETR model architecture; The KITTI augmented data set is input into the CoT-MonoDETR model framework for training to obtain the target detection model.
2. The monocular 3D target detection method according to claim 1, characterized in that: The multimodal feature fusion module is a multimodal feature fusion encoder, which is used to input the semantic image and the depth image into two ResNet50 pre-trained models respectively, and through the adaptive feature fusion module connected between the two ResNet50 models, the images of two different modalities can adaptively fuse different modal features.
3. The monocular 3D target detection method according to claim 1, characterized in that: The step of introducing the Fusion-ControlNet model architecture as a guidance module for conditional generation into the Stable Diffusion model architecture to obtain a data set expansion model includes: The weights of the trained Stable Diffusion model architecture are frozen, and the Fusion-ControlNet model architecture is trained to obtain the data set expansion model.
4. The monocular 3D target detection method according to claim 1, characterized in that: The step of obtaining a KITTI data set and using the data set expansion model to expand the KITTI data set to obtain a KITTI expanded data set includes: The dataset expansion model is used to generate a KITTI diffusion dataset using the KITTI dataset, and the annotations of the KITTI dataset are directly projected into the corresponding newly generated KITTI diffusion dataset through a projection method, thereby obtaining pseudo labels of new images and obtaining the final KITTI expansion dataset.
5. The monocular 3D target detection method according to claim 1, characterized in that: The loss function of the dataset expansion model is: ; in, represents the original image, represents the actual noise added, represents the image after adding noise, t represents the diffusion time step, Indicates text conditions, Indicates potential conditions, Represents a noisy estimate of the model output.
6. The monocular 3D target detection method according to claim 1, characterized in that: The loss function of the CoT-MonoDETR model architecture is: ; in, is the classification loss, is the geometric center point offset loss, is the size loss, is the rotation loss, is the speed loss, is the depth loss.
7. A monocular 3D target detection device, characterized in that: The device comprises: An acquisition module is used to acquire a detection image of a target to be detected; A detection module, used for inputting the target detection image into a pre-trained target detection model to obtain target detection information of the target detection image; Among them, the training process of the target detection model is: Add a multimodal feature fusion module to the ControlNet model architecture to establish the Fusion-ControlNet model architecture; The Fusion-ControlNet model architecture is introduced as a guidance module for conditional generation into the Stable Diffusion model architecture to obtain a data set expansion model. Obtain a KITTI dataset, and use the dataset expansion model to expand the KITTI dataset to obtain a KITTI expanded dataset; Use the Contextual transformer module to replace the self-attention mechanism in the MonoDETR model architecture to establish the CoT-MonoDETR model architecture; The KITTI augmented data set is input into the CoT-MonoDETR model framework for training to obtain the target detection model.
8. A readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method and device for generating Chinese character library based on Stable Diffusion
CN116975344A
Nuclear medicine imaging equipment target detection method based on improved DETR model
CN118486029A