3D Object Detection Method, Device, Equipment and Readable Storage Medium

By combining the DORN algorithm and the improved ResNet-50 network to generate a depth estimation graph, the loss function constraints two-dimensional and three-dimensional box relationships are constructed, which solves the problem of low robustness and correctness of the 3D object detection model and improves the detection accuracy.

CN116229448BActive Publication Date: 2025-07-11DONGFENG AUTOMOBILE COMPANY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211733172.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-07-11
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

The existing 3D object detection algorithm models are poorly robust and have low correctness, especially when detecting distant and small targets.

Method used

The depth estimation graph is generated through the DORN algorithm and its predicted target center point depth value, combined with the improved ResNet-50 network and RPN network, the target detection model is generated, and an adaptive local convolution kernel is generated using the depth estimation graph to construct the relationship between the loss function constraints of the two-dimensional enclosing box and the three-dimensional box.

Benefits of technology

The robustness and accuracy of three-dimensional target detection have been improved, especially when detecting semi-occluded objects, with car AP|R40 being increased by 5%, and pedestrians and other types of targets being increased by 1%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229448B_ABST
    Figure CN116229448B_ABST
Patent Text Reader

Abstract

The present application relates to a three-dimensional object detection method, device, equipment and readable storage medium, including processing an RGB image based on the DORN algorithm to obtain a depth estimation map and its predicted target center point depth value; training a preset neural network model with the RGB image, the depth estimation map and the predicted target center point depth value to generate an object detection model, where the local convolutional kernel in the object detection model is generated based on the depth estimation map, and the loss function in the object detection model is used to constrain the relationship between the regression two-dimensional bounding box and the three-dimensional box; inputting the image to be detected into the object detection model to obtain the target three-dimensional box. The present application combines the depth estimation map of the image to generate a local convolutional kernel adapted to the position of the target sample, so as to effectively capture context information and improve the robustness and accuracy of the three-dimensional object detection algorithm; at the same time, the geometric constraint supervision of the two-dimensional box and the three-dimensional box is constructed by using the target center depth value of the depth estimation map, so as to effectively improve the algorithm performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of autonomous driving target detection, and particularly to a three-dimensional target detection method, device, equipment and readable storage medium. Background Art

[0002] Autonomous driving mainly consists of three core systems, namely, an environment perception system, a behavior decision-making system and a motion control system. Among them, the environment perception system mainly senses the surrounding environmental conditions through sensors and collects corresponding data; the behavior decision-making system mainly processes the real-time input data and timely generates a reasonable and executable planned path; the motion control system issues a vehicle control instruction according to the current decision path to guide the vehicle to drive safely. Thus, the environment perception system is the premise of autonomous driving, and the accuracy and efficiency of the data it collects directly affect the operations of the behavior decision-making system and the motion control system. In the physical world, objects often contain geometric information such as length, width, height, and orientation angle. Therefore, 3D target detection aims to provide real information such as the size and pose of the target in the three-dimensional space based on 2D target detection. So, to ensure the safe driving of autonomous vehicles, the target detection in the environment perception layer will obtain the position information of the surrounding environment, and at the same time, the environment perception layer will provide the real situation of the target object on the road surface to the behavior decision-making layer.

[0003] In the related art, when implementing 3D target detection, the limitation of traditional two-dimensional convolution is often overcome by the assistance of depth information. However, since the convolution kernel in the traditional target detection model is a standard 3×3, the depth-guided local dilation convolution kernel cannot effectively capture the context information of the sample target for semi-occluded objects, resulting in poor robustness and low correctness of the algorithm model, thus causing poor performance in the perception results, especially for distant targets and small targets (pedestrians or cyclists). Summary of the Invention

[0004] This application provides a three-dimensional target detection method, device, equipment and readable storage medium to solve the problems of poor robustness and low correctness of the 3D target detection algorithm model in the related art.

[0005] In a first aspect, a three-dimensional target detection method is provided, including the following steps:

[0006] Process the RGB image based on the DORN algorithm to obtain a depth estimation map and its predicted target center point depth value;

[0007] Training a preset neural network model with an RGB image, a depth estimation map, and a predicted depth value of the target center point to generate an object detection model, where the local convolution kernels in the object detection model are generated based on the depth estimation map, and the loss function in the object detection model is used to constrain the relationship between the regression two-dimensional bounding box and the three-dimensional box;

[0008] Input the image to be detected into the object detection model to obtain the target three-dimensional box.

[0009] In some embodiments, the preset neural network model includes an upper branch feature extraction network and a lower branch convolution kernel generation network. The upper branch feature extraction network includes an improved ResNet-50 network, an RPN network, and a depth-guided convolution module.

[0010] In some embodiments, the improved ResNet-50 network includes 4 convolution modules composed of residual modules, and the lower branch convolution kernel generation network includes 3 convolution modules composed of residual modules.

[0011] In some embodiments, the training of the preset neural network model with an RGB image, a depth estimation map, and a predicted depth value of the target center point to generate an object detection model includes:

[0012] Using the depth estimation map as the input of the lower branch convolution kernel generation network to obtain a depth feature map;

[0013] Using the RGB image as the input of the improved ResNet-50 network, and the convolution modules in the improved ResNet-50 network generate an upper branch feature map based on the RGB image and the depth feature map;

[0014] The depth-guided convolution module fuses the depth feature map and the upper branch feature map to obtain a fused feature map;

[0015] Performing three-dimensional object detection training on the RPN network based on the fused feature map and the predicted depth value of the target center point to generate an object detection model.

[0016] In some embodiments, the loss function L total is:

[0017] L total =(1 - S t ) γ (L cls + L 2D + L 3D + L depth )

[0018]

[0019] where S t represents the score of the target classification, γ represents the focusing parameter of the loss function, and L cls represents the classification loss, and L 2D represents the regression loss of the 2D bounding box, and L 3D represents the regression loss of the 3D bounding box, and L depth represents the prediction loss of the depth estimation map, λ represents the weight, represents the SmoothL1 loss function, [z′] depth represents the predicted depth value of the target center point, and [z] 3D represents the true depth value of the target center point.

[0020] In a second aspect, a 3D object detection device is provided, including:

[0021] A processing unit configured to process the RGB image based on the DORN algorithm to obtain a depth estimation map and its predicted depth value of the target center point;

[0022] A training unit configured to train a preset neural network model through the RGB image, the depth estimation map, and the predicted depth value of the target center point to generate an object detection model, where the local convolution kernel in the object detection model is generated based on the depth estimation map, and the loss function in the object detection model is used to constrain the relationship between the regression 2D bounding box and the 3D box;

[0023] A detection unit configured to input the image to be detected into the object detection model to obtain the target 3D box.

[0024] In some embodiments, the preset neural network model includes an upper branch feature extraction network and a lower branch convolution kernel generation network, and the upper branch feature extraction network includes an improved ResNet-50 network, an RPN network, and a depth-guided convolution module.

[0025] In some embodiments, the loss function L total is:

[0026] L total = (1 - S t ) γ (L cls + L 2D + L 3D + L depth )

[0027]

[0028] where S t represents the score of the target classification, γ represents the focusing parameter of the loss function, and L cls represents the classification loss, and L 2DDenote the regression loss of the 2D bounding box, \(L\) 3D Denote the regression loss of the 3D bounding box, \(L\) depth Denote the prediction loss of the depth estimation map, \(\lambda\) denotes the weight Denote the SmoothL1 loss function, \([z']\) depth Denote the predicted depth value of the center point of the target, \([z]\) 3D Denote the true depth value of the center point of the target.

[0029] In a third aspect, a 3D object detection device is provided, including: a memory and a processor. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the foregoing 3D object detection method.

[0030] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing 3D object detection method is implemented.

[0031] The present application provides a 3D object detection method, device, equipment and readable storage medium, including processing an RGB image based on the DORN algorithm to obtain a depth estimation map and its predicted depth value of the target center point; training a preset neural network model through the RGB image, the depth estimation map and the predicted depth value of the target center point to generate an object detection model. The local convolution kernel in the object detection model is generated based on the depth estimation map, and the loss function in the object detection model is used to constrain the relationship between the regression 2D bounding box and the 3D box; inputting the image to be detected into the object detection model to obtain the target 3D box. The present application combines the depth estimation map of the image to generate a local convolution kernel adaptive to the position of the target sample to effectively capture context information, thereby improving the robustness and accuracy of the 3D object detection algorithm; at the same time, using the target center depth value of the depth estimation map to construct the geometric constraint supervision of the 2D box and the 3D box, so as to effectively improve the algorithm performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 It is a schematic flowchart of a 3D object detection method provided by an embodiment of the present application;

[0034] Figure 2 It is a schematic flowchart of the image data processing provided by an embodiment of the present application;

[0035] Figure 3 This is a schematic structural diagram of a three-dimensional object detection device provided by an embodiment of the present application;

[0036] Figure 4 This is a schematic structural diagram of a three-dimensional object detection device provided by an embodiment of the present application. Specific implementation manners

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0038] The embodiments of the present application provide a three-dimensional object detection method, device, equipment, and readable storage medium, which can solve the problems of poor robustness and low correctness of the 3D object detection algorithm model in the related art.

[0039] Figure 1 This is a three-dimensional object detection method provided by an embodiment of the present application, including the following steps:

[0040] Step S10: Process the RGB image based on the DORN algorithm (Depth Ordinal Regression Network) to obtain a depth estimation map and the predicted depth value of the target center point;

[0041] Exemplarily, in this embodiment, the DORN algorithm (i.e., the depth estimation algorithm for monocular images) is used to process the corresponding monocular image (i.e., the RGB image) to generate a depth estimation map and the depth value of the target center predicted from the depth estimation map. It should be noted that not only can the DORN algorithm be used to process the RGB image, but other monocular image depth estimation algorithms can also be used to process the RGB image. Which algorithm to specifically adopt can be determined according to actual needs and is not limited herein.

[0042] It should be understood that the depth estimation map obtained in this embodiment is used to guide the generation of convolution kernels. Each pixel in the two-dimensional image corresponding to the depth estimation map has a corresponding depth estimation value. Since the difficulty of regressing the three-dimensional bounding box mainly lies in the lack of accurate depth values in the monocular image, the depth value of the target sample center will be generated using the DORN algorithm in this embodiment.

[0043] It should be noted that in this embodiment, the KITTI dataset and the NuScenes dataset will be used to implement the training, validation, and testing of the model. Among them, the NuScenes dataset collects 1000 driving scenarios, which includes about 1.4 million camera images in 400,000 key frames, 390,000 LIDAR scan data, 1.4M RADAR scan data, and 14,000 object bounding boxes. The 40,000 point clouds and 1000 scenarios (i.e., including 850 scenarios for training and validation and 150 scenarios for testing) in its extension package nuScenes-lidarseg contain 1.4 billion annotated points; while the KITTI dataset is a computer vision algorithm evaluation dataset for autonomous driving scenarios. This dataset contains multiple scenarios such as urban areas, rural areas, and highways. The entire dataset consists of 389 pairs of stereo images and optical flow maps, 39.2km visual ranging sequences, and images of more than 200k 3D annotated objects. For 3D object detection, the labels are subdivided into car, van, truck, pedestrian, pedestrian(sitting), cyclist, tram, and misc. The above two datasets have been annotated and divided into training, validation, and testing datasets.

[0044] Step S20: Train a preset neural network model through an RGB image, a depth estimation map, and a predicted target center point depth value to generate an object detection model. The local convolutional kernel in the object detection model is generated based on the depth estimation map, and the loss function in the object detection model is used to constrain the relationship between the regression two-dimensional bounding box and the three-dimensional box;

[0045] Furthermore, the preset neural network model includes an upper branch feature extraction network and a lower branch convolutional kernel generation network. The upper branch feature extraction network includes an improved ResNet-50 network, an RPN network, and a depth-guided convolution module. Among them, the improved ResNet-50 network includes 4 convolutional modules composed of residual modules, and the lower branch convolutional kernel generation network includes 3 convolutional modules composed of residual modules.

[0046] Exemplarily, in this embodiment, the backbone in the preset neural network model mainly includes an upper branch feature extraction network and a lower branch convolution kernel generation network. The upper branch feature extraction network mainly includes an improved ResNet-50 network, an RPN network, and a depth-guided convolution module. Among them, the improved ResNet-50 network part is mainly used to extract relevant features, the RPN network further finely processes the generated anchor boxes, and the depth-guided convolution module guides the generation of convolution kernels for the feature extraction network. The convolution kernel can be a dilated convolution, and its receptive field can be allocated by an adaptive weight function so that it can more effectively extract the feature information of targets of different scales, thereby effectively solving the problem of feature extraction of different scales.

[0047] It can be understood that when the network layers built by the model are too deep, it will lead to an increase in computational overhead and a decrease in the accuracy of the neural network model. Therefore, to avoid the above phenomenon, the improved ResNet-50 network in this embodiment discards the last fully convolutional layer and pooling layer in the original ResNet-50 network. That is, the improved ResNet-50 network is divided into five stages and four convolutional modules: the first stage performs convolution and max-pooling operations on the input image; and each subsequent stage is composed of multiple residual modules with Bottleneck (i.e., bottleneck layer) structures, and each convolutional module is a convolutional layer for extracting effective information of samples. It should be noted that the parameter weights of each layer in the network can be pre-trained from the ImageNet classification dataset.

[0048] It should be understood that to reduce the amount of calculation, the network composed of the first three convolutional modules of ResNet-50 in this embodiment is used as the lower branch convolution kernel generation network, and the convolution kernels and their receptive fields of different pixels and channels of different images are different, which can be learned based on the depth map.

[0049] Furthermore, training the preset neural network model with the RGB image, the depth estimation map, and the predicted target center point depth value to generate a target detection model includes:

[0050] Using the depth estimation map as the input of the lower branch convolution kernel generation network to obtain a depth feature map;

[0051] Using the RGB image as the input of the improved ResNet-50 network, and the convolutional module in the improved ResNet-50 network generates an upper branch feature map based on the RGB image and the depth feature map;

[0052] The depth-guided convolution module fuses the depth feature map and the upper branch feature map to obtain a fused feature map;

[0053] Perform 3D object detection training on the RPN network based on the fused feature map and the depth value of the predicted target center point to generate an object detection model.

[0054] Exemplarily, in this embodiment, refer to Figure 2 As shown, use the depth estimation map as the input of the lower branch convolution kernel generation network. After being processed by the residual module in the lower branch convolution kernel generation network, a depth feature map H1 with a width of w1, a height of h1, and a channel of c1, a depth feature map H2 with a width of w2, a height of h2, and a channel of c2, and a depth feature map H3 with a width of w3, a height of h3, and a channel of c3 are obtained respectively.

[0055] Use the RGB image as the input of the improved ResNet-50 network. After being processed by the improved ResNet-50 network, an upper branch feature map F1 with a width of w1, a height of h1, and a channel of c1 is obtained; input the upper branch feature map F1 and the depth feature map H1 into the residual module of the improved ResNet-50 network for fusion, and then output an upper branch feature map F2 with a width of w2, a height of h2, and a channel of c2; then input the upper branch feature map F2 and the depth feature map H2 into the next residual module in the improved ResNet-50 network for fusion, and output an upper branch feature map F3 with a width of w3, a height of h3, and a channel of c3.

[0056] Input the upper branch feature map F3 and the depth feature map H3 into the depth-guided convolution module for fusion to obtain a fused feature map with a width of w4, a height of h4, and a channel of c4.

[0057] Then use the depth value of the predicted target center point as the depth loss branch and the fused feature map to input into the RPN network, so as to construct a new loss function through the RPN network, and further constrain the relationship between the regression 2D bounding box and the 3D box.

[0058] In this embodiment, use the 2D-3D single-stage detection of faster_RCNN as the detection head to detect the feature map output by the RPN network, and then the visualized 3D box can be output, thus completing the training of the preset neural network and generating an object detection model.

[0059] It can be seen that when the preset network model in this embodiment is trained, the inputs of the model are the RGB image and the depth estimation map, and the output of the model is the 3D position coordinate points of the objects in the image. Among them, the model in this embodiment uses forward network propagation, backward network feedback, and is optimized and iterated based on Stochastic Gradient Descent (SGD).

[0060] Further, the loss function L total is:

[0061] Ltotal = (1 - S t ) γ (L cls + L 2D + L 3D + L depth )

[0062]

[0063] where S t represents the score of the target classification, γ represents the focusing parameter of the loss function, L cls represents the classification loss, L 2D represents the regression loss of the 2D bounding box, L 3D represents the regression loss of the 3D bounding box, L depth represents the prediction loss of the depth estimation map, λ represents the weight, represents the SmoothL1 loss function, [z′] depth represents the predicted depth value of the target center point, [z] 3D represents the true depth value of the target center point.

[0064] Exemplarily, it can be understood that there are often some constraint conditions for the semantic information of corresponding pixels in the 2D image and the depth map, but such constraint conditions are not fully utilized in the existing object detection networks. In this embodiment, the geometric constraint relationship between the 3D information of the instance object and the 2D bounding box in the image is fully considered. To improve the accuracy rate of the 3D candidate boxes by the RPN network, the depth map instance object center point Z value generated by the lower branch convolution kernel network is input into the RPN network to construct a new loss function, thereby constraining the relationship between the regression 2D bounding box and the 3D box. This method strengthens the depth perception feature learning of the positive sample objects during the training process.

[0065] It should be noted that the detection rate of the 2D bounding box is usually very high, so most of the 3D bounding boxes are regressed based on the 2D bounding box using the perspective geometry relationship. And this embodiment will add a new loss function, that is, on the basis of the total loss, enabling the 2D bounding box to more accurately regress the 3D box. Specifically, the loss function in this embodiment includes the classification loss, the regression loss of the 2D bounding box, the regression loss of the 3D bounding box, and the target center loss based on the depth estimation map, that is, the total loss function L total is as follows:

[0066] L total = (1 - S t ) γ (L cls + L 2D + L 3D + L depth )

[0067] In the formula, S t represents the score of the target classification, γ represents the focusing parameter of the loss function, and L cls represents the classification loss, and L 2D represents the regression loss of the 2D bounding box, and L 3D represents the regression loss of the 3D bounding box, and L depth represents the prediction loss of the depth estimation map.

[0068] Among them, for the loss of the target center based on the depth estimation map, in this embodiment, the SmoothL1 loss function will be adopted, and the specific formula is as follows:

[0069]

[0070] In the formula, λ represents the weight, λ ∈ [0, 1], represents the SmoothL1 loss function, and [z′] depth represents the depth value of the target center point predicted from the depth estimation map, and [z] 3D represents the true value of the depth value of the target center point.

[0071] It should be understood that the depth prediction of the depth estimation map is accurate at small and medium distances, while the depth prediction at long distances is a reasonable estimate. Therefore, in this embodiment, a reverse S function will be used to define the weight λ at different distances, as shown in the following formula:

[0072]

[0073] In the formula, s is the distance threshold, which corresponds to the central symmetry point of the function, T is the function curvature, and dep i is the depth predicted by the RPN network.

[0074] Step S30: Input the image to be detected into the target detection model to obtain the target 3D box.

[0075] Exemplarily, in this embodiment, an adaptive local convolution kernel for the target sample position is generated in combination with the depth estimation map of the image to effectively capture context information, thereby improving the robustness and accuracy of the 3D object detection algorithm; at the same time, the geometric constraint supervision of the 2D box and the 3D box is constructed using the depth value of the target center of the depth estimation map, so as to effectively improve the algorithm performance. Therefore, when the image to be detected is input into the target detection model, an accurate target 3D box and its corresponding 3D position coordinate points can be output. In addition, experiments show that the accuracy of the 3D object detection results for semi-occluded objects in this embodiment has been significantly improved. Among them, for the 3D object detection of cars, the AP|R40 of each type has increased by about 5%, while for pedestrians and other types, it has increased by about 1%.

[0076] SeeFigure 3 As shown in the figure, an embodiment of the present application further provides a three-dimensional object detection device, including:

[0077] A processing unit, which is used to process the RGB image based on the DORN algorithm to obtain a depth estimation map and its predicted target center point depth value;

[0078] A training unit, which is used to train a preset neural network model through the RGB image, the depth estimation map, and the predicted target center point depth value to generate an object detection model. The local convolution kernel in the object detection model is generated based on the depth estimation map, and the loss function in the object detection model is used to constrain the relationship between the regression two-dimensional bounding box and the three-dimensional box;

[0079] A detection unit, which is used to input the image to be detected into the object detection model to obtain an object three-dimensional box.

[0080] Further, the preset neural network model includes an upper branch feature extraction network and a lower branch convolution kernel generation network. The upper branch feature extraction network includes an improved ResNet-50 network, an RPN network, and a depth-guided convolution module.

[0081] Further, the preset neural network model includes an upper branch feature extraction network and a lower branch convolution kernel generation network. The upper branch feature extraction network includes an improved ResNet-50 network, an RPN network, and a depth-guided convolution module.

[0082] Further, the improved ResNet-50 network includes 4 convolution modules composed of residual modules, and the lower branch convolution kernel generation network includes 3 convolution modules composed of residual modules.

[0083] Further, the training unit is specifically used for:

[0084] Taking the depth estimation map as the input of the lower branch convolution kernel generation network to obtain a depth feature map;

[0085] Taking the RGB image as the input of the improved ResNet-50 network, and the convolution modules in the improved ResNet-50 network generate an upper branch feature map based on the RGB image and the depth feature map;

[0086] The depth-guided convolution module fuses the depth feature map and the upper branch feature map to obtain a fused feature map;

[0087] Based on the fused feature map and the predicted target center point depth value, performing three-dimensional object detection training on the RPN network to generate an object detection model.

[0088] Further, the loss function Ltotal is:

[0089] L total =(1 - S t ) γ (L cls + L 2D + L 3D + L depth )

[0090]

[0091] Wherein, S t represents the score of the target classification, γ represents the focusing parameter of the loss function, L cls represents the classification loss, L 2D represents the regression loss of the 2D bounding box, L 3D represents the regression loss of the 3D bounding box, L depth represents the prediction loss of the depth estimation map, λ represents the weight, represents the SmoothL1 loss function, [z′] depth represents the predicted depth value of the target center point, [z] 3D represents the true depth value of the target center point.

[0092] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described device and each unit can refer to the corresponding processes in the foregoing embodiment of the 3D object detection method, and will not be described herein again.

[0093] The device provided in the above embodiment can be implemented in the form of a computer program, and the computer program can run on a 3D object detection device as Figure 4 shown.

[0094] The embodiment of the present application further provides a 3D object detection device, including: a memory, a processor, and a network interface connected through a system bus. At least one instruction is stored in the memory, and at least one instruction is loaded and executed by the processor to implement all or part of the steps of the foregoing 3D object detection method.

[0095] Among them, the network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 4 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0096] The processor can be a CPU, or it can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor, etc. The processor is the control center of the computer device, and connects various parts of the entire computer device through various interfaces and circuits.

[0097] The memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory, the processor realizes various functions of the computer device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as video playback function, image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as video data, image data, etc.). In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0098] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, all or part of the steps of the foregoing three-dimensional object detection method are realized.

[0099] The embodiments of the present application implement all or part of the foregoing processes, and may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the foregoing various methods may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0100] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, system, server, or computer program product. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0101] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0102] It should be noted that in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or system comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or system comprising such element.

[0103] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will conform to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A three-dimensional object detection method, characterized in that, It includes the following steps: Process the RGB image based on the DORN algorithm to obtain a depth estimation map and the depth value of the predicted target center point; Train a preset neural network model with the RGB image, the depth estimation map, and the depth value of the predicted target center point to generate a target detection model. The local convolution kernels in the target detection model are generated based on the depth estimation map, and the loss function in the target detection model is used to constrain the relationship between the regression two-dimensional bounding box and the three-dimensional box; Input the image to be detected into the target detection model to obtain a target three-dimensional box; Among them, the loss function is as follows: Wherein, represents the score of the target classification, represents the focusing parameter of the loss function, represents the classification loss, represents the regression loss of the two-dimensional bounding box, represents the regression loss of the three-dimensional bounding box, represents the prediction loss of the depth estimation map, represents the weight, represents the Smooth L1 loss function, represents the predicted depth value of the target center point, represents the true depth value of the target center point; Among them, the weights of different distances are defined by the reverse S function , the weight The calculation formula is as follows: Where s is the distance threshold corresponding to the central symmetry point of the function, and T is the function curvature, which is the depth predicted by the RPN network in the target detection model.

2. The three-dimensional object detection method according to claim 1, characterized in that: The preset neural network model includes an upper branch feature extraction network and a lower branch convolution kernel generation network. The upper branch feature extraction network includes an improved ResNet-50 network, an RPN network, and a depth-guided convolution module.

3. The three-dimensional object detection method according to claim 2, characterized in that: The improved ResNet-50 network includes 4 convolution modules composed of residual modules, and the lower branch convolution kernel generation network includes 3 convolution modules composed of residual modules.

4. The three-dimensional object detection method according to claim 3, wherein The training of the preset neural network model with the RGB image, the depth estimation map, and the depth value of the predicted target center point to generate a target detection model includes: Use the depth estimation map as the input of the lower branch convolution kernel generation network to obtain a depth feature map; Use the RGB image as the input of the improved ResNet-50 network. The convolution modules in the improved ResNet-50 network generate an upper branch feature map based on the RGB image and the depth feature map; The depth-guided convolution module fuses the depth feature map and the upper branch feature map to obtain a fused feature map; Perform three-dimensional object detection training on the RPN network based on the fused feature map and the depth value of the predicted target center point to generate a target detection model.

5. A three-dimensional object detection device, characterized in that, It includes: A processing unit for processing the RGB image based on the DORN algorithm to obtain a depth estimation map and the depth value of the predicted target center point; A training unit for training a preset neural network model with the RGB image, the depth estimation map, and the depth value of the predicted target center point to generate a target detection model. The local convolution kernels in the target detection model are generated based on the depth estimation map, and the loss function in the target detection model is used to constrain the relationship between the regression two-dimensional bounding box and the three-dimensional box; A detection unit for inputting the image to be detected into the target detection model to obtain a target three-dimensional box; Among them, the loss function is as follows: wherein, represents the score of the target classification, represents the focusing parameter of the loss function, represents the classification loss, represents the regression loss of the 2D bounding box, represents the regression loss of the 3D bounding box, represents the prediction loss of the depth estimation map, represents the weight, represents the Smooth L1 loss function, represents the predicted depth value of the center point of the prediction target, represents the true depth value of the center point of the target; Among them, the weights for different distances are defined by the reverse S function , the weight is calculated by the following formula: Where s is the distance threshold corresponding to the central symmetry point of the function, and T is the function curvature. is the depth predicted by the RPN network in the object detection model.

6. The three-dimensional object detection device according to claim 5, characterized in that: The preset neural network model includes an upper branch feature extraction network and a lower branch convolution kernel generation network. The upper branch feature extraction network includes an improved ResNet-50 network, an RPN network, and a depth-guided convolution module.

7. A three-dimensional object detection device, characterized in that, It includes: A memory and a processor. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the three-dimensional object detection method according to any one of claims 1 to 4.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the three-dimensional object detection method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Monocular 3D target detection method based on dynamic convolution

    CN114266900A