Training methods for parking object detection models and 3D object detection methods

By using a lightweight convolutional neural network and corner point correlation information training object detection model, the fisheye image is directly processed, solving the detection accuracy problem caused by fisheye camera image distortion, and achieving efficient 3D object detection.

CN120126101BActive Publication Date: 2025-08-29BEIJING YINWO AUTOMOBILE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510599873.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-29
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

In the prior art, when using a fisheye camera for 3D object detection, detection accuracy decreases due to image distortion, and computer vision processing will lose useful information.

Method used

A lightweight convolutional neural network is used as the backbone network, and the position information associated with the center point and corner point are trained, the fisheye image is directly processed, feature information is retained, and model parameters are adjusted to improve detection accuracy.

Benefits of technology

It improves the accuracy of target detection, adapts to complex environmental conditions, reduces the amount of calculation, is suitable for low-computing platforms, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126101B_ABST
    Figure CN120126101B_ABST
Patent Text Reader

Abstract

The present application provides a training method for a parking target detection model and a 3D target detection method, which relate to the field of target detection technology. The training method includes: obtaining fisheye image samples and sample labels; inputting the fisheye image samples into the target detection model, outputting predicted position information of a 2D bounding box associated with a center point, predicted position information of a 3D bounding box, and predicted position information associated with a corner point; and adjusting parameters of the target detection model based on the difference between the actual position information and the predicted position information of the 2D bounding box associated with the center point, the difference between the actual position information and the predicted position information of the 3D bounding box associated with the center point, and the difference between the actual position information and the predicted position information of the corner point. The present application can improve target detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of target detection technology, and in particular to a training method for a target detection model for parking and a 3D target detection method. Background Art

[0002] In the field of autonomous driving, 3D object detection in parking scenarios is a key technology. It helps the vehicle perceive its surroundings, identify and locate obstacles, vehicles, pedestrians, and other objects within the parking lot, and provides the necessary foundation for automated parking and obstacle avoidance. Fisheye cameras, with their wide viewing angle and ability to cover a large field of view, are well-suited for 3D object detection in parking scenarios. However, the significant distortion and nonlinearity of fisheye camera images pose challenges to 3D object detection based on these images.

[0003] Existing technologies typically use computer vision techniques, such as image distortion correction and feature extraction, to detect 3D objects. However, distortion correction of fisheye images can lose some useful information, affecting the accuracy of 3D object detection. Summary of the Invention

[0004] This application is proposed in view of at least one of the above-mentioned technical problems existing in the prior art, and this application can improve the accuracy of target detection.

[0005] In a first aspect, an embodiment of the present application provides a method for training a parking object detection model, comprising:

[0006] Obtain fisheye image samples and sample labels; wherein the sample labels include: real parameters of a 2D bounding box associated with the center point and real parameters of a 3D bounding box;

[0007] Inputting the fisheye image sample into an object detection model, and outputting predicted position information of a 2D bounding box associated with a center point, predicted position information of a 3D bounding box, and predicted position information associated with a corner point; wherein the object detection model comprises: a backbone network, a neck network, and a detection head, wherein the backbone network is a lightweight convolutional neural network;

[0008] Determining the true position information of the 2D bounding box associated with the center point based on the true parameters of the 2D bounding box associated with the center point;

[0009] Determining true position information of the 3D bounding box associated with the center point based on true parameters of the 3D bounding box associated with the center point;

[0010] Determining true position information associated with the corner point based on true parameters of a 2D bounding box associated with the center point and true parameters of a 3D bounding box associated with the center point;

[0011] Adjust the parameters of the target detection model based on the difference between the real position information and the predicted position information of the 2D bounding box associated with the center point, the difference between the real position information and the predicted position information of the 3D bounding box associated with the center point, and the difference between the real position information and the predicted position information associated with the corner point.

[0012] In a second aspect, an embodiment of the present application provides a 3D object detection method, including:

[0013] Obtain target fisheye images collected during parking;

[0014] Inputting the target fisheye image into a trained target detection model, and outputting position information of a 2D bounding box and a 3D bounding box associated with the center point; wherein the target detection model is trained using the method described in any of the above embodiments;

[0015] Parameters of the 3D bounding box associated with the center point are determined based on the position information of the 2D bounding box associated with the center point and the position information of the 3D bounding box associated with the center point.

[0016] In a third aspect, an embodiment of the present application provides a computer program product, which implements the method described in any of the above embodiments when the computer program / instructions are executed by a processor.

[0017] The training method of the target detection model for parking and the 3D target detection method provided in this application directly process the fisheye image through the target detection model without the need for dedistortion of the fisheye image. This can retain the feature information in the fisheye image, improve the accuracy of target detection, and adapt to complex environmental conditions. The target detection model uses a lightweight convolutional neural network as the backbone network, which can reduce the amount of calculation and is suitable for low-computing power platforms. In the process of training the target detection model, not only the parameters associated with the center point are used, but also the position information associated with the corner points. The position information associated with the corner points is used as a supervision item in the training process, which can improve the accuracy of target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 is a schematic diagram of a method for training a parking object detection model provided by an embodiment of the present application;

[0020] Figure 2This is a flowchart of a 3D object detection method provided by an embodiment of the present application;

[0021] Figure 3 1 is a schematic diagram of a training device for a parking object detection model provided by an embodiment of the present application;

[0022] Figure 4 3D target detection device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] To help those skilled in the art better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0024] like Figure 1 As shown, an embodiment of the present application provides a method for training a parking object detection model, comprising:

[0025] Step 101: Obtain fisheye image samples and sample labels.

[0026] The sample label includes the true parameters of the 2D bounding box associated with the center point and the true parameters of the 3D bounding box. The true parameters of the 2D bounding box associated with the center point can be obtained by manual annotation, while the true parameters of the 3D bounding box associated with the center point are the parameters of the target in the radar coordinate system.

[0027] The real parameters of the 2D bounding box associated with the center point include: the real coordinates of the center point of the 2D bounding box, the width and height of the 2D bounding box, where the width and height of the 2D bounding box correspond to the coordinates of the center point of the 2D bounding box.

[0028] The real parameters of the 3D bounding box associated with the center point include: the real coordinates of the center point of the 3D bounding box, length, width, height, and orientation. The length, width, height, and orientation of the 3D bounding box correspond to the coordinates of the center point of the 3D bounding box.

[0029] The fisheye image sample may be an original fisheye image or an original fisheye image obtained through processing. The original fisheye image is acquired by a fisheye camera in a parking scene.

[0030] Step 102: Input the fisheye image sample into the target detection model, and output the predicted position information of the 2D bounding box associated with the center point, the predicted position information of the 3D bounding box, and the predicted position information associated with the corner points.

[0031] Among them, the target detection model includes: backbone network, neck network and detection head, and the backbone network is a lightweight convolutional neural network.

[0032] A lightweight convolutional neural network (CNN) is a type of CNN designed to reduce computational complexity, parameter count, and memory usage. For example, the backbone network can be a depthwise separable CNN. Using a lightweight CNN significantly reduces computational effort, making this object detection model suitable for low-computing platforms.

[0033] Compared with existing target detection models, in addition to outputting the predicted position information associated with the center point, it can also output the predicted position information associated with the corner points.

[0034] Step 103: Determine the real position information of the 2D bounding box associated with the center point based on the real parameters of the 2D bounding box associated with the center point.

[0035] The predicted position information of the 2D bounding box associated with the center point, including the following information, as predicted by the target detection model: a heatmap of the predicted center point of the 2D bounding box, the 2D predicted offset between the predicted coarse coordinates of the 2D bounding box center point and the predicted 2D coordinates of the center point, and the predicted size of the 2D bounding box. This 2D prediction offset is caused by the loss of precision during the downsampling process of the target detection model. The predicted coarse coordinates of the 2D bounding box center point are generated by the target detection model after downsampling and are generally integers. The predicted 2D coordinates of the center point are the coordinates that the target should have been predicted to be, generally floating-point numbers.

[0036] The true position information of the 2D bounding box associated with the center point includes the following information: a true center point heat map of the 2D bounding box, the true 2D offset between the true coarse coordinates of the 2D bounding box center point and the true coordinates of the 2D bounding box center point, and the true size of the 2D bounding box. The true position information of the 2D bounding box associated with the center point can be calculated from the true parameters of the 2D bounding box associated with the center point.

[0037] Based on the predicted center point heat map of the 2D bounding box, the predicted rough coordinates of the center point of the 2D bounding box can be obtained.

[0038] Step 104: Determine the real position information of the 3D bounding box associated with the center point based on the real parameters of the 3D bounding box associated with the center point.

[0039] The predicted position information of the 3D bounding box associated with the center point, including the following predicted depth value, 3D predicted offset, predicted size, and predicted yaw angle of the 3D bounding box center point, as predicted by the object detection model. The 3D predicted offset refers to the offset between the predicted coarse coordinates of the 2D bounding box center point and the projected coordinates of the predicted 3D coordinates of the center point on the fisheye image sample.

[0040] The true position information of the 3D bounding box associated with the center point includes the following information, calculated from the true parameters of the 3D bounding box associated with the center point: the true depth value of the 3D bounding box center point, the true 3D offset, the true size of the 3D bounding box, and the true yaw angle. The true 3D offset refers to the offset between the true coarse coordinates of the 2D bounding box center point and the projected coordinates of the true coordinates of the 3D bounding box center point on the fisheye image sample.

[0041] Step 105: Determine the real position information associated with the corner point based on the real parameters of the 2D bounding box associated with the center point and the real parameters of the 3D bounding box associated with the center point.

[0042] Based on the true parameters of the 2D bounding box associated with the center point, the true parameters of the 2D bounding box associated with the corner point are obtained. For example, the true parameters of the 2D bounding box associated with the corner point include the true coordinates of the corner points of the 2D bounding box. Based on the true coordinates of the corner points of the 2D bounding box, a sample corner heat map can be obtained.

[0043] Based on the real parameters of the 3D bounding box associated with the center point, the real parameters of the 3D bounding box associated with the corner point are obtained. The real parameters of the 3D bounding box associated with the corner point are also parameters in the radar coordinate system. For example, the real parameters of the 3D bounding box associated with the corner point include the real coordinates of the 3D bounding box corner points and the depth values ​​of the sample corner points. In other words, based on the real parameters of the 3D bounding box associated with the center point, the real coordinates of the 3D bounding box corner points can be obtained.

[0044] The sample corner offset is obtained based on the projection coordinates of the corner coordinates in the sample corner heat map and the real coordinates of the corner points of the 3D bounding box on the fisheye image sample.

[0045] The real position information associated with the corner points includes: sample corner point heat map, sample corner point depth value and sample corner point offset.

[0046] The predicted position information associated with the corner point includes: a predicted corner point heat map, a predicted corner point depth value, and a predicted corner point offset; wherein the predicted corner point offset is the offset between the corner point coordinates in the predicted corner point heat map and the projected coordinates of the predicted 3D coordinates of the corner point on the fisheye image sample.

[0047] Step 106: Adjust the parameters of the target detection model based on the difference between the actual position information and the predicted position information of the 2D bounding box associated with the center point, the difference between the actual position information and the predicted position information of the 3D bounding box associated with the center point, and the difference between the actual position information and the predicted position information associated with the corner points.

[0048] Based on the difference between the real position information and the predicted position information of the 2D bounding box associated with the center point, the loss between the predicted center point heatmap of the 2D bounding box and the real center point heatmap of the 2D bounding box, the loss between the 2D predicted offset and the 2D real offset, and the loss between the predicted size of the 2D bounding box and the real size of the 2D bounding box can be calculated; based on the difference between the real position information and the predicted position information of the 3D bounding box associated with the center point, the loss between the predicted depth value of the 3D bounding box center and the real depth value of the 3D bounding box center, the loss between the 3D predicted offset and the 3D real offset, the loss between the predicted size of the 3D bounding box and the real size of the 3D bounding box, and the loss between the predicted yaw angle and the real yaw angle can be calculated; based on the difference between the real position information and the predicted position information associated with the corner point, the loss between the sample corner point heatmap and the predicted corner point heatmap, the loss between the depth value of the sample corner point and the predicted corner point depth value, and the loss between the sample corner point offset and the predicted corner point offset can be calculated. Based on the above losses, the parameters of the target detection model can be adjusted.

[0049] The embodiment of the present application directly processes the fisheye image through the target detection model, without the need for dedistortion processing of the fisheye image, and can retain the feature information in the fisheye image, improve the accuracy of target detection, and adapt to complex environmental conditions. The target detection model uses a lightweight convolutional neural network as the backbone network, which can reduce the amount of calculation and is suitable for low-computing power platforms. In the process of training the target detection model, not only the parameters associated with the center point are used, but also the position information associated with the corner points. The position information associated with the corner points is used as a supervision item in the training process, which can improve the accuracy of target detection.

[0050] In one embodiment of the present application, obtaining a fisheye image sample includes:

[0051] Obtain the original fisheye image collected during the parking process;

[0052] Calculate the affine transformation matrix based on the resolution of the original fisheye image and the preset sample resolution;

[0053] Perform affine transformation on the original fisheye image based on the affine transformation matrix to obtain fisheye image samples.

[0054] The resolution of the original fisheye image captured by the on-board fisheye camera is usually 1280×960. The image size is large, which affects the detection speed. In view of this, the embodiment of the present application performs an affine transformation on the original fisheye image and adjusts it to a smaller resolution (preset sample resolution), such as 512×384, which can not only ensure that the image ratio is not distorted, but also maintain a good zoom ratio, thereby improving the target detection speed.

[0055] In one embodiment of the present application, the actual parameters of the 3D bounding box associated with the center point are the parameters of the 3D bounding box associated with the center point of the target in the radar coordinate system.

[0056] The real parameters of the 3D bounding box associated with the center point include: the real coordinates, length, width, height, and orientation of the center point of the 3D bounding box.

[0057] The real position information of the 3D bounding box associated with the center point includes: the real depth value of the 3D bounding box center point, the real 3D offset, the real size of the 3D bounding box, and the real yaw angle.

[0058] Based on the true parameters of the 3D bounding box associated with the center point, determining the true position information of the 3D bounding box associated with the center point includes:

[0059] Based on the camera extrinsic parameters, the real coordinates of the center point of the target's 3D bounding box in the radar coordinate system are transformed to the camera coordinate system;

[0060] Based on the distortion coefficient, the target in the camera coordinate system is projected onto the fisheye image sample to obtain the projection coordinates of the center point of the 3D bounding box on the fisheye image sample;

[0061] The real rough coordinates of the center point of the 2D bounding box are subtracted from the projection coordinates of the real coordinates of the center point of the 3D bounding box on the fisheye image sample to obtain the 3D real offset.

[0062] Based on the real coordinates, length, width, height, and orientation of the center point of the 3D bounding box, the real depth value of the center point of the 3D bounding box, the real size of the 3D bounding box, and the real yaw angle can be obtained.

[0063] The true parameters of the 2D bounding box associated with the center point can be obtained by manual annotation, for example, by manually annotating based on fisheye image samples directly, or annotating based on the original fisheye image and then transforming the annotation result based on an affine transformation matrix.

[0064] In one embodiment of the present application, the object detection model is optimized based on MonoDLE, which is a monocular 3D object detection model based on the CenterNet architecture.

[0065] In practical applications, MobileNetV2 can be used as the backbone network instead of DLA34, and a serial bilinear interpolation upsampling structure can be used instead of DLA34UP. MobileNetV2 uses depthwise separable convolutions, which significantly reduces the number of model parameters and computational complexity, and better balances detection accuracy and efficiency. The serial bilinear interpolation upsampling structure uses a vertical stacking structure to complete feature map sampling, making it suitable for different hardware and applications in different scenarios.

[0066] Because the input fisheye image samples are small, the loss of target geometry can be avoided by controlling the feature map size during the backbone network feature extraction downsampling and the neck network feature upsampling. To this end, adjustments can be made to MobileNetV2 and the cascaded bilinear interpolation upsampling structure. For example, the downsampling ratio of MobileNetV2 can be halved, and the upsampling in the cascaded bilinear interpolation upsampling structure can be reduced to two layers, ensuring that the overall 4x downsampling remains unchanged.

[0067] Specifically, in one embodiment of the present application, the backbone network includes: a first input layer, multiple bottleneck layers, a standard convolutional layer, an average pooling layer, and a first output layer;

[0068] Among them, the bottleneck layer includes a 1×1 expansion layer, a 3×3 depth-separable convolution layer, and a 1×1 projection layer;

[0069] Among them, the stride of the expansion layer and the depth-wise separable convolutional layer in at least one bottleneck layer is 1.

[0070] The embodiment of the present application adjusts the step size of the expansion layer and the depth-wise separable convolution layer in the bottleneck layer to adjust the downsampling ratio of the backbone network to avoid losing geometric information.

[0071] In one embodiment of the present application, the neck network includes: a second input layer, two bilinear interpolation layers, a convolutional layer, an activation layer and a second output layer.

[0072] In the embodiment of the present application, two bilinear interpolation layers are set to ensure that the overall 4x downsampling remains unchanged.

[0073] Convolutional layers and activation layers are optional.

[0074] In actual application scenarios, if the target detection model is trained only on data related to the center point, the accuracy of the target detection results may be low.

[0075] In view of this, in one embodiment of the present application, the real position information associated with the corner point includes: a sample corner point heat map, a depth value of the sample corner point, and a sample corner point offset; wherein the sample corner point offset is the offset between the corner point coordinates in the sample corner point heat map and the projection coordinates of the real corner point coordinates of the 3D bounding box on the fisheye image sample;

[0076] The predicted position information associated with the corner point includes: a predicted corner point heat map, a predicted corner point depth value, and a predicted corner point offset; wherein the predicted corner point offset is the offset between the corner point coordinates in the predicted corner point heat map and the projected coordinates of the predicted 3D coordinates of the corner point on the fisheye image sample.

[0077] Specifically, when calculating the loss of the sample corner heatmap and the predicted corner heatmap, the FocalLoss is calculated point by point. Focal Loss is a loss function designed specifically to solve the class imbalance problem (ClassImbalance) in machine learning tasks.

[0078] The Laplace loss is calculated based on the depth value of the sample corner point and the depth value of the predicted corner point.

[0079] Based on the sample corner offset and the predicted corner offset, the corner offset loss is calculated.

[0080] Focal Loss can be combined with Laplace loss, corner offset loss, heatmap loss, 2D offset loss, 2D size loss, 3D depth loss, 3D offset loss, 3D yaw loss, and 3D size loss to adjust the parameters of the object detection model. For example, the combined loss is the sum of these losses and is used to adjust the parameters of the object detection model.

[0081] The embodiment of the present application combines the data associated with the corner points and the data associated with the center points to jointly train the target detection model and improve the quality of model training.

[0082] In one embodiment of the present application, adjusting parameters of the object detection model based on a difference between actual position information and predicted position information of a 2D bounding box associated with a center point, a difference between actual position information and predicted position information of a 3D bounding box associated with a center point, and a difference between actual position information and predicted position information associated with a corner point includes:

[0083] Adjust the parameters of the target detection model based on the difference between the actual position information and the predicted position information of the 2D bounding box associated with the center point, the difference between the actual position information and the predicted position information of the 3D bounding box associated with the center point, the difference between the sample corner point heat map and the predicted corner point heat map, the difference between the depth value of the sample corner point and the depth value of the predicted corner point, and the difference between the sample corner point offset and the predicted corner point offset.

[0084] The loss of sample corner offsets and predicted corner offsets is calculated using L1 loss. Since the projected coordinates of the predicted 3D coordinates of the corners on the fisheye image samples are floating-point numbers, and the corner coordinates in the predicted corner heat map in the predicted position information associated with the corners are integers, the embodiment of the present application introduces quantization residuals to make the target detection model aware of the precision loss in the downsampling stage, thereby improving the model training quality and the accuracy of target detection.

[0085] Existing detection heads of object detection models, such as MonoDLE's, output seven prediction branches. The detection head of this application builds on this by adding three more branches: predicted corner heatmaps, predicted corner depth values, and predicted corner offsets. This can be achieved by modifying the existing detection head structure, such as adding feature extraction structures and output layers corresponding to the branches, and determining the corresponding loss functions.

[0086] In one embodiment of the present application, the distance between the target and the vehicle in the fisheye image sample is no more than 30 meters.

[0087] Taking into account the effectiveness of parking perception and the accuracy of the target covered by the image size, the embodiment of the present application discards targets beyond 30 meters to prevent misleading model training and further improve the model training effect.

[0088] like Figure 2 As shown, the embodiment of the present application provides a 3D target detection method, including:

[0089] Step 201: Acquire a target fisheye image collected during the parking process.

[0090] Step 202: Input the target fisheye image into the trained target detection model, and output the position information of the 2D bounding box and the position information of the 3D bounding box associated with the center point.

[0091] The target detection model is trained using any of the methods described above.

[0092] The position information in the embodiment of the present application is consistent with the predicted position information in the aforementioned embodiment, that is, the position information of the 2D bounding box associated with the center point, including: a heat map of the center point of the 2D bounding box, a 2D offset between the rough coordinates of the center point of the 2D bounding box and the predicted 2D coordinates of the center point, and the size of the 2D bounding box.

[0093] The location information of the 3D bounding box associated with the center point, including the depth value of the 3D bounding box center point, the 3D offset, the size of the 3D bounding box, and the yaw angle. The 3D offset is the difference between the coarse coordinates of the 2D bounding box center point and the projected 3D coordinates of the center point on the target fisheye image.

[0094] It should be noted that, unlike the training phase, the target detection model does not output information associated with corner points during the inference phase. That is, it does not output the three newly added branches. These three branches are only used to improve the quality of model training and are not enabled during the inference phase. This ensures that the target detection model does not consume additional time during the inference phase, and maximizes the model's capabilities without consuming additional resources.

[0095] Step 203: Determine parameters of the 3D bounding box associated with the center point based on the position information of the 2D bounding box associated with the center point and the position information of the 3D bounding box associated with the center point.

[0096] The parameters of the 3D bounding box include: the 3D coordinates of the center point, the length, width, height and orientation of the 3D bounding box.

[0097] Determine the 3D coordinates of the center point based on the coarse coordinates of the 2D bounding box center point, the 3D offset, the depth value of the 3D bounding box center point, and the camera intrinsic parameters.

[0098] Based on the position information of the 3D bounding box associated with the center point, the length, width, height, and orientation of the 3D bounding box are determined.

[0099] The target detection model provided in this application has a simple structure and strong scalability. It also has a small number of network parameters and computational complexity, making it suitable for deployment and use on low-computing platforms. It directly processes fisheye images, avoiding tedious image processing steps such as dedistortion. It can be used directly for various fisheye images and has a wide range of applications. In the field of parking, it greatly improves the inaccuracy of 2D detection and positioning, while avoiding the use of high-cost equipment such as binocular cameras and lidars, which can reduce the cost of target detection and facilitate industry applications.

[0100] One embodiment of the present application provides a computer program product, which implements any of the above-mentioned method embodiments when the computer program / instructions are executed by a processor.

[0101] like Figure 3As shown, one embodiment of the present application provides a training device for a parking object detection model, comprising:

[0102] The acquisition module 301 is configured to acquire fisheye image samples and sample labels; wherein the sample labels include: real parameters of a 2D bounding box associated with a center point and real parameters of a 3D bounding box;

[0103] An input module 302 is configured to input a fisheye image sample into an object detection model and output predicted position information of a 2D bounding box associated with a center point, predicted position information of a 3D bounding box, and predicted position information associated with a corner point. The object detection model includes a backbone network, a neck network, and a detection head. The backbone network is a lightweight convolutional neural network.

[0104] Determination module 303 is configured to determine the real position information of the 2D bounding box associated with the center point based on the real parameters of the 2D bounding box associated with the center point; determine the real position information of the 3D bounding box associated with the center point based on the real parameters of the 3D bounding box associated with the center point; and determine the real position information associated with the corner point based on the real parameters of the 2D bounding box associated with the center point and the real parameters of the 3D bounding box associated with the center point.

[0105] The training module 304 is configured to adjust the parameters of the target detection model based on the difference between the actual position information and the predicted position information of the 2D bounding box associated with the center point, the difference between the actual position information and the predicted position information of the 3D bounding box associated with the center point, and the difference between the actual position information and the predicted position information associated with the corner point.

[0106] like Figure 4 As shown, one embodiment of the present application provides a 3D object detection device, comprising:

[0107] An acquisition module 401 is configured to acquire a target fisheye image collected during the parking process;

[0108] An input module 402 is configured to input the target fisheye image into a trained target detection model and output position information of a 2D bounding box and a 3D bounding box associated with the center point; wherein the target detection model is trained using the method of any of the above embodiments;

[0109] The determination module 403 is configured to determine parameters of the 3D bounding box associated with the center point based on the position information of the 2D bounding box associated with the center point and the position information of the 3D bounding box associated with the center point.

[0110] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0111] It should also be understood that the memory mentioned in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAMbus RAM (DR RAM).

[0112] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) is integrated into the processor.

[0113] It should be noted that the memory described herein is intended to include, but not be limited to, these and any other suitable types of memory.

[0114] In addition to the data bus, the bus may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, various buses are labeled as buses in the figure.

[0115] It should also be understood that the first, second, third, fourth and various numerical numbers involved in this document are only distinctions made for the convenience of description and are not intended to limit the scope of this application.

[0116] It should be understood that the term "and / or" in this document simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0117] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.

[0118] In various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0119] Those skilled in the art will appreciate that the various illustrative logical blocks (ILBs) and steps described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0120] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0121] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0122] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0123] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk).

[0124] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A training method for a parking object detection model, characterized in that: include: Obtain fisheye image samples and sample labels; wherein the sample labels include: real parameters of a 2D bounding box associated with the center point and real parameters of a 3D bounding box; Inputting the fisheye image sample into an object detection model, outputting predicted position information of a 2D bounding box associated with a center point, predicted position information of a 3D bounding box, and predicted position information associated with a corner point; wherein the object detection model comprises: a backbone network, a neck network, and a detection head, wherein the backbone network is a lightweight convolutional neural network; Determine the real position information of the 2D bounding box associated with the center point based on the real parameters of the 2D bounding box associated with the center point; determine the real position information of the 3D bounding box associated with the center point based on the real parameters of the 3D bounding box associated with the center point; determine the real position information associated with the corner point based on the real parameters of the 2D bounding box associated with the center point and the real parameters of the 3D bounding box associated with the center point; Adjusting parameters of the object detection model based on a difference between actual position information and predicted position information of a 2D bounding box associated with the center point, a difference between actual position information and predicted position information of a 3D bounding box associated with the center point, and a difference between actual position information and predicted position information associated with a corner point; Wherein, the target detection model is used to identify obstacles in the parking lot; The real parameters of the 2D bounding box include: the real coordinates of the center point of the 2D bounding box, the width and height of the 2D bounding box corresponding to the real coordinates of the center point; The real parameters of the 3D bounding box include: the real coordinates of the center point of the 3D bounding box, and the length, width, height, and orientation corresponding to the real coordinates of the center point; The predicted position information of the 2D bounding box includes: a heat map of the predicted center point of the 2D bounding box, the 2D predicted offset between the predicted coarse coordinates of the center point of the 2D bounding box and the predicted 2D coordinates of the center point, and the predicted size of the 2D bounding box. The predicted coarse coordinates are generated by the object detection model after downsampling. The parameters contained in the actual position information of the 2D bounding box correspond to the parameters contained in the predicted position information of the 2D bounding box; The predicted position information of the 3D bounding box includes: the predicted depth value of the 3D bounding box center point, the predicted 3D offset, the predicted size of the 3D bounding box, and the predicted yaw angle; the 3D prediction offset is the offset between the predicted rough coordinates of the 2D bounding box center point and the projected coordinates of the predicted 3D coordinates of the center point on the fisheye image sample; The parameters contained in the actual position information of the 3D bounding box correspond to the parameters contained in the predicted position information of the 3D bounding box; The real position information associated with the corner points includes: a sample corner point heat map, a depth value of the sample corner points, and a sample corner point offset; the sample corner point offset is the offset between the corner point coordinates in the sample corner point heat map and the projected coordinates of the 3D bounding box corner point's real coordinates on the fisheye image sample; based on the real parameters of the 3D bounding box, the real coordinates of the corner points of the 3D bounding box associated with the corner points and the depth value of the sample corner points are obtained; based on the real parameters of the 2D bounding box, the real coordinates of the corner points of the 2D bounding box are obtained, and the sample corner point heat map is obtained based on the real coordinates of the corner points; The predicted position information associated with the corner point includes: a predicted corner point heat map, a predicted corner point depth value, and a predicted corner point offset; the predicted corner point offset is the offset between the corner point coordinates in the predicted corner point heat map and the projected coordinates of the predicted 3D coordinates of the corner point on the fisheye image sample.

2. The method according to claim 1, wherein Get fisheye image samples, including: Obtain the original fisheye image collected during the parking process; Calculating an affine transformation matrix based on the resolution of the original fisheye image and a preset sample resolution; Performing an affine transformation on the original fisheye image based on the affine transformation matrix to obtain the fisheye image sample.

3. The method according to claim 1, wherein The true parameters of the 3D bounding box associated with the center point are the parameters of the 3D bounding box associated with the center point of the target in the radar coordinate system.

4. The method according to claim 1, wherein The backbone network includes: a first input layer, multiple bottleneck layers, a standard convolutional layer, an average pooling layer and a first output layer; The bottleneck layer includes a 1×1 expansion layer, a 3×3 depth-separable convolution layer, and a 1×1 projection layer; Wherein, the stride of the expansion layer and the depthwise separable convolutional layer in at least one of the bottleneck layers is 1.

5. The method according to claim 1, wherein The neck network includes: a second input layer, two bilinear interpolation layers, a convolutional layer, an activation layer and a second output layer.

6. The method according to claim 5, wherein Adjusting parameters of the object detection model based on a difference between actual position information and predicted position information of a 2D bounding box associated with the center point, a difference between actual position information and predicted position information of a 3D bounding box associated with the center point, and a difference between actual position information and predicted position information associated with a corner point, including: Adjust the parameters of the target detection model based on the difference between the actual position information and the predicted position information of the 2D bounding box associated with the center point, the difference between the actual position information and the predicted position information of the 3D bounding box associated with the center point, the difference between the sample corner point heat map and the predicted corner point heat map, the difference between the depth value of the sample corner point and the depth value of the predicted corner point, and the difference between the sample corner point offset and the predicted corner point offset.

7. The method according to claim 1, wherein The distance between the target and the vehicle in the fisheye image sample is no more than 30 meters.

8. A 3D object detection method, characterized in that: include: Obtain target fisheye images collected during parking; Inputting the target fisheye image into a trained target detection model, and outputting position information of a 2D bounding box and a 3D bounding box associated with the center point; wherein the target detection model is trained by the method according to any one of claims 1 to 7; Parameters of the 3D bounding box associated with the center point are determined based on the position information of the 2D bounding box associated with the center point and the position information of the 3D bounding box associated with the center point.

9. A computer program product, characterized in that The method comprises one or more computer instructions, which implement the method according to any one of claims 1 to 8 when executed by a processor.

Citation Information

Patent Citations

  • Target detection model training method, and target detection method and device

    CN112329873A

  • Parking control method, obstacle recognition model training method and device

    CN114802261A

  • Parking space detection method and device, vehicle and storage medium

    CN114841963A

  • Parking space detection method and device, vehicle and storage medium

    CN114926416A

  • Target detection model training method, target detection method and device

    CN116543143A