A novel monocular vision 3D object detection method based on key point constraints
The method addresses suboptimal keypoint selection and task branch collaboration in monocular 3D target detection by predicting 2D key points from 3D bounding box projections and integrating loss functions, improving detection accuracy and reducing complexity.
Patent Information
- Application Number
- CN202210846973.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-06
AI Technical Summary
In the existing monocular vision 3D object detection technology, the key point selection is not good, the detection branches are insufficient, and the post-processing process of geometric constraints is complicated, resulting in insufficient detection accuracy and robustness.
A prediction branch task with the projection point of the 2D key points in the 6 surface centers of the target 3D external frame is designed as a 2D key point projection point at the image plane. Through a multi-task collaborative training method, multiple branch tasks are trained together using a loss function, including the loss terms of the 2D key point and the 3D position, to improve detection performance.
It improves the accuracy and robustness of monocular visual 3D object detection, simplifies the computational complexity, and achieves real-time and feasibility.
Smart Images

Figure CN115205654B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of 3D object detection, and in particular to a monocular vision 3D object detection method based on key point constraints and multi-task collaboration. Background Art
[0002] 3D object detection technology plays an important role in the perception task of autonomous driving. Its core tasks include locating dynamic objects around an autonomous driving vehicle, estimating the object category, spatial size, and orientation, as Figure 2 shown. However, how to save costs while ensuring safety remains a problem that has not been fully solved yet. Among various perception sensors, the monocular vision camera has become the most widely used sensor in autonomous driving technology due to its low cost, small size, and light weight. However, how to use this data without depth information for 3D object detection and maintain high accuracy and robustness remains a huge challenge.
[0003] The existing object detection technologies have the following disadvantages:
[0004] 1. The key points selected by the existing technologies are not optimal. Usually, 8 corner points of the 3D bounding box of the object are selected, and these corner points do not fall on the object main body pixels and cannot effectively express the object features.
[0005] 2. The detection branches of the existing technologies are not sufficient, and multiple detection branches cannot be collaboratively trained, making it difficult for the model to learn more accurate features.
[0006] 3. The post-processing process introduced when the existing technologies fuse geometric constraints is too complex, with a high computational complexity, and the real-time performance and feasibility of the algorithm are low. Summary of the Invention
[0007] The purpose of the present invention is to overcome the deficiencies of the existing technologies and provide a new monocular vision 3D object detection method based on key point constraints. A prediction branch task with the projection points of the centers of 6 faces of the 3D bounding box of the object on the image plane as 2D key points is designed. At the same time, a loss function is designed based on the results predicted by multiple branch tasks, so that this loss function affects multiple branch tasks simultaneously during the training process, thereby realizing collaborative training between multiple branch tasks and improving the performance of the object detection model. In addition, the method also designs the loss between the predicted 2D key points and the 2D key points calculated based on the predicted 3D positions, thereby collaboratively training the 2D key point prediction branch task and other branch tasks and improving the detection performance.
[0008] The purpose of the present invention is achieved by the following technical solutions:
[0009] A new monocular vision 3D object detection method based on key point constraints, comprising the following steps:
[0010] Step 1: Data preprocessing. Use a monocular camera to collect monocular images, and perform image preprocessing and label preprocessing on the monocular images;
[0011] Step 2: Feature extraction. Build a convolutional neural network, pass the preprocessed monocular images through the convolutional neural network for digital feature extraction, and extract information for the network branch tasks based on the digital features to generate the final output result of the model;
[0012] Step 3: Target information decoding. Decode the final output result of the model to obtain the center point position and length, width, and height attributes of the target;
[0013] Step 4: Loss function calculation. Design a loss function according to the information of the network branch tasks, and add a loss term l between the predicted 2D key points and the 2D key points calculated based on the predicted 3D positions in the loss function 2d-3d , and use the loss function and the preprocessed monocular images to perform multi-task branch collaborative training and model optimization on the convolutional neural network to obtain a trained 3D object detection model;
[0014] Step 5: Model deployment. After converting the format and compiling the trained 3D object detection model, deploy it in the AI computing device for online inference of the model.
[0015] Specifically, the image preprocessing process in Step 1 specifically includes: for multiple monocular images collected by the monocular camera, first uniformly convert the shapes of all monocular images to [W, H, 3] through affine transformation; then normalize the RGB value of each pixel in the monocular image from [0, 255] to between [0, 1];
[0016] Specifically, the label preprocessing process in Step 1 specifically includes the following sub-steps:
[0017] S11, 2D center point label generation. First, initialize a zero matrix H with a shape of [W / 4, H / 4, C]; then calculate the floating-point coordinates (kf j , yf j ) and integer coordinates (xi j , yi j ) of each target in the monocular image after 4 times downsampling; finally, assign values to the elements in the zero matrix H according to the following formula:
[0018]
[0019] where, σ jis the standard deviation determined by the length and width of the target 2D box; after assigning values to the all-zero matrix H according to the above formula, the peak in the all-zero matrix H corresponds to the position of the target center point in the monocular image, and the 2D center point label of the target in the monocular image is obtained;
[0020] S12, Generation of 2D center point offset label. First, initialize an all-zero matrix O_2D with a shape of [W / 4, H / 4, 2]; then assign values to the all-zero matrix O_2D according to the following formula:
[0021] O_2D(xi j ,yi j ,0) = xf j -xi j
[0022] O_2D(xi j ,yi j ,1) = yf j -yi j
[0023] After the assignment is completed, the offset between the position of the target center point corresponding to the peak in the all-zero matrix H and the actual center point position of the target in the monocular image is obtained, and the 2D center point offset label of the target is generated;
[0024] S13, Generation of 2D length and width label. First, initialize an all-zero matrix S with a shape of [W / 4, H / 4, 2]; then assign values to the matrix S according to the following formula to generate the 2D length and width label:
[0025] S(xi j ,yi j ,0) = w j
[0026] S(xi j ,yi j ,1) = h j
[0027] where w j 、h j represent the length and width of each target respectively;
[0028] S14, Generation of 3D center point projection offset label. First, initialize an all-zero matrix O_3D with a shape of [W / 4, H / 4, 2]; then, assign values to the matrix O_3D according to the following formula to generate the 3D center point projection offset label:
[0029] O_3D(xi j ,yi j ,0) = dx j
[0030] O_3D(xi j ,yij , 1) = dy j
[0031] where dx j and dy j represent the offsets of the projection of the 3D center point of each target in the monocular image from the center point of the target 2D box in the x and y directions, respectively;
[0032] S15, 3D center point depth label generation. In the monocular image, the depth d of the target 3D center point is obtained by calculating the ratio of the dimensions of the 3D box in the pixel coordinate system to those in the predicted world coordinate system; then, the depth d is converted to the absolute depth d0 through the formula d0 = 1 / σ(d) - 1, where σ is the Sigmoid function;
[0033] S16, 3D length, width, and height label generation. Based on the three-dimensional length, width, and height information of the target in the monocular image, calculate the ratio of the three-dimensional length, width, and height information to the average dimensions;
[0034] S17, 3D orientation label generation. Adopt a regression method based on MultiBin to predict the residual of the target orientation relative to the bin center;
[0035] S18, 2D key point label generation. Use the face center point as the key point to establish a multi-task learning network to generate labels for 6 key points for each target; for the i-th 3D target box, assume its orientation is R i (θ), and its 3D coordinates are and the length, width, and height are D i = [l i , w i , h i T , and the homogeneous coordinates of the 6 face center points of the target can be expressed as:
[0036]
[0037]
[0038]
[0039] A rotation matrix R is established according to the angle of the target's change around the y-axis of the 3D coordinates. The expression of R is as follows:
[0040]
[0041] Given the camera intrinsic matrix K, after projecting the cube face center point into the image coordinate system, the coordinates of the 2D key points are calculated by the following formula
[0042]
[0043] Specifically, step two specifically includes:
[0044] Construct a convolutional neural network, which includes a backbone network and a detection head branch. The backbone network uses an improved DLA-34 network; the detection head branch includes a 2D center point branch 2D center point offset branch 2D length and width 3D center point projection offset 3D length, width and height 3D orientation 3D center point depth and 2D key point offset
[0045] Input the monocular image that has undergone image preprocessing and label preprocessing into the backbone network to extract highly abstract digital features; the detection head branch further extracts the information corresponding to the branch tasks based on the digital features, and takes the 8 tensors output by the detection head branch as the final output result of the convolutional neural network.
[0046] Specifically, step three specifically includes: decoding the 8 tensors output by the detection head branch. First, extract the peaks of the predicted heatmaps for each category and retain the top 100 peaks; denote as the set of n center points detected and the center points belonging to category c in it; where the coordinates represent the approximate position of the target 2D center point, and the length, width and height attributes of the target can also be obtained through the coordinates Similarly. For the 8 tensors output by the 8 branches of the detection head, the first two dimensions of each tensor are w / 4 x h / 4, and the third dimension is the attribute information. Substituting the coordinates (x, y) into the first two dimensions gives the information in the third dimension, that is, the target attributes given by this branch. Using these coordinates to index the results of other tensors in turn can obtain the other attributes of the target.
[0047] Specifically, step four specifically includes: first designing corresponding loss functions based on the 8 network branch information extracted in step two. The loss functions include the 2D center point prediction loss function l n , the 2D center point offset prediction loss function l 2Doff , the 2D length and width prediction loss l 2D , the 3D center point projection offset prediction loss function l 3Doff , the 3D center point depth prediction loss function l dep , the 3D length, width and height prediction loss l 3D , the 3D orientation angle prediction loss function l ori and the 2D key point prediction loss function lkp ; Then, add a loss term l between the 2D key points calculated based on the predicted 3D positions and the predicted 2D key points to the loss function. 2d-3d ; Finally, use the preprocessed monocular images to perform multi-task branch collaborative training on the convolutional neural network. Meanwhile, during the training process, use the loss function to optimize the parameters of the convolutional neural network, and use the finally trained convolutional neural network as the 3D object detection model.
[0048] Specifically, the fifth step specifically includes:
[0049] Model format conversion: Convert the 3D object detection model format into a model format supported by the model deployment tool.
[0050] Model compilation: The deep learning model is compiled into a format supported by the AI computing device at this stage, and a representative calibration image is used to quantize the model.
[0051] Model inference: Load and compile the code and data files generated after model compilation in the AI computing device, generate an executable program and execute it to achieve online model inference.
[0052] Advantages of the present invention:
[0053] 1. The present invention designs a prediction branch task for monocular vision 3D object detection, with the projection points of the centers of the 6 faces of the object's 3D bounding box on the image plane as the 2D key points.
[0054] 2. The present invention designs a multi-task collaborative training method. Design a loss function based on the prediction results of multiple branch tasks, so that the loss function affects multiple branch tasks simultaneously during the training process, thereby achieving collaborative training between multiple branch tasks and improving the model performance.
[0055] 3. The present invention designs a multi-task collaborative training method for monocular vision 3D object detection tasks. Design the loss between the predicted 2D key points and the 2D key points calculated based on the predicted 3D positions, so as to collaboratively train the 2D key point prediction branch task and other branch tasks and improve the detection performance. Description of the drawings
[0056] Figure 1 is the flowchart of the method steps of the present invention;
[0057] Figure 2 is an example diagram of monocular vision 3D object detection;
[0058] Figure 3 is the flowchart of the multi-task collaborative 3D object detection of the present invention;
[0059] Figure 4It is a schematic diagram of the difference between the center points of the 2D box and the 3D box;
[0060] Figure 5 It is a sample diagram of the position of the target 2D key points;
[0061] Figure 6 It is a schematic diagram of the backbone network structure. Detailed implementation manners
[0062] For a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, the technical solutions of the present invention are described in detail below. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments, and should not be construed as limiting the scope of implementation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0063] Embodiment 1:
[0064] In this embodiment, as Figure 1 shown, a novel monocular vision 3D object detection method based on key point constraints includes the following steps:
[0065] Step 1: Data preprocessing, using a monocular camera to collect monocular images, and performing image preprocessing and label preprocessing on the monocular images;
[0066] Step 2: Feature extraction, constructing a convolutional neural network, passing the preprocessed monocular images through the convolutional neural network for digital feature extraction, and extracting information for the network branch tasks based on the digital features to generate the final output result of the model;
[0067] Step 3: Target information decoding, decoding the final output result of the model to obtain the center point position and length, width, and height attributes of the target;
[0068] Step 4: Loss function calculation, designing a loss function according to the information of the network branch tasks, and adding a loss term l2d-3d between the predicted 2D key points and the 2D key points calculated based on the predicted 3D positions in the loss function. Using the loss function and the preprocessed monocular images to perform multi-task branch collaborative training and model optimization on the convolutional neural network to obtain a trained 3D object detection model;
[0069] Step 5: Model deployment, after converting the format and compiling the trained 3D object detection model, deploying it in an AI computing device for online inference of the model.
[0070] Regarding the existing defects in monocular vision 3D object detection, the current mainstream and the closest technical solution to the present invention is CenterNet and its derivative algorithms. The technical route of CenterNet is as follows: After the original image is preprocessed, feature extraction is performed through the backbone network. The extracted features are directly or indirectly output through different detection branches with various attributes of the 3D detection box (i.e., object category, center point coordinates, length, width, height, and orientation angle).
[0071] The derivative algorithms of CenterNet believe that the combination of a deep neural network and geometric constraints is required to jointly estimate the appearance and spatial-related information of the 3D detection box. For example, RTM3D regards the 8 corner points of the 3D detection box as key points, then further adds a key point prediction branch in the detection branch to predict the 2D coordinates of the corner points in the image, and finally uses geometric constraint conditions to solve the optimization problem to obtain the final 3D detection box. Compared with detection networks such as yolo, ssd, and faster_rcnn that rely on a large number of anchors, CenterNet is an anchor-free object detection network and has advantages in both speed and accuracy. In addition to the detection task, CenterNet can also be used for limb recognition or 3D object detection, etc. Therefore, CenterNet proposes three backbone network structures, namely Resnet-18, DLA-34, and Hourglass-104.
[0072] In this embodiment, as Figure 3 shown, the technical implementation process of the method is specifically as follows:
[0073] Step 1: Data preprocessing
[0074] Preprocessing of the input image. For an input image of any shape, in order to ensure that the size of the final output feature is fixed, it is transformed into a unified shape [W, H, 3] through affine transformation. In addition, for better feature learning, the RGB value of each pixel needs to be normalized from [0, 255] to between [0, 1]. For image input, multiple images in the time series of a monocular camera, multiple images obtained by multiple monocular cameras, or multiple images in the time series of multiple monocular cameras can be used for object detection.
[0075] Preprocessing of labels. Since the data in the annotation files of most existing datasets is not necessarily the object that the network needs to directly predict. Taking the nuScenes dataset [4] as an example, for each object in this dataset, it includes its category, coordinates of the center point in the world coordinate system, length, width, height, orientation, and coordinates of the 2D box in the pixel coordinate system. However, these labels need to be preprocessed as described below:
[0076] (1) 2D center point label generation. First, initialize a zero matrix H with a shape of [W / 4, H / 4, C]. Then, calculate the floating-point coordinates (xf j , yf j ) and integer coordinates (xi j , yi j ) of each target after 4x downsampling. Finally, assign values to the elements in the matrix according to the following formula.
[0077]
[0078] where σ j is the standard deviation determined by the length and width of the target 2D box. In fact, after assigning values according to the above formula, a peak in the matrix H corresponds to the approximate position of the center point of a target.
[0079] (2) 2D center point offset label generation. Since the peak in the matrix H actually corresponds to the coordinates after rounding the position of the target center point, there is a certain error from the actual coordinates. Therefore, to accurately locate the position of each target, it is also necessary to predict the offset between the actual center point position of the target and the distance of the peak in the matrix H. The specific method is as follows: First, initialize a zero matrix O_2D with a shape of [W / 4, H / 4, 2]. Then, assign values to the matrix O_2D according to the following formula.
[0080] O_2D(xi j , yi j , 0) = xf j - xi j
[0081] O_2D(xi j , yi j , 1) = yf j - yi j
[0082] (3) 2D length and width label generation. The model can directly predict the length and width of the target. First, initialize a zero matrix S with a shape of [W / 4, H / 4, 2]. Then, assign values to the matrix S according to the following formula. Where w j , h j represent the length and width of each target respectively.
[0083] S(xi j , yi j , 0) = w j
[0084] S(xi j , yi j , 1) = h j
[0085] (4) Generation of 3D center point projection offset labels. As Figure 4 shown, there may be a certain offset between the 2D center point of the target and the projection of the 3D center point.
[0086] Therefore, in order to accurately estimate the projection position of the 3D center point, the offset between it and the 2D center point can be estimated. The specific method is the same as above. First, initialize a zero matrix O_3D with a shape of [W / 4, H / 4, 2]. Then, assign values to the matrix O_3D according to the following formula. Where dx j , dy j respectively represent the offsets in the x and y directions of the projection of the 3D center point of each target in the image and the center point of its 2D box.
[0087] O_3D(xi j , yi j , 0) = dx j
[0088] O_3D(xi j , yi j , 1) = dy j
[0089] (5) Generation of 3D center point depth labels. In order to calculate the position of the 3D center point of the target in 3D space, the depth of the 3D center point also needs to be calculated. In depth prediction, the present invention is based on an uncertainty-based depth prediction method. First, the depth can be obtained by calculating the ratio of the dimensional information of the 3D box in the pixel coordinate system to its corresponding predicted world coordinate system. Considering that depth is a scalar and not easy to regress, we first transform the network output depth d into absolute depth d0 through the formula d0 = 1 / σ(d) - 1, where σ is the Sigmoid function.
[0090] (6) Generation of 3D length, width, and height labels. The information of the 3D length, width, and height of the target is also very important. The output of this branch is To make the dimensional regression here more stable, instead of directly regressing the specific values of its dimensions, we use the ratios of them to the average dimensions for regression.
[0091] (7) Generation of 3D orientation labels. The present invention adopts a regression method based on MultiBin. Specifically, each orientation interval is represented by several overlapping bins. The neural network assigns several scalars to each bin to determine in which bin's orientation range the vehicle orientation is, and can predict the residual of the orientation relative to the bin center. Here, after regression through eight bins, the size of the output is
[0092] (8) 2D Key Point Label Generation. The present invention proposes to use the center point of the face as the key point and establish a multi-task learning network (the panoramic driving perception network YOLOP). That is, in the subsequent object detection based on key point prediction, key point labels are required. Therefore, for each object, labels of 6 key points (corresponding to the 6 faces of each cube) are generated, as shown in Figure 4. For the i-th 3D object box, assuming its orientation is R i (θ), and its 3D coordinates are and the length, width, and height are D i =[l i , w i , h i T . Therefore, the homogeneous coordinates of the center points of the 6 faces of the object can be expressed as:
[0093]
[0094]
[0095]
[0096] In autonomous driving, we usually consider the road surface to be flat. Therefore, only one parameter of the rotation matrix R that considers the angle of the object's change around the y-axis is considered. Assuming that the parameter 'rotation_y' collected in the above first step is r y , the expression of R can be represented as follows:
[0097]
[0098] Given the camera intrinsic matrix K, the coordinates of the 2D key points after the center points of the cube faces are projected into the image coordinate system can be calculated by the following formula.
[0099]
[0100] The final effect diagram is as shown in Figure 5 .
[0101] Step 2: Feature Extraction
[0102] Feature extraction includes two parts in total: the backbone network and the detection head branch. The backbone network is responsible for representing the preprocessed image as highly abstract digital features through a convolutional neural network; the detection head branch further extracts information for the branch task based on this feature and generates the final output result of the model. As shown in Figure 6 , the backbone network is processed using an improved DLA-34 backbone network. In addition, backbone networks such as MLP, CNN, ResNet, and Transformer can also be used to extract the digital features of the image.
[0103] Input an RGB three-channel image I of W×H×3, and the output feature image will obtain a four-fold downsampling. In order to utilize key point information for geometric constraints and uncertainty for depth prediction optimization, the following information will be learned using a deep neural network.
[0104] The detection head branch contains a total of 8 branches, namely the 2D center point branch 2D center point offset branch 2D length and width 3D center point projection offset 3D length, width and height 3D orientation 3D center point depth 2D key point offset They correspond one-to-one with the preprocessed label data.
[0105] Step Three: Post-processing
[0106] This process can be understood as the inverse process of label preprocessing in Step One. Its purpose is to decode the tensors of the 8 branches output by the detection head into the center point position, length, width, and other attributes of the target.
[0107] The output tensor of the 2D center point branch has 3 dimensions. The first 2 dimensions are w / 4 x h / 4, which is a heat map, where the peak represents the target center position; the 3rd dimension is the category C (there are a total of C categories); that is, the output tensor of the 2D center point branch is a total of C heat maps, and each heat map represents the distribution of the center positions of the targets belonging to this category. Select the peak as the center position of the target predicted for this category. First, extract the peaks of the predicted heat maps of each category (that is, the places where its value is greater than its eight surrounding neighborhoods), and retain the first 100 peaks. Denote as the set of n center points detected the part of category c in whose coordinates
[0108] represent the approximate position of the target 2D center point. Other attributes of the target can also be obtained through this coordinate. For the 8 tensors output by the 8 branches of the detection head, the first two dimensions of each tensor are w / 4 x h / 4, and the third dimension is the attribute information (i.e., the category). After substituting the coordinates (x, y) into the first two dimensions, the information of the third dimension is obtained, that is, the target attribute given by this branch. Using this coordinate to index the results of other tensors in turn can obtain other attributes of the target.
[0109] Based on the branch information of the above multi-task collaborative 3D object detection neural network (convolutional neural network), the present invention relates to the following loss functions: 2D center point prediction loss function ln , 2D center point offset prediction loss function l 2Doff , 2D length and width prediction loss l 2D , 3D center point projection offset prediction loss function l 3Doff , and the 3D center point depth prediction loss function l dep , 3D length, width and height prediction loss l 3D , 3D orientation angle prediction loss function l ori , and the 2D key point prediction loss function l kp .
[0110] In addition, the present invention adds a loss term l between the predicted 2D key points and the 2D key points calculated based on the predicted 3D position 2d-3d The calculation of this loss uses the results of 2D center point prediction, 2D key point prediction, 3D center point depth prediction, 3D length, width and height prediction, and 3D orientation angle prediction. Therefore, this loss will collaboratively train multiple task branches. Through this collaborative training between multiple branches, it is possible to effectively introduce inter-branch constraints and improve detection performance. In addition, 3D key points can also be used to achieve collaborative training.
[0111] In summary, the results of 3D object detection in this multi-task branch are as follows:
[0112] L=w n l n +w 2Doff l 2Doff +w 2D l 2D +w ori l ori +w 3D l 3D +w 3Doff l 3Doff +w dep l dep +w kp l kp +w 2d-3d l 2d-3d
[0113] Among them, l n Consistent with the focal loss used in CenterNet, l 2D The depth-guided l-1 loss and the angle loss l are used. ori MultiBin loss is used, and the other loss functions all use L1-loss.
[0114] Step 5: Model deployment
[0115] To apply the model to an actual scenario, it is also necessary to deploy the trained model in an AI computing device. This invention takes TI's TDA4VM platform as an example to illustrate the model deployment process. Deploying a trained model includes three steps: model format conversion, model compilation, and inference.
[0116] Model format conversion: According to the model formats supported by the model deployment tool, for example, the Texas Instruments Deep Learning Runtime (TIDL-RT) deployment tool only supports models in onnx, caffe, and TensorFlow formats. The model needs to be converted to a supported format in advance.
[0117] Model compilation: The deep learning model is compiled into a format supported by the AI computing device at this stage. Representative calibration images are also used to quantize the model in this stage to reduce the model's computational workload.
[0118] Model inference: After model compilation, code and data files are generated. These files can be loaded and compiled in the AI computing device to generate an executable program. Executing this program can achieve online inference of the model. Deployment hardware can use other AI computing platforms such as NVIDIA XAVIER NX / AGX, NVIDIA ORIN, Raspberry PI, Khadas VIM3 / VIM4, Huawei Hi3559A, Qualcomm SA8155P / SA8195P / SA8295P, etc. for model deployment.
[0119] The method of this embodiment has the following advantages:
[0120] 1. This invention designs a prediction branch task for monocular vision 3D object detection, with the projection points of the centers of the 6 faces of the object's 3D bounding box on the image plane as 2D key points.
[0121] 2. This invention designs a multi-task collaborative training method. Based on the results predicted by multiple branch tasks, a loss function is designed to make this loss function affect multiple branch tasks simultaneously during the training process, thereby achieving collaborative training between multiple branch tasks and improving the model performance.
[0122] 3. This invention designs a multi-task collaborative training method for the monocular vision 3D object detection task. A loss is designed between the predicted 2D key points and the 2D key points calculated based on the predicted 3D positions, thereby co-training the 2D key point prediction branch task and other branch tasks to improve the detection performance.
[0123] 4. This invention designs a method for deploying the model on an actual AI computing device.
[0124] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A novel monocular vision 3D object detection method based on key point constraints, characterized in that, It includes the following steps: Step 1: Data preprocessing. Use a monocular camera to collect monocular images, and perform image preprocessing and label preprocessing on the monocular images. Step 2: Feature extraction. Build a convolutional neural network, pass the preprocessed monocular images through the convolutional neural network for digital feature extraction, and extract information for network branch tasks based on the digital features to generate the final output result of the model. Step 3: Target information decoding. Decode the final output result of the model to obtain the center point position, length, width, and height attributes of the target. Step 4: Loss function calculation. Design a loss function according to the information of the network branch tasks, and add a loss term l between the predicted 2D key points and the 2D key points calculated based on the predicted 3D positions in the loss function. 2d-3d Use the loss function and the preprocessed monocular image to perform multi-task branch collaborative training and model optimization on the convolutional neural network to obtain a trained 3D object detection model. Step 5: Model deployment. After converting the format and compiling the trained 3D object detection model, deploy it in the AI computing device for online model inference.
2. A novel monocular vision 3D object detection method based on key point constraints according to claim 1, characterized in that, The image preprocessing process in Step 1 specifically includes: For multiple monocular images collected by the monocular camera, first use affine transformation to uniformly convert the shapes of all monocular images to [W, H, 3]; then normalize the RGB value of each pixel in the monocular image from [0, 255] to between [0, 1].
3. A novel monocular vision 3D object detection method based on key point constraints according to claim 1, characterized in that, The label preprocessing process in Step 1 specifically includes the following sub-steps: S11, 2D center point label generation. First, initialize a zero matrix H with the shape of [W / 4, H / 4, C]; then calculate the floating-point coordinates (xf j , yf j ) and integer coordinates (xi j , yi j ) of each target in the monocular image after 4x downsampling; finally, assign values to the elements in the zero matrix H according to the following formula: where σ j is the standard deviation determined by the length and width of the target 2D box; after assigning values to the all-zero matrix H according to the above formula, the peak value in the all-zero matrix H corresponds to the position of the target center point in the monocular image, and the 2D center point label of the target in the monocular image is obtained; S12, Generation of 2D center point offset label. First, initialize a zero matrix O_2D with a shape of [W / 4, H / 4, 2]; then assign values to the zero matrix O_2D according to the following formula: O_2D(xi j ,yi j ,0)=xf j -xi j O_2D(xi j ,yi j ,1) = yf j -yi j After the assignment is completed, obtain the offset between the center point position corresponding to the peak in the zero matrix H and the true center point position of the target in the monocular image, and generate the 2D center point offset label of the target. S13, Generation of 2D length and width labels. First, initialize a zero matrix S with a shape of [W / 4, H / 4, 2]; then assign values to the matrix S according to the following formula to generate 2D length and width labels: S(xi j ,yi j ,0)=w j S(xi j ,yi j ,1)=h j where w j and h j represent the length and width of each target, respectively; S14, Generation of 3D center point projection offset label. First, initialize a zero matrix O_3D with a shape of [W / 4, H / 4, 2]; then, assign values to the matrix O_3D according to the following formula to generate the 3D center point projection offset label: O_3D(xi j ,yi j ,0) = dx j O_3D(xi j ,yi j ,1)=dy j where dx j and dy j respectively represent the offsets in the x - direction and y - direction between the projection of the 3D center point of each target in the monocular image and the center point of the target 2D box. S15, Generation of 3D center point depth label. In the monocular image, obtain the depth d of the 3D center point of the target by calculating the ratio of the dimensions of the 3D box in the pixel coordinate system to the ratio in the predicted world coordinate system; then convert the depth d to the absolute depth d0 through the formula d0 = 1 / σ(d) - 1, where σ is the Sigmoid function. S16, Generation of 3D length, width, and height labels. Calculate the ratio of the three-dimensional length, width, and height information of the target in the monocular image to the average dimension. S17, Generation of 3D orientation label. Use the regression method based on MultiBin to predict the residual of the target orientation relative to the bin center. S18, 2D key point label generation. Using the center point of the face as the key point, a multi-task learning network is established to generate labels of 6 key points for each target. For the i-th 3D target box, assuming its orientation is R i (θ), the 3D coordinates are and the length, width, and height are D i =[l i , w i , h i T , the homogeneous coordinates of the center points of the 6 faces of the target can be expressed as: Establish a rotation matrix R according to the angle of the target rotating around the y-axis of the 3D coordinate. The expression of R is as follows: Given the camera internal parameter matrix K, after projecting the center point of the cube face into the image coordinate system, calculate the coordinates of the 2D key points by the following formula 4. A novel monocular vision 3D object detection method based on key point constraints according to claim 1, characterized in that, Step 2 specifically includes: Construct a convolutional neural network, which includes a backbone network and a detection head branch. The backbone network uses an improved DLA-34 network; the detection head branch includes a 2D center point branch 2D center point offset branch 2D length and width 3D center point projection offset 3D length, width and height 3D orientation 3D center point depth and 2D key point offset Input the monocular image that has undergone image preprocessing and label preprocessing into the backbone network to extract highly abstract digital features; the detection head branch further extracts the information corresponding to the branch tasks based on the digital features, and uses the 8 tensors output by the detection head branch as the final output result of the convolutional neural network.
5. A novel monocular vision 3D object detection method based on key point constraints according to claim 1, characterized in that, The specific steps of Step 3 include: decoding the 8 tensors output by the detection head branch. First, extract the peaks of the predicted heatmaps for each category and retain the top 100 peaks; denote as the set of n detected center points in category c in the set; where the coordinates represent the approximate position of the target 2D center point, and the length, width, and height attributes of the target can also be obtained through the coordinates 6. A novel monocular vision 3D object detection method based on key point constraints according to claim 1, characterized in that The specific steps of Step 4 include: First, design corresponding loss functions based on the 8 network branch information extracted in Step 2. The loss functions include 2D center point prediction loss function l n , 2D center point offset prediction loss function l 2Doff , 2D length and width prediction loss l 2D , 3D center point projection offset prediction loss function l 3Doff , 3D center point depth prediction loss function l dep , 3D length, width and height prediction loss l 3D , 3D orientation angle prediction loss function l ori and 2D key point prediction loss function l kp ; Then add a loss term l 2d-3d between the 2D key points calculated based on the predicted 3D position and the predicted 2D key points in the loss function; Finally, use the preprocessed monocular image to perform multi-task branch collaborative training on the convolutional neural network, and at the same time use the loss function to optimize the parameters of the convolutional neural network during the training process, and use the finally trained convolutional neural network as the 3D object detection model.
7. A novel monocular vision 3D object detection method based on key-point constraints according to claim 1, characterized in that Step 5 specifically includes: Model format conversion. Convert the format of the 3D object detection model to the model format supported by the model deployment tool. Model compilation: In this stage, the deep learning model is compiled into a format supported by the AI computing device, and the model is quantized using representative calibration images; Model inference: Load and compile the code and data files generated after model compilation in the AI computing device to generate an executable program and execute it to achieve online model inference.
Citation Information
Patent Citations
Three-dimensional target detection method and device and storage medium
CN111126269A
Monocular image three-dimensional target detection method based on depth information estimation
CN113436239A
Cited By
Monocular three-dimensional target detection method, system and device based on Lyapunov optimization theory and multi-task learning and storage medium
CN121582744A