A 3D Object Detection Method and System Based on Plane Constraint and Position Constraint

By introducing normal vectors and gradient constraints into the depth estimation model, and adopting a two-stage training strategy, combined with 3D detection network, the problems of pseudo-point cloud shape distortion and position offset are solved, and the accuracy and accuracy of 3D object detection are improved.

CN116091574BActive Publication Date: 2025-07-18XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310028861.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-07-18
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

In the existing 3D object detection method based on image, depth estimation errors lead to shape distortion and position offset of pseudo-point clouds, affecting pseudo-point cloud detection performance.

Method used

The plane constraints and position constraints are used to improve the shape characteristics of pseudo-point clouds through normal vectors and gradient constraints, and the depth estimation model is optimized through two-stage training strategies, combined with 3D detection network for end-to-end training, and alternately trained using pseudo-point cloud box labels and GT labels.

Benefits of technology

It significantly improves the shape structure characteristics and position prediction accuracy of pseudo-point clouds, reduces depth estimation errors, and improves the performance of 3D detection models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116091574B_ABST
    Figure CN116091574B_ABST
Patent Text Reader

Abstract

The present invention discloses a 3D object detection method and system based on plane constraint and position constraint. The input data is an RGB image, and a depth map is obtained by training using a depth estimation model ForeSeE; the obtained depth map is segmented using an instance segmentation mask, and the obtained foreground part is converted into foreground point cloud; taking the obtained foreground point cloud as the center, a pseudo point cloud box label is generated with the same size as the GT detection box; the parameters of the depth estimation model are frozen, and a 3D detection network is trained. The pseudo point cloud label is used as the training label to complete the first-stage training. The parameters of the 3D detection network F-PointNet are frozen, and the depth estimation model is trained. The GT detection box is used as the label of the 3D detector to train the depth estimation network to complete the second-stage training; the first stage and the second stage are alternately trained so that the 3D detection network F-PointNet can correctly predict the pseudo point cloud position at all times. The present invention can not only significantly improve the depth estimation effect, make the contour more prominent, but also improve the performance of the 3D detection model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to a 3D target detection method and system based on plane constraints and position constraints. Background Art

[0002] 3D object detection is an important task in the fields of autonomous driving and robot obstacle avoidance. Its purpose is to obtain the position and volume information of surrounding objects in three-dimensional space. 3D object detection can be divided into point cloud-based algorithms and image-based detection algorithms according to the different forms of input data. Although the accuracy of image-based detection algorithms lags behind that of pure point cloud detection algorithms, image-based detection algorithms such as monocular 3D detection are still a hot research topic in academia and industry due to the advantages of cameras such as high resolution, low cost, and convenient deployment. In recent years, some scholars have proposed a monocular detection method in the form of pseudo point cloud. The pseudo point cloud detection algorithm decouples monocular detection into two separate modules: depth estimation and pure point cloud 3D detection. This method first performs depth estimation, then converts the depth map into a pseudo point cloud, and finally uses the pseudo point cloud as input to train the pure point cloud detection model. The pseudo point cloud detection algorithm can improve the accuracy of monocular detection with the help of high-precision pure point cloud detection algorithms. The depth estimation module in the algorithm can be pre-trained with the help of large-scale data sets, which can improve generalization and adapt to more complex and changing scenes. The 3D detection module can also flexibly select high-precision detection models according to actual scene requirements.

[0003] The difficulty of the pseudo point cloud detection method lies in depth estimation. Monocular depth estimation itself is an ill-posed problem, and its predicted depth map is usually inaccurate. Compared with the real point cloud, the objects in the pseudo point cloud are usually severely distorted in shape and accompanied by position offset, which leads to a decrease in 3D detection performance. Based on the above description, the depth estimation problems in the pseudo point cloud detection method can be summarized into the following two categories:

[0004] First, the blur of depth estimation causes serious distortion of the shape of objects in the pseudo point cloud. Since traditional depth estimation usually focuses on reducing the pixel-level error of depth estimation rather than optimizing the depth structure. The predicted depth map is usually blurred inside the object and around the contour. The blur in the predicted depth map will cause the deformation of the pseudo point cloud. There are shape distortions of objects in the pseudo point cloud, and there are tailing phenomena around the contour. This makes it difficult for the 3D detection network to learn effective features from the distorted pseudo point cloud during training. Therefore, the 3D detection network may produce a large number of false detection results during the prediction stage. In recent years, some post-processing methods have been proposed to solve the problem of pseudo point cloud distortion. These methods usually use instance segmentation or redesign pseudo point cloud sparse schemes to reduce tailing point clouds. However, additional processing operations may make the overall model complicated and unsuitable for real-time applications. In addition, the distortion inside the object cannot be well handled.

[0005] Second, the depth estimation error causes a deviation in the predicted position of the object. It is very difficult to estimate the absolute distance of an object from a single RGB image. Especially as the distance increases, the depth labels become sparse, and the depth estimation error becomes more serious. The depth estimation error will cause the position of the object in the pseudo-point cloud to shift, thus interfering with the 3D detection result. Some recent methods propose to solve this problem through the joint training of a depth estimation model and a 3D detection network. However, due to the position shift problem, some GT labels may not accurately correspond to the predicted position of the object in the pseudo-point cloud. Using these GT labels in the joint training will interfere with the training of the 3D detection network, making it unable to learn the correct position of the object in the pseudo-point cloud, thus reducing the performance of 3D detection.

[0006] Therefore, how to solve the two problems of object shape distortion and position shift in the pseudo-point cloud under monocular depth estimation has become the key to the pseudo-point cloud detection method. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a 3D object detection method and system based on plane constraint and position constraint for solving the technical problems of pseudo-point cloud shape distortion and prediction position error caused by depth estimation error, so as to improve the performance of pseudo-point cloud 3D object detection.

[0008] The present invention adopts the following technical solutions:

[0009] A 3D object detection method based on plane constraint and position constraint includes the following steps:

[0010] S1. Input the RGB image data, and use the depth estimation model ForeSeE to train to obtain a depth map;

[0011] S2. Use the instance segmentation mask to segment the depth map obtained in step S1, and convert the obtained foreground part into a foreground point cloud;

[0012] S3. Centering on the foreground point cloud obtained in step S2, generate a pseudo-point cloud box label with the same size as the GT detection box;

[0013] S4. Freeze the parameters of the depth estimation model, train the 3D detection network, use the pseudo-point cloud label obtained in step S3 as the training label to complete the first-stage training, freeze the parameters of the 3D detection network F-PointNet, train the depth estimation model, and use the GT detection box as the label of the 3D detector to train the depth estimation network to complete the second-stage training; alternately train the first stage and the second stage to make the 3D detection network F-PointNet correctly predict the pseudo-point cloud position at all times.

[0014] Specifically, in step S1, the loss function loss of depth is trained and predicted using the depth estimation model ForeSeE wcel as follows:

[0015]

[0016] where wcel_ fg , wcel_ bg are the pixel-level cross-entropy loss functions for the foreground and background respectively; α is the weight of the foreground loss.

[0017] Specifically, in step S1, the normal vector constraint loss normal is calculated as follows:

[0018]

[0019] where N is the number of valid point groups within the 2D bounding box of the object, is the predicted normal vector, is the true normal vector.

[0020] Specifically, in step S1, the gradient constraint is calculated as follows:

[0021]

[0022] where N is the number of valid point groups within the 2D bounding box of the object, is the predicted horizontal gradient difference, is the true horizontal gradient difference, is the predicted vertical gradient difference, is the true vertical gradient difference.

[0023] Specifically, in step S2, the following conversion formula is used to convert the foreground depth to the foreground point cloud:

[0024]

[0025]

[0026] where u, v are pixel coordinates, f x , f y is the camera focal length, c x , c y is the pixel coordinate of the image center point, and z is the predicted depth.

[0027] Specifically, in step S3, the pseudo-point cloud box is a detection box centered on the pseudo-point cloud with the same size as the GT detection box. When there is a deviation in the GT box, the pseudo-point cloud box represents the position of the pseudo-point cloud. When the ratio value is greater than 0.25, the center point of the GT box is used as the center point of the pseudo-point cloud box; when the ratio value is less than 0.25, the average value of the positions of all the pseudo-point clouds of the object is used as the center point of the pseudo-point cloud box.

[0028] Furthermore, the center position center of the pseudo-point cloud box pseudo is calculated as follows:

[0029]

[0030] where Num GT is the number of point clouds of a pseudo-point cloud object within the GT box, Num all is the total number of point clouds of a pseudo-point cloud object, center GT is the center position of the GT box, mean is the center position of the foreground point cloud, and thh is the threshold of the ratio.

[0031] Specifically, in step S4, the loss function loss1 of the first stage det is as follows:

[0032] loss1 det = PointNetLoss(Box pred , Pseudolabel)

[0033] where FPointNetLoss is the original loss function of the detection network F-PointNet, Box pred is the predicted 3D box, and Pseudo is the pseudo-point cloud label.

[0034] Specifically, in step S4, the loss function loss of the second stage all is as follows:

[0035] loss all = 1*oss wcel + 2*oss normal + 3*oss gradient + 4*loss2 del

[0036] where λ1 = 6, λ2 = λ3 = 1, λ4 = 0.001, loss wcel is the loss function, loss normal is the normal vector constraint, loss gradient is the gradient constraint, and loss2 del is the 3D detection function.

[0037] In a second aspect, an embodiment of the present invention provides a 3D object detection system based on plane constraint and position constraint, including:

[0038] An estimation module that inputs an RGB image of data and uses a depth estimation model ForeSeE to train and obtain a depth map;

[0039] A conversion module that uses an instance segmentation mask to segment the depth map obtained by the estimation module and converts the obtained foreground part into a foreground point cloud;

[0040] A labeling module that generates a pseudo-point cloud box label centered on the foreground point cloud obtained by the conversion module and having the same size as the GT detection box;

[0041] A prediction module that freezes the parameters of the depth estimation model, trains a 3D detection network, uses the pseudo-point cloud label obtained by the labeling module as a training label to complete the first-stage training, freezes the parameters of the 3D detection network F-PointNet, trains the depth estimation model, and uses the GT detection box as a label for the 3D detector to train the depth estimation network to complete the second-stage training; alternately train the first stage and the second stage so that the 3D detection network F-PointNet correctly predicts the position of the pseudo-point cloud at all times.

[0042] Compared with the prior art, the present invention has at least the following beneficial effects:

[0043] A 3D object detection method based on plane constraint and position constraint. Aiming at the problem of pseudo-point cloud shape distortion, plane constraints (random normal vector and gradient constraints) are proposed, which can significantly improve the shape of the pseudo-point cloud. The normal vector constraint is selected to enhance the shape structure features of the object in the flat area; the gradient constraint is selected to highlight the edges of the object in the depth map and reduce the trailing phenomenon in the pseudo-point cloud. Considering that the depth labels used during training are sparse and irregularly distributed, a method of randomly sampling points within the 2D box of the object is adopted to construct the normal vector and gradient constraints. After adopting the normal vector and gradient constraints, the depth estimation effect is significantly improved. The object contour in the depth map is more obvious. The shape structure features of the object in the pseudo-point cloud are strengthened, and the trailing phenomenon in the edge part is reduced; for the problem of depth estimation error, position constraints (end-to-end training + pseudo-point cloud box label + two-stage training strategy) are proposed, which can significantly reduce the pseudo-point cloud position prediction error. A 3D detection network is added behind the depth estimation model for end-to-end joint training. The depth estimation model for 3D detection is optimized by adding the information of the 3D detection box additionally. Since the position deviation of the pseudo-point cloud will also interfere with the training of the 3D detection network, it is proposed to use the pseudo-point cloud box (the detection box generated with the center of the pseudo-point cloud) as the training label. Further, a two-stage training method is proposed for the pseudo-point cloud label. In the first stage, the depth estimation model is frozen, and only the pseudo-point cloud box label is used to train the 3D detection network to correctly identify the objects in the scene. In the second stage, the 3D detection network is frozen, and only the GT label is used to train the depth estimation model. After adopting the two-stage training strategy, the depth estimation error of the objects in the middle and far distances is significantly reduced, thereby reducing the interference during the training of the 3D detection model and improving the performance of the detection model.

[0044] Furthermore, using the depth estimation model ForeSeE can obtain a high-precision pixel-level chromaticity map. The loss function adopted calculates the foreground object and the background separately and gives a high weight to the foreground. It can make the model focus on the prediction of the foreground object and improve the prediction accuracy of the foreground object.

[0045] Furthermore, adopt the normal vector constraint loss normal It will enhance the prediction accuracy of the internal structure of the object in the depth map, thereby enhancing the structure features of the foreground point cloud and improving the 3D detection accuracy.

[0046] Furthermore, adopt the normal vector constraint loss gradient It will enhance the prediction accuracy of the object contour, thereby reducing the trailing phenomenon at the edge of the foreground point cloud and improving the 3D detection accuracy.

[0047] Furthermore, using the conversion formula can accurately convert the foreground depth into the foreground point cloud and can do so without the background depth, improving the calculation efficiency.

[0048] Furthermore, pseudo point cloud boxes are generated centered on the previous scene point cloud. When there is a large error in the current scene depth prediction, the GT label will deviate from the previous scene point cloud; while the pseudo point cloud box can accurately represent the position of the previous scene point cloud, thus reducing the interference to the subsequent 3D detection training.

[0049] Furthermore, the center position center of the pseudo point cloud box is determined according to the ratio value. pseudo The ratio value can reflect the accuracy of the foreground depth prediction. When ratio > 0.25, the foreground depth prediction is accurate, and the center of the GT label is used as the center of the pseudo point cloud label; otherwise, the center of the previous scene point cloud is used as the center of the pseudo point cloud label. Such a setting is more accurate.

[0050] Furthermore, the loss function loss1 in the first stage det is used to train the 3D detection network F-Pointnet. The pseudo point cloud label is adopted in the loss function, which enables F-PointNet to learn the correct position of the foreground pseudo point cloud.

[0051] Furthermore, the loss function loss in the first stage all can optimize the depth estimation model ForeSee with the help of F-PointNet. One item loss2 in the loss function det adopts the GT label. In this way, when there is an error between the position of the current scene point cloud and the GT label, the value of the loss function of the 3D detection F-PointNet will increase. Subsequently, in the error backpropagation link, it affects the parameter update of the depth estimation network, making the previous scene point cloud closer to the GT label, thereby optimizing the depth estimation model.

[0052] It can be understood that the beneficial effects of the second aspect above can refer to the relevant descriptions in the first aspect above, and will not be elaborated here.

[0053] To sum up, the present invention can not only significantly improve the depth estimation effect with more prominent contours; the shape features of the pseudo point cloud are more obvious, but also significantly reduce the depth estimation error of medium and far objects, thereby improving the performance of the 3D detection model.

[0054] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is the overall framework diagram of the present invention;

[0056] Figure 2 is the input-output diagram of the depth estimation of the present invention. Among them, (a) is the input RGB image, and (b) is the output depth map;

[0057] Figure 3It is a display diagram of planar constraints (normal vector and gradient) adopted by the present invention. Among them, (a) is a schematic diagram of constructing normal vector constraints by randomly sampling points, and (b) is a schematic diagram of constructing gradient constraints by randomly sampling points;

[0058] Figure 4 It is the semantic segmentation adopted by the present invention;

[0059] Figure 5 It is a schematic diagram of the pseudo point cloud box label of the present invention;

[0060] Figure 6 It is a schematic diagram of the two-stage training strategy of the present invention. Among them, (a) is a schematic diagram of the first-stage training, and (b) is a schematic diagram of the second-stage training;

[0061] Figure 7 It is a comparison data graph of the 3D detection results of the present invention and various methods;

[0062] Figure 8 It is an objective comparison graph of the present invention before and after adding normal vector constraints on the depth map and point cloud. Among them, (a) is the depth map before adding normal vector constraints, and (b) is the depth map after adding normal vector constraints.

[0063] Figure 9 It is an objective comparison graph of the present invention before and after adding gradient constraints on the depth map and point cloud. Among them, (a) is the depth map before adding gradient constraints, and (b) is the depth map after adding gradient constraints.

[0064] Figure 10 It is an objective comparison graph of the present invention on the 3D detection results before and after successively adding normal vector and gradient constraints. Among them, (a) is the depth map before adding normal vector and gradient constraints, (b) is the depth map after adding normal vector constraints, and (c) is the depth map after adding normal vector and gradient constraints.

[0065] Figure 11 It is an objective comparison graph of the present invention on the point cloud position prediction results before and after two-stage training. Among them, (a) is the point cloud map without adding the two-stage training strategy, and (b) is the point cloud map after adding the two-stage training strategy. Detailed implementation manners

[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0067] In the description of the present invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0068] It should also be understood that the terms used in the specification of the present invention are merely for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0069] It should be further understood that the term " / and" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally indicates that the contextually related objects have an "or" relationship.

[0070] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present invention to describe preset ranges, etc., these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0071] Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" may be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0072] Structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of the various regions and layers shown in the figures and their relative sizes and positional relationships are merely exemplary and may deviate in practice due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes and relative positions according to actual needs.

[0073] Traditional depth estimation models are not well optimized for 3D detection, resulting in serious distortion (deformation, trailing) of the shape of existing pseudo-point clouds and position deviation. The present invention provides a 3D object detection method based on plane constraint and position constraint, by adding plane constraint (normal vector gradient and gradient constraint) and position constraint (end-to-end training + pseudo-point cloud box label + two-stage training strategy) in the training of the pseudo-point cloud detection method; thus, it can significantly improve the shape of the pseudo-point cloud, reduce the position deviation of the pseudo-point cloud, and ultimately improve the 3D detection accuracy, thereby being able to improve the two problems of shape distortion of the pseudo-point cloud and prediction position error caused by depth estimation error in the pseudo-point cloud detection method.

[0074] To address the problem of shape distortion of pseudo-point clouds, it is proposed to additionally add normal vector and gradient constraints during the training of the depth estimation model, thereby improving the shape characteristics of the pseudo-point cloud and reducing the trailing phenomenon. For the flat area of the object, normal vector constraint is selected to enhance the shape structure characteristics of the object on the pseudo-point cloud; for the edge area of the object, gradient constraint is selected to highlight the edge of the object in the depth map and reduce the trailing phenomenon in the pseudo-point cloud. Since the valid values in the actual depth label are very sparse and irregularly distributed, the traditional gradient operator is not applicable to this kind of data, so a random sampling method is used to calculate the normal vector and gradient constraints.

[0075] To address the problem of prediction position error of pseudo-point clouds, a 3D detection network is added behind the depth estimation model for end-to-end joint training. The depth estimation model for 3D detection is optimized by additionally adding the information of 3D detection boxes. And it is proposed to better conduct end-to-end training through a two-stage training strategy combined with pseudo-point cloud box labels (centered on the pseudo-point cloud). In the first stage, the depth estimation model is frozen, and only the 3D detection network is trained using the pseudo-point cloud box label to correctly identify the objects in the pseudo-point cloud. In the second stage, the 3D detection network is frozen, and only the depth estimation model is trained using the GT label.

[0076] Please refer to Figure 1 , a 3D object detection method based on plane constraint and position constraint of the present invention includes the following steps:

[0077] S1, Monocular depth estimation

[0078] S101, Input the RGB image of the data as Figure 2 (a), use the depth estimation model ForeSeE to train and predict the depth, and the predicted depth map is as Figure 2 (b);

[0079] The pixel-level loss function used during training is the weighted cross-entropy loss (wcel). The losses of the foreground and background are calculated separately using the 2D detection box, and the final pixel-level loss function is obtained by weighted summation. The calculation formula is as follows:

[0080]

[0081] Among them, wcel_ fg and wcel_ bg are the pixel-level cross-entropy loss functions of the foreground and background respectively; α is the weight of the foreground loss. In the experiment, α is set to 0.7, then the depth estimation model will focus on foreground prediction.

[0082] S102. Enhance the structural features of the pseudo-point cloud object by additionally adding normal vectors in the depth estimation stage;

[0083] In addition to the single pixel-level constraint loss wcel , it is also necessary to enhance the shape of the pseudo-point cloud object by adding normal vector constraints to reduce the phenomenon of pseudo-point cloud distortion; as Figure 3 (a), randomly sample points inside the object 2D box to calculate the virtual normal vector features.

[0084] The calculation formula of the normal vector constraint is as follows:

[0085]

[0086] Among them, N is the number of effective point groups inside the object 2D box, and each group consists of 3 randomly selected points. n is the normal vector of the plane formed by these 3 points.

[0087] S103. Improve the phenomenon of pseudo-point cloud object edge trailing by adding gradient constraints.

[0088] Enhance the shape of the pseudo-point cloud object by adding normal vector constraints to reduce the phenomenon of pseudo-point cloud distortion.

[0089] Such as Figure 3 (b), calculate the gradient by randomly sampling points inside and on the edge of the object 2D box. The calculation formula of the gradient constraint is as follows:

[0090]

[0091] Among them, N is the number of effective point pairs inside the object 2D box, and each pair consists of 2 randomly selected points. gu and gv are the horizontal and vertical image gradient differences calculated from these 2 points, and the calculation formula is as follows:

[0092]

[0093] Among them, u and v are the pixel coordinates in the horizontal and vertical directions, and z is the depth value obtained through depth estimation.

[0094] S2. Generation of pseudo-point cloud

[0095] S201. Use an instance segmentation network to accurately distinguish foreground objects in the RGB image;

[0096] Input the RGB image as Figure 2 (a). Use an instance segmentation network to accurately distinguish the foreground and background of the RGB image, and obtain the instance segmentation mask of the foreground. The schematic diagram of the instance segmentation result is as Figure 4 .

[0097] S202. Use the foreground instance segmentation mask to extract the corresponding foreground depth;

[0098] Use the foreground instance segmentation mask to extract the corresponding foreground part from the panoramic depth prediction map obtained in step S201. The subsequent 3D detection model only needs to use the depth corresponding to the foreground object.

[0099] S203. Use the following conversion formula to convert the foreground depth into foreground point cloud:

[0100]

[0101] where u and v are pixel coordinates, f x , f y is the camera focal length, c x , c y is the pixel coordinate of the image center point, and z is the predicted depth.

[0102] S3. Generation of pseudo point cloud box labels

[0103] The pseudo point cloud box is a detection box centered on the pseudo point cloud and having the same size as the GT detection box. As Figure 5 shown, when there is a deviation in the GT box, the pseudo point cloud box can correctly represent the position of the pseudo point cloud. The pseudo point cloud box can correctly represent the correct position of the pseudo point cloud. Using the pseudo point cloud box as a label can reduce interference in training.

[0104] The calculation formula for the center position of the pseudo point cloud box is as follows:

[0105]

[0106] where Num GT represents the number of point clouds of a pseudo point cloud object in the GT box, and Num all represents the total number of point clouds of a pseudo point cloud object.

[0107] When the ratio value is large, the deviation between the object and the GT box is small, and at this time, the center point of the GT box can be directly used as the center point of the pseudo point cloud box. When the ratio value is small, the deviation between the object and the GT box is large, and at this time, the average value of the positions of all pseudo point clouds of the object is used as the center point of the pseudo point cloud box.

[0108] S4. Two-stage training strategy

[0109] S401. The first stage of the two-stage training strategy;

[0110] As Figure 6 shown in (a), in the training of the first stage, the parameters of the depth estimation model are frozen, and only the 3D detection network is trained. In the first stage, the pseudo point cloud boxes are used as labels to train the 3D detection network F-PointNet. The pseudo point cloud boxes can enable the 3D detector to learn the correct positions of the pseudo point clouds. In the first stage, the loss function of F-PointNet is directly adopted, and only the GT labels in it are replaced with the pseudo point cloud box labels.

[0111] loss1 det = PointNetLoss(Box pred , Pseudolabel)

[0112] S402. The second stage of the two-stage training strategy;

[0113] As Figure 6 shown in (b), in the second stage, the parameters of the 3D detection network F-PointNet are frozen, and only the depth estimation model is trained, and the GT detection boxes are used as the labels of the 3D detector to train the depth estimation network. Since the parameters of the 3D detector are frozen, the detector can only predict the real positions of the pseudo point clouds. When there is an error between the positions of the pseudo point clouds and the GT detection boxes, the value of the loss function of the 3D detection will become larger. Subsequently, in the error backpropagation process, it affects the parameter update of the depth estimation network, making the pseudo point clouds closer to the GT detection boxes.

[0114] The overall loss function of the second stage is as follows:

[0115] loss all = 1*oss wcel + 2*oss normal + 3*oss gradient + 4*loss2 del

[0116] Among them, λ1 = 6, λ2 = λ3 = 1, λ4 = 0.001.

[0117] loss2 det is defined as follows:

[0118] loss2 det = PointNetLoss(Box pred , GTlabel)

[0119] S403. Alternately train these two stages to ensure that the 3D detection network F-PointNet can always correctly predict the positions of the pseudo point clouds.

[0120] Through the method of alternating training of the pseudo-point cloud box with the 3D detection model and the depth estimation model, the purpose of optimizing depth estimation by means of 3D detection is achieved.

[0121] In another embodiment of the present invention, a 3D object detection system based on plane constraint and position constraint is provided. This system can be used to implement the above-mentioned 3D object detection method based on plane constraint and position constraint. Specifically, the 3D object detection system based on plane constraint and position constraint includes an estimation module, a conversion module, a label module, and a prediction module.

[0122] Among them, the estimation module inputs the RGB image data and uses the depth estimation model ForeSeE to train and obtain a depth map;

[0123] The conversion module uses the instance segmentation mask to segment the depth map obtained by the estimation module and converts the obtained foreground part into a foreground point cloud;

[0124] The label module takes the foreground point cloud obtained by the conversion module as the center and generates a pseudo-point cloud box label with the same size as the GT detection box;

[0125] The prediction module freezes the parameters of the depth estimation model, trains the 3D detection network, uses the pseudo-point cloud label obtained by the label module as the training label to complete the first stage of training, freezes the parameters of the 3D detection network F-PointNet, trains the depth estimation model, and uses the GT detection box as the label of the 3D detector to train the depth estimation network to complete the second stage of training; alternately train the first stage and the second stage so that the 3D detection network F-PointNet can correctly predict the pseudo-point cloud position at all times.

[0126] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiments of the present invention can be used for the operation of the 3D object detection method based on plane constraint and position constraint, including:

[0127] Input the RGB image data, and use the depth estimation model ForeSeE to train to obtain a depth map; use the instance segmentation mask to segment the obtained depth map, and convert the obtained foreground part into foreground point cloud; take the obtained foreground point cloud as the center, and generate a pseudo point cloud box label with the same size as the GT detection box; freeze the parameters of the depth estimation model, train the 3D detection network, and use the pseudo point cloud label as the training label to complete the first stage of training. Freeze the parameters of the 3D detection network F-PointNet, train the depth estimation model, and use the GT detection box as the label of the 3D detector to train the depth estimation network to complete the second stage of training; alternately train the first stage and the second stage to make the 3D detection network F-PointNet correctly predict the position of the pseudo point cloud at all times.

[0128] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. Moreover, one or more instructions suitable for being loaded and executed by a processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory (Non-Volatile Memory), such as at least one disk memory.

[0129] One or more instructions stored in the computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the 3D object detection method based on plane constraint and position constraint in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor to perform the following steps:

[0130] Input the RGB image data, and use the depth estimation model ForeSeE to train to obtain a depth map; use the instance segmentation mask to segment the obtained depth map, and convert the obtained foreground part into foreground point cloud; generate a pseudo-point cloud box label with the obtained foreground point cloud as the center and the same size as the GT detection box; freeze the parameters of the depth estimation model, train the 3D detection network, use the pseudo-point cloud label as the training label to complete the first-stage training, freeze the parameters of the 3D detection network F-PointNet, train the depth estimation model, use the GT detection box as the label of the 3D detector to train the depth estimation network, and complete the second-stage training; alternately train the first stage and the second stage so that the 3D detection network F-PointNet can correctly predict the position of the pseudo-point cloud at all times.

[0131] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0132] The advantages of the present invention are described below by comparing the result graphs.

[0133] The main functions of the present invention are reflected in two aspects: the main advantage is that it can enhance the structural features of the pseudo-point cloud and improve the prediction accuracy of the pseudo-point cloud position, thereby improving the 3D object detection performance.

[0134] Please refer to Figure 8 , after adding the normal vector, the shape features of the point cloud of the nearby vehicle become more obvious and are very close to the features of the real point cloud.

[0135] Please refer to Figure 9 , before adding the gradient constraint. The contour of the object in the depth map is not clear; the edge transition between the object and the ground and the background is blurred. After adding the gradient constraint, the contour of the object on the depth map becomes more prominent; the edge transition with the ground and the background is more obvious; the trailing point cloud at the edge of the object disappears in the BEV point cloud map.

[0136] Please refer to Figure 10 , after successively adding the normal vector and the gradient constraint, the misdetection boxes in the 3D detection results gradually decrease.

[0137] Figure 7 The results of show that the detection results of the present invention exceed some recent monocular detection methods under both the 〖AP〗_3D moderate and hard models.

[0138] Please refer to Figure 11 , after adopting two-stage training, the huge deviation between the pseudo-point cloud position and the actual position (GT box) is significantly reduced, indicating that two-stage training can more effectively optimize the depth estimation model.

[0139] In summary, for the 3D object detection method and system based on plane constraint and position constraint of the present invention, the prediction effect of the depth estimation model is significantly improved and the contour is more prominent; thereby enhancing the structural features of the point cloud, reducing the point cloud trailing phenomenon and improving the point cloud position prediction accuracy; ultimately improving the 3D object detection performance.

[0140] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of distinguishing each other and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be repeated here.

[0141] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0142] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0143] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal and method can be implemented in other ways. For example, the device / terminal embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0144] The unit described as the separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0145] In addition, each functional unit in various embodiments of the present invention may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0146] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0147] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1means for the functions specified in one or more boxes.

[0148] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 means for the functions specified in one or more boxes.

[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 means for the functions specified in one or more boxes.

[0150] The above content is only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.

Claims

1. A 3D object detection method based on plane constraint and position constraint, characterized in that Including the following steps: S1. Input the RGB image data, obtain the depth map by training with the depth estimation model ForeSeE, and use the depth estimation model ForeSeE to train and predict the loss function of the depth as follows: Among them, , are the pixel-level cross-entropy loss functions for the foreground and background, respectively; is the weight of the foreground loss; Normal vector constraint The calculation is as follows: Among them, is the number of valid points within the 2D bounding box of the object, is the predicted normal vector, is the ground-truth normal vector; S2. Use the instance segmentation mask to segment the depth map obtained in step S1, and convert the obtained foreground part into foreground point cloud; S3. Taking the foreground point cloud obtained in step S2 as the center, generate a pseudo point cloud box label with the same size as the GT detection box; S4. Freeze the parameters of the depth estimation model, train the 3D detection network, use the pseudo point cloud label obtained in step S3 as the training label to complete the first-stage training, freeze the parameters of the 3D detection network F-PointNet, train the depth estimation model, and use the GT detection box as the label of the 3D detector to train the depth estimation network to complete the second-stage training; alternately train the first stage and the second stage so that the 3D detection network F-PointNet can correctly predict the position of the pseudo point cloud at all times; The loss function of the first stage is as follows: Among them, to detect the original loss function of the network F-PointNet, to predict the 3D bounding box, is the pseudo-point cloud label; The loss function of the second stage is as follows: wherein, = 6, = = 1, = 0.001, is the loss function, is the normal vector constraint, is the gradient constraint, is the 3D detection function.

2. The 3D object detection method based on plane constraint and position constraint according to claim 1, wherein, In step S1, the gradient constraint is calculated as follows: Among them, is the number of valid points within the 2D bounding box of the object, is the predicted horizontal gradient difference, is the true horizontal gradient difference, is the predicted vertical gradient difference, is the true vertical gradient difference.

3. The 3D object detection method based on plane constraint and position constraint according to claim 1, wherein In step S2, use the following conversion formula to convert the foreground depth into foreground point cloud: Among them, , is the pixel coordinate, , is the camera focal length, , is the pixel coordinate of the image center point, is the predicted depth.

4. The 3D object detection method based on plane constraint and position constraint according to claim 1, wherein In step S3, the pseudo-point cloud box is a detection box centered on the pseudo-point cloud and having the same size as the GT detection box. When there is a deviation in the GT box, the pseudo-point cloud box represents the position of the pseudo-point cloud. When the value is greater than 0.25, the center point of the GT box is used as the center point of the pseudo-point cloud box; when the value is less than 0.25, the average value of all the pseudo-point cloud positions of the object is used as the center point of the pseudo-point cloud box.

5. The 3D object detection method based on planar constraint and position constraint according to claim 4, wherein Pseudo point cloud frame center position The calculation is as follows: Among them, is the number of point clouds of a pseudo point cloud object in GT the box, is the total number of point clouds of a pseudo point cloud object, is the center position of the GT box, is the center position of the foreground point cloud, is ratio the threshold of.

6. A 3D object detection system based on plane constraint and position constraint, characterized in that, Including: An estimation module, which takes an input RGB image, obtains a depth map by training with a depth estimation model ForeSeE, and uses the depth estimation model ForeSeE to train and predict the loss function of depth is as follows: Among them, , are the pixel-level cross-entropy loss functions for the foreground and background respectively; is the weight of the foreground loss; Normal vector constraint The calculation is as follows: Among them, is the number of valid point groups within the 2D bounding box of the object, is the predicted normal vector, is the ground truth normal vector; A conversion module that uses the instance segmentation mask to segment the depth map obtained by the estimation module and converts the obtained foreground part into foreground point cloud; A label module that generates a pseudo point cloud box label with the same size as the GT detection box with the foreground point cloud obtained by the conversion module as the center; A prediction module that freezes the parameters of the depth estimation model, trains the 3D detection network, uses the pseudo point cloud label obtained by the label module as the training label to complete the first-stage training, freezes the parameters of the 3D detection network F-PointNet, trains the depth estimation model, and uses the GT detection box as the label of the 3D detector to train the depth estimation network to complete the second-stage training; alternately train the first stage and the second stage so that the 3D detection network F-PointNet can correctly predict the position of the pseudo point cloud at all times; The loss function of the first stage is as follows: Among them, is to detect the original loss function of the network F-PointNet, is to predict the 3D bounding box, is the pseudo-point cloud label; The loss function of the second stage is as follows: Among them, = 6, = = 1, = 0.001, is the loss function, is the normal vector constraint, is the gradient constraint, is the 3D detection function.

Citation Information

Patent Citations

  • Monocular image-oriented three-dimensional object detection method based on three-dimensional reconstruction

    CN110689008A

  • Mechanical arm grabbing control method of deep reinforcement learning DDPG algorithm based on visual information

    CN115464659A