A Monocular 3D Vehicle Detection Method Based on Dense Association
By constructing a monocular 3D vehicle detection method with dense association, the network is trained using Gaussian hybrid model and negative log likelihood loss function, 2D-3D association is generated, and through bottom-up instance segmentation and clustering, the problems of high data labeling cost and unreliable association of occlusion areas in the prior art are solved, and accurate vehicle identification and positioning are achieved.
Patent Information
- Application Number
- CN202111405543.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-11-24
AI Technical Summary
The existing monocular 3D vehicle detection methods have problems such as high data labeling cost, inability to fully combine target detection and instance segmentation, and inability to adaptively remove unreliable associations of occlusion areas, resulting in a decrease in positioning accuracy.
By constructing a monocular 3D vehicle detection method with dense associations, the network is trained using Gaussian hybrid model and negative log likelihood loss function, 2D-3D associations are generated, and the vehicle's position, angle and size information is obtained through bottom-up instance segmentation and clustering, avoiding the influence of additional annotation and occlusion areas.
Accurate vehicle identification and positioning is achieved, data labeling costs are reduced, positioning accuracy is improved, unreliable correlation problems of occlusion areas are solved, and the generalization ability of the method is enhanced.
Smart Images

Figure CN114119749B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and intelligent driving vehicles, and particularly to a monocular 3D vehicle detection method based on dense association. Background Art
[0002] Among the numerous sensors applied to intelligent vehicles, the camera, as a visual sensor, has the advantages of high resolution, low cost, and convenient deployment. Using the RGB image data obtained by the camera for 3D vehicle detection can replace the high-cost solution based on lidar in scenarios with slightly lower accuracy requirements. Using a single image for 3D vehicle detection, that is, monocular 3D vehicle detection, is one of the core technologies and has a wide range of demands in the field of intelligent vehicles.
[0003] The difficulty of monocular 3D vehicle detection lies in estimating the distance of the vehicle based solely on a 2D image. Currently, there are two main types of monocular 3D vehicle detection methods. One is to directly estimate the distance of the vehicle through a deep network, and the other is to construct a 2D-3D association and indirectly estimate the distance of the vehicle through geometric reasoning. Among them, the former often has problems such as relying on specific scenarios and camera internal parameters, and has poor generalization performance. The latter is more robust to data migration under different scenarios and camera internal parameters and has better practicability. However, existing methods still have some problems, which are mainly reflected as follows:
[0004] First, some methods require additional manual annotations, such as key points, vehicle 3D models, etc., when training the model, which increases the cost of data annotation;
[0005] Second, existing methods generally require separate object detection or instance segmentation modules. First, complete the detection, and then generate a 2D-3D association and perform geometric reasoning, without fully combining the two;
[0006] Third, existing methods often form 2D-3D associations using a fixed number of key points or regional grids, and cannot adaptively remove unreliable associations in the occluded areas of the vehicle, which easily reduces the positioning accuracy of some occluded vehicles. Summary of the Invention
[0007] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a monocular 3D vehicle detection method based on dense association, which can accurately identify and locate vehicles in a traffic scene.
[0008] The purpose of the present invention can be achieved through the following technical solutions:
[0009] The present invention provides a monocular 3D vehicle detection method based on dense association for an autonomous driving vehicle to identify and locate vehicles in a traffic scene, including the following steps:
[0010] S1: Collect a single front view image through an in-vehicle camera;
[0011] S2: Calculate the actual 2D coordinates of each pixel point in the front view image in the camera coordinate system;
[0012] S3: Process the front view image to sequentially obtain multi-scale features, a high-resolution feature map, and the probability distribution of the 3D coordinate vectors of each pixel point on the high-resolution feature map described by a Gaussian mixture model. Process the probability distribution of the 3D coordinate vectors of each pixel point into the probability distribution of the dynamic 3D coordinates of each pixel point in each local coordinate system. During training, project the 3D distribution into the probability distribution of 2D coordinates in the camera coordinate system, and train the network using the negative log-likelihood loss function to minimize the ghosting error, minimizing the negative log-likelihood of the actual 2D coordinates of each pixel point under the 2D coordinate probability distribution, thereby enabling each pixel point to generate a set of 2D-3D associations;
[0013] S4: Set up a first network branch. According to the Gaussian mixture model, determine the unique target vehicle corresponding to each pixel point, and cluster the center positions of the unique target vehicles corresponding to each pixel point to achieve bottom-up instance segmentation, thereby dividing the 2D-3D associations constructed in S3 into dense 2D-3D associations of each vehicle;
[0014] S5: Construct a PnP problem from the dense 2D-3D associations and solve it to obtain the position and angle of the target vehicle;
[0015] S6: According to the instance segmentation result of S4, set up a second network branch to obtain the size of the unique target vehicle corresponding to each pixel point, and combine the position and angle of the target vehicle obtained in S5 to obtain a vehicle 3D detection box containing position, angle, and size information.
[0016] Preferably, the target local coordinate system is a coordinate system established with the center point of the bottom surface of each target vehicle as the origin, the front of each target vehicle as the x-axis, the lower part of each target vehicle as the y-axis, and the left side of each target vehicle as the z-axis.
[0017] Preferably, S3 includes the following steps:
[0018] S3.1: Process the front view image through a residual network and a feature pyramid network in sequence to obtain the multi-scale features of the front view image;
[0019] S3.2: Perform deformable convolution, bilinear interpolation resampling, and splicing processing on the multi-scale features in sequence to obtain a multi-scale fused high-resolution feature map;
[0020] S3.3: Output the 3D coordinate vectors of each pixel on the high-resolution feature map through a branch network composed of convolutional layers, and use a Gaussian mixture model to describe the probability distribution of the 3D coordinate vectors of each pixel;
[0021] S3.4: Extract the regional features of each target vehicle from the multi-scale features, obtain the probability distribution of the dynamic 3D coordinates of each pixel in each local coordinate system according to the Gaussian mixture model in S3.3, convert the probability distribution of the dynamic 3D coordinates of each pixel in each local coordinate system into the probability distribution of 2D coordinates in the camera coordinate system, and train the network using the negative log-likelihood loss function to minimize the ghosting error, that is, minimize the negative log-likelihood of the actual 2D coordinates of each pixel under the 2D coordinate probability distribution, thereby enabling each pixel to generate a set of 2D-3D associations.
[0022] Preferably, using a Gaussian mixture model to describe the probability distribution of the 3D coordinate vectors of each pixel is specifically:
[0023]
[0024] In the formula, S is the number of pre-set Gaussian mixture models, φ i is the mixing weight of the i-th Gaussian mixture model, ∑ i is the covariance matrix of the i-th Gaussian mixture model, μ i is the mean of the i-th Gaussian mixture model, φ i , ∑ i , μ i are all variables output by the network, is the probability density estimate of x 3D , x 3D is a set of coordinate vectors in the target local coordinate system.
[0025] Preferably, the expression for projecting the probability distribution of the dynamic 3D coordinates of each pixel in each local coordinate system into the probability distribution of 2D coordinates in the camera coordinate system is:
[0026] [x cam y cam z cam T =Rx 3D +t
[0027]
[0028] In the formula, R and t are respectively the rotation matrix and displacement vector for converting the local coordinate system to the camera coordinate system, and the intermediate variables x cam , y cam , z cam are respectively the 3D coordinates in the camera coordinate system, x 3D is a set of coordinate vectors in the target local coordinate system, x 2D is a set of coordinate vectors in the transformed camera coordinate system.
[0029] Preferably, the formula for training the network using the negative log-likelihood loss function is:
[0030]
[0031]
[0032] In the formula, is the weight normalization parameter, satisfying to dynamically balance the weights of the loss function, is the actual 2D coordinate vector of each pixel point, is the true value of the 2D coordinate is the negative log-likelihood under the probability distribution density function of the transformed 2D coordinates, where is the covariance matrix of the i-th 2D Gaussian mixture model, is the mean of the i-th 2D Gaussian mixture model, φ i is the static mixing weight of the i-th Gaussian mixture model, ψ i is the dynamic mixing weight of the i-th Gaussian mixture model.
[0033] Preferably, the S4 includes the following steps:
[0034] S4.1: Divide the 2D space region occupied by each target vehicle in the front view image on the front view image, and set the first network branch to regress the offset of the 2D position of the geometric center point of the target vehicle to which each pixel point belongs relative to the pixel point position in the high-resolution feature map, so that each pixel point can locate the center position of the target vehicle to which it belongs;
[0035] S4.2: Determine the target vehicle corresponding to each Gaussian mixture model through the minimum reprojection error criterion, and then obtain the actual center position of the target vehicle corresponding to the Gaussian mixture model
[0036] S4.3: Calculate the offset between the center position of the target vehicle to which each pixel point belongs and the actual center position of the target vehicle corresponding to the Gaussian mixture model, and find the minimum offset, and then determine the unique target vehicle corresponding to each pixel point; and train the first network branch through the smooth L1 loss function, so that the center position x of the unique target vehicle corresponding to each pixel point ctr and the actual center position of the unique target vehicle has the minimum offset;
[0037] S4.4: Set up a foreground network branch to segment foreground pixels on the high-resolution feature map, train this network branch using the cross-entropy loss function, and obtain the foreground network branch;
[0038] S4.5: According to the foreground network branch and the Gaussian mixture model established in S3, cluster the central positions of the unique target vehicles corresponding to each pixel point to achieve bottom-up instance segmentation, and further divide the 2D-3D associations constructed in S3.4 into dense 2D-3D associations of each vehicle.
[0039] Preferably, the formula for obtaining the position and angle of the target vehicle in S5 is:
[0040]
[0041] In the formula, β and t are respectively the yaw angle and displacement offset of the target vehicle after initialization, μ 2D , are respectively the parameters of the 2D Gaussian mixture model and are functions of β and t, β * , t * are respectively the position and angle of the target vehicle optimized by β and t.
[0042] Preferably, before executing S5, use the EPnP algorithm to initialize the yaw angle and displacement offset of the target vehicle.
[0043] Preferably, S6 includes the following steps:
[0044] S6.1: Set up a second network branch, find the size of the target vehicle corresponding to each pixel point according to the same rule as S4.2, train the second network branch using the smooth L1 loss function, and output the size of the target vehicle corresponding to each pixel point of the high-resolution feature map;
[0045] S6.2: According to the instance segmentation result of S4.4, determine the size of the unique target vehicle corresponding to each pixel point;
[0046] S6.3: According to the position and angle of the target vehicle in S5 and the size of the target vehicle obtained in S6.2, obtain a vehicle 3D detection frame containing position, angle, and size information.
[0047] Compared with the prior art, the present invention has the following advantages:
[0048] 1. The monocular 3D vehicle detection method based on dense association designed by the present invention does not need to adopt the 3D geometric information of the vehicle including key points and 3D models. Instead, by constructing 2D-3D associations and training the network by minimizing the reprojection error, it enables the prediction of the 3D coordinates corresponding to each pixel point, avoiding the problem in the prior art that additional manual annotation is required when training the model, increasing the data annotation cost.
[0049] 2. The present invention solves the problem in the prior art that after detection is completed first and then 2D-3D associations are generated for geometric reasoning, the two cannot be fully combined, by means of predicting 2D-3D association points and clustering them to obtain object-level information.
[0050] 3. The present invention divides 2D-3D association points belonging to different target vehicles through clustering. The number of association points finally obtained for each target vehicle is determined according to the actual situation, and each pixel point cannot belong to two target vehicles at the same time. Therefore, it can solve the problem in the prior art that due to the inability to adaptively remove unreliable associations in the occluded areas of vehicles, the positioning accuracy of some occluded vehicles decreases. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a schematic flow chart of a monocular 3D vehicle detection method based on dense association according to this embodiment;
[0052] Figure 2 is a schematic diagram of a specific embodiment of the network structure used in this embodiment;
[0053] Figure 3 is a schematic diagram of a specific embodiment of the definition of the target local coordinate system in this embodiment;
[0054] Figure 4 is a schematic diagram of the relationship between the camera coordinate system and the target local coordinate system in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0056] This embodiment provides a monocular 3D vehicle detection method based on dense association for an autonomous vehicle to identify and locate vehicles in a traffic scene, including the following steps:
[0057] S1: Collect a single front view image through an in-vehicle camera;
[0058] S2: Calculate the actual 2D coordinates of each pixel point in the front view image in the camera coordinate system through the camera intrinsic matrix;
[0059]
[0060] Wherein, is the actual 2D coordinate of each pixel point in the camera coordinate system, K is the camera internal parameter matrix, and (u, v) is the pixel index coordinate, that is, the pixel coordinate of the u-th column and the v-th row in the front view image.
[0061] S3: Process the front view image, sequentially obtain multi-scale features, high-resolution feature maps, and the probability distribution of the 3D coordinate vectors of each pixel point on the high-resolution feature map described by the Gaussian mixture model. Process the probability distribution of the 3D coordinate vectors of each pixel point into the probability distribution of the dynamic 3D coordinates of each pixel point in each local coordinate system, and project it into the probability distribution of the 2D coordinates in the camera coordinate system during training. Then, use the negative log-likelihood loss function to train the network to minimize the ghosting error, minimize the negative log-likelihood of the actual 2D coordinates of each pixel point under the 2D coordinate probability distribution, and thus generate a set of 2D-3D associations for each pixel point.
[0062] The target local coordinate system is a coordinate system established with the autonomous driving vehicle as the origin. Refer to Figure 3 As shown, the target local coordinate system is a coordinate system established with the center point of the bottom surface of each target vehicle as the origin, the front of each target vehicle as the x-axis, the bottom of each target vehicle as the y-axis, and the left of each target vehicle as the z-axis.
[0063] The variables output by the Mixture Density Networks (MDN) are the parameters of the Gaussian Mixture Model (Gaussian Mixture Model), which include the means, covariances, and their mixing weights of n Gaussian mixture models in total.
[0064] S3.1: Process the front view image sequentially through the residual network and the feature pyramid network to obtain the multi-scale features of the front view image.
[0065] Use the residual network as the backbone network to extract the image features of the front view image, and obtain multi-scale features by passing the image features through the feature pyramid network; the resolutions of the multi-scale features are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image, and the channel dimension is 256.
[0066] S3.2: Perform deformable convolution, bilinear interpolation resampling, and splicing processing on the multi-scale features in sequence to obtain a multi-scale fused high-resolution feature map.
[0067] Process the multi-scale features through 3x3 deformable convolution, resample each level of features to the original Figure 1 / 4 size by bilinear interpolation, and splice them in the channel direction to obtain a multi-scale fused high-resolution feature map; the channel dimension is 512.
[0068] S3.3: Output the 3D coordinate vectors of each pixel on the high-resolution feature map through a branch network composed of convolutional layers, and use a Gaussian mixture model to describe the probability distribution of the 3D coordinate vectors of each pixel.
[0069] Use a Gaussian mixture model to describe the probability distribution of x 3D :
[0070]
[0071] In the formula, S is the number of pre-set Gaussian mixture models, and φ i is the mixing weight of the i-th Gaussian mixture model, ∑ i is the covariance matrix of the i-th Gaussian mixture model, μ i is the mean of the i-th Gaussian mixture model, φ i , ∑ i , μ i are all variables output by the network. is the probability distribution of x 3D , and x 3D is a set of coordinate vectors [x, y, z] of each pixel in the target local coordinate system. T .
[0072] Specifically, x 3D is completely learned by the network and does not necessarily have strong physical meaning. Ideally, the x 3D predicted by the network should satisfy the projection constraint, that is, the x 2D obtained in step S3.4.2 should be consistent with its corresponding actual 2D coordinates in the camera coordinate system. What the network actually predicts is not a single x 3D , but the probability distribution of x 3D , which can be described by the three parameters φ i , ∑ i , μ i . The branch network is composed of convolutional layers and maps the high-resolution feature map to φ i , ∑ i , μ i .
[0073] As can be seen from the above formula, the branch network outputs S groups of φ i , ∑ i , μ i . Among them, φ i needs to ensure that the sum is 1, so a softmax layer is required at the output end; the matrix ∑ i needs to ensure symmetry and positive definiteness. Therefore, LDL decomposition needs to be performed on the matrix:
[0074] ∑ = LDL T
[0075] D = exp diag[d1 d2 d3]
[0076]
[0077] Where D is the symmetric positive definite matrix after the LDL decomposition of matrix ∑, L is the unit lower triangular matrix, d1, d2, and d3 are the three parameters on the diagonal of matrix D respectively, and l1, l2, and l3 are the parameters in matrix L respectively.
[0078] After the LDL decomposition, it can be ensured that ∑ is symmetric positive definite. At this time, the network only needs to output six parameters d1, d2, d3, l1, l2, and l3. Therefore, the output layer dimension of the covariance is 6.
[0079] S3.4: Extract the regional features of each target vehicle from the multi-scale features, obtain the probability distribution of the dynamic 3D coordinates of each pixel point in each local coordinate system according to the Gaussian mixture model in S3.3, convert the probability distribution of the dynamic 3D coordinates of each pixel point in each local coordinate system into the probability distribution of 2D coordinates in the camera coordinate system, and train the network using the negative log-likelihood loss function to minimize the ghosting error, that is, to minimize the negative log-likelihood of the actual 2D coordinates of each pixel point under the 2D coordinate probability distribution, thereby enabling each pixel point to generate a set of 2D-3D associations.
[0080] S3.4.1: Add a region convolutional network (R-CNN) as an auxiliary branch to extract the regional features of each target vehicle from the multi-scale features and output the pixel boxes of the target vehicles in the front view image, that is, the target boxes. If there is overlap between the target boxes, the pixels in the overlapping region have a weight for each target box, and this weight is the dynamic mixing weight ψ i . Output the dynamic mixing weight ψ of each pixel point in the region of each target vehicle through this branch i , and then obtain the probability distribution of the dynamic 3D coordinates of each pixel point in each local coordinate system according to the Gaussian mixture model in S3.3.
[0081]
[0082] Where S is the number of pre-set Gaussian mixture models, φ i is the mixing weight of the i-th Gaussian mixture model, ∑ i is the covariance matrix of the i-th Gaussian mixture model, μ i is the mean of the i-th Gaussian mixture model, φ i , ∑ i , μ i are all variables output by the network, is the probability density estimate of x 3D , ψi is the dynamic mixing weight of each pixel point in the area of each target vehicle, x 3D is a set of coordinate vectors [x, y, z] of each pixel point in the target local coordinate system T .
[0083] S3.4.2: Project the probability distribution of the dynamic 3D coordinates of each pixel point in each target local coordinate system into the probability distribution of the 2D coordinates in the camera coordinate system.
[0084] [x cam y cam z cam T = Rx 3D + t
[0085]
[0086] In the formula, R and t are respectively the rotation matrix and displacement vector for the conversion from the target local coordinate system to the camera coordinate system. The intermediate variables x cam , y cam , z cam are respectively the 3D coordinates in the camera coordinate system.
[0087] For the Gaussian mixture distribution, the method of local linearization is used to calculate the parameters of the transformed 2D Gaussian mixture model.
[0088] The specific parameter transformation method is as follows: The transformation method of the mean μ i is the same as that of x in the above formula 3D , that is, first perform the pose transformation of Rμ + t to obtain and then normalize it by dividing by the Z-axis coordinate to obtain the mean vector μ 2D of the 2D Gaussian mixture model. The projection transformation of the covariance ∑ 2D of the 2D Gaussian mixture model is as follows:
[0089]
[0090] where [:2, :2] means taking the first two rows and two columns of the 3×3 matrix.
[0091] S3.4.3: Train the network using the negative log-likelihood loss function to minimize the ghosting error, that is, to minimize the negative log-likelihood of the actual 2D coordinates of each pixel point under the 2D coordinate probability distribution, and thus obtain the 2D-3D association.
[0092] The goal of network training is to minimize the reprojection error, that is, to minimize the negative log-likelihood of the actual 2D coordinates of each pixel point under the 2D coordinate probability distribution. Specifically, the network is trained using the negative log-likelihood loss function:
[0093]
[0094]
[0095] wherein, is a weight normalization parameter, satisfying and is used to dynamically balance the weights of the loss function, is the actual 2D coordinate vector of each pixel point, is the true value of the 2D coordinate is the negative log-likelihood under the probability distribution density function of the transformed 2D coordinate, where is the covariance matrix of the i-th 2D Gaussian mixture model, is the mean of the i-th 2D Gaussian mixture model, φ i is the static mixing weight of the i-th Gaussian mixture model, ψ i is the dynamic mixing weight of the i-th Gaussian mixture model.
[0096] S4: Set the first network branch. According to the Gaussian mixture model, determine the unique target vehicle corresponding to each pixel point, and cluster the central positions of the unique target vehicles corresponding to each pixel point to achieve bottom-up instance segmentation, thereby converting the 2D-3D association constructed in S3 into a dense 2D-3D association.
[0097] S4.1: Divide the 2D space region occupied by each target vehicle in the front view image on the front view image, and set the first network branch to regress the offset of the 2D position of the geometric center point of the target vehicle to which each pixel point belongs relative to the pixel point position in the high-resolution feature map, so that each pixel point can locate the central position of the target vehicle to which it belongs.
[0098] Since different Gaussian mixture models of the same pixel point in the Gaussian mixture model in S3.3 may be assigned to different target vehicles, K central offsets need to be output correspondingly for K Gaussian mixture models to distinguish these target vehicles; each pixel point in S4.1 corresponds to several target vehicles.
[0099] S4.2: Determine the target vehicle corresponding to each Gaussian mixture model through the minimum reprojection error criterion, and then obtain the actual central position of the target vehicle corresponding to the Gaussian mixture model
[0100] S4.3: Calculate the offset between the central position of the target vehicle to which each pixel point belongs and the actual central position of the target vehicle corresponding to the Gaussian mixture model, and find the minimum offset, thereby determining the unique target vehicle corresponding to each pixel point; and train the first network branch through the smooth L1 loss function so that the central position x of the unique target vehicle corresponding to each pixel pointctr has the smallest offset from the actual center position of the only target vehicle .
[0101] S4.4: Set up a foreground network branch to segment foreground pixels on the high-resolution feature map, train this network branch through the cross-entropy loss function, and obtain the foreground network branch.
[0102] As an alternative implementation, the method for obtaining the target value of the cross-entropy loss function includes: semantic segmentation annotation of the image, and using the vehicle 2D box as a rough foreground label;
[0103] S4.5: According to the foreground network branch and the Gaussian mixture model established in S3.3, cluster the center positions of the only target vehicle corresponding to each pixel point to achieve bottom-up instance segmentation, and further convert the 2D-3D association constructed in S3.4 into a dense 2D-3D association.
[0104] Specifically, first select all foreground pixel points through the foreground network branch, and take the Gaussian mixture model parameters μ, ∑ with the largest mixing weight φ in S3.3, and the center position x of the only target vehicle corresponding to each pixel point i , and the formula is: ctr ,
[0105]
[0106] Cluster the center positions x of the only target vehicle corresponding to each pixel point ctr to achieve bottom-up instance segmentation, and further divide the 2D-3D association constructed in S3.4 into dense 2D-3D associations of each vehicle.
[0107] As an alternative implementation, use the DBSCAN algorithm to cluster the center points of all foreground pixel points.
[0108] S5: Construct a PnP problem from the dense 2D-3D association and solve it to obtain the position and angle of the target vehicle;
[0109]
[0110] In the formula, β and t are respectively the yaw angle and displacement offset of the target vehicle after initialization, and optimize and solve according to the above formula, μ 2D , are respectively the parameters of the 2D Gaussian mixture model, and are functions of β and t, β * , t * are respectively the position and angle of the optimized target vehicle. Since x 2D uses R(β) and t during pose transformation, therefore, x2D It is a function of β and t.
[0111] This optimization problem is to find the vehicle angle and position with the minimum reprojection error under the Mahalanobis distance metric, so as to achieve 3D positioning of the vehicle.
[0112] Before performing S5, the yaw angle β and displacement offset t of the target vehicle are initialized using the EPnP algorithm, and then the Levenberg-Marquardt algorithm is used to solve the non-linear least squares problem described by the above formula to obtain the optimal solutions β * , t * .
[0113] S6: According to the instance segmentation result of S4, set the second network branch, obtain the size of the unique target vehicle corresponding to each pixel point, and combine the position and angle of the target vehicle obtained in S5 to obtain a 3D detection box of the vehicle containing position, angle, and size information.
[0114] S6.1: Set the second network branch, find the size of the target vehicle corresponding to each pixel point according to the same rules as S4.2, train the second network branch through the smooth L1 loss function, and output the size of the target vehicle corresponding to each pixel point of the high-resolution feature map;
[0115] Specifically, each pixel point corresponds to several target vehicles and the sizes of the target vehicles.
[0116] S6.2: Determine the size of the unique target vehicle corresponding to each pixel point according to the instance segmentation result of S4.4.
[0117] S6.3: According to the position and angle of the target vehicle in S5 and the size of the target vehicle obtained in S6.2, obtain a 3D detection box of the vehicle containing position, angle, and size information.
[0118] The above description of the embodiments is for the convenience of those of ordinary skill in the art to understand and use the invention. Obviously, those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made by those skilled in the art without departing from the scope of the present invention according to the disclosure of the present invention should be within the protection scope of the present invention.
Claims
1. A monocular 3D vehicle detection method based on dense association, which is used for an autonomous vehicle to identify and locate vehicles in a traffic scene, and is characterized in that, Including the following steps: S1: Collect a single forward-looking image through an in-vehicle camera; S2: Calculate the actual 2D coordinates of each pixel point in the forward-looking image in the camera coordinate system; S3: Process the forward-looking image to sequentially obtain multi-scale features, a high-resolution feature map, and the probability distribution of the 3D coordinate vectors of each pixel point on the high-resolution feature map described by a Gaussian mixture model. Process the probability distribution of the 3D coordinate vectors of each pixel point into the probability distribution of the dynamic 3D coordinates of each pixel point in each local coordinate system. During training, project the 3D distribution into the probability distribution of 2D coordinates in the camera coordinate system, and use the negative log-likelihood loss function to train the network to minimize the ghosting error, so that the negative log-likelihood of the actual 2D coordinates of each pixel point under the 2D coordinate probability distribution is minimized, thereby enabling each pixel point to generate a set of 2D-3D associations; S4: Set a first network branch. According to the Gaussian mixture model, determine the unique target vehicle corresponding to each pixel point, and cluster the center positions of the unique target vehicles corresponding to each pixel point to achieve bottom-up instance segmentation, so that the 2D-3D associations constructed in S3 are divided into dense 2D-3D associations of each vehicle; S5: Construct a PnP problem from the dense 2D-3D associations and solve it to obtain the position and angle of the target vehicle; S6: According to the instance segmentation result of S4, set a second network branch to obtain the size of the unique target vehicle corresponding to each pixel point, and combine the position and angle of the target vehicle obtained in S5 to obtain a vehicle 3D detection frame containing position, angle, and size information; The local coordinate system is a coordinate system established with the center point of the bottom surface of each target vehicle as the origin, the front of each target vehicle as the x-axis, the lower part of each target vehicle as the y-axis, and the left side of each target vehicle as the z-axis; S3 includes the following steps: S3.1: Process the forward-looking image sequentially through a residual network and a feature pyramid network to obtain multi-scale features of the forward-looking image; S3.2: Perform deformable convolution, bilinear interpolation resampling, and splicing processing on the multi-scale features in sequence to obtain a multi-scale fused high-resolution feature map; S3.3: Output the 3D coordinate vectors of each pixel point on the high-resolution feature map through a branch network composed of convolutional layers, and use a Gaussian mixture model to describe the probability distribution of the 3D coordinate vectors of each pixel point; S3.4: Extract the regional features of each target vehicle from the multi-scale features, obtain the probability distribution of the dynamic 3D coordinates of each pixel point in each local coordinate system according to the Gaussian mixture model in S3.3, convert the probability distribution of the dynamic 3D coordinates of each pixel point in each local coordinate system into the probability distribution of 2D coordinates in the camera coordinate system, and use the negative log-likelihood loss function to train the network to minimize the ghosting error, that is, to minimize the negative log-likelihood of the actual 2D coordinates of each pixel point under the 2D coordinate probability distribution, thereby enabling each pixel point to generate a set of 2D-3D associations; S4 includes the following steps: S4.1: Divide the 2D spatial regions occupied by each target vehicle in the front-view image, and set up the first network branch to regress the offset of the 2D position of the geometric center point of the target vehicle to which each pixel belongs relative to the pixel position in the high-resolution feature map, so that the center position of the target vehicle to which each pixel belongs can be located; S4.2: Determine the target vehicle corresponding to each Gaussian mixture model according to the minimum reprojection error criterion, and then obtain the actual center position of the target vehicle corresponding to the Gaussian mixture model S4.3: Calculate the offset between the center position of the target vehicle to which each pixel belongs and the actual center position of the target vehicle corresponding to the Gaussian mixture model, find the minimum offset, and then determine the unique target vehicle corresponding to each pixel; and train the first network branch through the smooth L1 loss function so that the center position x of the unique target vehicle corresponding to each pixel ctr and the actual center position of the unique target vehicle has the minimum offset; S4.4: Set up the foreground network branch to segment foreground pixels on the high-resolution feature map, train this network branch through the cross-entropy loss function, and obtain the foreground network branch; S4.5: According to the foreground network branch and the Gaussian mixture model established in S3, cluster the center positions of the unique target vehicles corresponding to each pixel point to achieve bottom-up instance segmentation, and further divide the 2D-3D associations constructed in S3.4 into dense 2D-3D associations of each vehicle.
2. The monocular 3D vehicle detection method based on dense association according to claim 1, characterized in that The specific method of using the Gaussian mixture model to describe the probability distribution of the 3D coordinate vectors of each pixel point is as follows: where S is the number of pre-set Gaussian mixture models, φ i is the mixing weight of the i-th Gaussian mixture model, ∑ i is the covariance matrix of the i-th Gaussian mixture model, μ i is the mean of the i-th Gaussian mixture model, φ i , ∑ i , μ i are all variables output by the network, is the probability density estimate of x 3D , x 3D is a set of coordinate vectors in the target local coordinate system.
3. The monocular 3D vehicle detection method based on dense association according to claim 2, wherein, The expression for projecting the probability distribution of the dynamic 3D coordinates of each pixel point in each local coordinate system into the probability distribution of 2D coordinates in the camera coordinate system is: [x cam y cam z cam T = Rx 3D + t wherein, R and t are respectively the rotation matrix and displacement vector for converting the local coordinate system to the camera coordinate system, and the intermediate variables x cam , y cam , z cam are respectively the 3D coordinates in the camera coordinate system, x 3D is a set of coordinate vectors in the target local coordinate system, and x 2D is a set of coordinate vectors in the converted camera coordinate system.
4. The monocular 3D vehicle detection method based on dense association according to claim 3, wherein The formula for training the network using the negative log-likelihood loss function is: Wherein, is the weight normalization parameter, satisfying which is used to dynamically balance the weights of the loss function, is the actual 2D coordinate vector of each pixel point, is the 2D coordinate ground truth is the negative log-likelihood under the converted 2D coordinate probability distribution density function, where is the covariance matrix of the i-th 2D Gaussian mixture model, is the mean of the i-th 2D Gaussian mixture model, φ i is the static mixing weight of the i-th Gaussian mixture model, ψ i is the dynamic mixing weight of the i-th Gaussian mixture model.
5. A monocular 3D vehicle detection method based on dense association according to claim 1, characterized in that, The formula for obtaining the position and angle of the target vehicle in S5 is: In the formula, β and t are respectively the yaw angle and displacement offset of the target vehicle after initialization, and μ 2D , are respectively the parameters of the 2D Gaussian mixture model and are functions of β and t. β * , t * are respectively the position and angle of the target vehicle optimized by β and t.
6. The monocular 3D vehicle detection method based on dense association according to claim 5, wherein Before executing S5, use the EPnP algorithm to initialize the yaw angle and displacement offset of the target vehicle.
7. A monocular 3D vehicle detection method based on dense association according to claim 5, characterized in that S6 includes the following steps: S6.1: Set up the second network branch, find the size of the target vehicle corresponding to each pixel point according to the same rule as S4.2, train the second network branch through the smooth L1 loss function, and output the size of the target vehicle corresponding to each pixel point of the high-resolution feature map; S6.2: Determine the size of the unique target vehicle corresponding to each pixel point according to the instance segmentation result of S4.4; S6.3: According to the position and angle of the target vehicle in S5 and the size of the target vehicle obtained in S6.2, obtain the vehicle 3D detection box containing position, angle, and size information.