An end-to-end bottom-up 3D object detection method for generating proposal boxes
Point cloud data is processed through Ball k-means clustering and Shell-based methods, combined with PointNet and multi-layer perceptron, fast and efficient three-dimensional object detection is achieved, solving the problem of slow convergence speed in the existing methods, and a high-precision three-dimensional detection box is generated.
Patent Information
- Application Number
- CN202111466037.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-03
AI Technical Summary
The existing three-dimensional object detection method is not effective in the way of directly reverting to the geometric center through the clustering center in autonomous driving, and the convergence speed is slow, making it difficult to detect three-dimensional objects quickly and accurately.
The original point cloud is processed by using the Ball k-means clustering algorithm, divided into multiple spherical clusters, point-by-point semantic features are extracted through PointNet encoding and decoder, and the geometric center of the suggestion box is calculated in combination with the Shell-based method, and fused with the semantic features through a multi-layer perceptron. Finally, Shell-based classification regression is performed through the total loss function to generate the final three-dimensional detection box.
It improves the accuracy and speed of three-dimensional object detection, reduces the computational complexity, and quickly generates high-precision three-dimensional detection frames through bottom-up and end-to-end object detection processes.
Smart Images

Figure CN114241225B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a bottom-up three-dimensional object detection method for end-to-end generating proposal boxes, belonging to the field of object detection in artificial intelligence. Background Art
[0002] Autopilot has become well-known in recent years, and scholars and companies at home and abroad have shown great enthusiasm for research in this field. Different from the object detection of traditional two-dimensional images: in the autopilot scenario, three-dimensional point cloud data is used as the input, and the 3D bounding boxes of labeled three-dimensional objects are well separated. This makes the detection of objects such as cars and pedestrians in three-dimensional space more realistic.
[0003] Traditionally, in two-dimensional space, such as fast-RCNN, SSD, and YOLOv3, two-dimensional information obtained from cameras is used as the input, and these network architectures show high accuracy in many image features. These methods provide good ideas for three-dimensional object detection. Prior to this, there were also many three-dimensional object detection methods such as PointNet and PointNet++. However, in these three-dimensional object detection methods, the way of directly regressing the geometric center is usually adopted. The effect of directly regressing the geometric center through the clustering center is not ideal, and it often requires multiple rounds of regression to obtain the geometric center, and the convergence speed is not fast enough. Summary of the Invention
[0004] In view of the deficiencies mentioned in the background art, the present invention provides an accurate and efficient bottom-up three-dimensional object detection method for end-to-end generating proposal boxes.
[0005] The technical solution of the present invention:
[0006] The present invention provides a bottom-up three-dimensional object detection method for end-to-end generating proposal boxes, and the steps are as follows:
[0007] (1) Processing the original point cloud through the Ball k-means clustering algorithm to obtain the clustering center, dividing the original point cloud into multiple spherical clusters, and obtaining the clustering center O and clustering radius R of each spherical cluster;
[0008] (2) Inputting the original point cloud in each spherical cluster into the PointNet encoder and decoder to obtain the per-point semantic features of each input point;
[0009] (3) Calculating the geometric center of the proposal box respectively through the Shell-based method, obtaining the foreground mask through foreground point semantic segmentation, and end-to-end generating the proposal box;
[0010] (4) Point cloud region pooling operation, separating local spatial points from the original point cloud and the foreground mask, performing spatial coordinate transformation, passing through a multi-layer perceptron (MLP), and fusing features with the semantic features;
[0011] (5) Perform PointNet encoding on the fused features;
[0012] (6) Through the total loss function perform Shell-based classification regression to obtain the final geometric center, predict the confidence level, suppress redundant 3D detection boxes, and obtain the final 3D detection boxes.
[0013] Further, in step (1), the specific steps of obtaining the clustering center by processing the original point cloud through the Ball k-means clustering algorithm are as follows:
[0014] For any spherical cluster C, it can be represented by (O, R), and the clustering center O can be expressed as:
[0015]
[0016] where N represents the number of point clouds in the spherical cluster C, i represents the counter, and p i represents the i-th arbitrary point in the spherical cluster C, and R represents the clustering radius of the spherical cluster;
[0017] Given two spherical clusters C i and C j , their clustering centers are represented as O i and O j ; R i represents the clustering radius of C i , if R i satisfies the following inequality:
[0018]
[0019] where C i and C j represent the i-th and j-th adjacent spherical clusters respectively, i and j represent the counters, and O i and O j represent the clustering centers of C i and C j respectively.
[0020] Further, in step (3), the calculation of the geometric center by the Shell - based method specifically includes the following steps: Taking the clustering center O as the center of the space, trisect the clustering radius R of the spherical cluster, where r1 = R / 3, r2 = 2R / 3, r3 = R. The concentric spheres with radii r1, r2, and r3 divide the space into three spatial domains (Shell1, Shell2, Shell3). Among them, Shell1 is presented as a sphere, and Shell2 and Shell3 are presented as spherical shells. The spherical coordinate radii r of these three spatial domains are represented by the following set: {0 ≤ r < r1, r1 ≤ r < r2, r2 ≤ r ≤ r3};
[0021] In the spherical coordinate system, the spherical coordinate representation of the clustering center O(x O , y O , z O ) is O(x O , y O , z O ) and The mapping relationship with
[0022] z O = r O cosα O
[0023] where represents the azimuth angle of the clustering center O in spherical coordinates, α O represents the zenith angle of the clustering center O in the spherical coordinate system, r O represents the radius of the clustering center O in spherical coordinates. Denote this mapping relationship as Then, mapping from Cartesian coordinates to the spherical coordinate system is represented by the following formula:
[0024]
[0025] Denote this mapping relationship as
[0026] Then Shell = (Shell1, Shell2, Shell3) is represented by the following formula:
[0027]
[0028] Δr, α, respectively represent the radius offset, zenith angle, and azimuth angle in the spherical coordinate system; r represents the radius of the spherical coordinate system in three spatial domains; for the divided spherical clusters, regarded as three different classification categories, in the designed deep neural network, through calculation and backpropagation, classify which one of Shell1, Shell2, and Shell3 the geometric center O' coordinates are in, and then through the regression method, calculate Δr, α, Then through the mapping relationship calculate the x, y, and z coordinates of the corresponding geometric center O' in the Cartesian coordinate system (x O′ , y O′ , z O′ ).
[0029] Furthermore, in step (6), the specific steps of fine-tuning the proposed box through the total loss function to obtain the three-dimensional detection box BBox are as follows:
[0030] (1) The classification of the Shell to which the geometric center O' belongs by the Shell-based method can be expressed as:
[0031]
[0032] where is the result after mapping the predicted value of the geometric center O' coordinates of its corresponding object, is the result after mapping the true value of the geometric center O' coordinates of its corresponding object, S represents the corresponding search range, and u generally refers to the x, y, and z coordinates;
[0033] Within the specified Shell, further position refinement is the residual with the true value along the x, y, and z axes can be expressed as:
[0034]
[0035] where and are the conversion values of the true values of the geometric center O' coordinates (x, y, z) in the Shell division, and R represents the clustering radius of the spherical cluster;
[0036] The loss of the Shell-based method
[0037]
[0038] where is the binary cross-entropy loss, represents the true value of which Shell the geometric center O' coordinates belong to, is the predicted value of which Shell the coordinates of the geometric center O' belong to, represents the Smooth-L1 loss, represents the true value of the residual part, represents the predicted value of the residual part;
[0039] (2) Confidence loss
[0040]
[0041]
[0042]
[0043]
[0044] Among them, G IoU represents the bounding box loss function, IoU represents the intersection over union of the BBox (Bounding Box) and the true value, A represents the volume of the smallest cuboid region enclosing the BBox and the true value, and u represents the volume of the union of the BBox and the true value; is the predicted confidence obtained by passing the i-th predicted value Conf through the Sigmoid function; T i represents the coincidence degree of the i-th BBox and the true value; K is the number of positive and negative samples;
[0045] (3) Total loss of all object detections is defined as follows:
[0046]
[0047] Among them, N pos represents the number of positive samples, λ conf , λ loc , λ cls , λ air , λ Shell respectively represent the balance coefficients of the confidence loss, localization loss, classification loss, orientation angle loss, and the balance coefficient of the Shell-based localization loss, respectively represent the confidence loss, localization loss, classification loss, orientation angle loss, and Shell-based localization loss.
[0048] Beneficial effects
[0049] (1) For the input original point cloud, the present invention first processes the point cloud data by using an unbounded fast adaptive accurate Ball k-means algorithm, divides the original point cloud data into multiple categories, each category has its own clustering center, calculates the clustering center and the clustering radius of the spherical cluster, thus, it is equivalent to dividing the three-dimensional space, greatly reducing the computational complexity and reducing the computational amount of sliding convolution;
[0050] (2) This method extracts features for each clustering center through PointNet, directly generates a proposal box, and achieves a bottom-up and end-to-end goal;
[0051] (3) In the fine-turn stage, this method adjusts the center point of the Bonding Box, designs a Shell-based model, adopts a strategy of classification first and then regression. For the clustering center obtained by the Ball k-means algorithm, it calculates the geometric center by using the Shell-based method. First, it uses the classification method to classify the object center into which Shell layer it belongs to, narrow down the range, and further calculates the geometric center by regression, effectively improving the network convergence speed. Brief Description of the Drawings
[0052] Figure 1 is a flowchart of a bottom-up three-dimensional object detection method for end-to-end generating proposal boxes according to the present invention;
[0053] Figure 2 is Figure 1 's framework diagram;
[0054] Figure 3 is an end-to-end generated proposal box for single-point prediction;
[0055] Figure 4 is a Shell-based classification and regression diagram;
[0056] Figure 5 is an implementation scheme of the bottom-up three-dimensional object detection method for end-to-end generating proposal boxes in the field of autonomous driving. Detailed Embodiment
[0057] The following further details the present application in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.
[0058] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The following will detail the present application with reference to the drawings and embodiments.
[0059] Figure 1 Shows an exemplary basic flowchart of a bottom-up 3D object detection method for end-to-end generating proposal boxes.
[0060] A bottom-up 3D object detection method for end-to-end generating proposal boxes according to the present invention, as Figure 2 shown, is divided into two-stage object detection. In the first stage, proposal boxes are output; in the second stage, the proposal boxes are fine-tuned to obtain the final 3D detection boxes BBox (Bounding Box).
[0061] Specifically, as Figure 2 and Figure 5 shown, the Velodyne HDL-64E lidar is used to collect the original point cloud data. With the original point cloud data PointCloud as the input, first, through the Ball k-means clustering algorithm, the PointCloud is divided into multiple spherical clusters, and the clustering center O and the clustering radius R of each spherical cluster are obtained. Further, the PointCloud in each spherical cluster is input into the PointNet encoder and decoder (PointNet encoder and decoder, hereinafter referred to as: PointNet codec). After passing through the PointNet codec, the feature vector of each input point can be obtained.
[0062] Further, post-processing tasks of obtaining the center point coordinates of the proposal box and foreground point semantic segmentation are respectively performed based on Shell. The center point coordinates of the proposal box are also the geometric center of the object.
[0063] For end-to-end generating proposal boxes, the obtained proposal boxes at this time still cannot meet the positioning accuracy, so fine-tuning in the next stage is required.
[0064] In the fine-turn stage, the input comes from the original point cloud, semantic features, foreground mask, and the proposal boxes proposal generated from single points in the previous stage. Therefore, this is a bottom-up process, as Figure 3 shown. The foreground mask is a mask for semantic segmentation, which filters out background points and retains foreground points.
[0065] Further, the coordinate transformation is performed on the original point cloud, and then through a multi-layer perceptron (MLP), feature fusion is performed with the semantic features. In this way, both local features and global features are included. Subsequently, the fused features are encoded by PointNet, and then based on Shell, classification and regression are performed to obtain the final geometric center, and the confidence can be further adjusted to obtain the object detection box.
[0066] The above algorithm is deployed and run on the processing unit of the Nuvo-5095GC industrial personal computer (IPC). For the final object detection box results, they are input into the autonomous driving logic processing system. Synchronously input along with these are such as: Global Positioning System (GPS), Inertial Measurement Unit (IMU), RGB color camera (camera), and other sensor data. The GPS provides the vehicle's longitude and latitude information and is input into the autonomous driving logic processing system; the IMU can provide richer information such as yaw angular velocity and angular acceleration and is input into the autonomous driving logic processing system; the camera is responsible for inputting image information into the recognition algorithm running on the IPC to identify traffic signs and signals, and finally inputs the recognition results into the autonomous driving logic processing system. The autonomous driving logic processing system outputs a comprehensive driving plan based on the input multi-modal information.
[0067] In some implementable ways of this embodiment, the Ball k-means clustering algorithm processes the point cloud to obtain the clustering centers as follows:
[0068] The point cloud data is a subset of points in the Euclidean space, and the discreteness of the data is naturally suitable for the clustering algorithm. The Ball k-means clustering algorithm is adopted in the present invention. By searching for adjacent spherical clusters and then dividing the spherical clusters, a series of spherical clusters are quickly and adaptively generated. The specific details are as follows. For any spherical cluster C, it can be represented by (O, R), and the clustering center O can be expressed as:
[0069]
[0070] where N represents the number of point clouds in the spherical cluster C, i represents the counter, and p i represents the i-th arbitrary point in the spherical cluster C, and R represents the clustering radius of the spherical cluster.
[0071] Given two spherical clusters C i and C j , whose clustering centers are represented as O i and O j ; R i represents the clustering radius of C i , if R i satisfies the following inequality:
[0072]
[0073] where C i and C j represent the i-th and j-th adjacent spherical clusters respectively, i and j represent the counters, and O i and O j represent C i and Cj The clustering center.
[0074] Using the nearest-neighbor spherical cluster can greatly reduce the distance calculation range of the points within a spherical cluster in the next iteration.
[0075] In some implementable ways of this embodiment, the Shell-based method calculates the geometric center as follows:
[0076] The clustering center O is obtained through the Ball k-means clustering algorithm. However, since the point cloud data is distributed on the surface of the object, the obtained clustering center is often not the geometric center (object center) required for target detection. Therefore, it is necessary to further obtain the geometric center O'. In the field of autonomous driving, the span of the point cloud data of the vehicle boundary reaches 3 to 5 meters. In this work, the Shell-based method is proposed to change the original "direct regression" into a "classification + regression" problem.
[0077] Shell-based: Taking the clustering center O as the center of the space, the radius R of the spherical cluster is divided into three equal parts, r1 = R / 3, r2 = 2R / 3, r3 = R. The concentric spheres with r1, r2, and r3 as the radii are divided into three spatial domains (Shell1, Shell2, Shell3), where Shell1 is presented as a sphere, and Shell2 and Shell3 are presented as spherical shells, as Figure 4 shown. The spherical coordinate radius r of these three spatial domains is represented by the following set: {0 ≤ r < r1, r1 ≤ r < r2, r2 ≤ r ≤ r3}.
[0078] In the spherical coordinate system, the spherical coordinates of the clustering center O(x O , y O , z O ) can be expressed as O(x O , y O , z O ) and The mapping relationship can be expressed by the following formula:
[0079] z O = r O cosα O
[0080] where represents the azimuth angle of the clustering center O in the spherical coordinates, α O represents the zenith angle of the clustering center O in the spherical coordinate system, and r O represents the radius of the clustering center O in the spherical coordinates. Denote this mapping relationship as Conversely, mapping from Cartesian coordinates to the spherical coordinate system can be expressed by the following formula:
[0081]
[0082] Denote this mapping relationship as
[0083] Furthermore, Shell = (Shell1, Shell2, Shell3) can be expressed by the following formula:
[0084]
[0085] Δr, α, respectively represent the radius offset, zenith angle, and azimuth angle in the spherical coordinate system; r represents the radius of the spherical coordinate system of the three spatial domains; for the divided spherical clusters, regarded as three different classification categories, in the designed deep neural network, through calculation and backpropagation, it can be classified which one of Shell1, Shell2, and Shell3 the geometric center O' coordinates are in. At this time, a rough result is obtained, and further through the regression method, Δr, α, Then the mapping relationship can calculate the x, y, and z coordinates (x O′ , y O′ , z O′ ) of the corresponding geometric center O' in the Cartesian coordinate system.
[0086] In some realizable ways of this embodiment, the BBox is adjusted through the total loss function as follows:
[0087] (1) For the Shell-based method, the classification of the geometric center O' belonging to the Shell can be expressed as:
[0088]
[0089] where, is the result after mapping the predicted value of the geometric center O' coordinates of its corresponding object, is the result after mapping the true value of the geometric center of its corresponding object, S represents the corresponding search range, and u generally refers to the x, y, and z coordinates.
[0090] Within the specified Shell, further position refinement of the residual from the true value along the x, y, and z axes can be expressed as:
[0091]
[0092] where, is the conversion value of the true value of the geometric center O' coordinates (x, y, z) in the Shell division, and R represents the clustering radius of the spherical cluster;
[0093] Thus, the loss function of the Shell-based method is expressed as follows:
[0094]
[0095] where, is the binary cross-entropy loss, represents the true value of which Shell the geometric center O' coordinates belong to, is the predicted value of which Shell the geometric center O' coordinates belong to, represents the Smooth-L1 loss, represents the true value of the residual part, represents the predicted value of the residual part.
[0096] (2) Confidence loss
[0097]
[0098]
[0099]
[0100]
[0101] where, G IoU represents the bounding box loss function, IoU represents the intersection over union of the BBox (Bounding Box) and the true value, A represents the volume of the smallest cubic region enclosing the BBox and the true value, and u represents the volume of the union of the BBox and the true value; is the predicted confidence obtained by passing the i-th predicted value Conf through the Sigmoid function; T i represents the coincidence degree of the i-th BBox and the true value; K is the number of positive and negative samples.
[0102] (3) The total loss of all object detections is defined as follows:
[0103]
[0104] where, N pos represents the number of positive samples, λ conf , λ loc , λ cls , λ dir , λ ShellThey respectively represent the balance coefficients of confidence loss, localization loss, classification loss, orientation angle loss, and the balance coefficient of shell-based localization loss. They respectively represent confidence loss, localization loss, classification loss, orientation angle loss, and shell-based localization loss.
[0105] As described above, it is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the embodiments of the present invention should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A bottom-up 3D object detection method for end-to-end generating proposal boxes, the steps are as follows: (1) Process the original point cloud through the Ball k-means clustering algorithm to obtain the clustering centers, divide the original point cloud into multiple spherical clusters, and obtain the clustering center O and clustering radius R of the spherical clusters; (2) Input the original point cloud in each spherical cluster into the PointNet encoder-decoder to obtain the point-level semantic features of each input point cloud; (3) Calculate the geometric center of the proposal box through the Shell-based method, and then through foreground point cloud semantic segmentation, obtain the foreground mask, and generate the proposal box end-to-end; The Shell - based method takes the clustering center O as the center of the space and divides the clustering radius R of the spherical cluster into three equal parts, where r1 = R / 3, r2 = 2R / 3, and r3 = R. Three concentric spheres with radii r1, r2, and r3 divide the space into three spatial domains (Shell1, Shell2, Shell3). Among them, Shell1 is presented as a sphere, and Shell2 and Shell3 are presented as spherical shells. The spherical coordinate radius r of these three spatial domains is represented by the following set: {0 ≤ r < r1, r1 ≤ r < r2, r2 ≤ r ≤ r3}. For the divided spherical cluster, it is regarded as three different classification categories. In the designed deep neural network, through calculation and backpropagation, it is classified which of Shell1, Shell2, and Shell3 the geometric center O' is in. Then, through the regression method, Δr and α are calculated. Then, through the mapping relationship the x, y, and z coordinates (x O' , y O' , z O' ) of the corresponding geometric center O' in the Cartesian coordinate system are calculated; (4) Point cloud region pooling operation, separate the local space points from the original point cloud and the foreground mask, perform spatial coordinate transformation, pass through a multi-layer perceptron (MLP), and fuse the features with the semantic features; (5) Encode the fused features by PointNet; (6) Through the total loss function Perform Shell-based classification regression to obtain the final geometric center, prediction confidence, suppress redundant 3D detection boxes, and obtain the final 3D detection boxes.
2. The three-dimensional object detection method according to claim 1, characterized in that, In step (1), the specific steps of processing the original point cloud through the Ball k-means clustering algorithm to obtain the clustering centers are as follows: For any spherical cluster C, it can be represented by (O, R), and the clustering center O can be expressed as: where N represents the number of point clouds in the spherical cluster C, i represents the counter, and p i represents the i-th arbitrary point of the spherical cluster C, and R represents the clustering radius of the spherical cluster; Given two spherical clusters C i and C j , whose cluster centers are denoted as O i and O j ; R i represents the clustering radius of C i , if R i satisfies the following inequality: Among which C i and C j respectively represent the i-th and j-th adjacent spherical clusters, where i and j are counters, O i and O j respectively represent the clustering centers of C i and C j .
3. The three-dimensional object detection method according to claim 2, wherein In step (3), the specific steps of calculating the geometric center through the Shell-based method are as follows: In the spherical coordinate system, the spherical coordinate representation of the clustering center O(x O , y O , z O ) is expressed as The mapping relationship between O(x O , y O , z O ) and is represented by the following formula: Among them represents the azimuth angle of the clustering center O in spherical coordinates, α O represents the zenith angle of the clustering center O in the spherical coordinate system, r O represents the radius of the clustering center O in spherical coordinates. Denote this mapping relationship as Mapping from Cartesian coordinates to the spherical coordinate system is expressed by the following formula: Denote this mapping relationship as Then Shell = (Shell1, Shell2, Shell3) is expressed by the following formula: where Δr, α, respectively represent the radius offset, zenith angle, and azimuth angle in the spherical coordinate system; r represents the radius of the spherical coordinate system in three spatial domains.
4. The three-dimensional object detection method according to claim 3, wherein In step (6), the specific steps of fine-tuning the proposal box through the total loss function to obtain the 3D detection box BBox are as follows: (1) The classification of the Shell to which the geometric center O' belongs in the Shell-based method can be expressed as: Among them, is the result after mapping the predicted value of the coordinates of the geometric center O' of its corresponding object, is the result after mapping the true value of the coordinates of the geometric center O' of its corresponding object. S represents the corresponding search range, and u generally refers to the x, y, and z coordinates; Within the specified Shell, further position refinement of the residuals from the true values along the x, y, and z axes can be expressed as: Among them, and are the conversion values of the true values of the coordinates (x, y, z) of the geometric center O' in the Shell division, and R represents the clustering radius of the spherical cluster; The loss of the Shell-based method Among them, is the binary cross-entropy loss, represents the true value of which Shell the coordinates of the geometric center O' belong to, is the predicted value of which Shell the coordinates of the geometric center O' belong to, represents the Smooth-L1 loss, represents the true value of the residual part, represents the predicted value of the residual part; (2) Confidence loss Among them, G IoU represents the border loss function, IoU represents the intersection over union of the BBox (Bounding Box) and the ground truth, A represents the volume of the smallest cubic region enclosing the BBox and the ground truth, and u represents the volume of the union of the BBox and the ground truth; is the predicted value Conf i of the predicted confidence obtained through the Sigmoid function; T i represents the coincidence degree of the i-th BBox and the ground truth; K is the number of positive and negative samples; (3) The total loss of all object detections is defined as follows: Among them, N pos represents the number of positive samples, λ conf , λ loc , λ cls , λ dir , λ Shell respectively represent the balance coefficients of confidence loss, localization loss, classification loss, orientation angle loss, and Shell-based localization loss. respectively represent confidence loss, localization loss, classification loss, orientation angle loss, and Shell-based localization loss.