Three-dimensional target detection method, system and storage medium based on auxiliary task learning network

By designing a three-dimensional object detection method based on auxiliary task learning network in the field of autonomous driving, the problems of high calculation cost, long detection time and spatial information loss of monocular vision methods are solved, and three-dimensional object detection is achieved with high precision and real-time performance, which is suitable for autonomous driving application scenarios.

CN116704464BActive Publication Date: 2025-05-06SUZHOU UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310706306.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-14
Publication Date
2025-05-06
Estimated Expiration
2043-06-14

AI Technical Summary

Technical Problem

In the field of autonomous driving, existing monocular vision methods have problems such as high computational cost, long detection time, and low spatial information loss and positioning accuracy after downsampling.

Method used

A three-dimensional object detection method based on auxiliary task learning network is designed, point cloud sampling and grouping is performed through the set abstraction layer (sampling layer, grouping layer and point network layer), point cloud geometric structure information is supervised by voxel feature encoding and self-attention mechanism layer, and new candidate points are generated through the candidate generation layer, and finally data enhancement strategies are used to prevent overfitting.

Benefits of technology

It realizes high-precision three-dimensional object detection, good real-time performance, strong generalization ability, and can meet the online processing and high-precision requirements in autonomous driving application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116704464B_ABST
    Figure CN116704464B_ABST
Patent Text Reader

Abstract

The present invention provides a three-dimensional target detection method, system and storage medium based on an auxiliary task learning network, including: step one, set abstraction: design a set abstraction layer, the set abstraction layer is composed of a sampling layer, a grouping layer and a point network layer; step two: input the output results of D-FPS and FS in parallel into the voxel feature encoding and self-attention mechanism layer to form an auxiliary task learning network; step three: in the candidate generation layer, use the representative points in F-FPS as the initial center points, the initial center points are transferred to their corresponding instances under the supervision of their relative positions and the correction of the auxiliary network center point estimation, so as to generate new candidate points; step four, data enhancement. The beneficial effects of the present invention are: the target detection of the present invention has high accuracy, good real-time performance, strong generalization ability, and can meet the requirements of online processing and high accuracy of target detection in the application scenario of autonomous driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a three-dimensional target detection method, system and storage medium based on an auxiliary task learning network. Background Art

[0002] In recent years, autonomous driving technology has made great progress. Three-dimensional object detection plays an important role in autonomous driving applications such as path planning, motion prediction, and collision avoidance. At present, the camera on the vehicle side is already the standard for three-dimensional object detection, and LiDAR has become more and more popular because it can collect high-precision point cloud depth information. In the field of autonomous driving, monocular cameras can be used to obtain high-resolution scene information. At the same time, due to its simple structure and low cost, it has a wide range of applicability. Researchers have done a lot of research on the application of monocular vision methods in the field of autonomous driving. However, the algorithm based on visual images selects more key points and requires a large number of errors to be eliminated, which greatly increases the computational cost and detection time. Compared with cameras, LiDAR can provide accurate environmental depth information. At present, the single-stage method is prone to loss of spatial information during downsampling, and there is a problem of low positioning accuracy in the subsequent prediction process. Summary of the invention

[0003] The present invention provides a three-dimensional target detection method based on an auxiliary task learning network, comprising the following steps:.

[0004] Step 1: Perform set abstraction: Design a set abstraction layer, which consists of a sampling layer, a grouping layer, and a point network layer.

[0005] The sampling layer first selects a set of points from the input points, which define the centroid of the local region. Given an input point {x1, x2, x3, ..., x n}, use iterative farthest point sampling FPS, that is, Euclidean space farthest point sampling D-FPS to select a subset of points {x i1 , x i2 , x i3 , ..., x im}; Then use the fusion sampling strategy FS of F-FPS and D-FPS respectively Points are sampled, and these two sets are placed in layer A for grouping operation. F-FPS represents the farthest point sampling of the feature.

[0006] The grouping layer constructs a corresponding local area set by searching for neighboring points around the centroid;

[0007] The point network layer uses a micro-point network to encode local area patterns into feature vectors. Given an unordered point set {x1, x2, x3, ..., x n};when When you define a collection function It maps a set of points to a vector;

[0008] Step 2: Input the output results of D-FPS and FS into the voxel feature encoding and self-attention mechanism layer in parallel to form an auxiliary task learning network. The auxiliary task learning network can supervise the geometric structure information in the learning point cloud and guide the backbone network to learn the intermediate features at different stages, so as to obtain the fine-grained structural information of the point cloud;

[0009] Step 3: In the candidate generation layer, the representative points in F-FPS are used as the initial center points. The initial center points are transferred to their corresponding instances under the supervision of their relative positions and the correction of the auxiliary network center point estimation, thereby generating new candidate points.

[0010] Step 4, data augmentation: First, use the general neighboring distribution to construct virtual training examples; then, use Gaussian noise perturbation to offset the point cloud, and finally, randomly rotate and translate each bounding box and follow the average distribution Δθ1∈[-π / 4, +π / 4]; at the same time, add random transformations (Δx, Δy, Δz).

[0011] The present invention also provides a three-dimensional target detection system based on an auxiliary task learning network, comprising: a memory, a processor, and a computer program stored on the memory, wherein the computer program is configured to implement the steps of the three-dimensional target detection method of the present invention when called by the processor.

[0012] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program is configured to implement the steps of the three-dimensional target detection method described in the present invention when called by a processor.

[0013] The beneficial effects of the present invention are: the target detection of the present invention has high accuracy, good real-time performance, and strong generalization ability, and can meet the requirements of online processing and high accuracy of target detection in autonomous driving application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a framework diagram of the present invention;

[0015] Figure 2 is the voxel feature encoding map;

[0016] Figure 3 This is a diagram explaining the movement correction operation of the CG layer, moving the F-FPS point to the center of the instance. DETAILED DESCRIPTION

[0017] like Figure 1As shown, the present invention discloses a three-dimensional target detection method based on auxiliary task learning network (Auxiliary Task Learning-Single Shot Multibox Detector, referred to as ATL-SSD). The three-dimensional target detection method proposed in the present invention has high target detection accuracy, good real-time performance, and strong generalization ability, and can meet the requirements of online processing and high precision of target detection in the application scenario of autonomous driving.

[0018] The present invention designs an auxiliary task learning network composed of voxel feature encoding (VFE) and a self-attention layer. The network is used to supervise the geometric structure information in the learning point cloud, strengthen the context connection, and reduce the spatial information loss caused by downsampling. The present invention also proposes a method for estimating the center point of an object obtained using an auxiliary network. The predicted center point is made close to the instance center to achieve better aggregation, thereby improving the prediction accuracy. Finally, the present invention also proposes a new data enhancement strategy to prevent overfitting.

[0019] The three-dimensional object detection method of the present invention comprises the following steps:

[0020] Step 1: Perform set abstraction. A set abstraction (SA) layer is designed, which consists of three key layers: sampling layer, grouping layer and point network layer.

[0021] The sampling layer first selects a set of points from the input points, which define the centroid of the local region. Given the input points {x1, x2, x3, ..., x n}, use iterative farthest point sampling (FPS), that is, Euclidean farthest point sampling (D-FPS) to select a subset of points {x i1 , x i2 , x i3 , ..., x im Then, the fusion sampling strategies (Fusion Sampling, FS) of F-FPS and D-FPS are used to The two sets are sampled at the A layer and the following grouping operation is performed.

[0022] In order to consider the spatial distance and semantic information of each point, the internal points (positive points) in the instance are retained and those points on the background (negative points) are eliminated. The present invention uses feature distance as the criterion for FPS sampling to screen out many useless negative points. However, only using semantic feature distance as the only criterion, redundant points are retained in the same instance. In order to reduce redundancy and increase diversity, the present invention uses spatial distance and semantic feature distance as the criteria for FPS, expressed as C(A, B)=λL d (A,B)+L f (A, B), where L d (A, B) and L f (A, B) represents the spatial distance and characteristic distance between points A and B, and λ is the balance factor. This sampling method is called feature farthest point sampling (F-FPS).

[0023] The fusion sampling of F-FPS sampling and D-FPS sampling is applied in the SA layer to retain more positive points for localization while retaining enough negative points for classification.

[0024] The grouping layer constructs the corresponding local region set by finding "neighboring points" around the centroid. The input is a set of points of size N×(d+C) and a set of centroid coordinates of size N′×d. The size of the output group point set is N′×K×(d+C), each group corresponds to a local region and the number of centroid neighborhood points K. N is the number of points in this group, d is the depth feature, c is the dimensional feature, and k is the number of centroid neighborhood points.

[0025] The point network layer uses a micro-point network to encode local area patterns into feature vectors. Given an unordered point set {x1, x2, x3, ..., x n}.when When you define a collection function It maps a set of points to a vector. Among them, γ and h are usually multilayer perceptron (MLP) networks: f({x1, x2, x3, ..., x n )=γ(MAX{h(x i )}), where i=1,...,n.

[0026] Step 2: Auxiliary task learning network, which inputs the D-FPS and FS output results into the voxel feature encoding and self-attention mechanism layer in parallel.

[0027] In general, the downsampled convolutional features extracted from point clouds will inevitably lose structural details that are critical for accurate positioning. Therefore, the present invention proposes an auxiliary learning network that inputs the D-FPS and FS output results into the voxel feature encoding and self-attention mechanism layer in parallel. The network can supervise the learning of geometric structure information in point clouds and guide the backbone network to learn intermediate features at different stages, thereby obtaining fine-grained structural information of point clouds.

[0028] The voxel feature encoding stage is as follows Figure 2 As shown in Figure 1, the average value of all points in the voxel is obtained to obtain (Vx, Vy, Vz), and the feature of each point is upgraded to a 7-dimensional feature point Vin. Then, P with 7 features is i Input to the FC network to obtain point-by-point features and 7-dimensional feature points. After maximum pooling, the features obtained in the previous step are aggregated element by element to obtain local aggregate features. The point-by-point features and local aggregate features are concatenated to obtain fused point-by-point connection features, and the final point-by-point features are max-pooled to obtain voxel features.

[0029] In order to make full use of the sampling advantage and better fuse bird's-eye views of different scales, the present invention proposes a self-attention mechanism. By processing the farthest point sampling and fusion sampling feature maps in three-dimensional space in parallel, the effect of retaining detail information and increasing the receptive field is achieved.

[0030] Step 1: Calculate the similarity between the query and each key to obtain the weight. Step 2: Use a softmax function to normalize these weights, and finally perform a weighted summation of the weights and the corresponding key values ​​to obtain the Attention vector. i , V i |i=1,2,...,m} is mapped to the output. The query, each key, and each value are vectors. The output is the weighted sum of all values ​​in V. The weight is calculated by the query and each key. The calculation method is divided into three steps:

[0031] Step 1: Calculate and compare the similarity between Q and K. The calculation method is represented by f. Step 2: Perform Softmax operation on the obtained similarity and normalize it. Step 3: For the calculated weight α i , perform weighted sum calculation on all values ​​in V to get the Attention vector. The input query is q and the key dimension is d k , value dimension is d v. Calculate the dot product of the query and each key and divide by Finally, the Softmax function is applied to calculate the weights.

[0032] Step 3: Through the candidate generation layer. In the candidate generation layer (CG), since most of the representative points in D-FPS are negative sample points, only the representative points in F-FPS are used as the initial center points. Figure 3 As shown, these initial center points are transferred to their corresponding instances under the supervision of their relative positions and the correction of the auxiliary network center point estimation, thereby generating new candidate points. Among them, the transfer mechanism adopts the deep Hough voting method, given an input point cloud containing N points and corresponding x, y, and z coordinates. These points are sampled and deep features are learned through the backbone network, and a subset of M points is output as seed points. Each seed generates a vote independently through the voting module. Then, the votes are grouped into clusters and processed by the proposal module to generate the final proposal. Finally, the correction vector is adjusted through the auxiliary network learning parameters, and the correction transfer generates the final candidate point.

[0033] Candidate points closer to the center of an instance can often obtain more accurate positioning predictions. The present invention uses three-dimensional center labels to help the network distinguish the instances corresponding to candidate points. For each candidate point, its center label is defined in two steps. First, determine whether it is in instance l mask Secondly, a center label is drawn based on its 6 surface distances to the corresponding instance. The calculation formula of the center label is as follows:

[0034] Step 4: Data enhancement.

[0035] Most neural networks can usually minimize the average error during training. This learning rule is called empirical risk minimization. But this also brings some problems. This learning rule allows memorization of training data, which reduces generalization ability. When evaluating on examples outside the training distribution, the prediction will change dramatically. The present invention proposes to use data augmentation instead. First, a general neighbor distribution is used to construct virtual training examples. Among them, λ ~ Beta (α, α), α∈ (0, ∞). (x i y i ) and (x j ,y j) are two feature target vectors randomly extracted from the training data, λ∈[0,1]. The hybrid hyperparameter α controls the interpolation strength between the feature and target pairs, and when α→0, it returns to the empirical risk minimization principle. In the present invention, the hyperparameter α is set to 0.2. Then, the point cloud is offset using Gaussian noise perturbation instead of adding new points to the original points. Finally, each bounding box is randomly rotated and translated, and follows the average distribution Δθ1∈[-π / 4, +π / 4]. At the same time, random transformations (Δx, Δy, Δz) are added.

[0036] The present invention is suitable for three-dimensional target detection during automatic driving. The driving vehicle uses multiple sensors to collect target data to ensure classification and positioning accuracy.

[0037] The target detection method of the present invention has high accuracy, good real-time performance, and strong generalization capability, and can meet the requirements of online processing and high accuracy of target detection in autonomous driving application scenarios.

[0038] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A three-dimensional object detection method based on an auxiliary task learning network, characterized in that: The steps include: Step 1: Perform set abstraction: Design a set abstraction layer, which consists of a sampling layer, a grouping layer, and a point network layer. The sampling layer first selects a set of points from the input points, which define the centroid of the local region. Given an input point {x1, x2, x3, ..., x n }, use iterative farthest point sampling FPS, that is, Euclidean space farthest point sampling D-FPS to select a subset of points {x 1i , x i2 , x i3 , ..., x im }; Then use the fusion sampling strategy FS of F-FPS and D-FPS respectively Points are sampled, and these two sets are placed in layer A for grouping operation. F-FPS represents the farthest point sampling of the feature. The grouping layer constructs a corresponding local area set by searching for neighboring points around the centroid; The point network layer uses a micro-point network to encode local area patterns into feature vectors. Given an unordered point set {x1, x2, x3, ..., x n };when When you define a collection function It maps a set of points to a vector; Step 2: Input the output results of D-FPS and FS into the voxel feature encoding and self-attention mechanism layer in parallel to form an auxiliary task learning network. The auxiliary task learning network can supervise the geometric structure information in the learning point cloud and guide the backbone network to learn the intermediate features at different stages, so as to obtain the fine-grained structural information of the point cloud; Step 3: In the candidate generation layer, the representative points in F-FPS are used as the initial center points. The initial center points are transferred to their corresponding instances under the supervision of their relative positions and the correction of the auxiliary network center point estimation, thereby generating new candidate points. Step 4: Data augmentation: First, use the general neighbor distribution to construct virtual training examples; Then, the point cloud is offset using Gaussian noise perturbation. Finally, each bounding box is randomly rotated and translated following the average distribution Δθ1∈[-π / 4, +π / 4]; at the same time, a random transformation (Δx, Δy, Δz) is added.

2. The three-dimensional target detection method according to claim 1, characterized in that: In step 1, the spatial distance and the semantic feature distance are used as the FPS standard, expressed as C(A, B)=λL d (A,B)+L f (A, B), where L d (A, B) and L f (A, B) represents the spatial distance and characteristic distance between points A and B, λ is the balance factor, and this sampling method is called characteristic farthest point sampling.

3. The three-dimensional target detection method according to claim 1, characterized in that: In the grouping layer, a group of points of size N×(d+C) and a group of centroid coordinates of size N′×d are input, and the size of the output group point set is N′×k×(d+C), each group corresponds to a local area and the number of centroid neighborhood points K, N is the number of points in this group, d is the depth feature, c is the dimensional feature, and k is the number of centroid neighborhood points.

4. The three-dimensional target detection method according to claim 1, characterized in that: In the second step, at the voxel feature encoding stage, the average value of all points in the voxel is obtained to obtain (Vx, Vy, Vz), and the feature of each point is upgraded to a 7-dimensional feature point Vin, and then the P with 7 features is converted into a 7-dimensional feature point Vin. i The points are input into the FC network to obtain point-by-point features and 7-dimensional features; the obtained features are aggregated element by element through maximum pooling to obtain local aggregate features; the point-by-point features and local aggregate features are concatenated to obtain fused point-by-point connection features, and the final point-by-point features are subjected to maximum pooling to obtain voxel features.

5. The three-dimensional target detection method according to claim 1, characterized in that: In the step 2, the feature maps of the farthest point sampling and the fused sampling in the three-dimensional space are processed in parallel, including the following steps: Step 1: Calculate the similarity between the query and each key to obtain the weight; Step 2: Use the softmax function to normalize the weights, and finally perform weighted summation on the weights and the corresponding key values ​​to obtain the Attention vector, and convert the query (Q) and key-value pairs {K i , V i |i=1,2,…,m} is mapped to the output; where the query, each key, and each value are vectors; the output is the weighted sum of all values ​​in V.

6. The three-dimensional target detection method according to claim 5, characterized in that: In step 2, the weight is calculated by the query and each key. The calculation method is divided into three steps: Step 1: Calculate and compare the similarity between Q and k. The calculation method is represented by f. Step 2: Perform Sogtmax operation on the obtained similarity and normalize it; Step 3: For the calculated weight α i , perform weighted sum calculation on all values ​​in V to obtain the Attention vector; the input query is q and the key dimension is d k , value dimension is d v ; Calculate the dot product of the query and each key and divide by Finally, the Softmax function is applied to calculate the weights.

7. The three-dimensional target detection method according to claim 1, characterized in that: In step 3, the transfer mechanism uses a deep Hough voting method. Given an input point cloud containing N points and corresponding x, y, and z coordinates, the points are sampled and deep features are learned through the backbone network, and a subset of M points is output as seed points. Each seed generates a vote independently through the voting module. Then, the votes are grouped into clusters and processed by the proposal module to generate the final proposals; finally, the correction vector is adjusted through the auxiliary network learning parameters and the correction transfer produces the final candidate points.

8. The three-dimensional target detection method according to claim 1, characterized in that: In step 3, the 3D center label is used to help the network distinguish the instance corresponding to the candidate point. For each candidate point, its center label is defined through two steps: first, determine whether the candidate point is in instance l mask Secondly, a center label is drawn according to the 6 surface distances from the candidate point to the corresponding instance. The calculation formula of the center label is as follows:

9. A three-dimensional object detection system based on an auxiliary task learning network, characterized in that: include: A memory, a processor, and a computer program stored in the memory, wherein the computer program is configured to implement the steps of the three-dimensional object detection method according to any one of claims 1 to 8 when called by the processor.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the three-dimensional object detection method according to any one of claims 1 to 8 when called by a processor.

Citation Information

Patent Citations

  • Three-dimensional visual inspection method, system and device based on shape attention mechanism

    CN110879994A

  • Multi-view image three-dimensional reconstruction method based on attention mechanism

    CN111402405A