A low-power fast target detection method for 3D point cloud based on YOLO

Through the improved method based on YOLO network, combined with three-dimensional point cloud metadata preprocessing and two-dimensional mapping technology, the calculation burden problem of the existing three-dimensional point cloud object detection method on low-computing equipment is solved, and efficient and low-power three-dimensional point cloud object detection is achieved.

CN116189147BActive Publication Date: 2025-05-16DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310155654.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-23
Publication Date
2025-05-16
Estimated Expiration
2043-02-23

AI Technical Summary

Technical Problem

The existing three-dimensional point cloud object detection method has a high computing burden and is difficult to realize real-time detection on embedded devices with low computing power and low power consumption.

Method used

Based on the improvement of YOLO network, the three-dimensional point cloud data is mapped to the two-dimensional plane through three-dimensional point cloud metadata preprocessing (including point cloud cropping, downsampling and outlier point removal), and a complex angle regression layer is added to the YOLO network to improve detection efficiency.

Benefits of technology

It reduces the computing burden and prediction time, is suitable for low-power platforms, and improves the computing efficiency and scope of application of three-dimensional point cloud target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189147B_ABST
    Figure CN116189147B_ABST
Patent Text Reader

Abstract

The present invention proposes a three-dimensional point cloud low-power consumption fast target detection method based on YOLO, which belongs to the field of target detection. The target detection method comprises a point cloud metadata processing step, a BEV mapping step, an RGB filling step, and a network feature extraction and regression step. The point cloud metadata processing step performs basic processing on the input data to reduce data interference caused by problems such as sensor noise. The BEV mapping step maps traditional three-dimensional point cloud data to a BEV perspective and compresses three-dimensional information into a two-dimensional space. The RGB filling step fills the three RGB channels after normalizing the three types of information of the point cloud. On the basis of the YOLO network, E-RPN is added for expansion to complete feature extraction and regression tasks. Compared with traditional methods, the present invention improves detection efficiency, effectively reduces calculation burden, reduces computing platform computing power requirements, reduces prediction time, and greatly expands the scope and scenarios of application of three-dimensional point cloud target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning, and in particular relates to a three-dimensional point cloud low-power consumption fast target detection method based on YOLO. Background Art

[0002] 3D point cloud object detection has been widely used in autonomous driving, AR, VR and robotics. 3D point cloud information has richer geometric information than other modal data. With the growth of the market for acquisition equipment such as LiDAR, the threshold for obtaining 3D point clouds is gradually decreasing. 3D point cloud object detection methods are generally divided into three categories: multi-view methods that project 3D point clouds into 2D point clouds, voxel convolution methods based on representing scenes in voxel form, and methods that directly process 3D point cloud data.

[0003] At present, there are several problems with 3D point cloud data: the point cloud density is inconsistent. During the acquisition process, the laser radar has dense point clouds near and sparse point clouds far away; the point cloud is disordered, and the point cloud on the same object can be represented by two completely different 3D point cloud coordinate matrices; the point cloud resolution is low, and the 3D point cloud is a low-resolution sampling of 3D geometric shapes, so only one-sided geometric information can be obtained; there are various noises in the early acquisition sensors, etc.

[0004] Generally speaking, object detection in 3D point clouds has a large computational burden, which limits the application scenarios of 3D point cloud technology to system computing power and cannot be applied to many low-computing and low-power platforms such as embedded devices. Therefore, how to reduce the computational burden, improve computing efficiency, and reduce prediction time is an important research topic in this field. Summary of the invention

[0005] In view of the shortcomings of current three-dimensional point cloud target detection, the present invention obtains a three-dimensional point cloud target detection method based on the improvement of the YOLO network. The model is relatively simple, and the computational efficiency of the three-dimensional point cloud target detection algorithm is improved, the computational burden is reduced, the prediction time is reduced, the hardware requirements are relatively low, and the three-dimensional point cloud target detection function can be quickly completed on a low-power platform. The method is suitable for low-power and low-computing power platforms.

[0006] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is:

[0007] A YOLO-based three-dimensional point cloud low-power fast target detection method includes the following steps:

[0008] Step 1: 3D point cloud metadata processing is used to remove worthless sampling points in the original 3D point cloud data, which specifically includes the following sub-steps:

[0009] (1.1) Point cloud clipping: according to the target environment and the performance of the acquisition equipment, a clipping box is set to remove the point cloud data with a long distance and low value;

[0010] (1.2) Downsampling of point cloud: set a voxel grid of appropriate size and complete downsampling by voxel method;

[0011] (1.3) Outlier removal: Using the Gaussian distribution statistical characteristics of the point cloud, the sampling points with more than α times the standard deviation within the search radius are removed.

[0012] Step 2: Map the 3D point cloud data to BEV, and compress the 3D data into 2D space by mapping, which specifically includes the following sub-steps:

[0013] (2.1) Rasterizing the 3D point cloud information;

[0014] (2.2) Distribute the point cloud into a grid from a bird’s-eye view.

[0015] Step 3: Normalize the information obtained in step 2 and fill it into the RGB three channels, extract the features under the BEV perspective and match them with the RGB channels, which specifically includes the following sub-steps:

[0016] (3.1) Get the maximum height, maximum intensity and point cloud density in each grid respectively;

[0017] (3.2) Normalize the three types of information separately;

[0018] (3.3) Fill the three types of information into the RGB channels to match them with the network.

[0019] Step 4: Use the YOLO network to complete feature extraction and loss regression. Add a complex angle regression layer to the YOLO network to expand it and complete feature extraction and loss regression. Specifically, it includes the following sub-steps:

[0020] (4.1) Feature extraction, using a simplified YOLO-v4 network, extended by adding a complex angle regression layer;

[0021] (4.2) Loss regression, introduce the complex angle into the loss function and complete the loss function calculation.

[0022] Compared with the prior art, the present invention has the following beneficial effects:

[0023] Based on the traditional deep learning neural network for target detection, the present invention modifies the network structure and provides an efficient target detection method. Compressing three-dimensional information into two-dimensional space can effectively reduce the computing burden, improve detection efficiency, and reduce the computing power requirements of the computing platform. Compared with other network structures, it is simpler, reduces the computing power occupied by the computing platform, reduces the prediction time, and expands the scope and scenarios of three-dimensional point cloud target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a basic flow chart of the method of the present invention. DETAILED DESCRIPTION

[0025] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.

[0026] like Figure 1 As shown, a low-power consumption fast target detection method of a three-dimensional point cloud based on YOLO in this embodiment includes the following steps:

[0027] Step 100: Obtain the 3D point cloud metadata of the laser radar.

[0028] Among them, the laser radar imaging system mainly measures the distance of the object through laser, and uses the control system and scanning system to adjust the laser emission angle and position to form an image. Among them, the laser radar ranging mainly calculates the distance between the two through the laser flight time. After the laser is emitted, the timing circuit will start, and the timing will end after the echo signal is received. By calculating the time difference between emission and reception, the target distance is calculated. Therefore, the coordinates with the laser radar as the origin are obtained according to the distance, horizontal angle and vertical angle. Each sampling point consists of a set of three-dimensional coordinates and a reflection intensity.

[0029] Step 101: 3D point cloud metadata preprocessing.

[0030] Among them, firstly, according to the influence of factors such as algorithm characteristics, sampling environment and sampling equipment performance, the cropping box is set to 80 meters in length, 40 meters in width and 3 meters in height, and the point cloud data with low value and far away are eliminated, so as to reduce unnecessary calculations; small voxels with length, width and height of 1 cm are divided in three-dimensional space, and the point set falling in each voxel is obtained. A sampling point is taken in each voxel to replace the original point set, so as to complete the point cloud downsampling; according to the characteristic of Gaussian distribution of point cloud distribution symbols, the number of the nearest points analyzed by sampling points in the point cloud is set to K, and the distances from all points of the point cloud to the sampling points are calculated. If the distance from a point to the sampling point exceeds the average distance by more than α times the standard deviation, this point is identified as an outlier and needs to be eliminated.

[0031] Step 102: Map the three-dimensional point cloud data to the BEV perspective.

[0032] Among them, the three-dimensional point cloud data is first rasterized, and the grid resolution is set to 8 cm. All point clouds are mapped to a two-dimensional plane from a bird's-eye view, so as to obtain a two-dimensional point cloud map from the BEV perspective.

[0033] Step 103: Fill to network RGB channel.

[0034] Among them, the maximum height, maximum intensity and point cloud density of each grid are first obtained respectively. The formula is as follows, and then the three types of information are normalized respectively, and the obtained three-channel data are filled into the RGB channel.

[0035] definition is a set of point clouds projected onto a specific grid from the BEV perspective. Describing the mapping function to a particular grid, then:

[0036]

[0037]

[0038]

[0039] Among them, z g Indicates the maximum height, z b Represents the maximum intensity, I function represents the intensity of a single point cloud, z r represents the normalized point cloud density within the grid, and N is the number of points mapped to a specific grid from the BEV perspective.

[0040] Step 104: Use the YOLO network to complete feature extraction and loss regression.

[0041] The overall network features are similar to those of the YOLO-v4 network, and the feature extraction stage is the same as that of the YOLO-v4 network, that is, the CSP-DarkNet53 network is used. After the feature extraction network, a complex angle regression layer is added to the output layer to decode the features output by the network into the three-dimensional spatial coordinates, size, category probability, and orientation angle of the target.

[0042] The size of the complex angle regression layer is determined by the size and shape of the input point cloud network (in this embodiment, a 32*16*75 regression network is used, i.e., 32*16 grids are divided, and each grid provides 5 predictions). Each prediction contains t x ,t y 、c x 、c y ,t w ,t l etc. prediction parameters.

[0043] b x=σ(t x )+c x

[0044] b y =σ(t y )+c y

[0045]

[0046]

[0047] b φ =arctan2(t Im ,t Re )

[0048] Among them, the predicted center point t x , t y The sigmoid function is used to normalize the relative position of each grid; the σ function indicates that the actual offset is obtained through the relative position; c x , x y is the grid index position on the output feature map; t w , t l The offset relative to the anchor box is obtained by characterizing the logarithmic function; p w , p l is the length and width of the anchor box; t Im , t Re The real and imaginary parts of the predicted complex angle are used to calculate the heading angle b by inverse tangent. φ .

[0049] Among them, b x is the x coordinate of the center point of the target three-dimensional space, b y is the y coordinate of the center point of the target three-dimensional space, b w is the width of the target three-dimensional space, b l is the length of the target three-dimensional space, b φ is the orientation angle of the target in three-dimensional space.

[0050] Among them, complex angle regression, the target's heading angle b φ The corresponding regression parameter t Im and t Re It is calculated that t Im and t Re They correspond to the real part and imaginary part of the complex number respectively. Using complex numbers can effectively avoid singularities.

[0051] Among them, the loss function is:

[0052] L=L Yolo +L Euler

[0053] Among them, L Yolo is the loss function of YOLO itself, L Euler is the loss function of the complex angle regression layer.

[0054] Generally speaking, it is hoped that the learning rate will be larger in the early stage of training so that the network converges quickly, and smaller in the later stage of training so that the network converges better to the optimal solution.

[0055] This embodiment is verified on the KITTI dataset. Compared with networks such as VoxelNet, the average detection accuracy is basically the same, but the model inference time and network size are greatly reduced. This method can achieve a real-time inference speed of 4.7fps on the NVIDIA TX2 low-power embedded platform.

[0056] In summary, the present invention provides a low-power fast target detection method for three-dimensional point clouds based on YOLO. After obtaining the three-dimensional point cloud metadata from the laser radar, the invalid sampling points are reduced by cutting and removing outliers, and then the three-dimensional point cloud is mapped to a two-dimensional plane, which greatly reduces the amount of network calculation. On the basis of the mature YOLO network, a complex convolution layer is added, and finally an efficient point cloud target detection network is formed, and the singularity problem caused by single angle estimation is avoided. The principle of this method is simple and clear, the calculation burden is small, the prediction time is short, and it can effectively expand the application scenarios of three-dimensional point cloud target detection, and has broad application value and market prospects.

Claims

1. A low-power fast target detection method for three-dimensional point cloud based on YOLO, characterized in that: The following steps are involved: Step 1: 3D point cloud metadata processing is used to remove worthless sampling points in the original 3D point cloud data, which specifically includes the following sub-steps: (1.1) Point cloud clipping: according to the target environment and the performance of the acquisition equipment, a clipping box is set to remove the point cloud data with long distance and low value; (1.2) Downsampling of point cloud, setting voxel grid and completing downsampling by voxel method; (1.3) Outlier removal: using the Gaussian distribution statistical characteristics of the point cloud, the sampling points with more than α times the standard deviation within the search radius are removed; Step 2: Map the 3D point cloud data to BEV, and compress the 3D data into 2D space by mapping, which specifically includes the following sub-steps: (2.1) Rasterizing the 3D point cloud information; (2.2) Distribute the point cloud into a grid from a bird’s eye view; Step 3: Normalize the information obtained in step 2 and fill it into the RGB three channels, extract the features under the BEV perspective and match them with the RGB channels, which specifically includes the following sub-steps: (3.1) Get the maximum height, maximum intensity and point cloud density in each grid respectively; (3.2) Normalize the three types of information separately; (3.3) Fill the three types of information into the RGB channels to match them with the network; Step 4: Use the YOLO network to complete feature extraction and loss regression. Add a complex angle regression layer to the YOLO network to expand it and complete feature extraction and loss regression. Specifically, it includes the following sub-steps: (4.1) Feature extraction, using the YOLO-v4 network, extended by adding a complex angle regression layer; (4.2) Loss regression, introduce the complex angle into the loss function and complete the loss function calculation.

2. According to the YOLO-based three-dimensional point cloud low-power fast target detection method according to claim 1, it is characterized in that: In step three, definition is a set of point clouds projected onto a specific grid from the BEV perspective. Describing the mapping function to a particular grid, then: Among them, z g Indicates the maximum height, z b represents the maximum intensity, the I function represents the intensity of a single point cloud, z r represents the normalized point cloud density within the grid, and N is the number of points mapped to a specific grid from the BEV perspective.

3. A YOLO-based three-dimensional point cloud low-power fast target detection method according to claim 1 or 2, characterized in that: In step four, The overall network features are similar to those of the YOLO-v4 network. The feature extraction stage is the same as that of the YOLO-v4 network, that is, the CSP-DarkNet53 network is used. After the feature extraction network, a complex angle regression layer is added to the output layer to decode the features output by the network into the three-dimensional spatial coordinates, size, category probability and orientation angle of the target. The size of the complex angle regression layer is determined by the size and shape of the input point cloud network, and each prediction contains t x ,t y 、c x 、c y ,t w ,t l The prediction parameters of b x =σ(t x )+c x b y =σ(t y )+c y b φ =arctan2(t Im ,t Re ) Among them, the predicted center point t x , t y The sigmoid function is used to normalize the relative position of each grid; the σ function indicates that the actual offset is obtained through the relative position; c x , c y is the grid index position on the output feature map; t w , t l The offset relative to the anchor box is obtained by characterizing the logarithmic function; p w , p l is the length and width of the anchor box; t Im , t Re The real and imaginary parts of the predicted complex angle are used to calculate the heading angle b by inverse tangent. φ ; Among them, b x is the x coordinate of the center point of the target three-dimensional space, b y is the y coordinate of the center point of the target three-dimensional space, b w is the width of the target three-dimensional space, b l is the length of the target three-dimensional space, b φ is the orientation angle of the target in three-dimensional space; Among them, complex angle regression, the target's heading angle b φ The corresponding regression parameter t Im and t Re It is calculated that t Im and t Re Corresponding to the real and imaginary parts of the complex number respectively; Among them, the loss function is: L=L Yolo +L Euler Among them, L Yolo is the loss function of YOLO itself, L Euler is the loss function of the complex angle regression layer.

Citation Information

Patent Citations

  • Visual inspection positioning system and method thereof

    CN113465505A

  • Point cloud pedestrian distance risk detection method, system, device and medium

    CN115100741A