A 3D target detection method, apparatus, terminal, and storage medium
By using a two-stage 3D target detection algorithm based on the attention mechanism of the original point cloud mesh, the problems of poor position detection accuracy and low detection efficiency in the existing technology are solved, and more efficient and accurate 3D target detection is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2026-03-13
AI Technical Summary
Among existing 3D target detection methods, voxel-based methods have poor position detection accuracy, while point-based methods have low detection efficiency, resulting in poor detection performance.
A two-stage 3D target detection algorithm based on the attention mechanism of the original point cloud grid is adopted. The algorithm acquires laser point cloud data, performs voxelization and 3D sparse convolutional layer processing to extract the region of interest, performs farthest point sampling and spatial gridding processing, and uses a multi-head attention mechanism to capture the dependency relationship between grid points to perform target category prediction and bounding box position regression.
It improves the accuracy and efficiency of 3D target detection in terms of location detection, thereby enhancing the detection effect.
Smart Images

Figure CN115311653B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer technology, specifically relating to a 3D target detection method, device, terminal, and storage medium, and particularly to a 3D target detection algorithm, device, terminal, and storage medium based on the spatial attention mechanism of the original point cloud. Background Technology
[0002] As a key technology for robotics and autonomous driving perception systems, 3D object detection technology has made rapid progress. Point clouds acquired by LiDAR (LiDAR radar) can be used for describing the 3D structure of objects, estimating their pose, and sensing spatial distance; therefore, LiDAR has become the most commonly used sensor in 3D object detection technology. 3D object detection technology based on raw point clouds aims to utilize point clouds acquired by LiDAR to detect the position, size, and orientation of targets such as vehicles and pedestrians in various scenes, thereby further enhancing scene understanding.
[0003] In relevant solutions, 3D object detection methods can be broadly categorized into voxel-based methods and point-based methods. Voxel-based methods divide the point cloud into a regular grid and then use mature 3D convolution for feature extraction. However, voxel-based methods lose precise location information of the point cloud during voxel feature encoding, resulting in poor location detection accuracy and creating a performance bottleneck for voxel-based 3D object detection models. Point-based methods, on the other hand, use the original point cloud for detection. Due to the large number of points, multi-level sampling and feature aggregation are required, and these methods are generally less efficient.
[0004] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The purpose of this invention is to provide a 3D target detection method, apparatus, terminal, and storage medium to address the problem that in related 3D target detection schemes, voxel-based 3D target detection methods have poor position detection accuracy, while point-based 3D target detection methods have low detection efficiency, resulting in poor 3D target detection performance. This invention achieves the goal of improving the position detection accuracy and efficiency of 3D target detection by setting a two-stage 3D target detection algorithm based on an attention mechanism of the original point cloud mesh, thereby enhancing the overall 3D target detection performance.
[0006] This invention provides a 3D target detection method, comprising: acquiring laser point cloud data containing a target object as the original 3D point cloud data of the target object; performing voxelization, 3D sparse convolutional layer, and RPN network processing on the original 3D point cloud data of the target object to obtain a region of interest (ROI) in the original 3D point cloud data of the target object, as the ROI of the target object; performing farthest point sampling and spatial meshing processing on the ROI of the target object to obtain local features of the center point of the target object; performing coordinate dimensionality upscaling and feature summation processing on the local features of the center point of the target object to obtain ROI features of the target object; and performing target category prediction and bounding box position regression processing on the ROI features of the target object to achieve 3D target detection of the target object.
[0007] In some embodiments, acquiring laser point cloud data containing the target object as the original three-dimensional point cloud data of the target object includes: acquiring laser point cloud data of the target object collected by a lidar as the original three-dimensional point cloud data of the target object; wherein the original three-dimensional point cloud data of the target object has a value range of a first predetermined range in the X-axis direction, a value range of a second predetermined range in the Y-axis direction, and a value range of a third predetermined range in the Z-axis direction; and / or, based on the original three-dimensional point cloud data of the target object, performing voxelization, 3D sparse convolutional layer, and RPN network processing to obtain the region of interest in the original three-dimensional point cloud data of the target object, as the region of interest. The region of interest (ROI) of the target object includes: performing voxelization processing on the original 3D point cloud data of the target object to obtain 3D voxels of the original 3D point cloud data of the target object; performing feature extraction using 4 layers of 3D sparse convolution on the 3D voxels of the original 3D point cloud data of the target object to obtain multiple scale spatial features of the original 3D point cloud data of the target object; and performing viewpoint compression on the multiple scale spatial features of the last layer of the 4 layers based on the multiple scale spatial features of the original 3D point cloud data of the target object, and then using a region proposal network to extract the ROI, thus obtaining the ROI in the original 3D point cloud data of the target object, which is taken as the ROI of the target object.
[0008] In some implementations, based on the region of interest (ROI) of the target object, farthest-point sampling and spatial meshing are performed to obtain the local features of the center point of the target object. This includes: dividing the ROI of the target object into a cylindrical structure; sampling the farthest point of the original 3D point cloud data of the target object within the cylindrical structure to obtain interest points in the ROI of the target object, which are used as interest sampling points of the target object; uniformly meshing the ROI of the target object to obtain multiple ROI meshes of the target object, which are used as multiple interest meshes of the target object; determining the center point of each interest mesh of the target object, and determining the relative distance between the center point of each interest mesh of the target object and the interest sampling points of the target object; and determining the local features of the center points of all interest meshes of the target object based on the center point of each interest mesh of the target object and the relative distance between the interest sampling points of the target object.
[0009] In some embodiments, dividing the region of interest (ROI) of the target object into a cylindrical structure based on the ROI of the target object includes: setting the ROI of the target object as a cylinder, and using the cylindrical structure containing the cylinder as the cylindrical structure after dividing the ROI of the target object; wherein, the base radius of the cylinder is... r for ,high h for ;in, , , These are the width, length, and height of the region of interest, respectively. and The parameters are: a set column expansion ratio parameter; and / or, based on the center point of each interest grid of the target object and the relative distance between the interest sampling points of the target object, determining the local features of the center points of all interest grids of the target object, including: modeling and unifying the positional coding coordinate scale processing of the spatial position of the corresponding interest grids of the target object based on the center point of each interest grid of the target object and the relative distance between the interest sampling points of the target object, to obtain the positional features of the center points of each interest grid of the target object; performing dimensionality-up processing on the center points of each interest grid of the target object based on the positional features of the center points of each interest grid of the target object, to obtain the set of positional features of the center points of all interest grids of the target object within a set radius in a set spherical region; obtaining the set of feature expressions of the center points of all interest grids of the target object at different radius scales by changing the radius size of the sphere to which the set spherical region belongs, based on the set of positional feature expressions of the center points of all interest grids of the target object at different radius scales; and stitching together the features at different radius scales based on the set of feature expressions of the center points of all interest grids of the target object, to obtain the local features of the center points of all interest grids of the target object.
[0010] In some implementations, based on the center point of each interest grid of the target object and the relative distance between the interest sampling points of the target object, the spatial position of the corresponding interest grid of the target object is modeled and uniformly coded with coordinate scales to obtain the positional features of the center point of each interest grid of the target object. This includes: calculating the positional features of the center point of each interest grid of the target object according to the following formula based on the center point of each interest grid of the target object and the relative distance between the interest sampling points of the target object:
[0011] ;
[0012] ;
[0013] in, It is the positional feature of the center point of each interest grid of the target object. This is a feature transformation function used to map the features of the relative distance to a high-dimensional feature space using a feedforward neural network. and The relative distance between the interest sampling point of the target object and the center point of each interest grid of the target object. Additional features for the interest sampling points of the target object.
[0014] In some implementations, based on the local features of the center point of the target object, coordinate dimensionality upscaling and feature summation are performed to obtain the region of interest (ROI) features of the target object. This includes: using a three-layer feedforward neural network to upscale the coordinates of the center point of the target object based on the local features of the center point, and aggregating the features of different radius scales of the local features of the center point of the target object using a max-pooling function; using a feedforward neural network to adjust the dimension of the local features of the center point of the target object after the dimensionality upscaling and aggregation, and summing the positional encoding features of the local features of the center point of the target object and the local features of different radius scales to obtain the center point features of all the grids of interest of the target object; based on the center point features of the grids of interest of the target object, using an attention mechanism to capture the dependencies between the center points of different grids of interest among the center points of all the grids of interest of the target object, assigning corresponding weights to the center point features of different grids of interest among the center points of all the grids of interest of the target object according to the dependencies, so as to obtain the association relationship between the center point features of all the grids of interest of the target object and the ROI of the target object; and using a multi-head attention mechanism to determine the ROI features of the target object based on the association relationship between the center point features of all the grids of interest of the target object and the ROI of the target object.
[0015] In some implementations, based on the region of interest (ROI) features of the target object, target category prediction and bounding box position regression processing are performed to achieve 3D target detection of the target object. This includes: inputting the ROI features of the target object into a preset detection head, performing classification and regression processing of the 3D target detection bounding boxes of the target object, and determining the detection model loss of the 3D target detection bounding boxes; as the detection model loss of the 3D target detection bounding boxes of the target object decreases, determining the 3D target detection bounding boxes of the target object, thus achieving 3D target detection of the target object; wherein, the detection model loss of the 3D target detection bounding boxes of the target object includes: region proposal network loss and refinement stage loss; the region proposal network loss includes: confidence loss of the 3D target detection bounding boxes of the target object, and position regression loss of the 3D target detection bounding boxes of the target object.
[0016] In conjunction with the above method, another aspect of the present invention provides a 3D target detection device, comprising: an acquisition unit configured to acquire laser point cloud data containing a target object, as the original three-dimensional point cloud data of the target object; a detection unit configured to perform voxelization, 3D sparse convolutional layer, and RPN network processing based on the original three-dimensional point cloud data of the target object to obtain a region of interest in the original three-dimensional point cloud data of the target object, as the region of interest of the target object; the detection unit is further configured to perform farthest point sampling and spatial meshing processing based on the region of interest of the target object to obtain local features of the center point of the target object; the detection unit is further configured to perform coordinate dimensionality increase and feature summation processing based on the local features of the center point of the target object to obtain features of the region of interest of the target object; the detection unit is further configured to perform target category prediction and bounding box position regression processing based on the features of the region of interest of the target object to achieve 3D target detection of the target object.
[0017] In conjunction with the above-described device, the present invention further provides a terminal comprising: the 3D target detection device described above.
[0018] In conjunction with the above method, the present invention further provides a storage medium comprising a stored program, wherein the program, when running, controls the device where the storage medium is located to execute the 3D target detection method described above.
[0019] Therefore, the solution of this invention obtains laser point cloud data containing the target object as the original 3D point cloud data, performs voxelization and 3D sparse convolutional layer processing on the original 3D point cloud data to extract the region of interest, performs farthest point sampling and spatial grid encoding processing on the region of interest to obtain feature points of interest, and then uses the features of the region of interest to predict the target category and regress the bounding box position to achieve 3D target detection. Thus, by setting a two-stage 3D target detection algorithm based on the attention mechanism of the original point cloud grid, the position detection accuracy and detection efficiency of 3D target detection can be improved, which is conducive to improving the detection effect of 3D target detection.
[0020] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention.
[0021] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating an embodiment of the 3D target detection method of the present invention;
[0023] Figure 2 This is a flowchart illustrating an embodiment of the method of the present invention, which involves voxelization, 3D sparse convolutional layers, and RPN network processing based on the original 3D point cloud data of the target object.
[0024] Figure 3 This is a flowchart illustrating an embodiment of the method of the present invention, which performs farthest-point sampling and spatial meshing based on the region of interest of the target object.
[0025] Figure 4 This is a flowchart illustrating an embodiment of the method of the present invention for determining the local features of the center points of all interest grids based on the relative distance between the center point of each interest grid and the interest sampling points of the target object;
[0026] Figure 5 This is a flowchart illustrating an embodiment of the method of the present invention, which performs coordinate dimension increase and feature summation based on local features of the center point of the target object.
[0027] Figure 6 This is a flowchart illustrating an embodiment of the method of the present invention, which performs target category prediction and bounding box position regression processing based on the region of interest features of the target object.
[0028] Figure 7 This is a schematic diagram of the structure of an embodiment of the 3D target detection device of the present invention;
[0029] Figure 8 This is a flowchart illustrating an embodiment of a 3D target detection algorithm based on a spatial attention mechanism of raw point clouds according to the present invention.
[0030] Figure 9 This is a schematic diagram of region of interest sampling in a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud according to the present invention;
[0031] Figure 10 This is a schematic diagram of multi-scale spatial feature aggregation in a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud according to the present invention;
[0032] Figure 11 This is a schematic diagram of point feature encoding in a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud according to the present invention;
[0033] Figure 12 This is a schematic diagram of grid attention feature weighting in a 3D target detection algorithm based on the original point cloud spatial attention mechanism of the present invention, wherein (a) is a schematic diagram of the gridded region of interest, and (b) is a schematic diagram of different grids having different feature weights after attention calculation;
[0034] Figure 13The above are schematic diagrams of the detection effect in multiple scenarios of an embodiment of a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud according to the present invention, wherein (a) is a schematic diagram of the detection effect in the first scenario, (b) is a schematic diagram of the detection effect in the second scenario, and (c) is a schematic diagram of the detection effect in the third scenario.
[0035] Figure 14 This diagram illustrates the comparison of the detection performance of a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud in this invention with other related algorithms. (a) shows the detection performance of the SECOND algorithm (i.e., a target detection algorithm based on three-dimensional point clouds), (b) shows the detection performance of the PointPillars algorithm (i.e., a 3D target detection algorithm based on laser point clouds), and (c) shows the detection performance of the 3D target detection algorithm based on the spatial attention mechanism of the original point cloud.
[0036] Referring to the accompanying drawings, the reference numerals in the embodiments of the present invention are as follows:
[0037] 102 - Acquisition unit; 104 - Detection unit. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0039] Considering that among the relevant 3D target detection methods, voxel-based 3D target detection methods have poor position detection accuracy, while point-based 3D target detection methods have low detection efficiency, the detection effect of the relevant 3D target detection methods is poor.
[0040] Furthermore, considering that 3D object detection methods can be categorized into single-stage and two-stage object detection paradigms, single-stage object detection directly predicts anchor boxes using extracted point cloud features. While fast, it suffers from low accuracy. Two-stage object detection, on the other hand, uses Region Proposal Networks (RPNs) to generate Regions of Interest (RoIs) that may contain target objects. Further feature extraction is then performed on these RoIs to determine the object's category, location, size, and orientation, generating more refined candidate boxes. This results in higher accuracy for two-stage object detection methods.
[0041] With the continuous development of 3D object detection algorithms, a trend in some schemes is to design more promising feature pooling methods during the two-stage refinement process. Through analysis of classic detection techniques in some schemes, we have identified some factors affecting the performance of 3D object models, such as:
[0042] (1) Compared with the single-stage method, the two-stage method can retain more spatial information of the point cloud due to the existence of the detection head structure, thereby improving the detection accuracy of the model;
[0043] (2) Selecting an appropriate receptive field size has a positive impact on the two-stage feature extraction, and it is not necessary to sample and aggregate features from the entire point cloud space;
[0044] (3) Enhancing the spatial location encoding of points is beneficial to improving model performance;
[0045] (4) The Transformer (i.e., a model that uses attention mechanism to improve model training speed) structure can learn point cloud features more effectively and calculate the contribution of different features to the features of the region of interest through attention mechanism.
[0046] Therefore, the present invention proposes a new 3D target detection method, specifically a two-stage 3D target detection algorithm based on the attention mechanism of the original point cloud mesh. The implementation process of the present invention will be illustrated below.
[0047] According to embodiments of the present invention, a 3D target detection method is provided, such as... Figure 1 The diagram shows a flowchart of an embodiment of the method of the present invention. The 3D target detection method may include steps S110 to S150.
[0048] In step S110, laser point cloud data containing the target object is acquired as the original three-dimensional point cloud data of the target object.
[0049] In some implementations, step S110, which involves acquiring laser point cloud data containing the target object as the original three-dimensional point cloud data of the target object, includes: acquiring laser point cloud data of the target object collected by a lidar as the original three-dimensional point cloud data of the target object.
[0050] The original three-dimensional point cloud data of the target object has a first set range in the X-axis direction, a second set range in the Y-axis direction, and a third set range in the Z-axis direction. The first set range is such as [0.0m, 70.4m], the second set range is such as [-40.0m, 40.0m], and the third set range is such as [-3.0m, 1.0m].
[0051] Figure 8 This is a flowchart illustrating an embodiment of a 3D object detection algorithm based on a spatial attention mechanism of raw point clouds according to the present invention. Figure 8 In this context, "Point Cloud" refers to the cloud of points. "Points of Interest" refers to the point cloud of regions of interest. "Proposal to Grid" refers to gridding the regions of interest. "Spatial Geometry Features" refers to spatial geometric features. "Multi-scale Local Feature" refers to multi-scale local features. "Detect Head" refers to the detection head. "Grid-Wise RoI Pooling" refers to grid pooling. "Confidence" refers to the confidence level. "FFN" refers to a feedforward neural network. "Box Refinement" refers to box regression. "Position Embedding" refers to position embedding. "Multi-Head Self-Attention" refers to multi-head self-attention. "3D Voxel-based Backbone" refers to a 3D backbone network. "RPN" refers to a region proposal network. Figure 8 As shown, the implementation flow of a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud provided by the present invention includes:
[0052] Step 1: Input the raw 3D point cloud data obtained by LiDAR.
[0053] Specifically: acquire laser point cloud data containing the target object as the data to be detected, that is, the original three-dimensional point cloud data containing the laser point cloud data of the target object. The original three-dimensional point cloud data containing the laser point cloud data of the target object is limited to the value range of the point cloud in the X-axis direction as [0.0m, 70.4m], the value range in the Y-axis direction as [-40.0m, 40.0m], and the value range in the Z-axis direction as [-3.0m, 1.0m].
[0054] In step S120, based on the original three-dimensional point cloud data of the target object, voxelization, 3D sparse convolutional layer, and RPN network processing are performed to obtain the region of interest in the original three-dimensional point cloud data of the target object, which is used as the region of interest of the target object.
[0055] In some implementations, step S120 involves voxelization, 3D sparse convolutional layers, and RPN network processing based on the original 3D point cloud data of the target object to obtain the region of interest in the original 3D point cloud data of the target object. For the specific process of obtaining the region of interest of the target object, please refer to the following exemplary description.
[0056] The following is combined Figure 2The diagram illustrates an embodiment of the method of the present invention, which involves voxelization, 3D sparse convolutional layer, and RPN network processing based on the original 3D point cloud data of the target object. It further explains the specific process of voxelization, 3D sparse convolutional layer, and RPN network processing based on the original 3D point cloud data of the target object in step S120, including steps S210 to S230.
[0057] Step S210: Based on the original three-dimensional point cloud data of the target object, voxelization processing is performed to obtain three-dimensional voxels of the original three-dimensional point cloud data of the target object.
[0058] Step S220: Based on the three-dimensional voxels of the original three-dimensional point cloud data of the target object, feature extraction is performed using 4 layers of 3D sparse convolution to obtain multiple scale spatial features of the original three-dimensional point cloud data of the target object.
[0059] Step S230: Based on the multiple scale spatial features of the original 3D point cloud data of the target object, the multiple scale spatial features of the last layer of the 4 layers are compressed from the perspective, and the region of interest is extracted using a region proposal network to obtain the region of interest in the original 3D point cloud data of the target object, which is then used as the region of interest of the target object.
[0060] like Figure 8 As shown, the implementation flow of a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud provided by the present invention further includes:
[0061] Step 2: The input raw 3D point cloud data is processed through voxelization and 3D sparse convolutional layers for feature extraction. The data is then input into the RPN network (i.e., Region Generation Network) to extract the region of interest. The specific steps include the following exemplary steps.
[0062] Step 21: Based on the original 3D point cloud data containing the laser point cloud data of the target object, voxelize the point cloud to obtain 3D voxels of the original 3D point cloud data. For example: set the size of the voxel blocks in the X, Y, and Z directions to 0.05m, 0.05m, and 0.1m, respectively, and the number of voxel blocks in the three directions to 1408, 1600, and 40, respectively, and set the number of points in each voxel to no more than 5.
[0063] Step 22: Extract features from the three-dimensional voxels of the original three-dimensional point cloud data using 4 layers of 3D sparse convolution to obtain spatial features of the point cloud at multiple scales.
[0064] Step 23: Based on the spatial features of the point cloud at multiple scales, compress the last layer of spatial features to a bird's-eye view and input it into the Region Proposal Network (RPN) to extract the region of interest. Here, a bird's-eye view is a camera position that uses the perspective of a bird flying in the sky as the camera's viewpoint.
[0065] In step S130, based on the region of interest of the target object, farthest point sampling and spatial gridding are performed to obtain the local features of the center point of the target object.
[0066] In some implementations, the specific process of performing farthest point sampling and spatial meshing based on the region of interest of the target object in step S130 to obtain the local features of the center point of the target object is described in the following exemplary description.
[0067] The following is combined Figure 3 The schematic diagram shown is a flowchart of an embodiment of the method of the present invention, which performs farthest point sampling and spatial gridding based on the region of interest of the target object. It further illustrates the process of performing farthest point sampling and spatial gridding based on the region of interest of the target object in step S130, including steps S310 to S340.
[0068] Step S310: Based on the region of interest (ROI) of the target object, the ROI of the target object is divided into a cylindrical structure. Inside the cylindrical structure, the farthest point of the original 3D point cloud data of the target object is sampled to obtain the interest points in the ROI of the target object, which are used as the interest sampling points of the target object.
[0069] Step S320: Based on the region of interest of the target object, uniformly mesh the region of interest of the target object to obtain multiple region of interest meshes of the target object, which serve as multiple interest meshes of the target object.
[0070] Step S330: Determine the center point of each interest grid of the target object, and determine the relative distance between the center points of each interest grid of the target object and the interest sampling points of the target object.
[0071] In some implementations, step S330, which divides the region of interest of the target object into a cylindrical structure based on the region of interest of the target object, includes: setting the region of interest of the target object as a cylinder based on the region of interest of the target object, and using the cylindrical structure containing the cylinder as the cylindrical structure after dividing the region of interest of the target object.
[0072] Wherein, the bottom radius of the cylinder r for ,high h for .in , , These are the width, length, and height of the region of interest, respectively. and The set column expansion ratio parameter.
[0073] Step S340: Based on the center point of each interest grid of the target object and the relative distance between the interest sampling points of the target object, determine the local features of the center points of all interest grids of the target object.
[0074] like Figure 8 As shown, the implementation flow of a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud provided by the present invention includes:
[0075] Step 3: Next, the region of interest is divided into a cylindrical structure, and the farthest point is sampled inside using the original point cloud to obtain Points of Interest. The specific steps include the following examples.
[0076] Step 31: Set the sampling space of the region of interest as a cylinder. Specifically, set the sampling space of the region of interest as a cylinder with a bottom radius of... r for ,high h for ,in , , These are the width, length, and height of the region of interest, respectively. and The set column expansion ratio parameter. Figure 9 This is a schematic diagram illustrating the region of interest sampling in a 3D object detection algorithm based on a spatial attention mechanism of raw point clouds, as described in this invention. Figure 9 As shown, the sampling region obtained from the region of interest sampling can be designed as a cylindrical structure. This cylindrical structure allows for filtering of point clouds containing objects on a car, such as the point cloud of a car parked under a tree, thus ensuring effective filtering.
[0077] Step 32: Based on the extracted regions of interest, use farthest point sampling to sample each region of interest to obtain the points of interest for each region of interest.
[0078] Among them, farthest point sampling is a very commonly used sampling algorithm. Because it can guarantee uniform sampling of samples, it is widely used. For example, PointNet++ in the 3D point cloud deep learning framework performs FPS sampling on sample points and then clusters them as receptive fields. The 3D object detection network VoteNet performs FPS sampling on scattered points obtained by voting and then clusters them. In the 6D pose estimation algorithm PVN3D, it is used to select 8 feature points of the object for voting and calculate the pose.
[0079] In this way, sampling is performed on points within the region of interest using the farthest point sampling method, thus fully preserving the shape features of the point cloud within the region.
[0080] like Figure 8 As shown, the implementation flow of a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud provided by the present invention includes:
[0081] Step 4: Divide the region of interest into a uniform spatial grid, and encode the region of interest by taking the center point of the grid. This includes encoding multi-scale local spatial features and point cloud spatial coordinates. In grid-wise pooling, the two are concatenated and then attention is encoded. The specific steps include the following exemplary steps.
[0082] Step 41: Perform uniform meshing on the region of interest, setting the number of meshes to 6×6×6, so that each region of interest contains 216 meshes.
[0083] Step 42: Next, define the center point of each grid as... ,in Calculate the number of grid cells for each region of interest, and then calculate the center point of each grid cell. to sampling point p i relative distance ;
[0084]
[0085] In some implementations, step S340 determines all the interests of the target object based on the center point of each interest grid of the target object and the relative distance between the interest sampling points of the target object.
[0086] For a detailed explanation of the local features of the center point of the mesh, please refer to the following exemplary description.
[0087] The following is combined Figure 4 The illustrated flowchart shows an embodiment of the method of the present invention for determining the local features of the center points of all interest grids based on the relative distance between the center point of each interest grid and the interest sampling point of the target object. The flowchart further illustrates the specific process of determining the local features of the center points of all interest grids based on the relative distance between the center point of each interest grid and the interest sampling point of the target object in step S340, including steps S410 to S440.
[0088] Step S410: Based on the center point of each interest grid of the target object and the relative distance between the interest sampling points of the target object, the spatial position of the corresponding interest grid of the target object is modeled and uniformly coded with coordinate scale to obtain the positional features of the center point of each interest grid of the target object.
[0089] In some implementations, step S410 involves modeling and unifying the positional coding coordinate scale of the spatial location of the corresponding interest grids of the target object based on the center point of each interest grid and the relative distance between the interest sampling points of the target object, to obtain the positional features of the center point of each interest grid of the target object. This includes: calculating the positional features of the center point of each interest grid of the target object according to the following formula based on the relative distance between the center point of each interest grid of the target object and the interest sampling points of the target object:
[0090]
[0091] in, It is the positional feature of the center point of each interest grid of the target object. This is a feature transformation function used to map the features of the relative distance to a high-dimensional feature space using a feedforward neural network. The distance of the interest sampling point of the target object to each interest grid of the target object. The relative distance between the center points is an additional feature of the interest sampling points of the target object.
[0092] Specifically, see Figure 8 In the example shown, each grid center point Location features The calculation is as follows;
[0093]
[0094] in, Here, a feedforward network (FFN) is used as the feature transformation function to map the distance features to a high-dimensional feature space. For point Euclidean distance from the center point of each grid Additional features for points.
[0095] Step S420: Based on the positional features of the center point of each interest grid of the target object, perform dimensionality-upgrading processing on the center point of each interest grid of the target object to obtain a set of positional features of the center points of all interest grids of the target object within a set radius in a set spherical region.
[0096] Step S430: Based on the set of positional features of the center points of all interest grids of the target object within a set radius in a set spherical region, by changing the radius of the sphere to which the set spherical region belongs, obtain the set of feature representations of the center points of all interest grids of the target object at different radius scales.
[0097] Step S440: Based on the feature representation set of the center points of all interest grids of the target object at different radius scales, the features at different radius scales are spliced together to obtain the local features of the center points of all interest grids of the target object.
[0098] like Figure 8 As shown, the implementation flow of a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud provided by the present invention further includes:
[0099] Step 43: Use the center point of each grid to sampling point p i relative distance The spatial location of grid points is explicitly modeled, the coordinate scale of the location codes is standardized, and finally the center point of each grid is obtained. Location features .
[0100] Step 44: Next, extract multi-scale local features of the grid points. Specifically, this can be done by: for each grid center point... Query the points within a spherical region of radius r around the grid center point, and apply PointNet to each point to increase its dimensionality, thereby obtaining the feature set of all points within a specified radius of the grid center point. .
[0101] in The number of points within that radius, such as Figure 12 As shown. Figure 12 This is a schematic diagram of grid attention feature weighting in a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud according to the present invention. (a) is a schematic diagram of the gridded region of interest, and (b) is a schematic diagram showing that different grids have different feature weights after attention calculation. Figure 12 The diagram illustrates the weighted representation of grid attention features, showing that different grid points contribute differently to the features of the region of interest. In this invention, an attention mechanism is used to model the features of grid points, fully considering their contribution to the target features, thereby extracting more complex point cloud spatial features.
[0102] To satisfy the permutation invariance requirement, a max pooling function is used to aggregate the feature set to obtain the features of the center point within that radius. :
[0103] .
[0104] in, For aggregation functions, vector concatenation is used here. Aggregate functions It is used to concatenate multi-head attention features. Figure 10 This is a schematic diagram illustrating multi-scale spatial feature aggregation in a 3D object detection algorithm based on a spatial attention mechanism of raw point clouds, as described in this invention. Figure 10 As shown, in the multi-scale local features aggregated by grid center points, feature aggregation is performed on points within multiple radii. In the scheme of this invention, by dividing the point cloud space into a uniform grid and using the grid center points to represent point cloud features, it is beneficial to improve the detection accuracy of occlusion.
[0105] Step 45: Next, by changing the radius of the sphere, we can obtain the feature representation of the center point at different scales.
[0106] Step 46: Finally, the multi-scale features are concatenated to obtain the final local features of the center point. f g :
[0107] .
[0108] In the solution of this invention, point cloud sampling and multi-scale local feature aggregation are performed in the second stage to preserve the spatial information of the target and avoid the problem of low detection efficiency caused by complex feature extraction in the original point cloud scene. Thus, this solves the problem in some solutions where the second-stage refinement of 3D target detection algorithms based on the original point cloud does not fully utilize the local features and contextual dependencies of points, resulting in poor detection performance for occluded targets and thus affecting detection accuracy.
[0109] In step S140, based on the local features of the center point of the target object, coordinate dimensionality increase and feature summation are performed to obtain the region of interest features of the target object.
[0110] In some implementations, the specific process of obtaining the region of interest features of the target object by performing coordinate dimensionality increase and feature summation based on the local features of the center point of the target object in step S140 is illustrated in the following exemplary description.
[0111] The following is combined Figure 5 The diagram illustrates an embodiment of the method of the present invention, which performs coordinate dimensionality increase and feature summation processing based on the local features of the center point of the target object. It further explains the specific process of performing coordinate dimensionality increase and feature summation processing based on the local features of the center point of the target object in step S140, including steps S510 to S540.
[0112] Step S510: The local features of the center point of the target object include the coordinates of the center point of the target object. Based on the local features of the center point of the target object, a 3-layer feedforward neural network is used to increase the dimensionality of the center point coordinates of the target object, and a max pooling function is used to aggregate the local features of the center point of the target object at different radius scales.
[0113] Step S520: Using a feedforward neural network, adjust the dimension of the local features of the center point of the target object after the dimensionality increase and aggregation, and sum the position encoding features of the local features of the center point of the target object and the local features of different radius scales to obtain the center point features of all grids of interest of the target object.
[0114] Step S530: Based on the center point features of the grids of interest of the target object, an attention mechanism is used to capture the dependency relationship between the center points of different grids of interest among all the center points of the grids of interest of the target object. According to the dependency relationship, corresponding weights are assigned to the center point features of different grids of interest among all the center points of the grids of interest of the target object, so as to obtain the association relationship between the center point features of all the grids of interest of the target object and the region of interest of the target object.
[0115] Step S540: Based on the correlation between the center point features of all the interest grids of the target object and the region of interest of the target object, a multi-head attention mechanism is used to determine the region of interest features of the target object.
[0116] like Figure 8 As shown, the implementation flow of a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud provided by the present invention includes:
[0117] Step 5: Finally, to enhance spatial information, a residual structure is used to upgrade the coordinates to a higher-dimensional space and sum them with the attention features to obtain the final region of interest features. The specific steps include the following exemplary steps.
[0118] Step 51: Use a 3-layer FFN to increase the dimensionality of the aggregated coordinates, and aggregate the features at each scale using a max pooling function. FFN is used to transform the dimensions of the features.
[0119] Step 52: Finally, use FFN to adjust the final local features of the center point. The positional encoding features and multi-scale local features are summed in the given dimension to obtain the final grid center point features. f grid :
[0120] .
[0121] Step 53: Use an attention mechanism to capture the long-range dependencies between grid points. Assign different weights to the grid point features to capture more complex relationships between grid point features and regions of interest. Input features , ,and . This represents the local features of the grid center point, specifically the features obtained by aggregating the features of the grid point with its surrounding points. Empty grid features are not included in attention encoding; only their positional encoding is retained. Here, the original coordinate features of the grid center point are used. f pos As a positional encoding: This represents the location characteristics of the grid center point, which here refers to the characteristics calculated using the coordinates of the grid center point. Figure 11 This is a schematic diagram of point feature encoding in a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud according to the present invention. Figure 11 The coordinates of the grid center point are encoded, and spatial information enhancement of the grid point coordinates is performed using sampling points. In the scheme of this invention, it was found that feature enhancement of point coordinates has a positive impact on improving detection accuracy, thus a new point cloud coordinate enhancement method was designed.
[0122] Step 54: Use a multi-head attention mechanism to capture richer features of the region of interest. Multi-head attention features The calculation method is as follows:
[0123]
[0124] Among them, A i V represents the attention coefficient. i For the feature F calculated above i It was multiplied by a linearly changing matrix. K i Qi, V i The calculation of d is a general calculation method. q For feature F i The number of dimensions.
[0125] Step 55: Establish a residual-like channel between the grid spatial location encoding and the attention encoding, concatenate the spatial location encoding of points with the attention features to enrich the expressive power of the features, and after FFN processing, obtain the final region of interest features. f i :
[0126] .
[0127] In step S150, based on the region of interest features of the target object, target category prediction and bounding box position regression are performed to achieve 3D target detection of the target object.
[0128] This invention proposes a two-stage 3D object detection algorithm based on an attention mechanism of original point cloud meshes. By expanding the receptive field, aggregating multi-scale local features, and performing refined modeling of point coordinates, it fully preserves the spatial information of points and considers the complex relationship between mesh points and regions of interest to improve detection accuracy. The receptive field is the size of the region mapped back to the input image from the pixels in the feature map output by each layer of the convolutional neural network. This solves the problem that in related 3D object detection schemes, voxel-based methods suffer from poor position detection accuracy, while point-based methods suffer from low detection efficiency, resulting in poor detection performance.
[0129] In some implementations, step S160 involves predicting the target category and regressing the bounding box position based on the region of interest features of the target object, thereby realizing the 3D target detection of the target object. See the following exemplary description for details.
[0130] The following is combined Figure 6 The diagram illustrates an embodiment of the method of the present invention, which performs target category prediction and bounding box position regression based on the region of interest features of the target object. It further explains the specific process of performing target category prediction and bounding box position regression based on the region of interest features of the target object in step S160, including steps S610 to S620.
[0131] Step S610: Based on the region of interest features of the target object, input the region of interest features of the target object into a preset detection head, perform classification and regression processing of the 3D target detection box of the target object, and determine the detection model loss of the 3D target detection box of the target object.
[0132] In step S620, the loss of the detection model containing the 3D target detection box of the target object is variable; naturally, the smaller the loss of the detection model containing the 3D target detection box of the target object, the better. As the loss of the detection model containing the 3D target detection box of the target object decreases, the 3D target detection box of the target object is determined, thereby achieving 3D target detection of the target object.
[0133] The loss of the detection model containing the 3D target detection box of the target object includes: region proposal network loss and refinement stage loss. The region proposal network loss includes: confidence loss of the 3D target detection box of the target object, and position regression loss of the 3D target detection box of the target object.
[0134] like Figure 8As shown, the implementation flow of a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud provided by the present invention includes:
[0135] Step 6: Use the final region of interest features to predict the target category and regress the bounding box location, specifically including the following exemplary steps.
[0136] Step 61: Extract the final region of interest features. Input the detection head to classify and regress bounding boxes.
[0137] Step 62: The model's loss is divided into region proposal network loss. and the loss in the refinement stage Two parts, of which Including the confidence loss of the box and position regression loss .
[0138] The encoding format of the box is ,in , , The center point of the box , , , These represent the width, length, height, and orientation angle of the bounding box, respectively. The error between the ground truth bounding box and the candidate bounding box position. ;
[0139]
[0140] Among them, subscript Indicates the parameters of the ground truth bounding boxes in the training set, subscript Indicates the candidate box parameters. .
[0141] Step 63: Suggest network loss for the region. Calculate confidence loss using Focal Loss (i.e., the focus loss function). This is to balance the contribution of positive and negative samples to the loss.
[0142]
[0143] in, To predict confidence levels, This is the actual label value.
[0144] Step 63: Box position regression loss The loss function is calculated using the Smooth-L1 loss function.
[0145]
[0146] in, This represents the predicted residual value of the bounding box. To calculate the residual value of the predicted box's distance from the true box position, only positive samples are used to calculate the box position loss.
[0147] Step 64: Finally, obtain the total regional suggested network loss. loss:
[0148]
[0149] in, and These are the weighting coefficients for the loss, used to balance the classification and regression pairs. The extent of their contribution.
[0150] Similarly, refine the stage losses. Calculation method and regional suggestions for network loss Similarly, the final model loss is obtained. L loss as follows:
[0151]
[0152] To verify the effectiveness of the 3D target detection algorithm based on the original point cloud spatial attention mechanism proposed in this invention, it was validated using the publicly available autonomous driving dataset KITTI, and extensive ablation experiments were conducted. Experiments were performed on targets of three difficulty levels—easy, medium, and hard—on both the validation and test sets, and the average accuracy (AP) was used to measure the model performance.
[0153] Figure 13 This is a schematic diagram illustrating the detection effect of a 3D object detection algorithm based on the original point cloud spatial attention mechanism according to an embodiment of the present invention in multiple scenarios, wherein (a) is the detection in the first scenario.
[0154] The diagram shows the detection results in two scenarios: (b) is the detection result in the second scenario, and (c) is the detection result in the third scenario. Figure 13 To demonstrate the actual detection performance of the algorithm proposed in this invention, the KITTI autonomous driving dataset was used for testing.
[0155] Figure 14This diagram illustrates the comparison of the detection performance of a 3D target detection algorithm based on the spatial attention mechanism of the original point cloud in this invention with other related algorithms. (a) shows the detection performance of the SECOND algorithm (i.e., a target detection algorithm based on three-dimensional point clouds), (b) shows the detection performance of the PointPillars algorithm (i.e., a 3D target detection algorithm based on laser point clouds), and (c) shows the detection performance of the 3D target detection algorithm based on the spatial attention mechanism of the original point cloud. Figure 14 To compare the detection performance of the algorithm proposed in this invention with other mainstream classic algorithms, the visualization results show that the SECOND algorithm and the PointPillar algorithm have different degrees of false detection. For example, the point cloud on the left wall is more complex from the perspective of a BEV, causing the SECOND algorithm and the PointPillar algorithm to misdetect it as a car. However, the algorithm proposed in this invention shows better robustness, with a lower false recognition rate for complex targets, and achieves good experimental results.
[0156] The solution of this invention can effectively improve the detection performance of difficult-to-detect targets such as occluded objects in original point cloud scenes. Experiments were conducted on the publicly available 3D object detection dataset KITTI using the two-stage 3D object detection algorithm based on the original point cloud mesh attention mechanism proposed in this invention. The results show that the model proposed in this invention significantly improves the detection accuracy compared to other publicly available point cloud-based 3D object detection algorithms. Furthermore, the two-stage 3D object detection algorithm based on the original point cloud mesh attention mechanism proposed in this invention achieved competitive detection results in public testing on the official KITTI test set.
[0157] KITTI is a publicly available dataset for autonomous driving in related schemes, and is one of the most important datasets in the field of autonomous driving. It contains real images and point cloud data collected in scenarios such as urban areas, rural areas, and highways. The dataset contains 7481 training samples and 7518 test samples. For details, please refer to Tables 1 and 2 for some of the experimental data.
[0158] Table 1 compares the detection performance of vehicles on the KITTI test set with state-of-the-art methods. All results are calculated using the average accuracy at a 0.7 IoU threshold and R40 recall locations.
[0159]
[0160] Table 2 compares the vehicle detection performance with state-of-the-art methods on the KITTI validation set. All results are calculated using the mean accuracy at a 0.7 IoU threshold and R11 recall locations.
[0161]
[0162] The technical solution of this embodiment acquires laser point cloud data containing the target object as the original 3D point cloud data. After voxelization and 3D sparse convolutional layer processing of the original 3D point cloud data, the region of interest is extracted. Based on the region of interest, farthest point sampling and spatial grid encoding are performed to obtain feature points of interest. Then, the feature of the region of interest is used to predict the target category and regress the bounding box position, thereby realizing 3D target detection. Thus, by setting a two-stage 3D target detection algorithm based on the attention mechanism of the original point cloud grid, the position detection accuracy and detection efficiency of 3D target detection can be improved, which is conducive to improving the detection effect of 3D target detection.
[0163] According to embodiments of the present invention, a 3D target detection apparatus corresponding to a 3D target detection method is also provided. See also Figure 7 The diagram shows a structural schematic of an embodiment of the device of the present invention. The 3D target detection device may include an acquisition unit and a detection unit.
[0164] The acquisition unit 102 is configured to acquire laser point cloud data containing the target object, as the original three-dimensional point cloud data of the target object. The specific functions and processing of the acquisition unit 102 are described in step S110, and will not be repeated here.
[0165] The detection unit 104 is configured to perform voxelization, 3D sparse convolutional layer, and RPN network processing on the original 3D point cloud data of the target object to obtain the region of interest in the original 3D point cloud data of the target object, which is then used as the region of interest of the target object. The specific functions and processing of the detection unit 104 are described in step S120, and will not be repeated here.
[0166] The detection unit 104 is further configured to perform farthest point sampling and spatial meshing based on the region of interest of the target object to obtain local features of the center point of the target object. The specific functions and processing of this detection unit 104 are described in step S130, and will not be repeated here.
[0167] The detection unit 104 is further configured to perform coordinate dimensionality upscaling and feature summation processing based on the local features of the center point of the target object to obtain the region of interest features of the target object. The specific functions and processing of this detection unit 104 are described in step S140, and will not be repeated here.
[0168] The detection unit 104 is further configured to perform target category prediction and bounding box position regression processing based on the region of interest features of the target object, thereby realizing 3D target detection of the target object. The specific functions and processing of the detection unit 104 are described in step S150, and will not be repeated here.
[0169] This invention proposes a two-stage 3D object detection device based on an attention mechanism of original point cloud meshes. By expanding the receptive field, aggregating multi-scale local features, and performing refined modeling of point coordinates, it fully preserves the spatial information of points and considers the complex relationship between mesh points and regions of interest to improve detection accuracy. The receptive field is the size of the region mapped back to the input image from the pixels in the feature map output by each layer of the convolutional neural network. This solves the problem that in related 3D object detection schemes, voxel-based methods suffer from poor position detection accuracy, while point-based methods suffer from low detection efficiency, resulting in poor detection performance.
[0170] Since the processing and functions implemented by the device in this embodiment are basically the same as the embodiments, principles and examples of the aforementioned methods, any details not covered in the description of this embodiment can be found in the relevant descriptions in the aforementioned embodiments, and will not be repeated here.
[0171] The technical solution of this invention acquires laser point cloud data containing the target object as the original 3D point cloud data. After voxelization and 3D sparse convolutional layer processing of the original 3D point cloud data, the region of interest is extracted. Based on the region of interest, farthest point sampling and spatial grid encoding are performed to obtain feature points of interest. Then, the feature of the region of interest is used to predict the target category and regress the bounding box position, thereby realizing 3D target detection of the target object. This solves the problems of poor position detection accuracy of voxel-based 3D target detection methods and low detection efficiency of point-based 3D target detection methods. It has good detection accuracy and fast detection speed.
[0172] According to an embodiment of the present invention, a terminal corresponding to a 3D target detection device is also provided. This terminal may include the 3D target detection device described above.
[0173] Since the processing and functions implemented by the terminal in this embodiment are basically the same as the embodiments, principles and examples of the aforementioned device, any details not covered in this embodiment can be found in the relevant descriptions in the aforementioned embodiments, and will not be repeated here.
[0174] The technical solution of this invention obtains laser point cloud data containing the target object as the original three-dimensional point cloud data. After voxelization and 3D sparse convolutional layer processing of the original three-dimensional point cloud data, the region of interest is extracted. Based on the region of interest, farthest point sampling and spatial grid encoding are performed to obtain feature points of interest. Then, the feature of the region of interest is used to predict the target category and regress the bounding box position, thereby realizing 3D target detection of the target object. The detection accuracy is high and the detection process is relatively simple.
[0175] According to an embodiment of the present invention, a storage medium corresponding to a 3D target detection method is also provided, the storage medium including a stored program, wherein the program controls the device where the storage medium is located to execute the 3D target detection method described above when it is running.
[0176] Since the processing and functions implemented by the storage medium in this embodiment are basically the same as the embodiments, principles and examples of the aforementioned methods, any details not covered in this embodiment can be found in the relevant descriptions in the aforementioned embodiments, and will not be repeated here.
[0177] The technical solution of this invention obtains laser point cloud data containing the target object as the original three-dimensional point cloud data. After voxelization and 3D sparse convolutional layer processing of the original three-dimensional point cloud data, the region of interest is extracted. Based on the region of interest, farthest point sampling and spatial grid encoding are performed to obtain feature points of interest. Then, the feature of the region of interest is used to predict the target category and regress the bounding box position, thereby realizing 3D target detection of the target object. The false recognition rate is low for complex targets and the recognition efficiency is high.
[0178] In summary, it is readily understood by those skilled in the art that, without conflict, the aforementioned advantageous methods can be freely combined and superimposed.
[0179] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A 3D object detection method, characterized in that, The method comprises the following steps: acquiring laser point cloud data containing a target object as original three-dimensional point cloud data of the target object; based on the original three-dimensional point cloud data of the target object, voxelization, 3D sparse convolution layer and region proposal network processing are performed to obtain a region of interest in the original three-dimensional point cloud data of the target object as a region of interest of the target object; based on the region of interest of the target object, farthest point sampling and spatial gridding processing are performed to obtain a center point local feature of the target object; based on the center point local feature of the target object, coordinate dimension lifting and feature addition processing are performed to obtain a region of interest feature of the target object; comprising: based on the center point local feature of the target object, a 3-layer feedforward neural network is used to lift the dimension of the center point coordinates of the target object, and a max-pooling function is used to aggregate different radius scale features of the center point local feature of the target object; using a feedforward neural network, the dimension of the center point local feature of the target object after the lifting and the aggregation is adjusted, the position encoding feature and the different radius scale local feature of the center point local feature of the target object are added, and the center point feature of all the interest grids of the target object is obtained; based on the center point feature of the interest grid of the target object, an attention mechanism is used to capture the dependency relationship between different interest grids in the center points of all the interest grids of the target object, and according to the dependency relationship, corresponding weights are assigned to the center point features of different interest grids in the center points of all the interest grids of the target object, so as to obtain the association relationship between the center point features of all the interest grids of the target object and the region of interest of the target object; based on the association relationship between the center point features of all the interest grids of the target object and the region of interest of the target object, a multi-head attention mechanism is used to determine the region of interest feature of the target object; based on the region of interest feature of the target object, target category prediction and frame position regression processing of the target object are performed to realize 3D target detection of the target object.
2. The 3D object detection method of claim 1, wherein, Wherein, acquiring laser point cloud data containing a target object as original three-dimensional point cloud data of the target object comprises: acquiring laser point cloud data of the target object collected by a laser radar as original three-dimensional point cloud data of the target object; wherein, the value range of the original three-dimensional point cloud data of the target object in the X-axis direction is a first set range, the value range in the Y-axis direction is a second set range, and the value range in the Z-axis direction is a third set range; and / or, based on the original three-dimensional point cloud data of the target object, voxelization, 3D sparse convolution layer and region proposal network processing are performed to obtain a region of interest in the original three-dimensional point cloud data of the target object as a region of interest of the target object, comprising: based on the original three-dimensional point cloud data of the target object, voxelization processing is performed to obtain three-dimensional voxels of the original three-dimensional point cloud data of the target object; Based on the three-dimensional voxel of the original three-dimensional point cloud data of the target object, feature extraction is performed using 4 layers of 3D sparse convolution to obtain a plurality of scale space features of the original three-dimensional point cloud data of the target object; Based on the plurality of scale space features of the original three-dimensional point cloud data of the target object, after the plurality of scale space features of the last layer of the 4 layers are compressed in perspective, a region of interest is extracted using a region proposal network to obtain a region of interest in the original three-dimensional point cloud data of the target object as the region of interest of the target object.
3. The 3D object detection method of claim 1, wherein, Based on the region of interest of the target object, farthest point sampling and spatial gridding processing are performed to obtain the center point local feature of the target object, including: Based on the region of interest of the target object, the region of interest of the target object is divided into a column structure; inside the column structure, farthest point sampling is performed on the original three-dimensional point cloud data of the target object to obtain interest points in the region of interest of the target object as interest sampling points of the target object; Based on the region of interest of the target object, the region of interest of the target object is uniformly gridded to obtain a plurality of region of interest grids of the target object as a plurality of interest grids of the target object; The center point of each interest grid of the target object is determined, and the relative distance between the center point of each interest grid of the target object and the interest sampling point of the target object is determined; Based on the relative distance between the center point of each interest grid of the target object and the interest sampling point of the target object, the local feature of the center point of all interest grids of the target object is determined.
4. The 3D object detection method of claim 3, wherein, Wherein, Based on the region of interest of the target object, the region of interest of the target object is divided into a column structure, including: Based on the region of interest of the target object, the region of interest of the target object is set as a cylinder, and the column structure in which the cylinder is located is taken as the column structure after the region of interest of the target object is divided; wherein the radius of the bottom of the cylinder is r , , the height of the cylinder is h , ; wherein, are the width, length and height of the region of interest, respectively, and are the set cylinder dilation parameters; And / or, Based on the relative distance between the center point of each interest grid of the target object and the interest sampling point of the target object, the local feature of the center point of all interest grids of the target object is determined, including: Based on the relative distance between the center point of each interest grid of the target object and the interest sampling point of the target object, the spatial position of the corresponding interest grid of the target object is modeled and uniformly position coded coordinate scaled to obtain the position feature of the center point of each interest grid of the target object; Based on the position feature of the center point of each interest grid of the target object, the center point of each interest grid of the target object is processed by dimension lifting to obtain a position feature set of the center point of all interest grids of the target object within a set radius in a set spherical region; Based on the position feature set of the center point of all interest grids of the target object within the set radius in the set spherical region, a feature expression set of the center point of all interest grids of the target object at different radius scales is obtained by changing the radius of the spherical body to which the set spherical region belongs. Based on the feature expression set of the center points of all the interest grids of the target object at different radius scales, the features at different radius scales are spliced to obtain the local features of the center points of all the interest grids of the target object.
5. The 3D object detection method of claim 4, wherein, Based on the relative distances between the center points of each interest grid of the target object and the interest sampling points of the target object, the spatial positions of the corresponding interest grids of the target object are modeled and processed by uniform position coding coordinate scale to obtain the position features of the center points of each interest grid of the target object, including: Based on the relative distances between the center points of each interest grid of the target object and the interest sampling points of the target object, the position features of the center points of each interest grid of the target object are calculated according to the following formula: ; wherein, is a position feature of a center point of each interest grid of the target object, is a feature transformation function using a feedforward neural network to map the relative distance feature into a high-dimensional feature space, is a relative distance of an interest sampling point of the target object to a center point of each interest grid of the target object, is an additional feature of the interest sampling point of the target object.
6. The 3D object detection method of any one of claims 1-5, wherein, Based on the region of interest features of the target object, target category prediction and frame position regression processing of the target object are performed to realize 3D target detection of the target object, including: Based on the region of interest features of the target object, the region of interest features of the target object are input into a preset detection head to perform classification and regression processing of the 3D target detection frame of the target object, and determine the detection model loss of the 3D target detection frame of the target object; With the decrease of the detection model loss of the 3D target detection frame of the target object, the 3D target detection frame of the target object is determined to realize 3D target detection of the target object. The detection model loss of the 3D target detection frame of the target object includes: region proposal network loss and refinement stage loss; the region proposal network loss includes: confidence loss of the 3D target detection frame of the target object, and position regression loss of the 3D target detection frame of the target object.
7. A 3D object detection apparatus, characterized by comprising: including: An acquisition unit configured to acquire laser point cloud data containing a target object as original three-dimensional point cloud data of the target object; A detection unit configured to perform voxelization, 3D sparse convolution layer, and region proposal network processing based on the original three-dimensional point cloud data of the target object to obtain a region of interest in the original three-dimensional point cloud data of the target object as a region of interest of the target object; The detection unit is further configured to perform farthest point sampling and spatial gridding processing based on the region of interest of the target object to obtain center point local features of the target object. The detection unit is further configured to perform coordinate dimension lifting and feature addition processing based on the center point local features of the target object to obtain region of interest features of the target object; including: Based on the center point local features of the target object, a 3-layer feedforward neural network is used to lift the coordinates of the center points of the target object, and a max-pooling function is used to aggregate different radius scale features of the center point local features of the target object; A feedforward neural network is used to adjust the dimensions of the center point local features of the target object after the lifting and the aggregation, add the position coding features and different radius scale local features of the center point local features of the target object, and obtain the center point features of all the interest grids of the target object. Based on the center point features of the target object's interest grids, the attention mechanism is used to capture the dependency between different center points of the interest grids, and the corresponding weights are assigned to the center point features of different interest grids according to the dependency, so as to obtain the association between the center point features of all interest grids of the target object and the target object's interest region; Based on the association between the center point features of all interest grids of the target object and the target object's interest region, the multi-head attention mechanism is used to determine the target object's interest region feature. The detection unit is further configured to perform target class prediction and bounding box position regression of the target object based on the target object's interest region feature, and realize 3D target detection of the target object.
8. A terminal, characterized by comprising: Comprise: The 3D target detection device of claim 7.
9. A storage medium, characterized by The storage medium comprises a stored program, wherein the program controls the device where the storage medium is located to execute the 3D target detection method of any one of claims 1 to 6 when the program is running.
Citation Information
Patent Citations
Multimode data fusion-based three-dimensional target detection method
CN112347987A
Three-dimensional target detection method based on weighted sampling and multi-resolution feature extraction
CN113468994A
Target detection method and device based on pseudo point cloud, electronic equipment and storage medium
CN114842313A