Detection method based on end-to-end single-stage 3D target detection model

By introducing the DRF-SSD model into the point-based 3D object detection model, the PointNet++ backbone network and the hierarchical encoder-decoder module are used to solve the problem of discontinuous receptive fields, the extraction ability of object semantics and geometric information is improved, and the 3D object detection accuracy is achieved.

CN119942068APending Publication Date: 2025-05-06SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411966211.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing point-based 3D object detection method ignores information that does not occupy space, resulting in discontinuous receptive field problems, resulting in the loss of object semantics and geometric information.

Method used

An end-to-end single-stage 3D object detection model DRF-SSD is proposed, adopting a 3D backbone network in the PointNet++ style, and a hierarchical encoder-decoder module and a hybrid deep learning module are designed. Through these modules, point cloud features are extracted and integrated, candidate areas are generated and target objects are located.

Benefits of technology

It effectively solves the problem of discontinuous receptive fields, improves the extraction ability of object semantics and geometric information, improves 3D average accuracy (mAP3D), and performs better than many structure-based methods on the KITTI dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942068A_ABST
    Figure CN119942068A_ABST
Patent Text Reader

Abstract

The invention provides a detection method based on an end-to-end single-stage 3D target detection model. According to the model, a 3D trunk of a Point Net + + style is adopted to maintain high-speed reasoning, point-by-point features are projected to a plane in a Neck structure through designed hierarchical encoding-decoding (HED) and hybrid transformer (HT) modules, and local and global information is aggregated. The HED fills unoccupied space features through a convolutional layer, and local feature learning is enhanced; and the HT expands a receptive field by using the global learning ability of the transformer. The modules only process key points, keep the sparsity of feature maps and ensure the reasoning efficiency. And finally, the query features are reversely projected to the points, and after the features are enhanced, the features are input into a detection head for prediction. Experiments on a KITTI data set show that the 3D average precision (AP3D of the DRF-SSD is remarkably improved compared with that of a conventional method and is improved by 2.25%, 0.66% and 0.42% under simple, medium and difficult settings respectively, and the effectiveness of the method and the substantial improvement of other point-based detectors are proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to point-based 3D target detection, which is a method for solving the problem of discontinuous receptive field and the loss of semantic and geometric information of objects. Background Art

[0002] Point-based 3D object detection methods have significant advantages in fields such as intelligent transportation and autonomous driving due to their lightweight and fast inference speed. However, current advanced methods mainly focus on learning features from point clouds, ignoring the information of unoccupied spaces. This leads to the discontinuous receptive field (DRF) problem, resulting in the loss of semantic and geometric information of objects.

[0003] In LiDAR-based 3D object detection, the perception task is crucial because it provides information about the size and distance of objects. Detection models are mainly divided into structured and point-based methods. Structured methods learn dense features by stacking convolutional layers, which are suitable for sparse point clouds, but have high computational overhead and affect real-time performance. In contrast, point-based methods have significantly improved inference speed, such as IA-SSD and DBQ-SSD, which achieve high frame rates on the KITTI dataset. However, these methods can only extract features from scanned points and cannot obtain information from unoccupied space, resulting in insufficient local feature extraction, high uncertainty in object box regression, and limited ability to extract semantic information, which are attributed to discontinuous receptive fields. Summary of the invention

[0004] In view of the shortcomings of the prior art, the present invention provides a detection method based on an end-to-end single-stage 3D target detection model, which is a method for solving the problem of discontinuous receptive fields and the loss of object semantics and geometric information. It is used to solve the loss of object semantics and geometric information caused by discontinuous receptive fields. The present invention proposes a new end-to-end single-stage point-based model, called DRF-SSD. DRF-SSD adopts a PointNet++ style 3D backbone network to maintain fast reasoning capabilities.

[0005] The technical solution adopted by the present invention to achieve the above-mentioned purpose is: a detection method based on an end-to-end single-stage 3D object detection model, comprising the following steps:

[0006] 1) For the 3D point cloud collected by LiDAR, a feature map of the point-based backbone network is established;

[0007] 2) construct a hierarchical encoder-decoder module to extract and compress features to generate semantic information;

[0008] 3) Voting on point cloud features to generate candidate regions to locate the target object; Voting on point cloud features to generate candidate regions to locate the target object, and connecting features through a hybrid deep learning module to determine the prediction result;

[0009] 4) Obtain an end-to-end single-stage 3D object detection model through steps 1) to 3) and perform training;

[0010] 5) Collect 3D point clouds in real time and input them into the trained end-to-end single-stage 3D object detection model to obtain the detected object.

[0011] The feature graph of the point-based backbone network is represented as follows:

[0012]

[0013] Among them, i represents the stage of feature learning; Sampling is the downsampling operation; Group is the local feature aggregation, and Pool is the pooling layer; E θ (·) denotes the feature embedding layer.

[0014] The step of establishing a feature map of a point-based backbone network and sampling the sampling points comprises the following steps:

[0015] 1) Yes Use the farthest sampling downsampling method; Represent the feature maps of the first and second feature learning stages respectively;

[0016] 2) Yes Adopting the foreground point downsampling method with semantic embedding, following the IA-SSD method; Represent the feature maps of the 3rd and 4th feature learning stages respectively;

[0017] 3) Sampling points for the next stage It is expressed as follows:

[0018]

[0019] Among them, topK represents the first K sampling points, S represents the semantic extraction network, It is the feature map of the backbone network.

[0020] Constructing a layered encoder-decoder module involves the following steps:

[0021] 1) Based on the feature map output by the backbone network, extract the key point p3 and its features And input it into the voxel feature encoding module for normalization:

[0022]

[0023] Among them, u k 、v k represents the normalized two-dimensional coordinates, i.e., the voxelized output result of the key point p3, x min ,y min 、x max ,y nax They represent the minimum x value, the minimum y value, the maximum x value, and the maximum y value, respectively, and are used to represent the effective range of the plane projection of the 3D point cloud; W and H are the width and height of the feature map, respectively. k ,y k is the horizontal and vertical coordinates of the k-th frame point cloud projection;

[0024] 2) Each point has a known feature Collect the current feature map grid R through maximum pooling h,w Characteristics of the midpoint Form the corresponding 2D mesh features

[0025]

[0026] in, Indicates that the features projected in the same grid are processed along the channel for maximum pooling; k is the grid number; c represents the maximum value of each channel;

[0027] 3) For characterizing 2D mesh features Feature map Downsampling is performed to make the effective feature distance that was originally larger than the convolution kernel size closer to the kernel size;

[0028] Using the submanifold convolution layer, the effective features in the feature map are upsampled to maintain the same sparsity as the feature map before downsampling, and multi-scale feature values ​​are output.

[0029] The voting process of the point cloud features to generate candidate regions to locate the target object and the feature connection through the hybrid deep learning module to determine the prediction result includes the following steps:

[0030] (1) Construct a voting module to predict the target center p to which the key point p3 extracted based on the backbone network output feature map belongs v :

[0031]

[0032] Among them, Vote means voting, represents the feature map of the fourth feature learning stage, offset represents the voting result, p4 represents the feature point of the fourth feature learning stage, and n4 represents the number of sampling points in the fourth feature learning stage;

[0033] The voting point p v The position of is regarded as the RoI of the current scene, and then projected to the multi-scale feature value output by the hierarchical encoder-decoder module;

[0034] (2) Set the polling point p v The projected coordinates of are embedded into the object’s location query through the self-attention module to obtain local features

[0035]

[0036] Among them, Pool means pooling, Conv means convolution, and Padding means local area selection. It is the feature map of the third feature learning stage;

[0037] (3) Calculate through the self-attention module to obtain the 2D feature projection:

[0038]

[0039] Where m represents the index of the attention head, M represents the number of attention heads; W m and W′ m is a learnable weight; A mij is the attention weight; The value is used as the feature value q of the query i,j ,q i ,q j Represent the row and column values ​​of the eigenvalues ​​respectively;

[0040] (4) Use bilinear interpolation to project the 2D features back to the point:

[0041]

[0042] w i,j,k =(1-|u k -(|u k |+i)|)(1-|v k -(|v k |+j)|)

[0043] Among them, w i,j,k represents the bilinear interpolation weight, i represents the number of row values, j represents the number of column values, k is the grid number, It is an integrated feature map that integrates the multi-scale features of the hierarchical encoder-decoder module; |uk |+i and |v k |+j is used to limit the projection grid coordinates;

[0044] (5) Local features Attention results Q, return points These three features are connected together to determine the prediction results:

[0045]

[0046] Where Q = MultiHeadAttn(q i ).

[0047] The training is performed in an end-to-end multi-task manner, which includes the following steps:

[0048] 1) Use weighted cross entropy loss to manually emphasize rare categories, weighted cross entropy loss as follows:

[0049]

[0050] Among them, ∈ represents the hyperparameter, a c represents the weight of the cth class, y c Represents the true label determined by whether the point is within the bounding box. Represents the predicted probability, and C represents the number of target categories;

[0051] 2) Calculate the total loss

[0052]

[0053] in, They are sample loss, voting module loss, classification prediction loss, and voting point feature regression loss;

[0054] They represent position loss, size loss, angle range loss, angle resolution loss, and corner component loss respectively.

[0055] A detection system based on an end-to-end single-stage 3D object detection model, comprising:

[0056] The target detection model building module is used to build a feature map of the point-based backbone network for the three-dimensional point cloud collected by the lidar; build a hierarchical encoder-decoder module to extract and compress features to generate semantic information; perform voting on point cloud features to generate candidate regions to locate target objects; perform voting on point cloud features to generate candidate regions to locate target objects, and connect features through a hybrid deep learning module to determine the prediction results; and train the resulting end-to-end single-stage 3D target detection model;

[0057] The target detection module is used to collect 3D point clouds in real time and input the trained end-to-end single-stage 3D target detection model to obtain the detected target.

[0058] A detection device based on an end-to-end single-stage 3D object detection model comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement the detection method based on the end-to-end single-stage 3D object detection model when executing the computer program.

[0059] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a detection method based on an end-to-end single-stage 3D object detection model is implemented.

[0060] The present invention has the following beneficial effects and advantages:

[0061] 1. This paper identifies the discontinuous receptive field (DRF) problem in current point-based detectors and proposes the first model that attempts to solve this problem using a point-based backbone network;

[0062] 2. In the KITTI ranking, DRF-SSD of the present invention improves the 3D average precision (mAP3D) by 1.1% over the previous model, surpassing many structure-based methods and improving the safety baseline and application prospects of point-based detectors.

[0063] 3. The HED and HT modules designed by the present invention can effectively extract non-occupied space information from raw point cloud data, demonstrating excellent plug-and-play capabilities. They also provide significant performance improvements for other point-based models. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a schematic diagram of the overall framework of the method of the present invention;

[0065] Figure 2 is a schematic diagram of the design of a current single-stage 3D object detector according to the method of the present invention;

[0066] Figure 3 is a schematic diagram of a layered encoder-decoder module of the method of the present invention;

[0067] Figure 4 Schematic diagram of feature conversion learning pipeline of the method of the present invention. DETAILED DESCRIPTION

[0068] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0069] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0070] Point-based 3D object detection methods have significant advantages in fields such as intelligent transportation and autonomous driving due to their lightweight and fast inference speed. However, existing methods mainly focus on feature learning of point clouds and ignore the importance of unoccupied space, which leads to the discontinuous receptive field (DRF) problem, which in turn affects the integrity of the semantic and geometric information of objects. To solve this problem, this paper proposes a new end-to-end single-stage model DRF-SSD. The model adopts a PointNet++-style 3D backbone to maintain high-speed inference. The designed hierarchical encoding-decoding (HED) and hybrid transformer (HT) modules project point-by-point features to the plane in the Neck structure to aggregate local and global information. HED fills the unoccupied space features through convolutional layers to enhance local feature learning; HT uses the global learning ability of transformers to expand the receptive field. These modules only process key points, maintain the sparsity of feature maps, and ensure inference efficiency. Finally, the query features are back-projected onto the points, and the enhanced features are input into the detection head for prediction. Experiments on the KITTI dataset show that DRF-SSD achieves significant improvements over previous methods in 3D average precision (AP3D), by 2.25%, 0.66%, and 0.42% in simple, medium, and difficult settings, respectively, demonstrating the effectiveness of the approach and its substantial improvement over other point-based detectors.

[0071] like Figure 1 As shown, the present invention comprises the following steps:

[0072] Step 1: Establish a feature map based on the point-based backbone network work, according to the formula:

[0073]

[0074] Where i represents the stage of feature learning; Sampling is the downsampling operation; Group is the local feature aggregation, and Pool is the pooling layer; E θ (·) denotes the feature embedding layer, which expands the feature space of the point cloud. As the number of learning stages increases, the query radius of the local feature extractor becomes larger to increase the receptive field of the model.

[0075] like Figure 2As shown, following previous point-based design work, DRF-SSD follows the previous point-based work design in backbone design, as shown in Figure 3 The location of the sampling points, the number of sampling points, and the point-by-point features generated at each stage are represented as {p1, p2, p3, p4}, {n1, n2, n3, n4}, respectively. Before each stage, a multi-scale point network is used to extract local features, including three steps: point indexing, feature learning, and feature aggregation. The farthest sampling (FPS) downsampling method is used to The foreground point downsampling method with semantic embedding is adopted, following the IA-SSD method, and the sampling points of the next stage are expressed as follows:

[0076]

[0077] By using the above 3D backbone, an extremely sparse set of keypoints (usually 256 points) and their locally rich spatial and semantic information are obtained.

[0078] Step 2: Design a hierarchical encoder-decoder module such as Figure 3 Extract key points p3 and their features from the 3D backbone network output Then, they are voxelized into pillars and input into the voxel feature encoding module to normalize their features.

[0079]

[0080] where x min ,y min 、x max ,y max is the effective range of the point cloud; W and H are the dimensions of the feature map;

[0081] Then, the current feature map grid R is collected by max pooling h,w Characteristics of the midpoint Form the corresponding 2D mesh features

[0082]

[0083] For feature maps Downsampling is performed so that the effective feature distance that was originally larger than the convolution kernel size is closer to the kernel size. Then, submanifold convolution is used to maintain the sparsity of the feature map for feature interaction. In this way, the extensive use of conventional sparse convolution layers is avoided, thereby maintaining the sparsity of the feature map while allowing interactions between local features. After passing through the submanifold convolution layer, the valid features in the feature map are upsampled to maintain the same sparsity as the feature map before downsampling. In order to recover the detailed features lost during the downsampling process, a conventional sparse convolution layer is placed at the beginning of each ED block to obtain multi-scale feature representation, followed by a multi-layer encoder-decoder module for feature fusion. Finally, sparse structure-based features are obtained from the original point-by-point feature transformation. As Figure 4 shown.

[0084] The unoccupied areas are filled with features through sparse convolution, and the detailed features are restored through the multi-scale fusion encoder-decoder structure. However, the scope of these operations is limited and aims to enhance local features. Therefore, the Hybridtransformer module can make up for the shortcomings of the model in this regard.

[0085] Step 3: The hybrid deep learning module uses the Hybrid transformer module, as follows:

[0086] For Transformer-based detectors, the most important step is to gradually update and learn rich content queries. and location query The HT module constructs queries using the results output by the previous model. As mentioned before, the point-based 3D backbone can effectively extract the geometric information of point clouds.

[0087] A voting module is designed to predict the target center to which the key point belongs, according to the formula:

[0088]

[0089] Query the content As a voting point v The position of is considered as the RoI of the current scene, and then projected onto the feature map output by the Hierarchical encoder-decoder (HED) module. Then, a multi-layer perceptron (MLP) is used to embed the projected coordinates into the location query of the object. For content queries, a self-attention module is set for the voting points to learn their local features, according to the formula:

[0090]

[0091] After obtaining the object query, self-attention calculation can be performed according to the traditional structure.

[0092]

[0093] Where m represents the index of the attention head; W m With W m ′ is a learnable weight; A mij is the attention weight; The value is used as the feature value q of the query i,j ,q i ,q j Represent the row and column values ​​of the eigenvalues ​​respectively.

[0094] Globally, the features of a voting point are used to predict the properties of its regressed bounding box.

[0095] In the local area, bilinear interpolation is used to project 2D features back to the point.

[0096]

[0097] w i,j,k =(1-|u k -(|u k |+i)|)(1-|v k -(|v k |+j)|)

[0098] Among them, w i,j,k represents the bilinear interpolation weight, i represents the number of row values, j represents the number of column values, k is the grid number, It is an integrated feature map that integrates the multi-scale features of the hierarchical encoder-decoder module; |u k |+i and |v k |+j is used to limit the projection grid coordinates.

[0099] Local features Attention results Q, return points These three features are connected together to determine the prediction results:

[0100]

[0101] Where Q = MultiHeadAttn(q i ).

[0102] Step 4: The proposed DRF-SSD is trained in an end-to-end multi-task manner. For the semantic loss, the weighted cross entropy (WCE) loss is used to manually emphasize rare categories. The WCE loss can be formulated as:

[0103]

[0104] Among them, ∈ represents the hyperparameter, a c represents the weight of the cth class, y c Represents the true label determined by whether the point is within the bounding box. represents the predicted probability, and C represents the number of target categories.

[0105] In order to calculate the semantic loss, each point is assigned a weight, which is determined by the formula:

[0106]

[0107] Except for the sampling loss in the backbone, all other losses are calculated in the detection head. First, the voting point loss Calculated as the L1 loss between the foreground point and its corresponding object center; followed by the bounding box classification prediction loss It is calculated using the weighted cross entropy (WCE) loss. Next is the voting point feature regression loss By incorporating an additional intersection over union (IoU) prediction branch, missed detection and false detection are resolved. All losses involved in the model are as follows:

[0108]

[0109] in, They are sample loss, voting module loss, classification prediction loss, and voting point feature regression loss; They represent position loss, size loss, angle range loss, angle resolution loss, and corner component loss respectively.

[0110] Through hybrid deep learning modules, multiple features are combined, feature representation is enhanced through the Transformer mechanism, and local and global information are integrated;

[0111] Finally, the outputs of the hybrid deep learning module (HT module) and the HED module are fused and concatenated with the output of the voting module to finally output the target detection results through the detection head.

[0112] Through the above process, an end-to-end single-stage point-based model DRF-SSD is constructed and model training is performed; the sample set is pictures containing targets collected through dense beam lidar scene acquisition, and the labels are target categories;

[0113] In practical applications, the images collected by the lidar in real time are directly input into DRF-SSD to obtain the detected target location and category.

Claims

1. A detection method based on an end-to-end single-stage 3D object detection model, characterized in that: The following steps are involved: 1) For the 3D point cloud collected by LiDAR, a feature map of the point-based backbone network is established; 2) construct a hierarchical encoder-decoder module to extract and compress features to generate semantic information; 3) Voting on point cloud features to generate candidate regions to locate the target object; Voting on point cloud features to generate candidate regions to locate the target object, and connecting features through a hybrid deep learning module to determine the prediction result; 4) Obtain an end-to-end single-stage 3D object detection model through steps 1) to 3) and perform training; 5) Collect 3D point clouds in real time and input them into the trained end-to-end single-stage 3D object detection model to obtain the detected object.

2. The detection method based on an end-to-end single-stage 3D object detection model according to claim 1, characterized in that: The feature graph of the point-based backbone network is represented as follows: Among them, i represents the stage of feature learning; Sampling is the downsampling operation; Group is the local feature aggregation, and Pool is the pooling layer; E θ (·) denotes the feature embedding layer.

3. The detection method based on an end-to-end single-stage 3D object detection model according to claim 2, characterized in that: The step of establishing a feature map of a point-based backbone network and sampling the sampling points comprises the following steps: 1) Yes Use the farthest sampling downsampling method; Represent the feature maps of the first and second feature learning stages respectively; 2) Yes Adopting the foreground point downsampling method with semantic embedding, following the IA-SSD method; Represent the feature maps of the 3rd and 4th feature learning stages respectively; 3) Sampling points for the next stage It is expressed as follows: Among them, topK represents the first K sampling points, S represents the semantic extraction network, It is the feature map of the backbone network.

4. The detection method based on an end-to-end single-stage 3D object detection model according to claim 1, characterized in that ,Constructing the hierarchical encoder-decoder module includes the following steps: 1) Based on the feature map output by the backbone network, extract the key point p3 and its features And input it into the voxel feature encoding module for normalization: Among them, u k 、v k represents the normalized two-dimensional coordinates, i.e., the voxelized output result of the key point p3, x min ,y min 、x max ,y max They represent the minimum x value, the minimum y value, the maximum x value, and the maximum y value, respectively, and are used to represent the effective range of the plane projection of the 3D point cloud; W and H are the width and height of the feature map, respectively. k ,y k is the horizontal and vertical coordinates of the k-th frame point cloud projection; 2) Each point has a known feature Collect the current feature map grid R through maximum pooling h,w Characteristics of the midpoint Form the corresponding 2D mesh features in, Indicates that the features projected in the same grid are processed along the channel for maximum pooling; k is the grid number; c represents the maximum value of each channel; 3) For characterizing 2D mesh features Feature map Downsampling is performed to make the effective feature distance that was originally larger than the convolution kernel size closer to the kernel size; Using the submanifold convolution layer, the effective features in the feature map are upsampled to maintain the same sparsity as the feature map before downsampling, and multi-scale feature values ​​are output.

5. The detection method based on an end-to-end single-stage 3D object detection model according to claim 1, characterized in that: The voting process of the point cloud features to generate candidate regions to locate the target object and the feature connection through the hybrid deep learning module to determine the prediction result includes the following steps: (1) Construct a voting module to predict the target center p to which the key point p3 extracted based on the backbone network output feature map belongs v : Among them, Vote means voting, represents the feature map of the fourth feature learning stage, offset represents the voting result, p4 represents the feature point of the fourth feature learning stage, and n4 represents the number of sampling points in the fourth feature learning stage; The voting point p v The position of is regarded as the RoI of the current scene, and then projected to the multi-scale feature value output by the hierarchical encoder-decoder module; (2) Set the polling point p v The projected coordinates of are embedded into the object’s location query through the self-attention module to obtain local features Among them, Pool means pooling, Conv means convolution, and Padding means local area selection. It is the feature map of the third feature learning stage; (3) Calculate through the self-attention module to obtain the 2D feature projection: Where m represents the index of the attention head, M represents the number of attention heads; W m With W m ′ m is a learnable weight; A mij is the attention weight; The value is used as the feature value q of the query i,j ,q i ,q j Represent the row and column values ​​of the eigenvalues ​​respectively; (4) Use bilinear interpolation to project the 2D features back to the point: w i,j,k =(1-|u k -(|u k |+i)|)(1-|v k -(|v k |+j)|) Among them, w i,j,k represents the bilinear interpolation weight, i represents the number of row values, j represents the number of column values, k is the grid number, It is an integrated feature map that integrates the multi-scale features of the hierarchical encoder-decoder module; |u k |+i and |v k |+j is used to limit the projection grid coordinates; (5) Local features Attention results Q, return points These three features are connected together to determine the prediction results: Where Q = MultiHeadAttn(q i ).

6. The detection method based on an end-to-end single-stage 3D object detection model according to claim 1, characterized in that: The training is performed in an end-to-end multi-task manner, which includes the following steps: 1) Use weighted cross entropy loss to manually emphasize rare categories, weighted cross entropy loss as follows: Among them, δ represents the hyperparameter, a c represents the weight of the cth class, y c Represents the true label determined by whether the point is within the bounding box. represents the predicted probability, and C represents the category of the target; 2) Calculate the total loss in, They are sample loss, voting module loss, classification prediction loss, and voting point feature regression loss; They represent position loss, size loss, angle range loss, angle resolution loss, and corner component loss respectively.

7. A detection system based on an end-to-end single-stage 3D object detection model, characterized in that: include: The target detection model building module is used to build a feature map of the point-based backbone network for the three-dimensional point cloud collected by the lidar; Construct a hierarchical encoder-decoder module to extract and compress features to generate semantic information; perform voting on point cloud features to generate candidate regions to locate the target object; perform feature concatenation through a hybrid deep learning module to determine the prediction result; Train the obtained end-to-end single-stage 3D object detection model; The target detection module is used to collect 3D point clouds in real time and input the trained end-to-end single-stage 3D target detection model to obtain the detected target.

8. A detection device based on an end-to-end single-stage 3D object detection model, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement a detection method based on an end-to-end single-stage 3D target detection model as described in any one of claims 1-6 when executing the computer program.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, a detection method based on an end-to-end single-stage 3D object detection model as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • 3D target detection method, device and equipment based on self-attention mechanism

    CN117058472A