A 3D object detection method and system based on implicit representation in 3D modeling

By using an implicit representation method based on 3D modeling in 3D object detection, processing LiDAR point cloud data and generating high-quality 3D target bounding boxes, the problem of processing LiDAR point cloud in the prior art is solved, and the robustness and accuracy are improved.

CN114463737BActive Publication Date: 2025-05-02FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210107083.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-28
Publication Date
2025-05-02
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

The existing 3D object detection technology is difficult to effectively deal with the irregular, sparse and disordered characteristics of LiDAR point clouds, and it is difficult to directly inherit the mature network framework and algorithm structure of 2D object detection.

Method used

Using an implicit representation method based on 3D modeling, the features of the point cloud and voxel dimensions are extracted and combined to convert them into bird's-eye view features by preprocessing the LiDAR point cloud data. Then, use implicit functions to assign values ​​to points in the local three-dimensional space, generate target bounding boxes, and further adjust the bounding boxes through feature optimization.

Benefits of technology

It realizes the generation of high-quality 3D target bounding boxes without anchor boxes, which improves the robustness and accuracy of detection and can effectively handle the complexity of 3D scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463737B_ABST
    Figure CN114463737B_ABST
Patent Text Reader

Abstract

The present invention relates to a 3D target detection method and system based on implicit representation in 3D modeling, the method comprising collecting point cloud data and preprocessing to obtain preprocessed point cloud data; extracting corresponding features in the point cloud dimension and voxel dimension respectively according to the obtained preprocessed point cloud data, and combining and converting the two features into bird's-eye view features; performing coordinate and feature offset on each pixel point on the bird's-eye view feature map, screening and sampling the candidate center point with the maximum probability; using implicit functions to assign values ​​to all points contained in the surrounding local three-dimensional space with the candidate center point as a unit, and generating a target bounding box according to the assigned results; optimizing the bounding box by combining the features in the generated target bounding box. Compared with the prior art, the present invention has the advantages of fast speed, high accuracy, good robustness, etc., and is suitable for applications such as target detection and segmentation in three-dimensional scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual detection, and in particular to a 3D target detection method and system based on implicit representation in 3D modeling. Background Art

[0002] In recent years, object detection technology has attracted widespread attention in academia and industry, with applications spanning the currently popular fields of virtual reality, autonomous driving, and robotics. Object detection technology is primarily categorized into 2D and 3D object detection, depending on the task. 2D object detection is a fundamental and highly applicable task in vision. It identifies objects in an image and locates their locations at the pixel level.

[0003] With the rapid development of deep neural networks in computer vision, the reintroduction of convolutional neural networks has enabled unprecedented achievements in 2D object detection. However, localizing objects based solely on images presents many limitations in real-world applications. For example, in autonomous driving, the specific distance and orientation of the target object are required for more accurate spatial path planning and navigation. Consequently, 3D object detection has gradually emerged and flourished. 3D object detection builds on 2D object detection by adding rotational orientation, length, width, height, and center location to the target. In the field of 3D object detection, the most common algorithm uses point clouds generated by laser radar (LiDAR) sensors as input for further detection. Although LiDAR point clouds can capture precise distance measurements and geometric information about the surrounding environment, their irregular, sparse, and disordered nature makes them difficult to encode and hinders the direct inheritance of established network frameworks and algorithmic structures for two-dimensional (2D) object detection. Summary of the Invention

[0004] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a 3D target detection method and system based on implicit representation in 3D modeling.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] A 3D object detection method based on implicit representation in 3D modeling includes the following steps:

[0007] Collecting point cloud data collected from LiDAR and preprocessing it to obtain preprocessed point cloud data;

[0008] Based on the obtained pre-processed point cloud data, corresponding features are extracted in the point cloud dimension and voxel dimension respectively, and these two features are combined and converted into bird's-eye view features;

[0009] The coordinates and features of each pixel on the bird's-eye view feature map are offset, and the candidate center point with the highest probability is screened and sampled;

[0010] Use an implicit function to assign values ​​to all points in the local three-dimensional space around the candidate center point, and generate a target bounding box based on the assigned results;

[0011] The bounding box is optimized by combining the features within the generated target bounding box.

[0012] Furthermore, preprocessing point cloud data specifically includes the following steps:

[0013] According to the detection range, only the point cloud data within the range of the x, y and z axes are retained to form a point cloud set;

[0014] The point cloud collection is divided into three-dimensional grid voxels according to the voxel size in three-dimensional space. When the number of points contained in each voxel exceeds the set number, it is randomly sampled so that the number of points contained in each voxel grid does not exceed the set number.

[0015] Furthermore, pre-processing the point cloud data to extract corresponding features in the point cloud dimension and voxel dimension specifically includes the following steps:

[0016] The preprocessed point cloud data is passed through a multi-layer perceptron to obtain the point feature vector;

[0017] The obtained point feature vector is fed into multiple voxel feature extraction layers to obtain initial features for each grid voxel;

[0018] The obtained point feature vector and the obtained initial feature are fused and sent to an MLP layer to obtain the features of the point cloud dimension;

[0019] The obtained initial features are sent to multiple 3D sparse convolution blocks to obtain voxel-dimensional features; the voxel-dimensional features are compressed along the z-axis and sent to the multi-scale 2D convolution layer to obtain 2D bird's-eye view features.

[0020] Furthermore, screening and sampling the candidate center point with the maximum probability on the bird's-eye view feature map specifically includes the following steps:

[0021] Adjust each pixel on the bird's-eye view feature map toward its true center point, that is, (bev) Feed it into an MLP layer to generate the center offset and feature offset for each pixel on the bird's-eye view feature. By adding the offset, the expression of the candidate center is

[0022] p (ctr) =p (ofs) +p (bev)

[0023] f (ctr) =f (ofs) +f (bev)

[0024] Among them, p (bev) and f (bev) Represents the coordinates and features of each pixel on the bird's-eye view feature map, p (ctr) and f (ctr) Represent the coordinates and features of the candidate center point, p (ofs) represents the center offset, f (ofs) Indicates feature offset;

[0025] The candidate centers obtained after migration are screened and sampled for quality, and the 3D center distance is used as the quality measure.

[0026] Furthermore, generating the target bounding box specifically includes the following steps:

[0027] A sampling strategy is used for the candidate center point to obtain the point cloud coordinates and features in the local three-dimensional space around it, where the sampling points include original points and virtual points;

[0028] Use implicit functions to assign values ​​to each sampled point in the local three-dimensional space. The assignment is expressed as Specifically, the implicit function generates a kernel conditioned on the candidate center, which is convolved with the sampling point to dynamically adjust the assignment result of the sampling point; similarly, the assignment result of each sampled original point is calculated The assignment results of the two types of sampling points based on the candidate center point are collectively referred to as

[0029] The sampling points in the local three-dimensional space are filtered according to the set threshold and the assignment result, and the target boundary is generated based on the filtered points.

[0030] Furthermore, using the sampling strategy to obtain the point cloud coordinates and features in the local three-dimensional space around it specifically includes the following steps:

[0031] Step 1: Given a candidate center point Draw a ball with a radius of r to obtain the local space around it, and randomly select m original points from the space as sampling points;

[0032] Step 2: For each sampled original point Collect its corresponding point-based features f (point) and marked as

[0033] Step 3: Set a series of virtual points Evenly placed at the candidate center points around;

[0034] Step 4: At the virtual point In the same way, m virtual points are randomly sampled;

[0035] Step 5: For the virtual points obtained by sampling, the K-nearest neighbor algorithm is used to calculate the voxel features. The virtual point features are obtained by interpolation;

[0036] Step 6: The interpolated virtual point features are fed into an MLP layer for encoding, and the virtual point coordinates and features are marked as and

[0037] Furthermore, the sampling points in the local three-dimensional space are filtered according to the set threshold and the assignment result, and then the target boundary is generated according to the filtered points. Specifically, the following steps are included:

[0038] Step 1, If the value of is higher than the set threshold, it is considered as a point inside the target area, otherwise it is considered as a point outside the target area;

[0039] Step 2: Generate the size of the bounding box: Use the minimum enclosing rectangle to generate an axis-parallel bounding box that fits all internal points;

[0040] Step 3: Generate the direction of the bounding box: reduce the direction space from [0,2π] to Then divide it into multiple different angles, calculate the distance from the sampling point to the surface within the target border point, select the bounding box with the smallest distance, and the corresponding angle is used as the angle r of the bounding box a ; At the same time, by comparing the length l of the bounding box a and width a , empirically correct the direction range to [0,π], and its expression is:

[0041]

[0042] Among them, r a represents the angle of the bounding box, l a Indicates the length of the bounding box, w a Indicates the width of the bounding box.

[0043] Furthermore, the process of optimizing the generated target bounding box includes the following steps:

[0044] Reuse implicit values ​​to refine and adjust the bounding box by aggregating features of internal sample points and suppressing the influence of features from external points. Specifically, multiple grid points are uniformly sampled within each bounding box, and then a point set abstraction layer is used to aggregate internal point features and voxel features at each grid point position.

[0045] The features of all grid points are concatenated and fed into the detection head; the detection head is constructed with three parallel branches for classification confidence prediction, direction prediction, and box boundary refinement, respectively.

[0046] Furthermore, for the three parallel branches of the detection head, each branch has four MLP layers with 256 channels, and all branches share the first two layers.

[0047] A 3D object detection system based on implicit representation in 3D modeling, comprising:

[0048] The point cloud data preprocessing unit is used to collect and preprocess the point cloud data collected from the LiDAR to obtain preprocessed point cloud data;

[0049] The root point cloud feature extraction unit is used to extract corresponding features in the point cloud dimension and voxel dimension based on the pre-processed point cloud data, and combine the two features to convert them into bird's-eye view features;

[0050] The target center point sampling unit is used to perform coordinate and feature offset on each pixel point on the bird's-eye view feature map, and to screen and sample the candidate center point with the highest probability;

[0051] An implicit target boundary generation unit is used to assign values ​​to all points in the local three-dimensional space around the candidate center point using an implicit function, and generate a target bounding box based on the assigned results;

[0052] The candidate region integration unit is used to optimize the bounding box by combining the features within the generated target bounding box.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] 1. This paper assigns values ​​to points within a local three-dimensional space and distinguishes internal and external points based on the assigned values, thereby generating a target bounding box based on the internal points. Compared with the traditional hyperparameter definition of the target bounding box, the implicit representation method has good robustness.

[0055] 2. The present invention uses an implicit function to assign values ​​to all points in the local three-dimensional space around the candidate center point, and generates a target bounding box based on the assigned results. Therefore, when predicting the bounding box, there is no need to rely on any anchor point box that needs to be fine-tuned.

[0056] 3. The present invention adopts implicit representation in 3D modeling, that is, assigning values ​​to points in the local three-dimensional space, generating a target bounding box based on the assigned results, and optimizing the bounding box by combining the features in the generated target bounding box. Therefore, it has the advantages of fast speed, high accuracy, and good robustness. It can effectively apply segmentation tasks to target detection and improve the understanding and analysis of 3D scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 This is a flowchart of the first embodiment of the present invention.

[0058] Figure 2 This is a schematic diagram of the boundary generation of the implicit target boundary generation unit in the first embodiment of the present invention.

[0059] Figure 3 It is a structural diagram of embodiment 2 of the present invention.

[0060] Figure 4 It is a flowchart of the second embodiment of the present invention.

[0061] Figure 5 This is a flowchart of the second embodiment of the present invention.

[0062] Figure 6 It is a structural diagram of embodiment 3 of the present invention. DETAILED DESCRIPTION

[0063] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0064] Example 1

[0065] like Figure 1 As shown, this embodiment provides a 3D object detection method based on implicit representation in 3D modeling, and the specific steps are as follows:

[0066] Step S1: collect point cloud data collected from LiDAR and preprocess it to obtain preprocessed point cloud data.

[0067] Step S2: Based on the obtained pre-processed point cloud data, corresponding features are extracted in the point cloud dimension and voxel dimension respectively, and these two features are combined and converted into bird's-eye view features.

[0068] Step S3: perform coordinate and feature offset on each pixel point on the bird's-eye view feature map, and screen and sample the candidate center point with the maximum probability.

[0069] Step S4: Use an implicit function to assign values ​​to all points in the local three-dimensional space around the candidate center point, and generate a target bounding box based on the assigned results.

[0070] Step S5: Optimize the bounding box by combining the features within the generated target bounding box.

[0071] 1. The specific expansion of step S1 is as follows:

[0072] Step S1-1, according to the detection range setting, only retains the point cloud set that meets the set range in the x, y and z axis directions. The point cloud set that meets the detection range includes the following steps: Step S1-1-a, read the point cloud data Among them, (x i ,y i ,z i ) is the three-dimensional space coordinate value, r i is the intensity value of the point, and N is the total number of point clouds; step S1-1-b, removes out-of-bounds point cloud data according to the ranges set on the x, y, and z axes.

[0073] Step S1-2: The entire point cloud set is divided into three-dimensional grid voxels according to the voxel size of the three-dimensional space, and only a maximum of 5 point cloud points are sampled in each voxel. In particular, according to the voxel size (x s ,y s ,z s ) The point cloud data is meshed. When the number of points contained in each three-dimensional grid voxel exceeds 5, random sampling is performed so that the number of points contained in each voxel grid does not exceed 5. At this time, the point cloud data set is represented as P1.

[0074] In this example, the x, y, and z axes are set to [0,70.4]m, [-40,40]m, and [-3,1]m, respectively. The voxel size is set to [0.05,0.05,0.1]m.

[0075] 2. The specific development of step S2 is as follows:

[0076] In step S2-1, the point cloud set P1 obtained in step S1 is passed through a multi-layer perceptron (MLP) to obtain a point feature vector.

[0077] Step S2-2: Send the obtained point feature vector to multiple voxel feature extraction layers to obtain initial features for each grid voxel

[0078] Step S2-3, the point feature vector obtained in S2-1 and the initial feature obtained in S2-2 are fused and sent to an MLP layer to obtain the point-based feature f (point) .

[0079] Step S2-4, the initial features obtained in S2-2 Feed it into multiple 3D sparse convolution blocks to obtain voxel-based multi-scale features

[0080] Step S2-5, voxel features Compress along the z-axis and feed into the multi-scale 2D convolution layer to obtain 2D bird's-eye view features Where H, W and C represent the length, width and feature dimension of the bird's-eye view feature, respectively.

[0081] In this example, the two voxel feature extraction layers in step S2-2 have 32 and 64 channels, respectively. The output channels of the 3D sparse convolution block in step S2-4 are 32, 32, 64, 64, and 128, respectively. The multi-scale 2D convolution layer structure in step S2-5 consists of two convolutional layers and two deconvolutional layers, and has an output channel of 128.

[0082] 3. The specific development of step S3 is as follows:

[0083] Step S3-1, the bird's-eye view feature f (bev) Each pixel on the image is adjusted towards its true center point, that is, the bird's-eye view feature f (bev) Feed it into an MLP layer to generate the center offset for each pixel on the bird’s-eye view feature and feature offset By adding the offset, the candidate center can be expressed as

[0084] p (cfr) =p (ofs) +p (bev)

[0085] f (ctr) =f (ofs) +f (bev)

[0086] Among them, p (bev) and f (bev) Represent the coordinates and features of each pixel on the bird's-eye view map.

[0087] Step S3-2: Perform quality screening and sampling on the candidate centers obtained after the migration, using the 3D center distance as the quality criterion:

[0088]

[0089] Among them, x f ,x b ,y l ,y r ,z t and zb Respectively represent the distances from the candidate center to the front, back, left, right, top, and bottom of the true target frame; s (ctrns) The closer the value is to 1, the closer the candidate midpoint is to the true target center. (ctrns) When it is 0, it means that the adjusted pixel is outside the target area. During the training and testing process, this value is calculated by adding the center feature f (ctr) It is fed into an MLP layer and a sigmoid nonlinear layer for prediction.

[0090] In this embodiment, a total of 512 optimal candidate center points are sampled.

[0091] 4. The specific development of step S4 is as follows Figure 2 As shown:

[0092] Step S4-1 uses a sampling strategy to obtain the point cloud coordinates and features in the surrounding local three-dimensional space based on the center point obtained in step S3. Step S4-1 further includes the following sub-steps:

[0093] Step S4-1-a, given a candidate center point Draw a ball with a radius of r to obtain the local space around it, and randomly select m original points from the space. The sampling point set is defined as:

[0094]

[0095] in, Indicates the candidate center point As the center, perform local space point sampling operation; represents the original point in the local space; r is the radius of the spherical local three-dimensional space;

[0096] Step S4-1-b, for each sampled original point Collect its corresponding point-based features f (point) and marked as

[0097] Step S4-1-c, a series of virtual points The grid size is S×S×S, and the spacing is (x s ,y s ,z s ) are evenly placed at the candidate center points around;

[0098] Step S4-1-d, in order to reduce the amount of calculation, at the virtual point In the same way, m virtual points are randomly sampled;

[0099] Step S4-1-e: For the virtual points obtained by sampling, in order to obtain the corresponding features, the K-nearest neighbor algorithm is used to extract the voxel features. The virtual point features are obtained by interpolation;

[0100] In step S4-1-f, the interpolated virtual point features are fed into an MLP layer for encoding. Similarly, the virtual point coordinates and features are marked as: and

[0101] Step S4-2 uses an implicit function to assign a value to each sampled point; whether a sample point belongs to a target area (i.e., within a frame) depends on its corresponding candidate center. The closer the Euclidean distance or feature distance between two points, the higher the probability that they belong to the same frame (target). Here, an implicit function is used to generate a kernel conditioned on the candidate center. This kernel is convolved with the sample point to dynamically adjust the assignment result of the sample point. The kernel here can be expressed as:

[0102]

[0103] The generated kernel θ k It is reshaped into the parameters of two convolutional layers with a channel number of 16. Taking the virtual sampling point as an example, its assignment can be expressed as:

[0104]

[0105] It can be seen that The value of the sampling point and the corresponding voxel features Similarly, the assignment result of each sampled original point can be calculated Based on the candidate center point The assignment results of the two types of sampling points are collectively referred to as

[0106] Step S4-3 selects sampling points in the local three-dimensional space according to the threshold value and generates the target boundary based on these points. This process includes the following sub-steps:

[0107] Step S4-3-a, according to the threshold setting, If the value of is higher than the threshold, it is considered as a point inside the target area, otherwise it is considered as a point outside the target area;

[0108] Step S4-3-b, generating the size of the bounding box: using the minimum enclosing rectangle to generate an axis-parallel bounding box that fits all interior points;

[0109] Step S4-3-c, generate the direction of the bounding box: reduce the direction space from [0,2π] to Then divide it into h = 7 different angles, calculate the distance from the sampling point to the surface within the target border point, select the bounding box with the smallest distance, and the corresponding angle is used as the angle r of the bounding box a At the same time, by comparing the length l of the bounding box a and width a , empirically correct the direction range to [0,π]:

[0110]

[0111] In this example, in S4-1, the radius r is set to 3.2m, m=256 points are randomly sampled, the grid size is set to S=10, and the spacing (x s ,y s ,z s )=(0.6,0.6,0.3)m.

[0112] 5. The specific implementation of step S5 is as follows:

[0113] Step S5-1 Reusing implicit values The bounding box is refined by aggregating the features of internal sample points and suppressing the influence of features of external points. Specifically, 6×6 grid points are uniformly sampled within each bounding box. Then, a point set abstraction layer is used to aggregate the internal point features and voxel features at each grid point position. and

[0114] Step S5-2 concatenates the features of all grid points and inputs them into the detection head. The detection head consists of three parallel branches, one for classification confidence prediction, one for orientation prediction, and one for box boundary refinement. Specifically, each branch has four MLP layers with 256 channels, and all branches share the first two layers.

[0115] Example 2

[0116] like Figure 3 As shown, this embodiment discloses a 3D target detection system based on implicit representation in 3D modeling, including a point cloud data preprocessing unit 101, a point cloud feature extraction unit 102, a target center point sampling unit 103, an implicit target boundary generation unit 104, a candidate area integration unit 105, a picture storage unit 106, an output display unit 107, a system communication unit 108 and a system control unit 109 for controlling the above-mentioned units.

[0117] The point cloud data preprocessing unit 101 is used to preprocess the obtained point cloud data to be analyzed to obtain preprocessed point cloud data. In this embodiment, point cloud data is a set of point coordinates generated by a laser radar in three-dimensional space. Point cloud data is the most commonly used data input form in 3D detection.

[0118] The point cloud feature extraction unit 102 extracts corresponding features in the point dimension and the voxel dimension respectively, combines the two features, and converts them into bird's-eye view features.

[0119] The target center sampling unit 103 offsets each pixel point on the bird's-eye view feature map and samples a candidate center point with the maximum probability.

[0120] The implicit object boundary generation unit 104 assigns values ​​to all points in a local three-dimensional space around a candidate center point using an implicit function, and generates a high-quality object boundary based on the assigned results.

[0121] The candidate region integration unit 105 optimizes the generated target bounding box by fusing the features of the sampling points within the bounding box.

[0122] The picture storage unit 106 is used to store the pictures of the detection output results. In this embodiment, the picture storage unit 106 stores the detection results optimized by the candidate region integration unit 106.

[0123] The output display unit 107 is used to display the test results received from the system communication unit 108, allowing users to complete corresponding human-computer interaction through these images. The image storage unit 106 and the output display unit 107 are a display device that is connected to a computing device, such as a computer, television, or mobile device.

[0124] The system communication unit 108 transmits the detection result stored in the screen storage unit 106 to the output display unit 107 .

[0125] Figure 4 and Figure 5 3D object detection system based on implicit representation in this embodiment. The 3D object detection process of the 3D object detection system based on implicit representation 100 includes the following steps:

[0126] In step T1 , the point cloud data preprocessing unit 101 performs data preprocessing on the data collected from the LiDAR to obtain preprocessed point cloud data, and then proceeds to step 2 .

[0127] In step T2, the point cloud feature extraction unit 102 extracts corresponding features in the point dimension and the voxel dimension respectively, combines the two features, and converts them into bird's-eye view features, and then enters step 3.

[0128] In step T3, the target center sampling unit 103 performs coordinate and feature offset on each pixel point on the bird's-eye view feature map, and samples the candidate center point with the maximum probability, and then enters step 4.

[0129] In step T4, the implicit target boundary generation unit uses an implicit function to assign values ​​to all points in the local three-dimensional space around the candidate center point, and generates a high-quality target boundary based on the assigned results, and then enters step 5.

[0130] In step T5, the candidate region integration unit optimizes the target bounding box by fusing the features of the sampling points within the target bounding box, and then enters the end state.

[0131] The system of the present invention has the advantages of high speed, high accuracy and good robustness. It introduces implicit representation into 3D object detection. It not only effectively improves the understanding of 3D scenes by combining segmentation and detection tasks, but also utilizes the inherent advantages of implicit representation to improve the robustness of the predicted bounding box without the need for any anchor box.

[0132] Example 3

[0133] like Figure 6 As shown, this embodiment discloses a 3D target detection device based on implicit representation in 3D modeling, which is composed of a computing device and a display device for processing external media data. The computing device is composed of a processor and a memory. The processor is a hardware processor for computing and running executable code. Common processors include a central processing unit (CPU) or a graphics computing processor (GPU); the memory is a non-volatile memory for storing executable code and various intermediate data and parameters so that the processor can execute the corresponding calculation process. The memory stores the relevant execution program code for running the point cloud data preprocessing unit 101, the point cloud feature extraction unit 102, the target center point sampling unit 103, the implicit target boundary generation unit 104, and the candidate area integration unit 105; the display device includes a picture storage unit 106 and an output display unit 107.

[0134] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.

Claims

1. A 3D object detection method based on implicit representation in 3D modeling, characterized in that: The following steps are involved: Collecting point cloud data collected from LiDAR and preprocessing it to obtain preprocessed point cloud data; According to the obtained pre-processed point cloud data, corresponding features are extracted in the point cloud dimension and voxel dimension respectively, and these two features are combined and converted into bird's-eye view features; The coordinates and features of each pixel point on the bird's-eye view feature map are offset, and the candidate center point with the highest probability is screened and sampled; Use an implicit function to assign values ​​to all points in the local three-dimensional space around the candidate center point, and generate a target bounding box based on the assigned results; Optimize the bounding box by combining the features within the generated target bounding box; Generating the target bounding box specifically includes the following steps: A sampling strategy is used for the candidate center point to obtain the point cloud coordinates and features in the local three-dimensional space around it, where the sampling points include original points and virtual points; Use an implicit function to assign a value to each point sampled in the local three-dimensional space. The assignment is expressed as Specifically, the implicit function generates a kernel conditioned on the candidate center, which is convolved with the sampling point to dynamically adjust the assignment result of the sampling point; similarly, the assignment result of each sampled original point is calculated The assignment results of the two types of sampling points based on the candidate center points are collectively referred to as The sampling points in the local three-dimensional space are selected according to the set threshold and the assignment result, and the target boundary is generated according to the selected points; Using the sampling strategy to obtain the point cloud coordinates and features in the local three-dimensional space around it specifically includes the following steps: Step 1: Given a candidate center point Draw a ball with a radius of r to obtain the local space around it, and randomly select m original points from the space as sampling points; Step 2: For each sampled original point Collect its corresponding point-based features f (point) and marked as Step 3: Convert a series of virtual points Evenly placed at the candidate center points around; Step 4: At the virtual point In the same way, m virtual points are randomly sampled; Step 5: For the sampled virtual points, the K-nearest neighbor algorithm is used to calculate the voxel features. The virtual point features are obtained by interpolation; Step 6: Send the interpolated virtual point features to an MLP layer for encoding, and mark the virtual point coordinates and features as and The sampling points in the local three-dimensional space are selected according to the set threshold and the assignment result, and then the target boundary is generated according to the selected points. Specifically, the following steps are included: Step 1, If the value of is higher than the set threshold, it is considered as a point inside the target area, otherwise it is considered as a point outside the target area; Step 2: Generate the size of the bounding box: Use the minimum enclosing rectangle to generate an axis-parallel bounding box that fits all internal points; Step 3: Generate the direction of the bounding box: Reduce the direction space from [0,2π] to Then divide it into multiple different angles, calculate the distance from the sampling point to the surface within the target border point, select the bounding box with the smallest distance, and the corresponding angle is used as the angle r of the bounding box a ; At the same time, by comparing the length l of the bounding box a and width a , empirically correct the direction range to [0,π], and its expression is: Among them, r a represents the angle of the bounding box, l a Indicates the length of the bounding box, w a Indicates the width of the bounding box.

2. The 3D object detection method based on implicit representation in 3D modeling according to claim 1, characterized in that: Preprocessing point cloud data specifically includes the following steps: According to the detection range, only the point cloud data within the range of the x, y and z axis directions are retained to form a point cloud set; The point cloud set is divided into three-dimensional grid voxels according to the voxel size in three-dimensional space. When the number of points contained in each voxel exceeds the set number, it is randomly sampled so that the number of points contained in each voxel grid does not exceed the set number.

3. The 3D object detection method based on implicit representation in 3D modeling according to claim 1, characterized in that: Preprocessing point cloud data to extract corresponding features in point cloud dimension and voxel dimension respectively includes the following steps: The preprocessed point cloud data is passed through a multi-layer perceptron to obtain the point feature vector; The obtained point feature vector is fed into multiple voxel feature extraction layers to obtain initial features for each grid voxel; The obtained point feature vector and the obtained initial feature are fused and sent to an MLP layer to obtain the features of the point cloud dimension; The obtained initial features are sent to multiple 3D sparse convolution blocks to obtain voxel-dimensional features; the voxel-dimensional features are compressed along the z-axis and sent to a multi-scale 2D convolution layer to obtain 2D bird's-eye view features.

4. The 3D object detection method based on implicit representation in 3D modeling according to claim 1, characterized in that: The specific steps of screening and sampling the candidate center point with the maximum probability on the bird's-eye view feature map include: Adjust each pixel on the bird's-eye view feature map toward its true center point, that is, (bev) Send it to an MLP layer to generate the center offset and feature offset for each pixel on the bird's-eye view feature. By adding the offset, the candidate center is expressed as p (ctr) =p (ofs) +p (bev) f (ctr) =f (ofs) +f (bev) Among them, p (bev) and f (bev) Respectively represent the coordinates and features of each pixel on the bird's-eye view feature map, p (ctr) and f (ctr) Represent the coordinates and features of the candidate center point, p (ofs) represents the center offset, f (ofs) Indicates feature offset; The candidate centers obtained after migration are screened and sampled for quality, and the 3D center distance is used as a quality measure.

5. The 3D object detection method based on implicit representation in 3D modeling according to claim 1, characterized in that: The process of optimizing the generated object bounding box includes the following steps: Reuse implicit values ​​to refine and adjust the bounding box by aggregating the features of internal sampled points and suppressing the influence of features of external points. Specifically, multiple grid points are uniformly sampled in each bounding box, and then a point set abstraction layer is used to aggregate the internal point features and voxel features at each grid point position. The features of all grid points are concatenated and fed into the detection head; the detection head is constructed with three parallel branches for classification confidence prediction, direction prediction, and box boundary refinement, respectively.

6. The 3D object detection method based on implicit representation in 3D modeling according to claim 5, characterized in that: For the three parallel branches of the detection head, each branch has four MLP layers with 256 channels, and all branches share the first two layers.

7. A 3D object detection system based on implicit representation in 3D modeling, characterized in that: include: The point cloud data preprocessing unit is used to collect and preprocess the point cloud data collected from the LiDAR to obtain preprocessed point cloud data; The root point cloud feature extraction unit is used to extract corresponding features in the point cloud dimension and the voxel dimension respectively according to the pre-processed point cloud data obtained, and combine and convert the two features into bird's-eye view features; The target center point sampling unit is used to perform coordinate and feature offset on each pixel point on the bird's-eye view feature map, and screen and sample the candidate center point with the highest probability; An implicit target boundary generation unit is used to assign values ​​to all points in the local three-dimensional space around the candidate center point using an implicit function, and generate a target bounding box according to the assigned results; A candidate region integration unit, used to optimize the bounding box by combining the features within the generated target bounding box; Generating the target bounding box specifically includes the following steps: A sampling strategy is used for the candidate center point to obtain the point cloud coordinates and features in the local three-dimensional space around it, where the sampling points include original points and virtual points; Use an implicit function to assign a value to each point sampled in the local three-dimensional space. The assignment is expressed as Specifically, the implicit function generates a kernel conditioned on the candidate center, which is convolved with the sampling point to dynamically adjust the assignment result of the sampling point; similarly, the assignment result of each sampled original point is calculated The assignment results of the two types of sampling points based on the candidate center points are collectively referred to as The sampling points in the local three-dimensional space are selected according to the set threshold and the assignment result, and the target boundary is generated according to the selected points; Using the sampling strategy to obtain the point cloud coordinates and features in the local three-dimensional space around it specifically includes the following steps: Step 1: Given a candidate center point Draw a ball with a radius of r to obtain the local space around it, and randomly select m original points from the space as sampling points; Step 2: For each sampled original point Collect its corresponding point-based features f (point) and marked as Step 3: Convert a series of virtual points Evenly placed at the candidate center points around; Step 4: At the virtual point In the same way, m virtual points are randomly sampled; Step 5: For the sampled virtual points, the K-nearest neighbor algorithm is used to calculate the voxel features. The virtual point features are obtained by interpolation; Step 6: Send the interpolated virtual point features to an MLP layer for encoding, and mark the virtual point coordinates and features as and The sampling points in the local three-dimensional space are selected according to the set threshold and the assignment result, and the target boundary is generated according to the selected points, which specifically includes the following steps: Step 1, If the value of is higher than the set threshold, it is considered as a point inside the target area, otherwise it is considered as a point outside the target area; Step 2: Generate the size of the bounding box: Use the minimum enclosing rectangle to generate an axis-parallel bounding box that fits all internal points; Step 3: Generate the direction of the bounding box: Reduce the direction space from [0,2π] to Then divide it into multiple different angles, calculate the distance from the sampling point to the surface within the target border point, select the bounding box with the smallest distance, and the corresponding angle is used as the angle r of the bounding box a ; At the same time, by comparing the length l of the bounding box a and width a , empirically correct the direction range to [0,π], and its expression is: Among them, r a represents the angle of the bounding box, l a Indicates the length of the bounding box, w a Indicates the width of the bounding box.

Citation Information

Patent Citations

  • Three-dimensional modeling method, device and system based on implicit function and storage medium

    CN110033519A

  • Sign and lane creation for high definition maps used for autonomous vehicles

    CN111542860A