A 3D object detection method, system, and storage medium based on the center point

Through the three-dimensional object detection method based on the center point, the dynamic voxelization and semi-supervised training framework is used to solve the problems of rotational alignment and information loss in the existing technology, the feature map quality and detection frame representation are improved, the annotation cost is reduced, and the efficient three-dimensional object detection is achieved.

CN117058398BActive Publication Date: 2025-07-11UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310781067.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-28
Publication Date
2025-07-11
Estimated Expiration
2043-06-28

AI Technical Summary

Technical Problem

The existing three-dimensional object detection technology has problems with anchor rotation alignment, point cloud data voxelization information loss, feature extraction network failure to fuse high-level semantic features and low-level spatial features, inaccurate quality representation of detection frames, and high costs due to relying on a large amount of labeled data.

Method used

The three-dimensional object detection method based on the center point is adopted, dynamic voxelization, spatial semantic feature aggregation network and attention fusion network are used, combined with the semi-supervised training framework of the teacher-student model, and the detection model is optimized through a consistent regularization loss function to reduce the need for labeling data.

Benefits of technology

The rotational alignment and information loss problems are solved, the feature map quality is improved, the position representation of the detection box is optimized, the annotation cost is reduced, and the three-dimensional object detection performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058398B_ABST
    Figure CN117058398B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional object detection method, system, and storage medium based on a center point. The target image is sent to a trained three-dimensional detection model to output the detection result of the target image. The three-dimensional detection model includes a point cloud processing network, a feature extraction network, and a detection box prediction network connected in sequence. The feature extraction network uses a spatial semantic feature aggregation network as the feature extraction framework, and includes a spatial feature extraction network, a semantic feature extraction network, and an attention fusion network. The detection box prediction network includes a target category prediction branch, a detection box coordinate prediction branch, and a detection box position quality estimation prediction branch arranged in parallel. This three-dimensional object detection improves the prediction and detection results of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a three-dimensional object detection method, system, and storage medium based on a center point. Background Art

[0002] In recent years, with the rise of the concept of autonomous driving, the research on deep learning in the field of autonomous driving has received increasing attention. In autonomous driving algorithms, three-dimensional object detection models based on multi-modal data such as images and point clouds play a key role.

[0003] Current three-dimensional object detection models can be classified according to the input mode into: three-dimensional object detection models based on point clouds, three-dimensional object detection models based on images, and multi-modal three-dimensional object detection models based on image-point cloud fusion. Three-dimensional object detection models based on point clouds can be classified according to the encoding method of point clouds into: three-dimensional object detection models based on voxelization, three-dimensional object detection models based on BEV projection, and three-dimensional object detection models based on perspective views. The three-dimensional object detection model based on voxelization first processes three-dimensional point cloud data into voxel representations through the VoxelNet network or the PointPillar network, then extracts three-dimensional features through a backbone network, and finally performs three-dimensional object detection based on the feature map.

[0004] Currently, mainstream three-dimensional object detection models adopt an anchor-based detection method. Due to the preset size and direction of the anchors, there is a problem of rotational alignment between the anchors and the targets in turning scenarios. At the same time, in order to improve the detection performance, current various three-dimensional object detection models rely on a large amount of labeled training data to cover various long-tail scenarios, and the corresponding huge cost is brought by data annotation.

[0005] In summary, the existing three-dimensional object detection technologies have the following disadvantages: (1) The anchor-based three-dimensional object detection method has a problem of rotational alignment; (2) There is an information loss problem in the traditional voxelization process of point cloud data; (3) The general three-dimensional feature extraction network cannot well fuse the high-level semantic features and low-level spatial features of the point cloud feature map; (4) The object detector only uses the class score as the quality representation of the detection box, does not consider the position quality of the detection box, and cannot accurately represent the quality of the detection box; (5) The mainstream three-dimensional object detection methods rely on a large amount of labeled data, resulting in high costs. Summary of the Invention

[0006] Based on the technical problems existing in the background art, the present invention proposes a three-dimensional object detection method, system, and storage medium based on a center point, which improves the prediction and detection results of images.

[0007] A 3D object detection method based on the center point proposed by the present invention sends the target image to a trained 3D detection model to output the detection result of the target image;

[0008] The 3D detection model includes a point cloud processing network, a feature extraction network, and a detection box prediction network connected in sequence;

[0009] The point cloud processing network extracts the point cloud data of the image based on dynamic voxelization to obtain point cloud voxel data, and sends the point cloud voxel data to the feature extraction network;

[0010] The feature extraction network uses a spatial semantic feature aggregation network as the feature extraction framework, including a spatial feature extraction network, a semantic feature extraction network, and an attention fusion network. The input of the spatial feature extraction network is connected to the output of the point cloud processing network. After passing through three sequentially connected convolutional neural networks, initial spatial features are obtained. The input of the semantic feature extraction network is connected to the output of the spatial feature extraction network. After passing through three sequentially connected convolutional neural networks, initial semantic features are obtained. After deconvolving the initial semantic features and performing point-by-point accumulation with the initial spatial features, enhanced spatial features are obtained. The enhanced spatial features and the initial semantic features are input into the attention fusion network to output the final aggregated features;

[0011] The detection box prediction network includes a target category prediction branch, a detection box coordinate prediction branch, and a detection box position quality estimation prediction branch arranged in parallel. The inputs of the target category prediction branch, the detection box coordinate regression branch, and the detection box position quality estimation prediction branch are all connected to the output of the feature extraction network. The target category prediction branch consists of one convolutional network and predicts the target category score. The detection box coordinate prediction branch consists of one convolutional network and predicts the center point position of the detection box, the 3D size of the detection box, and the target orientation. The detection box position quality estimation prediction branch consists of two fully connected layers and an activation function. The output of the fully connected layer is used as the input of the activation function to predict the detection box position quality estimation.

[0012] Further, the training process of the 3D detection model is as follows:

[0013] Construct a training set and a validation set. The training set is a set of unlabeled point cloud data, and the validation set is a set of labeled point cloud data;

[0014] Perform data augmentation on the training set and the validation set, and extract all point features of the point cloud data in the augmented training set based on the dynamic voxelization in the point cloud processing network to obtain point cloud voxel data;

[0015] The voxel data of the point cloud is sequentially fed into the feature extraction network and the detection box prediction network, and the intersection over union between the detection box output by the 3D detection model and the ground truth box is predicted through the detection box position quality estimation prediction branch to optimize the quality of the 3D detection model;

[0016] The validation set is fed into the 3D detection model after weak data augmentation processing, and the 3D detection model is trained in a supervised manner to obtain a trained 3D detection model as the teacher model;

[0017] The training set is input into the teacher model after weak data augmentation processing to output the predicted value of the image. At the same time, the training set is input into the 3D detection model to be trained after strong data augmentation processing, and the 3D detection model is trained in a semi-supervised manner. At this time, the 3D detection model outputs the predicted value of the image as a student model;

[0018] The predicted value output by the teacher model is used as the pseudo-label, and the predicted value output by the student model is constrained based on the pseudo-label through the consistency regularization loss function to obtain a trained 3D detection model.

[0019] Furthermore, the formula of the consistency regularization loss function is as follows:

[0020]

[0021] where L represents the consistency regularization loss function, H and W represent the height and width of the feature map, i and j represent the positions of points on the feature map, C t ,B t represent the classification branch prediction result and the regression branch prediction result of the teacher model, C S ,B s represent the classification branch prediction result and the regression branch prediction result of the student model, and CE represents the cross-entropy loss function.

[0022] Furthermore, in the step of inputting the enhanced spatial feature and the initial semantic feature into the attention fusion network to output the final aggregated feature, it specifically includes:

[0023] The enhanced spatial feature and the initial semantic feature are respectively input into a convolutional network to obtain two single-channel feature maps;

[0024] The two single-channel feature maps are concatenated in dimension and then input into the Softmax function for normalization, and then split into two single-channel attention weight maps;

[0025] The two single-channel attention weight maps are respectively multiplied by the enhanced spatial feature and the initial semantic feature to obtain the attention mechanism-enhanced spatial feature and the attention mechanism-enhanced semantic feature;

[0026] The spatially enhanced feature with the attention mechanism and the semantically enhanced feature with the attention mechanism are accumulated point by point to obtain the final aggregated feature, and the final aggregated feature is fed into the detection box prediction network to output the detection result of the image.

[0027] Furthermore, in the data augmentation of the training set and the validation set, the data augmentation methods include horizontal and vertical flipping, global rotation, random offset, scale scaling, and GT-AUG;

[0028] The horizontal and vertical flipping rate is 0.5, the global rotation angle range is [-π / 4, π / 4], the random offset range is [-0.2, 0.2] m, and the scale range for point cloud data scaling is [0.95, 1.05];

[0029] Furthermore, the GT-AUG data augmentation method is specifically as follows: create an annotated target database where the labels of the target objects correspond one-to-one with the point cloud data. During the training process of the 3D detection model, different target categories are sampled at a certain ratio, and the sampled targets are added to the input point cloud data.

[0030] A 3D object detection system for feeding a target image into a trained 3D detection model to output the detection result of the target image;

[0031] The 3D detection model includes a point cloud processing network, a feature extraction network, and a detection box prediction network connected in sequence;

[0032] The point cloud processing network extracts the point cloud data of the image based on dynamic voxelization to obtain point cloud voxel data, and feeds the point cloud voxel data into the feature extraction network;

[0033] The feature extraction network uses a spatial semantic feature aggregation network as the feature extraction framework, including a spatial feature extraction network, a semantic feature extraction network, and an attention fusion network. The input of the spatial feature extraction network is connected to the output of the point cloud processing network, and after passing through three sequentially connected convolutional neural networks, the initial spatial feature is obtained; the input of the semantic feature extraction network is connected to the output of the spatial feature extraction network, and after passing through three sequentially connected convolutional neural networks, the initial semantic feature is obtained. After deconvolving the initial semantic feature, it is accumulated point by point with the initial spatial feature to obtain the enhanced spatial feature; the enhanced spatial feature and the initial semantic feature are input into the attention fusion network to output the final aggregated feature;

[0034] The detection box prediction network includes a target category prediction branch, a detection box coordinate prediction branch, and a detection box position quality estimation prediction branch that are set in parallel. The inputs of the target category prediction branch, the detection box coordinate regression branch, and the detection box position quality estimation prediction branch are all connected to the output of the feature extraction network. The target category prediction branch consists of a single convolutional network and predicts the target category scores. The detection box coordinate prediction branch consists of a single convolutional network and predicts the center point position of the detection box, the 3D size of the detection box, and the target orientation. The detection box position quality estimation prediction branch consists of two fully connected layers and an activation function. The output of the fully connected layer is used as the input of the activation function to predict the detection box position quality estimation.

[0035] A computer-readable storage medium stores a number of programs thereon, and the number of programs is used to be called by a processor and execute the above-mentioned center point-based three-dimensional object detection method.

[0036] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disc that can store program codes.

[0037] The advantages of a center point-based three-dimensional object detection method, system, and storage medium provided by the present invention are as follows: In the structure of the present invention, a center point-based three-dimensional object detection method, system, and storage medium are provided. Based on the recorded three-dimensional detection model, it is used to solve the rotation alignment problem existing in the anchor point-based center point-based three-dimensional object detection method, solve the information loss and randomness problems existing in the voxelization process of point cloud data, improve the quality of the three-dimensional point cloud feature map, better fuse the high-level semantic features and low-level spatial features of the point cloud feature map, estimate the position quality of the detection box to optimize the detection box quality representation, improve the three-dimensional object detection performance. Finally, a semi-supervised three-dimensional object detection framework is designed to alleviate the cost problem brought by manual annotation. Use a small amount of labeled data to assist in training the three-dimensional detection model with a large amount of unlabeled data, and use a small amount of labeled and single-scene data to train the three-dimensional detection model, so that it can achieve good detection performance on the validation test set containing rich scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flowchart of the present invention;

[0039] Figure 2 is a schematic diagram of a semi-supervised three-dimensional object detection framework of a teacher-student model;

[0040] Figure 3 Schematic diagrams of traditional voxelization and dynamic voxelization;

[0041] Figure 4 Schematic diagram of the model of the feature extraction network. Specific implementation manners

[0042] Next, the technical solution of the present invention will be described in detail through specific embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementations disclosed below.

[0043] As Figures 1 to 4 shown, a three-dimensional object detection method based on a center point proposed by the present invention conveys a target image to a trained three-dimensional detection model to output a detection result of the target image, wherein the three-dimensional object detection method represents a detection box output by the model using the center point position, the 3D size of the detection box, and the target orientation.

[0044] (A) The three-dimensional detection model includes a point cloud processing network, a feature extraction network, and a detection box prediction network connected in sequence.

[0045] (A1) The point cloud processing network extracts point cloud data of an image based on dynamic voxelization to obtain point cloud voxel data, and conveys the point cloud voxel data to the feature extraction network.

[0046] The point cloud processing network is improved based on the VoxelNet network model, and a dynamic voxelization method is adopted to replace the traditional voxelization method. The dynamic voxelization method abandons the step of presetting the tensor dimension in the traditional voxelization method. By extracting the features of all points in the point cloud, it avoids the loss of voxel information caused by random sampling, and at the same time avoids the redundant calculation brought by filling voxels with 0. Since a definite voxel embedding is generated, it overcomes randomness and makes the detection result more stable;

[0047] As Figure 3 shown, the traditional voxelization process is as follows: First, the point cloud data is spatially partitioned, where each small space is called a voxel. Then, the dimension of the preset tensor is K×T×F, where K is the maximum number of voxels, T is the maximum number of points that can be accommodated in each voxel, and F represents the feature dimension. Then, according to the spatial partition and the preset dimension, the points in the point cloud are voxel-matched and randomly sampled. The 13 points in the point cloud data are divided into 4 voxel blocks {V1, V2, V3, V4}, and K is preset to 3 and T is preset to 5, that is, for Figure 3For the point cloud in it, the maximum number of voxel features extracted is 3, and the maximum number of points that can be accommodated in each voxel is 5. First, randomly sample from four pixel blocks to obtain {V1, V2, V4}. At this time, the point cloud information in the entire V3 area is lost. Then, since the number of points in the V1 area exceeds T, randomly sample again from V1 to obtain T points. At this time, part of the point cloud information in the V1 area is lost, and random sampling leads to the uncertainty of the voxelization process. However, the dynamic voxelization process is as follows: First, divide the point cloud data in space, then without presetting the dimension of the tensor, extract all the points in the point cloud data into the feature tensor, and finally use the feature encoding technology to transform the point cloud features into a high-dimensional space, effectively solving the problems of point cloud information loss and randomness.

[0048] (A2) As Figure 4 shown, the feature extraction network uses the spatial semantic feature aggregation network as the feature extraction framework, including a spatial feature extraction network, a semantic feature extraction network, and an attention fusion network. The input of the spatial feature extraction network is connected to the output of the point cloud processing network, and after three sequentially connected convolutional neural networks, the initial spatial feature is obtained; the input of the semantic feature extraction network is connected to the output of the spatial feature extraction network, and after three sequentially connected convolutional neural networks, the initial semantic feature is obtained. After deconvolving the initial semantic feature, it is added point by point to the initial spatial feature to obtain the enhanced spatial feature; the enhanced spatial feature and the initial semantic feature are input into the attention fusion network to output the final aggregated feature.

[0049] The feature extraction network uses the spatial semantic feature aggregation network as the feature extraction framework to enhance the spatial semantic feature extraction ability, as Figure 4As shown in the figure, the spatial semantic feature aggregation network fully integrates the high-level semantic features and low-level spatial features of the point cloud feature map through two aggregations. The first aggregation: Set the number of channels of the spatial feature extraction network to be the same as the number of channels of the input feature map, keep the spatial feature dimension unchanged, double the number of channels of the semantic feature extraction network first, and then use the deconvolution network to restore its dimension to be the same as the dimension of the spatial feature after completion. Then, perform point-by-point addition to obtain the enhanced spatial feature. The second aggregation: Use the deconvolution network to perform upsampling on the semantic feature network for the semantic feature, and then use the fusion network of the attention mechanism to fuse the enhanced spatial feature and the upsampled semantic feature. The enhanced spatial feature and the upsampled semantic feature are used as fusion data and sent to the detection box prediction network to output the detection result of the image. The specific operation of the second aggregation is as follows: The enhanced spatial feature and the initial semantic feature are respectively input into a layer of convolutional network to obtain two single-channel feature maps; the two single-channel feature maps are concatenated in dimension and then input into the Softmax function for normalization, and then split into two single-channel attention weight maps; the two single-channel attention weight maps are respectively multiplied by the enhanced spatial feature and the initial semantic feature to obtain the attention mechanism-enhanced spatial feature and the attention mechanism-enhanced semantic feature; the attention mechanism-enhanced spatial feature and the attention mechanism-enhanced semantic feature are added point by point to obtain the final aggregated feature, and the final aggregated feature is sent to the detection box prediction network to output the detection result of the image.

[0050] (A3) The detection box prediction network includes a target category prediction branch, a detection box coordinate prediction branch, and a detection box position quality estimation prediction branch arranged in parallel. The inputs of the target category prediction branch, the detection box coordinate prediction branch, and the detection box position quality estimation prediction branch are all connected to the output of the feature extraction network (the final aggregated feature output). The target category prediction branch consists of a layer of convolutional network and predicts the target category score. The detection box coordinate prediction branch consists of a layer of convolutional network and predicts the center point position of the detection box, the 3D size of the detection box, and the target orientation. The detection box position quality estimation prediction branch consists of two fully connected layers and an activation function. The output of the fully connected layer is used as the input of the activation function to predict the detection box position quality estimation.

[0051] The detection box prediction network adopts center - point - based 3D detection box prediction. This detection box prediction network uses the center point to replace a large number of anchor points in the traditional method for predicting detection boxes. The regression branch includes the prediction of the center point position, the 3D size prediction of the detection box, and the target orientation prediction, which reduces the search space while maintaining the rotational invariance of the target. An IoU (Intersection over Union) prediction branch is designed at the head of the detection box prediction network. The designed IoU prediction branch is parallel to the classification branch and the regression branch in the 3D detection model. The IoU prediction branch consists of two fully - connected layers and a sigmoid activation function, which is used to predict the intersection over union between the detection box output by the 3D detection model and the ground - truth box, representing the position quality of the detection box, optimizing the quality representation of the detection box, and being used for the post - processing process. In addition, in this embodiment, a deformable convolutional network is introduced on the target category prediction branch and the detection box coordinate prediction branch respectively, and the ordinary convolutional network is replaced with the deformable convolutional network to enhance the model's ability to resist random uncertainties.

[0052] (B), as Figure 2 shown, in this embodiment, a semi - supervised 3D object detection framework based on the teacher - student model is then designed. Both the teacher model and the student model adopt the 3D detection model proposed in (A). The training of the semi - supervised network is constrained by the consistency regularization loss function between the prediction results of the teacher model and the student model, and the student model is used as the final 3D detection model.

[0053] (C) Under the combination of (A) and (B), the training process of the 3D detection model is as follows.

[0054] S1: Construct a training set and a validation set. The training set is a set of unlabeled point cloud data, and the validation set is a set of labeled point cloud data.

[0055] S2: Perform data augmentation on the training set and the validation set, and extract all point features of the point cloud data in the augmented training set based on dynamic voxelization in the point cloud processing network to obtain point cloud voxel data;

[0056] In the data augmentation of the training set and the validation set, the data augmentation methods include horizontal and vertical flipping, global rotation, random offset, scale scaling, and GT - AUG;

[0057] The horizontal and vertical flipping rate is 0.5, the global rotation angle range is [-π / 4, π / 4], the random offset range is [-0.2, 0.2] m, and the scale scaling range of the point cloud data is [0.95, 1.05];

[0058] The GT-AUG data augmentation method is specifically as follows: create an annotated target database where the labels of the target objects correspond one-to-one with the point cloud data. During the training process of the 3D detection model, different target categories are sampled according to a certain ratio. In a specific embodiment, vehicles, pedestrians, and bicycles are sampled at a ratio of 4:1:1, and the sampled targets are added to the input point cloud data to improve the performance of the 3D detection model.

[0059] S3: Sequentially feed the point cloud voxel data into the feature extraction network and the detection box prediction network. Predict the intersection over union (IoU) between the detection boxes output by the 3D detection model and the ground truth boxes through the detection box position quality estimation prediction branch to optimize the quality of the 3D detection model.

[0060] Among them, the processing processes of the feature extraction network and the detection box prediction network for the point cloud voxel data are detailed in (A).

[0061] S4: Feed the validation set into the 3D detection model after weak data augmentation to train the 3D detection model in a supervised manner, and obtain the trained 3D detection model as the teacher model.

[0062] Specifically: Use a small amount of labeled point cloud data as input to train the teacher model in a supervised manner. In addition, weak data augmentation includes horizontal and vertical flipping, global rotation, random offset, and scale scaling.

[0063] S5: Feed the training set into the teacher model after weak data augmentation to output the predicted values of the images. At the same time, feed the training set into the 3D detection model to be trained after strong data augmentation to train the 3D detection model in a semi-supervised manner. At this time, the 3D detection model outputs the predicted values of the images as the student model.

[0064] Specifically: Use a large amount of unlabeled point cloud data as input to train the student model in a semi-supervised manner. First, feed the unlabeled point cloud data into the teacher model after weak data augmentation, and feed the unlabeled point cloud data into the student model after strong data augmentation. The two are fed synchronously. In addition, strong data augmentation includes horizontal and vertical flipping, global rotation, random offset, scale scaling, and GT-AUG.

[0065] Using a small amount of labeled data to assist in training the 3D detection model with a large amount of unlabeled data, and using a small amount of labeled and single-scene data to train the 3D detection model to achieve good detection performance on the validation and test sets containing rich scenes can alleviate the cost problem brought by manual annotation.

[0066] S6: Use the predicted values output by the teacher model as pseudo-labels, and constrain the predicted values output by the student model based on the pseudo-labels through the consistency regularization loss function to obtain the trained 3D detection model.

[0067] The consistency regularization loss function ensures the consistency between the student model and the teacher model by constraining the student model to learn the dense prediction results of the teacher model. The specific formula is as follows:

[0068]

[0069] Where, L represents the consistency regularization loss function, H and W represent the height and width of the feature map, i and j represent the positions of points on the feature map, C t ,B t represents the prediction results of the classification branch and the regression branch of the teacher model, C S ,B s represents the prediction results of the classification branch and the regression branch of the student model, and CE represents the cross-entropy loss function.

[0070] (D) The student model (corresponding to the 3D detection model) obtains the trained student model through the above training process, and uses the student model as the final 3D detection model for use.

[0071] Before the student model is actually used, it is necessary to test and verify the accuracy of the model. At this time, a test set is constructed, which is a set of labeled point cloud data. The test set is fed into the trained student model to obtain the target detection results output by the student model. The detection results are compared with the actual results. If it is within the preset threshold range, it means that the student model is trained and can be put into actual use. If it is not within the preset threshold range, it means that the student model needs further training until the comparison result is within the preset threshold range.

[0072] (E) The student model is put into actual use

[0073] The target image is fed into the trained student model to output the detection results of the target image.

[0074] In this embodiment, the 3D detection model described in (A) is first proposed to solve the rotation alignment problem existing in the center point-based 3D object detection method based on anchors, solve the information loss and randomness problems existing in the voxelization process of point cloud data, improve the quality of 3D point cloud feature maps and the position quality of 3D prediction boxes, and improve 3D object detection performance; then a semi-supervised 3D object detection framework described in (B) is designed to alleviate the cost problem caused by manual annotation, use a small amount of labeled data to assist in training the 3D detection model with a large amount of unlabeled data, and use a small amount of labeled and single-scene data to train the 3D detection model to achieve good detection performance on the validation test set containing rich scenes; through the training in (C) and the testing in (D), the output accuracy of the student model in (E) is improved.

[0075] As described above, it is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. A three-dimensional object detection method based on the center point, characterized in that, Feed the target image into the trained 3D detection model to output the detection result of the target image; The 3D detection model includes a point cloud processing network, a feature extraction network, and a detection box prediction network connected in sequence; The point cloud processing network extracts the point cloud data of the image based on dynamic voxelization to obtain point cloud voxel data, and feeds the point cloud voxel data into the feature extraction network; The feature extraction network uses a spatial semantic feature aggregation network as the feature extraction framework, including a spatial feature extraction network, a semantic feature extraction network, and an attention fusion network. The input of the spatial feature extraction network is connected to the output of the point cloud processing network, and after passing through three sequentially connected convolutional neural networks, the initial spatial features are obtained; The input of the semantic feature extraction network is connected to the output of the spatial feature extraction network. After passing through three sequentially connected convolutional neural networks, the initial semantic features are obtained. After deconvolving the initial semantic features, they are added point by point to the initial spatial features to obtain enhanced spatial features; The enhanced spatial features and the initial semantic features are input into the attention fusion network to output the final aggregated features; The detection box prediction network includes a target class prediction branch, a detection box coordinate prediction branch, and a detection box position quality estimation prediction branch arranged in parallel. The inputs of the target class prediction branch, the detection box coordinate regression branch, and the detection box position quality estimation prediction branch are all connected to the output of the feature extraction network. The target class prediction branch consists of one convolutional network and predicts the target class score. The detection box coordinate prediction branch consists of one convolutional network and predicts the center point position of the detection box, the 3D size of the detection box, and the target orientation. The detection box position quality estimation prediction branch consists of two fully connected layers and one activation function. The output of the fully connected layer is used as the input of the activation function to predict the detection box position quality estimation.

2. The three-dimensional object detection method based on the center point according to claim 1, wherein, The training process of the 3D detection model is as follows: Construct a training set and a validation set. The training set is a set of unlabeled point cloud data, and the validation set is a set of labeled point cloud data; Perform data augmentation on the training set and the validation set, and based on the dynamic voxelization in the point cloud processing network, extract all point features of the point cloud data in the augmented training set to obtain point cloud voxel data; Feed the point cloud voxel data into the feature extraction network and the detection box prediction network in sequence, and predict the intersection over union between the detection box output by the 3D detection model and the ground truth box through the detection box position quality estimation prediction branch to optimize the quality of the 3D detection model; Feed the validation set into the 3D detection model after weak data augmentation to train the 3D detection model in a supervised manner to obtain the trained 3D detection model as the teacher model; Feed the training set into the teacher model after weak data augmentation to output the predicted value of the image. At the same time, feed the training set into the 3D detection model to be trained after strong data augmentation to train the 3D detection model in a semi-supervised manner. At this time, the 3D detection model, as the student model, outputs the predicted value of the image; Take the predicted values output by the teacher model as pseudo-labels, and constrain the predicted values output by the student model based on the pseudo-labels through a consistency regularization loss function to obtain a trained 3D detection model.

3. The three-dimensional object detection method based on the center point according to claim 2, wherein, The formula of the consistency regularization loss function is as follows: Among them, represents the consistency regularization loss function, and represent the height and width of the feature map, represents the position of a point on the feature map, , represent the prediction results of the classification branch and the regression branch of the teacher model, , represent the prediction results of the classification branch and the regression branch of the student model, and CE represents the cross-entropy loss function.

4. The 3D object detection method based on the center point according to claim 1, characterized in that When inputting the enhanced spatial feature and the initial semantic feature into the attention fusion network and outputting the final aggregated feature, it specifically includes: Input the enhanced spatial feature and the initial semantic feature into a convolutional network respectively to obtain two single-channel feature maps; Concatenate the two single-channel feature maps in dimension and then input them into the Softmax function for normalization, and then split them into two single-channel attention weight maps; Multiply the two single-channel attention weight maps by the enhanced spatial feature and the initial semantic feature respectively to obtain the attention mechanism-enhanced spatial feature and the attention mechanism-enhanced semantic feature; Pointwise accumulate the attention mechanism-enhanced spatial feature and the attention mechanism-enhanced semantic feature to obtain the final aggregated feature, and send the final aggregated feature to the detection box prediction network to output the detection result of the image.

5. The three-dimensional object detection method based on the center point according to claim 2, wherein In the data augmentation of the training set and the validation set, the data augmentation methods include horizontal and vertical flipping, global rotation, random offset, scale scaling, and GT-AUG; The horizontal and vertical flipping rate is 0.5, the global rotation angle range is [-π / 4, π / 4], the random offset range is [-0.2, 0.2] m, and the scale scaling range of the point cloud data is [0.95, 1.05].

6. The three-dimensional object detection method based on the center point according to claim 5, wherein The GT-AUG data augmentation method is specifically: create an annotated target database, where the labels of the target objects correspond one-to-one with the point cloud data. During the training process of the 3D detection model, different target categories are sampled according to a certain proportion, and the sampled targets are added to the input point cloud data.

7. A three-dimensional object detection system, characterized in that, Used to send the target image into the trained 3D detection model to output the detection result of the target image; The 3D detection model includes a point cloud processing network, a feature extraction network, and a detection box prediction network connected in sequence; The point cloud processing network extracts the point cloud data of the image based on dynamic voxelization to obtain point cloud voxel data, and sends the point cloud voxel data to the feature extraction network; The feature extraction network uses a spatial semantic feature aggregation network as the feature extraction backbone, including a spatial feature extraction network, a semantic feature extraction network, and an attention fusion network. The input of the spatial feature extraction network is connected to the output of the point cloud processing network, and after three sequentially connected convolutional neural networks, the initial spatial feature is obtained; The input of the semantic feature extraction network is connected to the output of the spatial feature extraction network, and after three sequentially connected convolutional neural networks, the initial semantic feature is obtained. After deconvolving the initial semantic feature, it is pointwise accumulated with the initial spatial feature to obtain the enhanced spatial feature; Input the enhanced spatial feature and the initial semantic feature into the attention fusion network to output the final aggregated feature; The detection box prediction network includes a target class prediction branch, a detection box coordinate prediction branch, and a detection box position quality estimation prediction branch that are arranged in parallel. The inputs of the target class prediction branch, the detection box coordinate regression branch, and the detection box position quality estimation prediction branch are all connected to the output of the feature extraction network. The target class prediction branch consists of a single convolutional network and predicts the target class scores. The detection box coordinate prediction branch consists of a single convolutional network and predicts the center point position of the detection box, the 3D size of the detection box, and the target orientation. The detection box position quality estimation prediction branch consists of two fully connected layers and an activation function. The output of the fully connected layer is used as the input of the activation function to predict the detection box position quality estimation.

8. A computer-readable storage medium, characterized in that, A number of programs are stored on the computer-readable storage medium, and the number of programs is used to be called by the processor and execute the center point-based three-dimensional object detection method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Target detection model training method and device, electronic equipment and storage medium

    CN111241964A

  • Traffic target identification method and system

    CN114627437A