3D Object Detection Method with Multimodal Input and Spatial Partitioning
Through the three-dimensional object detection method of multimodal input and spatial division, combining point cloud and color image feature extraction, effective grouping and fusing features, the problem of point cloud data lacking texture and color information is solved, and the accuracy and real-time detection are improved.
Patent Information
- Application Number
- CN202111173391.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-10-08
AI Technical Summary
In the existing three-dimensional object detection methods, the lack of texture and color information of the original point cloud data leads to low recognition accuracy, and the grouping of no objects in the three-dimensional spatial grouping leads to wasted computing resources, affecting the real-time and accuracy of detection.
The multimodal input combined with spatial division method is used, and the original point cloud data and RGB three-channel color images are used as inputs. Features are extracted through PointNet and VGG16, point cloud packets of possible objects are filtered, dimensionality is reduced, and feature vectors are fused, and the border loss function is used to train the network.
It improves the accuracy and real-time nature of three-dimensional object detection, reduces the amount of computing, makes up for the problem of missing information, and improves the accuracy of classification and detection.
Smart Images

Figure CN114118125B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of object detection in artificial intelligence, and particularly relates to a three-dimensional object detection method with multi-modal input and space division. Background Art
[0002] With the advent of the era of artificial intelligence, people begin to pursue a more intelligent lifestyle. As an essential means of transportation in life, automobiles urgently need an intelligent driving system as an assistant to provide convenience for drivers. The primary challenge in intelligent driving is to understand the real world, and the most effective and commonly used way to enable an automobile to understand the world is through computer vision (CV). Intelligent driving is a technology with high requirements for safety and is mainly affected by two aspects: (1) accuracy; (2) real-time performance.
[0003] Previous work mostly used raw point cloud data as input and extracted three-dimensional features through a deep convolutional network. However, such a method has some drawbacks, that is, due to the overly single input information, the raw point cloud data lacks texture and color information. Therefore, only using raw point cloud data as input will result in low recognition accuracy. In addition, in previous work, some people tried to use the method of dividing the three-dimensional space and adopted CNN (Convolutional Neural Network) for feature extraction. However, due to the sparsity of the raw point cloud and the characteristic that most points are distributed on the surface of objects, there are no objects in most three-dimensional space groups. Therefore, extracting three-dimensional features in this way will cause a waste of a large amount of computing resources. To improve the real-time performance and accuracy of detection, it is of extremely important research significance to improve the existing problems of three-dimensional object detection. Summary of the Invention
[0004] Aiming at the deficiencies in the prior art, the present invention proposes a brand-new three-dimensional object detection method with multi-modal input and space division. The goal of this method is to improve the accuracy and real-time performance of three-dimensional object detection, and the following targeted strategies are respectively proposed. (1) This method uses multi-modal input and directly extracts the color and texture information of the image through the two-dimensional convolutional neural network VGG16 to make up for the lack of information caused by only using raw point cloud data as input. (2) This method still uses the method of dividing the three-dimensional space into several groups, but before extracting features, these groups are screened first, and only the three-dimensional groups with objects are used to extract features and perform subsequent predictions, and the remaining three-dimensional groups are directly discarded, which significantly improves the operation efficiency of the present invention. (3) The post-processing task directly performs regression prediction information and draws a BBox (Bonding Box), which belongs to a single-stage object detection method.
[0005] To achieve the above objectives, the present invention adopts the following technical solutions:
[0006] A three-dimensional object detection method for multi-modal input and spatial division, comprising the following steps:
[0007] (1) Using the original point cloud data and RGB three-channel color images as multi-modal input;
[0008] (2) Divide the space of the original point cloud data, index the point cloud groups row by row and column by column. Each point cloud group is a subset of the original point cloud data. Randomly sample each point cloud group, sample K points from each point cloud group, input them into PointNet to extract features, obtain K feature vectors, and reduce the dimension of these K feature vectors through a max pooling layer to obtain K local-global feature vectors;
[0009] (3) Split the RGB three-channel color image, index the slices row by row, and input them into the two-dimensional feature extractor VGG16 to extract only the shallow related features of the texture color of the 8th layer, and obtain K color texture feature vectors extracted from the RGB three-channel color image;
[0010] (4) Fuse the K local-global feature vectors and the K color texture feature vectors to obtain a fused feature vector;
[0011] (5) The fused feature vector passes through a fully connected layer to obtain the output prediction result. According to the confidence level, draw the BBox (Bonding Box) to complete the post-processing task.
[0012] To optimize the above technical solutions, the specific measures taken also include:
[0013] Furthermore, extract the features of the original point cloud data: Divide the space of the original point cloud data, index row by row and column by column. There are a total of W*H*D point cloud groups, and the index numbers are {0, 1, 2...W*H*D}, where W, H, and D respectively represent the number of spatial divisions in the width direction, height direction, and depth direction;
[0014] According to the distribution of points, screen out the point cloud groups where objects may exist, and remove the point cloud groups where there are no objects. If a group does not contain points, it will be removed and is not responsible for predicting objects;
[0015] Randomly sample the remaining point cloud groups, sample K points from each point cloud group, input them into PointNet to obtain K feature vectors, and then reduce the dimension of these K feature vectors in the depth direction through a max pooling layer to obtain K 1024-dimensional local-global feature vectors.
[0016] Further, extract the shallow relevant features of the colors of the RGB three-channel color image: Divide the RGB three-channel color image, index it row by row, with a total of W*H slices, and the slice numbers are {0, 1, 2...W*H}, where W and H respectively represent the number of spatial divisions in the width direction and the height direction;
[0017] Input each slice into the two-dimensional feature extractor VGG16, and only extract the shallow relevant features of the texture colors of the 8th layer to obtain K color texture feature vectors extracted from the RGB three-channel color image.
[0018] Further, predict the fused feature vectors, complete the post-processing task, and train the entire network through the loss function to obtain the output prediction result:
[0019] For each point cloud group with extracted features, it is necessary to judge the possibility of containing the target to be detected within this group, through the confidence loss to measure:
[0020]
[0021]
[0022]
[0023]
[0024] Among them, G IoU represents the bounding box loss function, IoU represents the intersection over union of the BBox (Bounding Box) and the ground truth, A c represents the volume of the smallest cubic region enclosing the BBox and the ground truth, and u represents the union volume of the BBox and the ground truth; is the predicted confidence obtained by the i-th predicted value C i through the Sigmoid function; O i represents the coincidence degree of the i-th BBox and the ground truth; represents the confidence loss, and N is the number of positive and negative samples;
[0025] Total loss function is defined as follows:
[0026]
[0027] Among them, N pos is the number of positive samples, λ conf , λ loc , λ cls , λ dir respectively represent the balance coefficients of the confidence loss, the localization loss, the classification loss, and the orientation angle loss, They respectively represent confidence loss, localization loss, classification loss, and orientation angle loss; finally, the entire network is trained through the loss function.
[0028] The beneficial effects of the present invention are as follows:
[0029] (1) Divide the original point cloud data space into W*H*D groups, and sample K points from each point cloud group, which can reduce the computational amount and offset the influence of the difference (dense near and sparse far) brought by the distance in the radar's acquisition of the original point cloud data; in addition, before extracting features, first screen out the point cloud groups where objects may exist according to the distribution of points, and remove the point cloud groups where there are no objects, significantly reducing the computational amount;
[0030] (2) Adopt multi-modal input and introduce the shallow features of the RGB three-channel color image to make up for the deficiency of the existing PointNet that only inputs the original point cloud data and causes information losses such as color and texture, and improve the accuracy of classification and detection;
[0031] (3) Divide the original point cloud data space into multiple groups, extract features from each group through PointNet, convert the original point cloud data into structured data, provide a basis for further fusion with the feature vectors extracted from the RGB three-channel color image, and after structuring the data, the idea of two-dimensional object detection can be adopted to predict each group. Description of the Drawings
[0032] Figure 1 It is a schematic diagram of the method of the present invention. Detailed Embodiment
[0033] Now, the present invention will be further described in detail with reference to the accompanying drawings.
[0034] It should be noted that the terms such as "up", "down", "left", "right", "front", "back", etc. cited in the invention are only for the convenience of description and are not used to limit the scope of implementation of the present invention. The change or adjustment of their relative relationship, without substantial change in the technical content, should also be regarded as the scope of implementation of the present invention.
[0035] Such as Figure 1 , the present invention provides a three-dimensional object detection method with multi-modal input and spatial division, including the following steps:
[0036] Use the original point cloud data and the RGB three-channel color image as multi-modal input;
[0037] Partition the space of the original point cloud data, index the point cloud groups row by row and column by column. Each point cloud group is a subset of the original point cloud data. Randomly sample each point cloud group, sample K points from each point cloud group, input them into PointNet to extract features, obtain K feature vectors, and reduce the dimensionality of these K feature vectors through a max pooling layer to obtain K local-global feature vectors;
[0038] Slice the RGB three-channel color image, index the slices row by row, and input them into the two-dimensional feature extractor VGG16. Only extract the shallow relevant features of the texture color of the 8th layer to obtain K color texture feature vectors extracted from the RGB three-channel color image;
[0039] Fuse the K local-global feature vectors and the K color texture feature vectors to obtain the fused feature vectors;
[0040] The fused feature vectors pass through a fully connected layer to obtain the output prediction results. According to the confidence level, draw the BBox (Bonding Box) to complete the post-processing task.
[0041] Explanation from the theoretical basis:
[0042] (1) Extract the features of the original point cloud data: Partition the space of the original point cloud data, index row by row and column by column. There are a total of W*H*D point cloud groups, and the index numbers are {0, 1, 2...W*H*D}, where W, H, and D respectively represent the number of spatial partitions in the width direction, height direction, and depth direction;
[0043] According to the distribution of points, screen out the point cloud groups where objects may exist, and remove the point cloud groups where there are no objects. If a group does not contain points, it will be removed and is not responsible for predicting objects;
[0044] Randomly sample the remaining point cloud groups, sample K points from each point cloud group, input them into PointNet to obtain K feature vectors, and then reduce the dimensionality of these K feature vectors in the depth direction through a max pooling layer to obtain K 1024-dimensional local-global feature vectors.
[0045] (2) Extract the shallow relevant features of the color of the RGB three-channel color image: Slice the RGB three-channel color image, index row by row. There are a total of W*H slices, and the index numbers are {0, 1, 2...W*H}, where W and H respectively represent the number of spatial partitions in the width direction and height direction;
[0046] Input each slice into the two-dimensional feature extractor VGG16, and only extract the shallow relevant features of the texture color of the 8th layer to obtain K color texture feature vectors extracted from the RGB three-channel color image.
[0047] (3)Fuse the K local-global feature vectors and the K color texture feature vectors to obtain the fused feature vectors; introduce multi-modal input. After feature fusion, 1) it can effectively solve the classification accuracy problem caused by the lack of texture and color information in the original point cloud data; 2) it can effectively solve the problem of inaccurate positioning of the bounding box due to the lack of three-dimensional information in the image data.
[0048] (4) Loss function: Predict the fused feature vectors, complete the post-processing tasks, and train the entire network through the loss function to obtain the predicted output results. For each point cloud group with extracted features, it is necessary to judge the possibility of containing the target to be detected within the group, and measure it through the confidence loss as follows:
[0049]
[0050]
[0051]
[0052]
[0053] Among them, G IoU represents the bounding box loss function, IoU represents the intersection-over-union ratio of the BBox (Bounding Box) and the ground truth, A C represents the volume of the smallest cubic region enclosing the BBox and the ground truth, and u represents the union volume of the BBox and the ground truth; is the predicted confidence obtained by the i-th predicted value C i through the Sigmoid function; O i represents the coincidence degree of the i-th BBox and the ground truth; represents the confidence loss, N is the number of positive and negative samples, and ln represents the logarithmic function;
[0054] Each 3D BBox is represented by x, y, z, w, l, h, θ, where x, y, z represent the three-dimensional coordinates of the object center, w, l, h represent the width, length, and height dimension data, and θ represents the horizontal direction angle in the radar coordinate system. Then the parameters to be learned in the detection box regression task are the offsets of these 7 variables:
[0055]
[0056]
[0057]
[0058] Δθ = sin(θgt -θ a )
[0059] Δx, Δy, Δz, Δw, Δl, Δh, Δθ respectively represent the relative offsets of the BBox from their corresponding true values, where x gt , y gt , z gt , w gt , l gt , h gt , θ gt represent the true values of the three-dimensional coordinates, width, length, height, and horizontal angle, and x a , y a , z a , w a , l a , h a , θ a respectively represent the predicted values of the three-dimensional coordinates, width, length, height, and horizontal angle;
[0060] Classification loss Adopts the multi-class cross-entropy loss function, defined as follows:
[0061]
[0062] where t k represents the true value of the k-th correct classification, and y k represents the predicted result of the output of the k-th neural network;
[0063] Total loss function Is defined as follows:
[0064]
[0065] where N pos is the number of positive samples, and λ conf , λ loc , λ cls , λ dir respectively represent the balance coefficients of the confidence loss, localization loss, classification loss, and orientation angle loss, respectively represent the confidence loss, localization loss, classification loss, and orientation angle loss; finally, the entire network is trained through the loss function.
[0066] The present invention divides the original point cloud data space into W*H*D groups, and samples K points from each point cloud group, which can reduce the computational amount and offset the influence of the difference (dense near and sparse far) brought by the distance in the original point cloud data obtained by the radar. In addition, before extracting features, first, according to the distribution of points, the point cloud groups where objects may exist are screened out, and the point cloud groups where no objects exist are removed, significantly reducing the computational amount. The present invention adopts multi-modal input and introduces the shallow features of the RGB three-channel color image, making up for the deficiency of the existing PointNet which only inputs the original point cloud data and causes information loss such as color and texture, and improving the accuracy of classification and detection. The present invention divides the original point cloud data space into multiple groups, extracts features from each group through PointNet, converts the original point cloud data into structured data, provides a basis for further fusion with the feature vectors extracted from the RGB three-channel color image, and after structuring the data, the idea of two-dimensional object detection can be adopted to predict each group.
[0067] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art, several improvements and retouches made without departing from the principle of the present invention should be regarded as the protection scope of the present invention.
Claims
1. A three-dimensional object detection method for multimodal input and spatial division, characterized in that, It includes the following steps: (1) Using the original point cloud data and the RGB three-channel color image as multi-modal inputs; (2) Dividing the space of the original point cloud data, indexing the point cloud groups row by row and column by column. Each point cloud group is a subset of the original point cloud data. Randomly sampling each point cloud group, sampling K points from each point cloud group, inputting them into PointNet to extract features, obtaining K feature vectors, and reducing the dimension of these K feature vectors through a max pooling layer to obtain K local-global feature vectors; (3) Splitting the RGB three-channel color image, indexing the slices row by row, inputting them into the two-dimensional feature extractor VGG16, and only extracting the shallow relevant features of the texture color of the 8th layer to obtain K color texture feature vectors extracted from the RGB three-channel color image; (4) Fusing the K local-global feature vectors and the K color texture feature vectors to obtain the fused feature vectors; (5) The fused feature vectors pass through a fully connected layer to obtain the output prediction results. According to the confidence level, draw the BBox to complete the post-processing task; In step 3, when extracting the shallow relevant features of the color of the RGB three-channel color image, the RGB three-channel color image is split and indexed row by row. There are a total of W*H slices, and the slice numbers are {0, 1, 2...W*H}, where W and H respectively represent the number of spatial divisions in the width direction and the height direction; Inputting each slice into the two-dimensional feature extractor VGG16, and only extracting the shallow relevant features of the texture color of the 8th layer to obtain K color texture feature vectors extracted from the RGB three-channel color image; In step 2, when extracting the features of the original point cloud data, the space of the original point cloud data is divided and indexed row by row and column by column. There are a total of W*H*D point cloud groups, and the group numbers are {0, 1, 2...W*H*D}, where W, H, and D respectively represent the number of spatial divisions in the width direction, the height direction, and the depth direction; According to the distribution of the points, screening out the point cloud groups where objects may exist, removing the point cloud groups where no objects exist. If a group does not contain points, it is removed and is not responsible for predicting objects; Randomly sampling the remaining point cloud groups, sampling K points from each point cloud group, inputting them into PointNet to obtain K feature vectors, and then reducing the dimension of these K feature vectors in the depth direction through a max pooling layer to obtain K 1024-dimensional local-global feature vectors.
2. The three-dimensional object detection method for multi-modal input and spatial division according to claim 1, wherein In step 5, predicting the fused feature vectors to complete the post-processing task, and training the entire network through a loss function to obtain the output prediction results; For each point cloud grouping with features extracted, it is necessary to determine the possibility that the grouping contains the target to be detected, through confidence loss to measure: Among them, G IoU represents the bounding box loss function, IoU represents the intersection over union of the BBox and the ground truth, A C represents the volume of the smallest cubic region enclosing the BBox and the ground truth, and u represents the volume of the union of the BBox and the ground truth; is the predicted value C of the i-th i predicted confidence obtained through the Sigmoid function; O i represents the coincidence degree of the i-th BBox and the ground truth; represents the confidence loss, and N is the number of positive and negative samples; Total loss function is defined as follows: Among them, N pos is the number of positive samples, λ conf , λ loc , λ cls , λ dir respectively represent the balance coefficients of confidence loss, localization loss, classification loss, and orientation angle loss, respectively represent confidence loss, localization loss, classification loss, and orientation angle loss; finally, the entire network is trained through the loss function.
Citation Information
Patent Citations
Vehicle detection method based on monocular vision and laser radar fusion
CN111291714A
Power transmission line abnormal target detection method based on improved YOLOv3
CN111444809A
Unmanned aerial vehicle target detection method in laser point cloud
CN111738214A
Reinforcing steel bar cluster classification method and device, electronic equipment and storage medium
CN112287992A