A 3D object detection method and device for unmanned systems based on multi-modal fusion
Through the multimodal fusion of three-dimensional object detection method, two-dimensional images and three-dimensional point cloud data are used to improve the recognition ability of distant objects in unmanned driving systems, solve the problem of weak distant target modeling capabilities, and improve detection accuracy and efficiency.
Patent Information
- Application Number
- CN202410292746.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-03-14
AI Technical Summary
The ability to recognize distant objects in unmanned driving systems is weak, resulting in weak modeling capabilities of distant objects, affecting the recognition effect.
The three-dimensional object detection method of multimodal fusion is adopted. By acquiring two-dimensional image data and three-dimensional point cloud data, the point cloud feature extraction network, timing dynamic convolution fusion module, adaptive gated fusion module, point cloud density fusion module and density confidence prediction module are used for feature fusion and detection, thereby improving the recognition ability of distant targets.
It improves the recognition ability of distant objects, reduces misjudgment behavior, enhances the driving performance of unmanned driving systems in complex environments, and reduces the occurrence of traffic accidents.
Smart Images

Figure CN118172632B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target detection, and particularly relates to a three-dimensional target detection method and device for an unmanned system based on multi-modal fusion. Background Art
[0002] In recent years, with the vigorous development of artificial intelligence technology, the intelligent requirements for various unmanned systems have been continuously improved. Among them, as a typical unmanned system, the driverless system has shown great potential in reducing traffic congestion, improving road safety and optimizing traffic flow efficiency. This system enables vehicles to intelligently perceive the surrounding environment and achieve safe driving with little human intervention, thus reducing human errors and improving road safety. However, achieving fully autonomous driving is still a daunting task, mainly because the complex road environment poses a huge challenge to intelligent perception.
[0003] In the driverless scenario, lidar emits lasers at certain intervals at different vertical angles, resulting in a larger point spacing of the point cloud of distant targets than that of the point cloud of nearby targets. Targets at the same height may obtain more laser points nearby, and the distant point cloud is sparser, which will lead to relatively weak modeling ability of distant targets, thus resulting in weak recognition ability of distant objects. Summary of the Invention
[0004] The present application aims at the problem of weak recognition ability of distant objects in an unmanned system, and provides a three-dimensional target detection method and device for an unmanned system based on multi-modal fusion.
[0005] In a first aspect, a three-dimensional target detection method for an unmanned system based on multi-modal fusion is provided, including the following steps:
[0006] Obtain two-dimensional image data and three-dimensional point cloud data of the road scene to be detected;
[0007] Use the two-dimensional image data and three-dimensional point cloud data of the road scene to be detected as the input of a pre-constructed multi-modal fusion detection model; the multi-modal fusion detection model includes:
[0008] A point cloud feature extraction network for extracting three-dimensional point cloud features of the three-dimensional point cloud;
[0009] The sequential dynamic convolution fusion module discretizes the camera frustum space to generate grid points, converts the grid points into three-dimensional world space coordinates using camera parameters, and uses a three-dimensional position embedding encoder to generate the three-dimensional position perception feature of the previous frame based on the three-dimensional world space coordinates and the two-dimensional image features of the previous frame. The three-dimensional position perception feature of the previous frame is used as the supplementary position perception feature of the current frame after pose transformation; the first fused frame three-dimensional point cloud feature is generated based on the supplementary position perception feature of the current frame and the three-dimensional point cloud feature of the current frame; the preliminary fusion feature is obtained based on the first fused frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame;
[0010] The multi-modal fusion detection model outputs a detection result based on the preliminary fusion feature.
[0011] Optionally, the multi-modal fusion detection model further includes an image feature extraction network for extracting two-dimensional image features of two-dimensional image data;
[0012] In the dynamic convolution module, obtaining the preliminary fusion feature based on the first fused frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame includes:
[0013] Using sub-manifold sparse convolution to predict the importance score of each dimension of the first fused frame three-dimensional point cloud feature and generate an importance weight distribution map;
[0014] If the weight of the first fused frame three-dimensional point cloud feature is greater than or equal to a set importance discrimination threshold, the feature is considered an important feature;
[0015] Performing sparse convolution calculation on the important features to obtain the second fused frame three-dimensional point cloud feature;
[0016] Stitching and fusing the second fused frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame to obtain the preliminary fusion feature.
[0017] Optionally, performing the sparse convolution calculation on the important features to obtain the second fused frame three-dimensional point cloud feature includes:
[0018] After filling the corresponding extended positions of the important features with zero vectors of the same dimension, using sub-manifold sparse convolution for feature extraction to obtain the second fused frame three-dimensional point cloud feature.
[0019] Optionally, the multi-modal fusion detection model further includes an adaptive gating fusion module,
[0020] The adaptive gating fusion module is used to obtain a gating fusion feature and a gating image feature using the preliminary fusion feature and the two-dimensional image feature; project the gating feature onto the gating image feature and perform a convolution operation to obtain a gating image enhancement feature.
[0021] Optionally, the multi-modal fusion detection model further includes a point cloud density fusion module;
[0022] The point cloud density fusion module is used to calculate the point centroid of the gated fusion feature; obtain a reference point in the image plane according to the point centroid of the gated fusion feature;
[0023] weight a group of gated image enhancement features around the reference point to generate an aggregated image feature;
[0024] fuse the aggregated image feature and the gated fusion feature to obtain a fused cross-modal feature;
[0025] perform density-aware regional grid pooling on the fused cross-modal feature to obtain a fused cross-modal flattened feature.
[0026] Optionally, the multi-modal fusion detection model further includes a density confidence prediction module, and the density confidence prediction module includes a shared feed-forward network, a bounding box feed-forward network, and a confidence feed-forward network;
[0027] After the shared feed-forward network encodes the fused cross-modal flattened feature, it is respectively input into the bounding box feed-forward network and the confidence feed-forward network;
[0028] The bounding box feed-forward network outputs a bounding box and the centroid of the final bounding box according to the encoding result;
[0029] The confidence feed-forward network predicts the confidence according to the encoding result, the centroid of the final bounding box, and the original number of points of the final bounding box;
[0030] Output the bounding box corresponding to the confidence higher than the set confidence threshold as the detection result.
[0031] Optionally, the multi-modal fusion detection model is optimized using a first loss function, and the first loss function is:
[0032]
[0033] where L is the total loss, is a balance hyperparameter, L RPN is the loss for generating candidate target boxes, L reg is the bounding box regression loss, L DC is the density confidence prediction loss.
[0034] In a second aspect, a three-dimensional object detection device for an unmanned system based on multi-modal fusion is provided, including:
[0035] A data acquisition unit for acquiring two-dimensional image data and three-dimensional point cloud data;
[0036] A model construction unit for constructing a multi-modal fusion detection model, where the multi-modal fusion detection model includes:
[0037] A point cloud feature extraction network for extracting three-dimensional point cloud features of a three-dimensional point cloud;
[0038] A temporal dynamic convolution fusion module discretizes the camera frustum space of a two-dimensional image to generate grid points, converts the grid points into three-dimensional world space coordinates using camera parameters, and uses a three-dimensional position embedding encoder to generate a three-dimensional position perception feature of the previous frame based on the three-dimensional world space coordinates of the previous frame and the two-dimensional image features of the previous frame. The three-dimensional position perception feature of the previous frame is used as a supplementary position perception feature of the current frame after pose transformation; a first fusion frame three-dimensional point cloud feature is generated based on the supplementary position perception feature of the current frame and the three-dimensional point cloud feature of the current frame; a preliminary fusion feature is obtained based on the first fusion frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame; the multi-modal fusion detection model obtains a detection result based on the preliminary fusion feature;
[0039] A detection unit for using the two-dimensional image data and three-dimensional point cloud data of a road scene to be detected as inputs to a pre-constructed multi-modal fusion detection model and outputting a detection result.
[0040] In a third aspect, an electronic device is provided, where the electronic device includes:
[0041] A processor;
[0042] A memory for storing executable instructions executable by the processor;
[0043] The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the method described in the first aspect above.
[0044] In a fourth aspect, a computer-readable storage medium is provided, where the computer-readable storage medium stores a computer program, and the computer program is used to execute the method described in the first aspect above.
[0045] Advantageous effects: In this method, the temporal dynamic convolution fusion module is used to solve the problem of difficult recognition of distant targets. The three-dimensional position perception feature of the previous frame is used to supplement the three-dimensional point cloud feature of the current frame after pose transformation, which is equivalent to supplementing some other points in a point cloud coordinate system to improve the detection accuracy; important features are screened and retained to ensure that important feature points are still retained in the case of fewer foreground points for distant targets, and the time relationship between different frames and the movement trajectory of objects are used to improve the recognition ability of distant objects. Moreover, since only important feature points are retained, the detection efficiency is ensured while maintaining the accuracy. Description of the Drawings
[0046] The following further describes the present invention in detail with reference to the accompanying drawings and specific embodiments.
[0047] Figure 1 It is a schematic flow chart of a three-dimensional object detection method for an unmanned system based on multi-modal fusion provided for this exemplary embodiment.
[0048] Figure 2 It is a schematic flow chart of another three-dimensional object detection method for an unmanned system based on multi-modal fusion provided for this exemplary embodiment.
[0049] Figure 3 It is a flow chart during the operation of a multi-modal fusion detection model constructed for this exemplary embodiment.
[0050] Figure 4 It is a schematic structural diagram of a multi-modal fusion detection model constructed for this exemplary embodiment.
[0051] Figure 5 It is a schematic diagram of the operation of a temporal dynamic convolution fusion module in this exemplary embodiment.
[0052] Figure 6 It is a schematic diagram of the operation of an adaptive gating fusion module in this exemplary embodiment.
[0053] Figure 7 It is a schematic diagram of the operation of a point cloud density fusion module in this exemplary embodiment.
[0054] Figure 8 It is a schematic diagram of the operation of a density confidence prediction module in this exemplary embodiment.
[0055] Figure 9 It is a schematic structural diagram of a three-dimensional object detection device for an unmanned system based on multi-modal fusion provided according to this exemplary embodiment;
[0056] Figure 10 It is a schematic diagram of an electronic device provided for this exemplary embodiment;
[0057] Figure 11 It is a schematic diagram of a computer-readable medium provided for this exemplary embodiment. Specific Embodiments
[0058] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0060] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations. In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0061] Embodiment
[0062] As Figure 1 shown, this embodiment provides a three-dimensional object detection method for an unmanned system based on multi-modal fusion, including the following steps:
[0063] S11. Obtain two-dimensional image data and three-dimensional point cloud data of the road scene to be detected;
[0064] S12. Use the two-dimensional image data and three-dimensional point cloud data of the road scene to be detected as the input of a pre-constructed multi-modal fusion detection model; the multi-modal fusion detection model includes: an image feature extraction network, a point cloud feature extraction network, a temporal dynamic convolution fusion module, an adaptive gating fusion module, a point cloud density fusion module, and a density confidence prediction module.
[0065] The point cloud feature extraction network is used to extract three-dimensional point cloud features of the three-dimensional point cloud data;
[0066] The image feature extraction network is used to extract two-dimensional image features of the two-dimensional image data;
[0067] The temporal dynamic convolution fusion module discretizes the camera frustum space to generate grid points, converts the grid points into three-dimensional world space coordinates using camera parameters, and uses a three-dimensional position embedding encoder to generate the three-dimensional position perception feature of the previous frame based on the three-dimensional world space coordinates and the two-dimensional image features of the previous frame. The three-dimensional position perception feature of the previous frame is used as the supplementary position perception feature of the current frame after pose transformation; the first fused frame three-dimensional point cloud feature is generated based on the supplementary position perception feature of the current frame and the three-dimensional point cloud feature of the current frame; the preliminary fusion feature is obtained based on the first fused frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame;
[0068] In the dynamic convolution module, obtaining the preliminary fusion feature based on the first fused frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame includes:
[0069] Using submanifold sparse convolution, predicting the importance score of each dimension of the first fused frame three-dimensional point cloud feature, and generating an importance weight distribution map;
[0070] If the weight of the first fused frame three-dimensional point cloud feature is greater than or equal to the set importance discrimination threshold, then this feature is considered an important feature;
[0071] Performing sparse convolution calculation on the important features to obtain the second fused frame three-dimensional point cloud feature; specifically: after filling the corresponding extended positions of the important features with 0 vectors of the same dimension, using submanifold sparse convolution for feature extraction to obtain the second fused frame three-dimensional point cloud feature.
[0072] Stitching and fusing the second fused frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame to obtain the preliminary fusion feature.
[0073] The multi-modal fusion detection model outputs the detection result based on the preliminary fusion feature, specifically including:
[0074] The adaptive gating fusion module is used to obtain the gating fusion feature and the gating image feature using the preliminary fusion feature and the two-dimensional image feature; and project the gating feature onto the gating image feature and perform a convolution operation to obtain the gating image enhancement feature.
[0075] The point cloud density fusion module is used to calculate the point centroid of the gating fusion feature; obtain the reference point in the image plane according to the point centroid of the gating fusion feature; weight a group of gating image enhancement features around the reference point to generate the aggregated image feature; fuse the aggregated image feature and the gating fusion feature to obtain the fused cross-modal feature; perform density-aware regional grid pooling on the fused cross-modal feature to obtain the fused cross-modal flat feature.
[0076] The density confidence prediction module includes a shared feedforward network, a bounding box feedforward network, and a confidence feedforward network. After encoding the fused cross-modal flattened features, the shared feedforward network inputs them into the bounding box feedforward network and the confidence feedforward network respectively. The bounding box feedforward network outputs a bounding box and the centroid of the final bounding box according to the encoding result. The confidence feedforward network predicts the confidence according to the encoding result, the centroid of the final bounding box, and the original number of points of the final bounding box. The bounding boxes corresponding to the confidence higher than the set confidence threshold are output as detection results.
[0077] The multi-modal fusion detection model is optimized using a first loss function, and the first loss function is:
[0078]
[0079] where L is the total loss, is a balance hyperparameter, L RPN is the loss for generating candidate target boxes, L reg is the bounding box regression loss, L DC is the density confidence prediction loss.
[0080] As Figure 2 shown, this embodiment provides another three-dimensional object detection method based on an unmanned system, including the following steps:
[0081] S20. Construct a multi-modal fusion detection model, including:
[0082] S201. Use a camera and a lidar to collect model-building two-dimensional image data and model-building three-dimensional point cloud data of an open road scene respectively;
[0083] S202. Perform data annotation on the model-building two-dimensional image data and the model-building three-dimensional point cloud data to obtain a multi-modal fusion object detection data set, and randomly divide the multi-modal fusion object detection data set into an 80% training set and a 20% test set according to a preset ratio;
[0084] S203. Construct an initial multi-modal fusion detection model; as Figure 4 shown, the multi-modal fusion detection model includes: an image feature extraction network, a point cloud feature extraction network, a temporal dynamic convolution fusion module, an adaptive gating fusion module, a point cloud density fusion module, a density confidence prediction module, a three-dimensional backbone network, a region candidate network, and a three-dimensional proposal box.
[0085] S204. Input the training set into the initial multi-modal fusion detection model to output a bounding box and a confidence. Then define a first loss function, and repeat step S14 to optimize the initial multi-modal fusion detection model until a trained multi-modal fusion detection model is finally obtained; the first loss function is:
[0086]
[0087] Where \(L\) is the total loss, \(\lambda\) is the balance hyperparameter, \(L_{gen}\) is the loss of generating candidate target boxes, \(L_{reg}\) is the bounding box regression loss, and \(L_{density}\) is the density confidence prediction loss.
[0088] S205. Input the test set into the trained multi-modal fusion detection model to obtain test results, and compare them with the original test set annotation results to further evaluate the performance of the multi-modal fusion detection model.
[0089] S21. Obtain the two-dimensional image data and three-dimensional point cloud data of the road scene to be detected.
[0090] Specifically, in an open road field, use a camera to obtain two-dimensional image data and a lidar to obtain three-dimensional point cloud data.
[0091] S22. Use the two-dimensional image data and three-dimensional point cloud data of the road scene to be detected as the input of the multi-modal fusion detection model pre-constructed in step S20; the pre-constructed multi-modal fusion detection model is the multi-modal fusion detection model trained in step S205. As Figure 3 shown, the detection process includes the following steps:
[0092] S221. Use the image feature extraction network to extract the two-dimensional image features of the two-dimensional image data;
[0093] S222. Use the point cloud feature extraction network to extract the three-dimensional point cloud features of the three-dimensional point cloud data;
[0094] S223. Use the temporal dynamic convolution fusion module to discretize the camera frustum space shared by all views to generate grid points, use the camera parameters to convert the grid points into three-dimensional world space coordinates, use the three-dimensional position embedding encoder to generate the three-dimensional position perception features of the previous frame according to the three-dimensional space coordinates and the two-dimensional image features of the previous frame, and after pose transformation, the three-dimensional position perception features are used as the supplementary position perception features of the current frame; generate the first fusion frame three-dimensional point cloud features according to the supplementary position perception features of the current frame and the three-dimensional point cloud features of the current frame; obtain the preliminary fusion features based on the first fusion frame three-dimensional point cloud features and the two-dimensional image features of the current frame; specifically, the temporal dynamic convolution fusion module includes a three-dimensional position embedding encoder and a submanifold sparse convolution; when the temporal dynamic convolution fusion module runs, as Figure 5 shown, it includes the following steps:
[0095] S2231. Discretize the camera frustum space shared by all views to generate grid points, use the camera parameters to convert the grid points of the current frame into spatial coordinates, generate the corresponding current three-dimensional world space coordinates, and the calculation formula is as follows:
[0096]
[0097] Among them, represents the three-dimensional world space coordinates of the current frame, and T a is the transformation matrix for converting the grid points of the current frame into the three-dimensional world space coordinates of the current frame, and U i ∈R 4×4 is the camera parameter conversion matrix of the i-th camera, represents the grid points of the current frame.
[0098] S2232. Obtain the two-dimensional image features and three-dimensional world space coordinates of the previous frame and input them into the three-dimensional position embedding encoder to generate the three-dimensional position perception features of the previous frame, and obtain the supplementary position perception features of the current frame through a certain pose transformation. The calculation formula in the three-dimensional position embedding encoder is:
[0099]
[0100] In the formula, is the three-dimensional position perception feature of the previous frame, is the position embedding encoder encoding function, is the two-dimensional image feature of the previous frame, is the three-dimensional world space coordinates of the previous frame;
[0101] The calculation formula for the pose transformation is:
[0102]
[0103] In the formula, T b is the transformation matrix for converting the three-dimensional coordinate space of the previous frame into the three-dimensional coordinate space of the current frame, is the three-dimensional world space coordinates of the current frame, is the three-dimensional world space coordinates of the previous frame.
[0104] S2233. The supplementary position perception feature of the current frame and the three-dimensional point cloud feature of the current frame are in the same three-dimensional coordinate system. The supplementary position perception feature of the current frame and the three-dimensional point cloud feature of the current frame are directly fused to form the three-dimensional point cloud feature of the first fusion frame.
[0105] S2234: Use the three-dimensional point cloud feature of the first fusion frame as the input of the submanifold sparse convolution. The submanifold sparse convolution predicts the importance scores of each dimension of the three-dimensional point cloud feature of the first fusion frame, and then uses the Sigmoid function to normalize the importance scores of each dimension to generate the importance weight distribution map.
[0106] During the training process, the ground truth is used for importance prediction supervision. Then, the center of the input voxel is projected into the point cloud space. If the center of the input voxel falls within the ground truth bounding box, the input is defined as an important feature, and the importance weight of the voxel is set to 1; otherwise, the importance weight of the voxel is 0. In this embodiment, the input voxel is a small cube region obtained by discretizing three-dimensional point cloud data.
[0107] To alleviate the problem of class imbalance, the focal cross-entropy loss function is used for loss calculation. The focal cross-entropy loss function is as follows:
[0108] L1 = -α t δlog(θ t );
[0109] δ = (1 - θ t ) γ ;
[0110] In the formula, L1 is the focal cross-entropy loss function, α t ∈[0, 1] is the weight factor, δ is the modulation factor, and θ t is the loss coefficient, γ ∈ [0, 5], and γ is an adjustable focal parameter.
[0111] Step S2235: If the weight of the three-dimensional point cloud feature of the first fusion frame is greater than or equal to the set importance discrimination threshold, the feature is considered an important feature. The determination process formula is as follows.
[0112]
[0113] In the formula, P imp is the input feature space of the important feature, is the importance weight of the input feature, τ is the importance discrimination threshold, p is the feature vector of the input feature space, and P in is the input feature space of the feature.
[0114] S2236: Perform sparse convolution calculation on the important feature to obtain the three-dimensional point cloud feature of the second fusion frame, specifically including:
[0115] Fill the corresponding extended positions of the important feature with 0 vectors of the same dimension. That is, after dynamically outputting the extended positions, use submanifold sparse convolution for feature extraction to obtain the three-dimensional point cloud feature of the second fusion frame.
[0116] The formula for dynamically outputting the extended position is as follows:
[0117]
[0118] In the formula, is the dynamically output extended position, and k is the feature vector of the dynamically extended feature space. is the importance weight of the peripheral expansion feature.
[0119] S2237: Feature splicing and fusion are performed on the two-dimensional image features of the current frame and the three-dimensional point cloud features of the second fusion frame after dynamic convolution, and finally the preliminary fusion features are output.
[0120] S224. As Figure 6 shown, in the adaptive gating module, the gating fusion feature and the gating image feature are obtained by using the preliminary fusion feature and the two-dimensional image feature. The gating feature is projected onto the gating image feature, and a convolution operation is performed to obtain the gating image enhanced feature. The calculation formulas for the gating fusion feature and the gating image feature are as follows:
[0121]
[0122]
[0123] In the formula, A is the gating fusion feature, F T is the preliminary fusion feature, F C is the two-dimensional image feature, × is the product of position elements, σ is the sigmoid activation function, is the concatenation operation, and B is the gating image feature.
[0124] In the adaptive gating module, the gating fusion feature is directly output. After the gating image enhanced feature is further enhanced by convolution, the gating image enhanced feature is output.
[0125] S225. The fusion cross-modal flat feature is obtained by using the point cloud density fusion module. As Figure 7 shown, it specifically includes:
[0126] S2251: Calculate the point centroid of the gating fusion feature. Specifically, calculate the non-empty voxel feature vector of the gating fusion feature, and obtain the non-empty voxel feature set of the gating fusion. Solve for the sum of the non-empty voxel feature vectors, and obtain the point centroid of the gating fusion feature according to the sum of the non-empty voxel feature vectors. The specific calculation formula is as follows:
[0127]
[0128]
[0129]
[0130] In the formula, F gate is the non-empty voxel feature set of the gating fusion, V gate is the voxel index, is the non-empty voxel feature vector of the gating fusion, N gate is the number of non-empty voxels in the gating fusion feature, M总 is the sum of non-empty voxel feature vectors is p at the spatial coordinates gate =(x gate , y gate z gate ), c gate is the point centroid of the gated fusion feature, |P(V gate )| is the set of points in the voxel index
[0131] S2252. Obtain a reference point in the image plane based on the point centroid of the gated fusion feature
[0132] Assign each point centroid of the gated fusion to its respective voxel index, and map the voxel index to the associated voxel feature in the sparse convolutional layer through a three-dimensional hash table. Calculate the reference point in the image plane, i.e., the gated image centroid, from the point centroid of each calculated gated fusion feature using the camera projection matrix. The calculation formula for the reference point is
[0133] p i =ρ·c gate ;
[0134] where ρ is the product of the camera internal matrix and the external matrix
[0135] S2253. Weight a set of gated image enhanced features around the reference point to generate an aggregated image feature
[0136] That is, generated by applying the learned offset to the gated image enhanced feature. Using the gated fusion feature as the query, the aggregated image feature as the key and value, fuse the aggregated image feature and the gated fusion feature through the cross-attention module to obtain the final fused cross-modal feature enhanced based on the aggregated image feature. The calculation formula is as follows
[0137]
[0138]
[0139] where is the aggregated image feature generated by the k-th sampling point, Δp omk is the sampling offset of the k-th sampling point in the m-th attention head is the probability density given the gated fusion feature and the aggregated image feature is the aggregated image feature, β1, β2 are learnable weight parameters, M is the number of self-attention heads, K is the total number of sampling points, Z omk is the attention weight of the k-th sampling point in the m-th attention head
[0140] S2254. Integrate the aggregated image features and the gated fusion features to obtain fused cross-modal features;
[0141] S2255. Perform density-aware regional grid pooling on the fused cross-modal features to obtain fused cross-modal flattened features.
[0142] S226. In the density confidence prediction module, obtain the bounding box and confidence according to the fused cross-modal flattened features.
[0143] As Figure 8 shown, the density confidence prediction module includes a shared feed-forward network, a bounding box feed-forward network, and a confidence feed-forward network;
[0144] The shared feed-forward network encodes the fused cross-modal flattened features, and the encoding results are respectively sent to the bounding box feed-forward network and the confidence feed-forward network;
[0145] The bounding box feed-forward network outputs the bounding box and the centroid of the final bounding box according to the encoding result; the bounding box includes the obstacle category and location;
[0146] The confidence feed-forward network predicts the confidence according to the encoding result, the centroid of the final bounding box, and the number of original points of the final bounding box; the number of original points of the final bounding box has been obtained in the data preprocessing stage.
[0147] Output the bounding boxes corresponding to the confidence higher than the set confidence threshold as the detection results.
[0148] The calculation formula of the confidence is as follows:
[0149] Y b = FFN([f b , c b , log(|N(b)|)]);
[0150] where f b is the output feature vector from the shared feed-forward neural network, c b is the centroid of the final bounding box, and |N(b)| is the number of original points of the final bounding box.
[0151] The three-dimensional backbone network converts input data (such as point cloud or voxel representation) into a feature representation; the region proposal network generates candidate boxes on the feature map and classifies and regresses these candidate boxes simultaneously; the three-dimensional proposal box selects some candidate boxes according to the output of the region proposal network, extracts the features corresponding to the candidate boxes on the feature map, and uses these features for object classification and localization, thus completing the object detection or object recognition task. Candidate boxes can be regarded as the preliminary predictions of the object detection algorithm, used to propose regions that may contain the object; while the bounding box is the final detection result, which is the precise localization of the object's position. Usually, the candidate boxes are sorted or filtered according to the confidence level, and the candidate boxes with a confidence level higher than a certain threshold are selected as the final bounding boxes for output, used to describe the position and size of the detected object.
[0152] A three-dimensional object detection method for unmanned systems based on multi-modal fusion provided in this embodiment can be evaluated in real time in the actual scenario of an open road, reduce traffic accidents caused by human errors, and improve the driving performance of vehicles in various complex situations, reducing the occurrence of misjudgment behaviors.
[0153] In this method, the temporal dynamic convolution fusion module is used to solve the problem of difficult recognition of distant objects. It uses the three-dimensional position perception features of the previous frame after pose transformation to supplement the three-dimensional point cloud features of the current frame, which is equivalent to supplementing some other points in a point cloud coordinate system to improve the detection accuracy; it screens and retains important features, ensuring that important feature points are still retained when there are fewer foreground points for distant objects, and uses the temporal relationship between different frames and the motion trajectory of objects to improve the recognition ability of distant objects. Moreover, since only important feature points are retained, the detection efficiency is ensured while maintaining the accuracy.
[0154] This method also integrates perception technologies of multiple modalities, which can further improve the refined three-dimensional perception ability in the actual scenario of an open road, thereby better improving the detection accuracy. In the adaptive gating fusion module, to solve the problem of key information loss in the combination of two heterogeneous feature maps, according to the relevance between the feature map and the object detection task, the feature maps are selectively combined. After the image and point cloud features are fused using the temporal dynamic convolution fusion module, the gated fusion feature and the gated image feature are fused using the attention mechanism for the feature fusion of the gating circuit. This design aims to effectively reduce the loss of key information in the relevant feature maps to improve the model performance.
[0155] In the point cloud density fusion module, to solve the problem of partial point cloud data and image data being affected by occlusion interference, the point cloud density change is considered through the point cloud density fusion module, enabling the network to pay more attention to the regions with large density changes when aggregating features, thereby improving the perception ability of occluded objects.
[0156] Please refer to Figure 9 , which shows a schematic diagram of a three-dimensional object detection device for an unmanned system based on multi-modal fusion provided by some embodiments of the present application. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiments. The device embodiments described below are merely illustrative.
[0157] As Figure 9 shown, the three-dimensional object detection device 900 for an unmanned system based on multi-modal fusion may include:
[0158] A data acquisition unit 901, configured to acquire two-dimensional image data and three-dimensional point cloud data;
[0159] A model construction unit 902, configured to construct a multi-modal fusion detection model, where the multi-modal fusion detection model includes: an image feature extraction network, a point cloud feature extraction network, a temporal dynamic convolution fusion module, an adaptive gating fusion module, a point cloud density fusion module, and a density confidence prediction module.
[0160] The point cloud feature extraction network is configured to extract three-dimensional point cloud features of the three-dimensional point cloud data;
[0161] The image feature extraction network is configured to extract two-dimensional image features of the two-dimensional image data;
[0162] The temporal dynamic convolution fusion module discretizes the camera frustum space of the two-dimensional image to generate grid points, converts the grid points into three-dimensional world space coordinates using camera parameters, and uses a three-dimensional position embedding encoder to generate a three-dimensional position perception feature of the previous frame based on the three-dimensional world space coordinates of the previous frame and the two-dimensional image features of the previous frame. The three-dimensional position perception feature of the previous frame is used as a supplementary position perception feature of the current frame after pose transformation; a first fusion frame three-dimensional point cloud feature is generated based on the supplementary position perception feature of the current frame and the three-dimensional point cloud feature of the current frame; a preliminary fusion feature is obtained based on the first fusion frame three-dimensional point cloud feature and the two-dimensional image features of the current frame; and the multi-modal fusion detection model obtains a detection result based on the preliminary fusion feature.
[0163] In the dynamic convolution module, a preliminary fusion feature is obtained based on the three-dimensional point cloud feature of the first fusion frame and the two-dimensional image feature of the current frame, including: using submanifold sparse convolution to predict the importance score of each dimension of the three-dimensional point cloud feature of the first fusion frame, and generating an importance weight distribution map; if the weight of the three-dimensional point cloud feature of the first fusion frame is greater than or equal to a set importance discrimination threshold, the feature is considered an important feature; performing sparse convolution calculation on the important features to obtain the three-dimensional point cloud feature of the second fusion frame; specifically: after filling the corresponding extended positions of the important features with 0 vectors of the same dimension, using submanifold sparse convolution to extract features to obtain the three-dimensional point cloud feature of the second fusion frame; splicing and fusing the three-dimensional point cloud feature of the second fusion frame and the two-dimensional image feature of the current frame to obtain a preliminary fusion feature.
[0164] The adaptive gating fusion module is used to obtain a gating fusion feature and a gating image feature by using the preliminary fusion feature and the two-dimensional image feature; project the gating feature onto the gating image feature, and perform a convolution operation to obtain a gating image enhancement feature.
[0165] The point cloud density fusion module is used to calculate the point centroid of the gating fusion feature; obtain a reference point in the image plane according to the point centroid of the gating fusion feature; weight a group of gating image enhancement features around the reference point to generate an aggregated image feature; fuse the aggregated image feature and the gating fusion feature to obtain a fused cross-modal feature; perform density-aware region grid pooling on the fused cross-modal feature to obtain a fused cross-modal flat feature.
[0166] The density confidence prediction module includes a shared feed-forward network, a bounding box feed-forward network, and a confidence feed-forward network; after encoding the fused cross-modal flat feature by the shared feed-forward network, it is respectively input into the bounding box feed-forward network and the confidence feed-forward network; the bounding box feed-forward network outputs a bounding box according to the encoding result, and the centroid of the final bounding box; the confidence feed-forward network predicts the confidence according to the encoding result, the centroid of the final bounding box, and the original number of points of the final bounding box; the bounding box corresponding to the confidence higher than the set confidence threshold is output as the detection result.
[0167] The detection unit 903 is used to take the two-dimensional image data and three-dimensional point cloud data of the road scene to be detected as the input of a pre-constructed multi-modal fusion detection model, and output a detection result.
[0168] In some embodiments of the embodiments of the present application, the device 900 provided by the embodiments of the present application has the same inventive concept and the same beneficial effects as the method provided by the foregoing embodiments of the present application.
[0169] Embodiments of the present application also provide an electronic device corresponding to the method provided in the foregoing embodiments. The electronic device may be an electronic device for a server, such as a server, including an independent server and a distributed server cluster, etc., to execute the above method; the electronic device may also be an electronic device for a client, such as a mobile phone, a laptop computer, a tablet computer, a desktop computer, etc., to execute the above method.
[0170] Please refer to Figure 10 , which shows a schematic diagram of an electronic device provided in some embodiments of the present application. As Figure 10 shown, the electronic device 40 includes: a processor 400, a memory 401, a bus 402, and a communication interface 403. The processor 400, the communication interface 403, and the memory 401 are connected through the bus 402; a computer program that can run on the processor 400 is stored in the memory 401, and when the processor 400 runs the computer program, it executes the foregoing method of the present application.
[0171] Among them, the memory 401 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 403 (which may be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.
[0172] The bus 402 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 401 is used to store programs, and after receiving an execution instruction, the processor 400 executes the programs. Any of the methods disclosed in the foregoing embodiments of the present application can be applied to or implemented by the processor 400.
[0173] The processor 400 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 400 or instructions in the form of software. The above-mentioned processor 400 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory 401, and the processor 400 reads the information in the memory 401 and combines its hardware to complete the steps of the above method.
[0174] The electronic device provided in the embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by it.
[0175] The embodiments of the present application also provide a computer-readable medium corresponding to the method provided in the foregoing embodiments. Please refer to Figure 11 , which shows that the computer-readable storage medium is an optical disc 50, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the foregoing method.
[0176] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.
[0177] The computer-readable storage medium provided in the above embodiments of the present application and the method provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by the application program stored in it.
[0178] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0179] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0180] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical functional division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed among each other can be through some communication interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0181] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0182] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0183] When the above functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0184] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of various embodiments of this application, and they should all be covered by the scope of the claims and the description of this application.
Claims
1. A three-dimensional object detection method for unmanned systems based on multimodal fusion, characterized in that, It includes the following steps: Obtain two-dimensional image data and three-dimensional point cloud data of the road scene to be detected by using a camera and a lidar respectively; Use the two-dimensional image data and three-dimensional point cloud data of the road scene to be detected as the input of a pre-constructed multi-modal fusion detection model; The multi-modal fusion detection model includes: A point cloud feature extraction network for extracting three-dimensional point cloud features of the three-dimensional point cloud data; A temporal dynamic convolution fusion module that discretizes the camera frustum space to generate grid points, converts the grid points into three-dimensional world space coordinates by using camera parameters, and uses a three-dimensional position embedding encoder to generate a three-dimensional position perception feature of the previous frame based on the three-dimensional world space coordinates and the two-dimensional image features of the previous frame. The three-dimensional position perception feature of the previous frame is used as the supplementary position perception feature of the current frame after pose transformation; Generate the first fused frame three-dimensional point cloud feature according to the supplementary position perception feature of the current frame and the three-dimensional point cloud feature of the current frame; Obtain the preliminary fusion feature based on the first fused frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame; The multi-modal fusion detection model outputs the detection result based on the preliminary fusion feature; The multi-modal fusion detection model further includes an image feature extraction network for extracting two-dimensional image features of the two-dimensional image data; In the temporal dynamic convolution module, obtaining the preliminary fusion feature based on the first fused frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame includes: Using sub-manifold sparse convolution to predict the importance score of each dimension of the first fused frame three-dimensional point cloud feature and generate an importance weight distribution map; If the weight of the first fused frame three-dimensional point cloud feature is greater than or equal to the set importance discrimination threshold, then consider this feature as an important feature; Perform sparse convolution calculation on the important features to obtain the second fused frame three-dimensional point cloud feature; Stitch and fuse the second fused frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame to obtain the preliminary fusion feature.
2. The three-dimensional object detection method for an unmanned system based on multimodal fusion according to claim 1, wherein, The performing sparse convolution calculation on the important features to obtain the second fused frame three-dimensional point cloud feature includes: After filling the corresponding extended positions of the important features with 0 vectors of the same dimension, use sub-manifold sparse convolution for feature extraction to obtain the second fused frame three-dimensional point cloud feature.
3. A three-dimensional object detection method for an unmanned system based on multimodal fusion according to claim 1, characterized in that, The multi-modal fusion detection model further includes an adaptive gating fusion module, The adaptive gating fusion module is used to obtain a gating fusion feature and a gating image feature by using the preliminary fusion feature and the two-dimensional image feature; Project the gating feature onto the gating image feature and perform a convolution operation to obtain a gating image enhancement feature.
4. A three-dimensional object detection method for an unmanned system based on multi-modal fusion according to claim 3, characterized in that The multi-modal fusion detection model further includes a point cloud density fusion module; The point cloud density fusion module is used to calculate the point centroid of the gating fusion feature; Obtain a reference point in the image plane according to the point centroid of the gating fusion feature; Weight a group of gating image enhancement features around the reference point to generate an aggregated image feature; Fuse the aggregated image feature and the gating fusion feature to obtain a fused cross-modal feature; Perform density-aware regional grid pooling on the fused cross-modal feature to obtain a fused cross-modal flattened feature.
5. The 3D object detection method for an unmanned system based on multi-modal fusion according to claim 3, characterized in that The multi-modal fusion detection model further includes a density confidence prediction module, which includes a shared feed-forward network, a bounding box feed-forward network, and a confidence feed-forward network; After encoding the fused cross-modal flat features, the shared feed-forward network inputs them into the bounding box feed-forward network and the confidence feed-forward network respectively; The bounding box feed-forward network outputs a bounding box and the centroid of the final bounding box according to the encoding result; The confidence feed-forward network predicts the confidence according to the encoding result, the centroid of the final bounding box, and the number of original points of the final bounding box; The bounding boxes corresponding to the confidence higher than the set confidence threshold are output as detection results.
6. The three-dimensional object detection method for an unmanned system based on multi-modal fusion according to claim 1, characterized in that The multi-modal fusion detection model is optimized using a first loss function, and the first loss function is: ; where \(L\) is the total loss, is the balance hyperparameter, is the loss for generating candidate target boxes, is the bounding box regression loss, is the density confidence prediction loss.
7. A three-dimensional object detection device for an unmanned system based on multimodal fusion, characterized in that, including: A data acquisition unit for acquiring two-dimensional image data and three-dimensional point cloud data; A model construction unit for constructing a multi-modal fusion detection model, and the multi-modal fusion detection model includes: A point cloud feature extraction network for extracting three-dimensional point cloud features of the three-dimensional point cloud; A temporal dynamic convolution fusion module that discretizes the camera frustum space to generate grid points, converts the grid points into three-dimensional world space coordinates using camera parameters, uses a three-dimensional position embedding encoder to generate the three-dimensional position perception feature of the previous frame according to the three-dimensional world space coordinates of the previous frame and the two-dimensional image features of the previous frame, and after pose transformation, the three-dimensional position perception feature of the previous frame is used as the supplementary position perception feature of the current frame; according to the supplementary position perception feature of the current frame and the three-dimensional point cloud feature of the current frame, a first fusion frame three-dimensional point cloud feature is generated; based on the first fusion frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame, a preliminary fusion feature is obtained; the multi-modal fusion detection model obtains a detection result based on the preliminary fusion feature; A detection unit for using the two-dimensional image data and three-dimensional point cloud data of the road scene to be detected as the input of a pre-constructed multi-modal fusion detection model and outputting a detection result; The multi-modal fusion detection model further includes an image feature extraction network, and the image feature extraction network is used to extract two-dimensional image features of the two-dimensional image data; In the temporal dynamic convolution module, obtaining a preliminary fusion feature based on the first fusion frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame includes: Using submanifold sparse convolution to predict the importance score of each dimension of the first fusion frame three-dimensional point cloud feature and generate an importance weight distribution map; If the weight of the first fusion frame three-dimensional point cloud feature is greater than or equal to the set importance discrimination threshold, then the feature is considered an important feature; Performing sparse convolution calculation on the important features to obtain a second fusion frame three-dimensional point cloud feature; Stitching and fusing the second fusion frame three-dimensional point cloud feature and the two-dimensional image feature of the current frame to obtain a preliminary fusion feature.
8. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing executable instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the executable instructions to implement the method described in any one of claims 1 to 6 above.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is used to execute the method described in any one of claims 1 to 6 above.
Citation Information
Patent Citations
3D target detection method and device based on multi-sensor fusion
CN115761723A