Feature fusion method of point cloud data and image data and related device
By extracting and fusing the features of point cloud data and image data, the problem of mismatch between point cloud and image data in autonomous driving is solved, and more accurate and robust 3D perception is achieved.
Patent Information
- Application Number
- CN202210333993.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-03-30
AI Technical Summary
In autonomous driving, the mismatch between point cloud data and image data leads to difficulty in fusion of lidar point cloud and camera images, affecting the accuracy and robustness of 3D perception.
By obtaining point cloud data and image data in the current field of view, point cloud features and multi-scale image features are extracted, and fusion is performed to generate fusion features. The specific steps include projecting point cloud features onto the plane of the image feature, obtaining a reference point, obtaining a target image feature from the multi-scale fusion image feature based on the reference point, and finally obtaining a fusion feature based on the point cloud features and the target image feature.
Through the use of fusion features, the feature fusion matching degree of point cloud data and image data is improved, so that the fusion features not only contain information of multi-scale image features, but also information of point cloud features, thereby improving the accuracy and robustness of 3D perception.
Smart Images

Figure CN114723793B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of feature fusion technology, and in particular to a feature fusion method of point cloud data and image data and a related device. Background Art
[0002] 3D perception has always been a key task in autonomous driving. Although point cloud-based methods have made great progress in 3D perception, its sparsity is a challenge for small objects and long-distance perception. Camera images contain dense texture and color information, but lack depth for accurate 3D positioning.
[0003] Therefore, effective fusion of point clouds and cameras is needed to achieve accurate and robust 3D perception.
[0004] Currently, the main challenge of fusion lies in the mismatch between lidar point clouds and camera images. Summary of the invention
[0005] The present application provides a method for fusion of features of point cloud data and image data and related devices, which can improve the matching degree of feature fusion of point cloud data and image data.
[0006] To solve the above technical problems, a technical solution adopted in the present application is: to provide a feature fusion method for point cloud data and image data, the method comprising: obtaining point cloud data and image data within the current field of view; extracting point cloud features based on the point cloud data; performing multi-scale feature extraction on the image data to obtain corresponding multiple image features; fusing multiple image features to obtain multi-scale fused image features; and obtaining fused features based on point cloud features and multi-scale fused image features.
[0007] The method further includes: projecting the point cloud features onto the plane of the image features to obtain reference points; obtaining target image features from multi-scale fused image features based on the reference points; and obtaining fused features based on the point cloud features and the target image features.
[0008] Among them, obtaining the target image feature from the multi-scale fusion image feature based on the reference point includes: obtaining the offset and attention weight corresponding to the reference point based on the reference point; and obtaining the target image feature based on the reference point, the offset and the attention weight.
[0009] The method includes: obtaining an offset feature corresponding to each offset based on a reference point and an offset; and obtaining a target image feature by weighted averaging each offset feature based on each attention weight.
[0010] The step of projecting the point cloud features onto the plane of the image features to obtain the reference points includes: performing scale-invariant projection based on calibration parameters to obtain the reference points.
[0011] Among them, obtaining the offset and attention weight corresponding to the reference point based on the reference point includes: dynamically enhancing the reference point feature to obtain the enhanced reference point feature; predicting the enhanced reference point feature to obtain the offset and attention weight corresponding to the reference point.
[0012] Among them, the reference point is dynamically enhanced to obtain enhanced reference point features, including: normalizing the point cloud features after MLP processing to obtain normalized point cloud features; normalizing the reference point features after convolution processing to obtain normalized convolution reference point features; fusing the normalized point cloud features and the normalized convolution reference point features to obtain enhanced reference point features.
[0013] The method includes: if the image data includes image data captured by at least two cameras; obtaining the reference point corresponding to each image data; averaging the target image features corresponding to each reference point to obtain the average target image features; and obtaining the fusion features based on the point cloud features and the average target image features.
[0014] To solve the above-mentioned technical problems, another technical solution adopted in the present application is: to provide an autonomous driving computing platform, which includes a processor and a memory coupled to the processor, the memory is used to store computer programs, and the processor is used to execute computer programs to implement the method provided by the above-mentioned technical solution.
[0015] To solve the above technical problems, another technical solution adopted in the present application is: providing a computer-readable storage medium, which is used to store a computer program, and when the computer program is executed by a processor, it is used to implement the method provided by the above technical solution.
[0016] The beneficial effects of the embodiments of the present application are as follows: Different from the prior art, the feature fusion method of point cloud data and image data provided by the present application includes: obtaining point cloud data and image data in the current field of view; extracting point cloud features based on point cloud data; performing multi-scale feature extraction on image data to obtain corresponding multiple image features; fusing multiple image features to obtain multi-scale fused image features; obtaining fused features based on point cloud features and multi-scale fused image features. In the above manner, point cloud features and multi-scale fused image features are fused, so that the obtained fused features not only have information on multi-scale image features, but also have information on point cloud features, thereby improving the matching degree of feature fusion of point cloud data and image data. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0018] Figure 1 It is a flow chart of an embodiment of a method for feature fusion of point cloud data and image data provided by the present application;
[0019] Figure 2 It is a structural schematic diagram of an embodiment of a fusion model provided by the present application;
[0020] Figure 3 It is a flowchart of another embodiment of the method for feature fusion of point cloud data and image data provided by the present application;
[0021] Figure 4 It is a flowchart of another embodiment of the method for feature fusion of point cloud data and image data provided by the present application;
[0022] Figure 5 It is a flowchart of another embodiment of the method for feature fusion of point cloud data and image data provided by the present application;
[0023] Figure 6 is a flowchart of an embodiment of step 56 provided by the present application;
[0024] Figure 7 It is a flowchart of an embodiment of step 561 provided by the present application;
[0025] Figure 8 It is a structural diagram of an embodiment of an autonomous driving computing platform provided by the present application;
[0026] Fig. 9 It is a structural schematic diagram of an embodiment of a computer-readable storage medium provided by the present application. DETAILED DESCRIPTION
[0027] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It will be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only some but not all structures related to the present application are shown in the drawings. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0028] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0029] See also Figure 1 , Figure 1 1 is a flow chart of an embodiment of a method for feature fusion of point cloud data and image data provided by the present application. The method comprises:
[0030] Step 11: Get the point cloud data and image data in the current field of view.
[0031] In some embodiments, an image sensor may be used to collect image data within the current field of view, and a radar sensor may be used to collect point cloud data of objects within the current field of view. For example, the image sensor and the radar sensor may be installed on a movable device. The device may be an automatic mobile device, such as a robot, an automatic driving vehicle, etc.
[0032] In some embodiments, the image sensor may be a camera. The radar sensor may be a lidar sensor, such as a mechanical lidar.
[0033] In one application scenario, while driving, an autonomous driving vehicle obtains image data within a current field of view through an image sensor disposed on the autonomous driving vehicle and obtains corresponding point cloud data within the current field of view through a radar sensor.
[0034] In some embodiments, as image data is acquired, the camera optimizes the calibration matrix in an iterative and dynamic manner in real time.
[0035] Step 12: Extract point cloud features based on point cloud data.
[0036] In some embodiments, a 3D feature extraction network may be used to extract features from point cloud data. The 3D feature extraction network consists of stacked basic blocks. In each basic block, the point-based method extracts and aggregates point-by-point features, while the voxel-based method performs sparse convolution. The entire 3D feature extraction network stacks basic blocks of different scales, which are referred to as stages in this application. Since only 3D features from valid points or voxels are fused, a sparse representation is used to describe the point cloud features. The 3D feature containing N voxels or points is defined as Where F represents the point cloud feature and P represents the corresponding coordinates.
[0037] Step 13: Perform multi-scale feature extraction on the image data to obtain corresponding multiple image features.
[0038] In some embodiments, a CNN architecture may be used to extract image features. For example, VGG and ResNet generate a feature pyramid from low to high. The feature pyramid is used to extract features of multiple scales from image data to obtain corresponding multiple image features.
[0039] Step 14: Fuse multiple image features to obtain multi-scale fused image features.
[0040] Among them, low-level features contain spatial details, while high-level features provide strong semantics. In order to utilize low-level details and high-level semantics, multi-level image features from ResNet50 can be fused. The image features from the four levels of ResNet50 are represented as The scales are 4, 8, 16, and 32 respectively. k represents the image acquired by the kth camera. Note that the image network here includes but is not limited to ResNet50, and multi-level image networks can be applied to this architecture.
[0041] Step 15: Obtain fusion features based on point cloud features and multi-scale fusion image features.
[0042] In some embodiments, point cloud features and multi-scale fused image features are fused one-to-many to obtain fused features. That is, one point cloud feature can correspond to the corresponding target image feature in each image feature. The fused features thus obtained have both image feature information and three-dimensional space point cloud information. Since multi-scale fused image features have different scales, the fused features can have contextual information of 3D geometry and image data.
[0043] In some embodiments, the 3D detection head may be further used to perform target object recognition on the fused features, so as to determine the target object in the image data and obtain the 3D coordinates of the object.
[0044] In an application scenario, after obtaining fusion features based on point cloud features and multi-scale fusion image features, it is necessary to use the fusion features as target point cloud data, extract point cloud features again, and perform point cloud features and multi-scale fusion image features to obtain fusion features. This process is repeated to obtain the final fusion features.
[0045] Combination Figure 2 The fusion model is explained as follows:
[0046] The fusion model includes an image feature extraction module, a point cloud feature extraction module and a dynamic attention fusion module.
[0047] After acquiring the point cloud data and image data, the image feature extraction module is used to extract features of multiple different scales on the image data to obtain corresponding multiple image features.
[0048] The point cloud feature extraction module is used to extract features from the point cloud data to obtain point cloud features.
[0049] The dynamic attention fusion module is used to fuse multiple image features to obtain multi-scale fused image features.
[0050] The dynamic attention fusion module is used to fuse the point cloud features and the multi-scale fusion image features to obtain the fusion features. After fusion, the fusion features have the information of multiple image features.
[0051] The point cloud feature extraction module extracts features from each pair of point cloud data, and the obtained point cloud features will be fused with multiple image features in the dynamic attention fusion module to further obtain fused features. As a result, the fused features have stronger image context information and 3D coordinate information.
[0052] like Figure 2 As shown, the image feature extraction module includes a first layer, a second layer, a third layer and a fourth layer, that is, the image data will be subjected to image feature extraction in the first layer, the second layer, the third layer and the fourth layer in sequence to obtain image features of corresponding scales.
[0053] The dynamic attention fusion module includes the first module, the second module, the third module and the fourth module. Image features of all scales are input into the first module, the second module, the third module and the fourth module.
[0054] The point cloud feature extraction module includes a first-stage module, a second-stage module, a third-stage module, and a fourth-stage module. In the first-stage module, feature extraction is performed on the point cloud data to obtain the first point cloud feature, and then the first point cloud feature is input into the first module. In the first module, multiple image features are fused to obtain multi-scale fused image features, and the first point cloud feature is fused with the multi-scale fused image feature in a one-to-many manner, and local attention learning representation is performed to further obtain the first fused feature.
[0055] The first fused feature is fed back to the first stage module, so that the first stage inputs the first fused feature to the second stage module.
[0056] In the second stage module, the first fusion feature is extracted to obtain the second point cloud feature, and then the second point cloud feature is input into the second module. In the second module, multiple image features are fused to obtain multi-scale fusion image features, and the second point cloud feature is fused with the multi-scale fusion image feature in a one-to-many manner, and local attention learning representation is performed to further obtain the second fusion feature.
[0057] The second fused feature is fed back to the second stage module, so that the second stage inputs the second fused feature to the third stage module.
[0058] In the third stage module, the second fusion feature is extracted to obtain the third point cloud feature, and then the third point cloud feature is input into the third module. In the third module, multiple image features are fused to obtain multi-scale fusion image features, and the third point cloud feature is fused with the multi-scale fusion image feature in a one-to-many manner, and local attention learning representation is performed to further obtain the third fusion feature.
[0059] The third fused feature is fed back to the third stage module, so that the third stage inputs the third fused feature to the fourth stage module.
[0060] In the fourth stage module, the third fusion feature is extracted to obtain the fourth point cloud feature, and then the fourth point cloud feature is input into the fourth module. In the fourth module, multiple image features are fused to obtain multi-scale fusion image features, and the fourth point cloud feature is fused with the multi-scale fusion image feature in a one-to-many manner, and local attention learning representation is performed to further obtain the fourth fusion feature.
[0061] The fourth fused feature at this time has the contextual information of 3D geometry and image data.
[0062] In this embodiment, point cloud data and image data in the current field of view are acquired; point cloud features are extracted based on the point cloud data; multi-scale feature extraction is performed on the image data to obtain corresponding multiple image features; multiple image features are fused to obtain multi-scale fused image features; and fusion features are obtained based on point cloud features and multi-scale fused image features, so that the obtained fusion features have not only multi-scale image feature information, but also point cloud feature information, thereby improving the matching degree of feature fusion of point cloud data and image data.
[0063] Furthermore, the image data is acquired by a camera, and the camera calibration method optimizes the calibration matrix in an iterative and dynamic manner. As a result of this dynamic optimization, the calibration matrix and the cross-modal feature pairs obtained therefrom are also dynamic. However, in the related art, the one-to-one mapping is obtained by calibration independent of the model and is fixed, which makes the point-based method unaware of the calibration and therefore lacks robustness to calibration errors. The calibration error caused by time and space synchronization is inevitable, and the method proposed in this application to fuse point cloud features and multiple image features to perform a one-to-many mapping between point cloud features and image features can adaptively learn the local context of the image, thereby reducing the impact of calibration errors on fused features.
[0064] See also Figure 3 , Figure 3 1 is a flow chart of another embodiment of the method for feature fusion of point cloud data and image data provided by the present application. The method comprises:
[0065] Step 31: Obtain the point cloud data and image data within the current field of view.
[0066] Step 32: Extract point cloud features based on the point cloud data.
[0067] Step 33: Perform multi-scale feature extraction on the image data to obtain corresponding multiple image features.
[0068] Step 31 to step 33 have the same or similar technical solutions as those in the above embodiment, and will not be described in detail here.
[0069] Step 34: Project the point cloud features onto the plane of the image features to obtain reference points.
[0070] Among them, the point cloud features are scale-invariantly projected based on the calibration parameters to obtain the reference points.
[0071] In some embodiments, the point cloud features are projected based on the image plane of the image data to obtain reference points.
[0072] Specifically, the calibration parameter may be a calibration matrix. The calibration matrix may be dynamically determined when the camera captures the image data. Then, the point cloud features are projected based on the calibration matrix to obtain the reference points.
[0073] Among them, the reference point after point cloud feature projection is a 2D point, and the reference point serves as the initial clue for subsequent fusion.
[0074] The calibration matrix of K cameras is defined as
[0075] The points in the point cloud feature are projected onto the 2D plane using the following formula:
[0076]
[0077] in, Indicates the concat operation. is the homogeneous coordinate corresponding to P. Then It can be normalized by its third dimension coordinate to obtain homogeneous coordinates Next, each reference point The coordinates of are normalized to the range [0, 1] by segmenting the image, so they remain consistent across different scales of image features.
[0078] Step 35: Fuse multiple image features to obtain multi-scale fused image features.
[0079] Step 36: Obtain target image features from the multi-scale fusion image features based on the reference points.
[0080] For example, prediction is performed based on the reference point to obtain the offset and attention weight corresponding to the reference point. The target image features are obtained based on the reference point, offset and attention weight.
[0081] Step 37: Obtain fusion features based on point cloud features and target image features.
[0082] In some embodiments, if the image data includes image data captured by at least two cameras, the image data is combined with Figure 4 To explain:
[0083] Step 41: Obtain the corresponding reference point in each image data.
[0084] Step 42: Averaging the target image features corresponding to each reference point to obtain an average target image feature.
[0085] It can be understood that, since the image data of all K cameras overlap, each point cloud data corresponds to a valid reference point in each image data.
[0086] The same point cloud data may correspond to multiple target image features. Therefore, the target image features corresponding to each reference point are averaged to obtain the average target image features.
[0087] In some embodiments, the target image features corresponding to each reference point may be averaged using the following formula:
[0088]
[0089] Among them, K * Indicates the number of valid reference points, the number of reference points on all images corresponding to a point in the point cloud data. represents the average target image feature, Indicates k * The target image features corresponding to the reference points.
[0090] Step 43: Obtain fusion features based on the point cloud features and the average target image features.
[0091] In this embodiment, a reference point is obtained by projecting the point cloud feature onto the plane of the image feature, a target image feature is obtained from the multi-scale fused image feature based on the reference point, and a fused feature is obtained according to the point cloud feature and the target image feature. The obtained fused feature has not only the information of the multi-scale image feature, but also the information of the point cloud feature, thereby improving the matching degree of the feature fusion of the point cloud data and the image data.
[0092] See also Figure 5 , Figure 51 is a flow chart of another embodiment of the method for feature fusion of point cloud data and image data provided by the present application. The method comprises:
[0093] Step 51: Obtain point cloud data and image data within the current field of view.
[0094] Step 52: Extract point cloud features based on the point cloud data.
[0095] Step 53: Perform multi-scale feature extraction on the image data to obtain corresponding multiple image features.
[0096] Step 54: Project the point cloud features onto the plane of the image features to obtain reference points.
[0097] The reference points are obtained by performing scale-invariant projection based on the calibration parameters.
[0098] Step 55: Fuse multiple image features to obtain multi-scale fused image features.
[0099] Steps 51 to 55 have the same or similar technical solutions as any of the above embodiments, and are not described in detail here.
[0100] Step 56: Obtain the offset and attention weight corresponding to the reference point based on the reference point.
[0101] In some embodiments, see Figure 6 , step 56 may be the following process:
[0102] Step 561: Perform dynamic reference point feature enhancement on the reference point to obtain enhanced reference point features.
[0103] In some embodiments, see Figure 7 , step 561 may be the following process:
[0104] Step 71: Normalize the point cloud features processed by MLP to obtain normalized point cloud features.
[0105] MLP (Multilayer Perceptron). The layers in a multilayer perceptron are fully connected. The bottom layer of a multilayer perceptron is the input layer, the middle layer is the hidden layer, and the last layer is the output layer.
[0106] Step 72: Normalize the reference point features after the convolution process to obtain normalized convolution reference point features.
[0107] Step 73: Fuse the normalized point cloud features and the normalized convolution reference point features to obtain enhanced reference point features.
[0108] In some embodiments, for each reference point The corresponding point cloud feature is f n ∈F. Obtain image features through bilinear interpolation:
[0109]
[0110] in, is the image feature corresponding to the nth reference point on the lth level image from the kth camera. For each camera view k, the reference point feature can be processed using 1×1 convolution and the point cloud feature can be processed using MLP to unify the feature channels of the image feature and point cloud feature corresponding to the reference point.
[0111] Then the enhanced reference point feature is obtained by the following formula:
[0112]
[0113] in, represents the concat operation, and LN represents the LayerNorm function.
[0114] Step 562: Predict the enhanced reference point features to obtain the offset and attention weight corresponding to the reference point.
[0115] Step 57: Obtain target image features based on reference points, offsets, and attention weights.
[0116] Before predicting the reference point feature, the reference point feature may be position-coded, such as cosine coding or sine coding.
[0117] Then is fed into two MLP layers to predict the reference point offset and the corresponding attention weights A total of D offset points and weights pointing in M directions on L levels of image features are predicted, with a total number of D×M×L.
[0118] By adding the offset to the coordinates of the reference point, the corresponding coordinate points can be determined, and these coordinate points correspond to the target image features. That is, one reference point corresponds to image features of multiple scales.
[0119] Step 58: Obtain fusion features based on point cloud features and target image features.
[0120] The features of all learned offset positions are sampled and weightedly added by the learned weights to obtain the fused features:
[0121]
[0122] in, Implemented by softmax.
[0123] In this embodiment, the reference point is obtained by projecting the point cloud feature onto the plane of the image feature, and the reference point is dynamically enhanced to obtain the enhanced reference point feature, so that the obtained fusion feature has not only the information of the multi-scale image feature, but also the information of the point cloud feature, thereby improving the matching degree of the feature fusion of the point cloud data and the image data. Furthermore, the reference point is dynamically enhanced to obtain the enhanced reference point feature, and the dynamic reference feature enhancement can be used to perceive the image calibration that is not related to the model to prevent mismatching.
[0124] In an application scenario, the method of the above embodiment can be applied to target recognition, for example, by projecting the coordinates of the 3D point cloud data onto the image plane in a scale-invariant manner to obtain a 2D reference point.
[0125] The multi-scale image features corresponding to the 2D reference points and the 3D sparse features are enhanced through dynamic reference point feature enhancement to obtain the reference point features.
[0126] Sparse fused features are obtained via dynamic cross-attention fusion.
[0127] Specifically, the reference point features are positionally encoded; then the encoded reference point features are input into two MLP layers to predict offsets and corresponding weights, where the offsets cover multiple directions and multiple image scales.
[0128] Then, the target image features corresponding to the reference point plus the predicted offset are obtained, weighted averaged using the predicted weights, and then added to the input 3D sparse features to obtain the 3D sparse fusion features.
[0129] Then the target object is identified based on the 3D sparse fusion features.
[0130] In this embodiment, the fusion process can be modularized, such as forming a dynamic cross-attention fusion module. The dynamic cross-attention fusion module learns a one-to-many mapping across modalities through a local attention module, thereby modeling dynamic calibration and obtaining robustness to calibration errors. This application uses offset and weight learning as the implementation of the local attention module, and the one-to-many mapping obtained by other local attention modules should also belong to the content covered by this application.
[0131] The dynamic cross-attention module can be a plug-and-play fusion module, and the entire network architecture is applicable to general image feature extraction and point cloud perception. This application uses 3D detectors and 2D detectors as examples, but the dynamic cross-attention module is also applicable to tasks such as 3D semantic segmentation, 3D panorama segmentation, and 3D object tracking. The image feature extraction network can also use the network corresponding to tasks such as image semantic segmentation and image instance segmentation.
[0132] See also Figure 8 , Figure 8 80 is a schematic diagram of an embodiment of an autonomous driving computing platform provided by the present application. The autonomous driving computing platform 80 includes a processor 81 and a memory 82 coupled to the processor 81, the memory 82 is used to store a computer program, and the processor 81 is used to execute the computer program to implement the following method:
[0133] Acquire point cloud data and image data within the current field of view; extract point cloud features based on the point cloud data; perform multi-scale feature extraction on the image data to obtain corresponding multiple image features; fuse multiple image features to obtain multi-scale fused image features; obtain fused features based on point cloud features and multi-scale fused image features.
[0134] It can be understood that the processor 81 is also used to execute computer programs to implement the methods provided by any of the above embodiments, which will not be elaborated here.
[0135] See also Fig. 9 , Fig. 9 1 is a schematic diagram of a computer readable storage medium according to an embodiment of the present application. The computer readable storage medium 90 is used to store a computer program 91. When the computer program 91 is executed by a processor, it is used to implement the following method:
[0136] Acquire point cloud data and image data within the current field of view; extract point cloud features based on the point cloud data; perform multi-scale feature extraction on the image data to obtain corresponding multiple image features; fuse multiple image features to obtain multi-scale fused image features; obtain fused features based on point cloud features and multi-scale fused image features.
[0137] It can be understood that when the computer program 91 is executed by the processor, it is also used to implement the method provided by any of the above embodiments, which will not be elaborated here.
[0138] In summary, the present application provides a technical solution to model dynamic cross-modal matching by using a dynamic cross-attention module with one-to-many mapping. In contrast to the one-to-one mapping in the related art, the present application uses the initial calibration as a clue to learn multiple offsets and weights in adjacent spaces and at adjacent feature levels. On the one hand, this one-to-many feature mapping adaptively learns local context, so a certain degree of calibration error can be tolerated. On the other hand, when the offset is zero, the one-to-many mapping degenerates into a one-to-one mapping, thereby including the latter. In addition, the present application provides a dynamic reference feature enhancement that incorporates the initial one-to-one feature pair into the offset prediction. Such a feature pair naturally perceives the calibration, while including 3D geometry and initial image context, and can serve as an information guide for the best match.
[0139] In this application, the entire network is referred to as the dynamic criss-cross attention network, which consists of the general image and lidar branches and the dynamic criss-cross attention module. Compared with some fusion methods that only use single-layer image features for fusion and require expensive 2D segmentation annotations, the dynamic criss-cross attention network uses multi-level image features for fusion and only requires more easily accessible 2D detection annotations. The choice of the lidar branch is flexible because point-based, grid-based, and voxel-based methods can all be easily adapted to the dynamic criss-cross attention module.
[0140] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only illustrative, for example, the division of the circuit or unit is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0141] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0142] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0143] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made according to the description and drawings of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A feature fusion method for point cloud data and image data, It is characterized in that The method comprises: Get the point cloud data and image data in the current field of view; Extracting point cloud features based on the point cloud data; Performing multi-scale feature extraction on the image data to obtain corresponding multiple image features; fusing the multiple image features to obtain a multi-scale fused image feature; and projecting the point cloud feature onto the plane of the image feature to obtain a reference point; Obtaining an offset and an attention weight corresponding to the reference point based on the reference point; Obtaining the target image feature based on the reference point, the offset, and the attention weight; Obtaining fusion features according to the point cloud features and the target image features; The fusion feature is used as the target point cloud data, point cloud features are extracted again based on the target point cloud data, and the step of obtaining the fusion feature according to the point cloud feature and the multi-scale fusion image feature is performed until the final fusion feature is obtained.
2. The method according to claim 1, It is characterized in that The method comprises: Obtaining an offset feature corresponding to each offset based on the reference point and the offset; and The target image feature is obtained by performing weighted averaging on each of the offset features based on each of the attention weights.
3. The method according to claim 1, It is characterized in that The step of projecting the point cloud feature onto the plane of the image feature to obtain a reference point comprises: The reference point is obtained by performing scale-invariant projection based on the calibration parameters.
4. The method according to claim 1, It is characterized in that The obtaining, based on the reference point, an offset and an attention weight corresponding to the reference point, comprises: Performing dynamic reference point feature enhancement on the reference point to obtain enhanced reference point features; The enhanced reference point feature is predicted to obtain the offset and attention weight corresponding to the reference point.
5. The method according to claim 4, It is characterized in that The step of performing dynamic reference point feature enhancement on the reference point to obtain enhanced reference point features includes: Normalize the point cloud features processed by MLP to obtain normalized point cloud features; Normalizing the reference point features after convolution processing to obtain normalized convolution reference point features; The normalized point cloud features and the normalized convolution reference point features are fused to obtain the enhanced reference point features.
6. The method according to claim 1, It is characterized in that The method comprises: If the image data includes image data captured by at least two cameras; Obtaining a reference point corresponding to each of the image data; Averaging the target image features corresponding to each of the reference points to obtain an average target image feature; The fusion feature is obtained according to the point cloud feature and the average target image feature.
7. An autonomous driving computing platform, It is characterized in that The autonomous driving computing platform includes a processor and a memory coupled to the processor, the memory is used to store a computer program, and the processor is used to execute the computer program to implement the method described in any one of claims 1-6.
8. A computer-readable storage medium, It is characterized in that The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the computer program is used to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Three-dimensional target detection method based on point cloud and image data fusion
CN114092780A
Novel feature layer data fusion method and system for unmanned driving and target detection method
CN114155414A