New feature layer data fusion method, system and target detection method for unmanned driving

By adopting layer-by-layer convolution and feature mapping based on the Krigin model in 3D object detection of driverless cars, combined with feature layer data fusion methods with multi-scale hierarchical fusion and attention mechanism, the problems of premature data fusion and information loss in the prior art are solved, and more efficient and accurate 3D object detection is achieved.

CN114155414BActive Publication Date: 2025-06-06JIANGSU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111376645.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-19
Publication Date
2025-06-06
Estimated Expiration
2041-11-19

AI Technical Summary

Technical Problem

The existing feature layer fusion method has problems such as premature data fusion, loss of information, and improper processing of information from different perspectives in the 3D object detection of driverless cars, resulting in low detection accuracy.

Method used

A new feature layer data fusion method is adopted to generate feature information of point clouds and images through layer-by-layer convolution and point cloud feature mapping based on the Krigin model, and superimpose and fusion through multi-scale hierarchical fusion and attention mechanism to generate the final fusion feature.

Benefits of technology

It effectively avoids the loss of data semantic information, improves the accuracy and efficiency of 3D object detection, and achieves faster and more accurate information fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155414B_ABST
    Figure CN114155414B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and target detection method for new feature layer data fusion for unmanned driving, including: S1, projecting the collected original point cloud to the bird's-eye view and the foreground view respectively, extracting features, and generating the final point cloud feature information after feature point mapping; S2, extracting features from the information collected by the camera, and generating the final image information through multi-scale hierarchical fusion; S3, superimposing and fusing the generated point cloud feature information and image feature information; S4, voxelizing the original point cloud at multiple scales as additional point cloud information; S5, using the image features after feature extraction as additional image information through spatial pyramid pooling; S6, splicing and fusing the features of S3, S4 and S5 to generate the final fusion feature, and performing three-dimensional target detection. The present invention ensures the integrity of semantic information, reduces the information loss in the fusion process, improves the speed of algorithm operation, and improves the accuracy of road target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of environmental perception of driverless cars, and specifically refers to a feature layer sensor fusion method, system and 3D target detection method. Background Art

[0002] The unmanned driving system mainly consists of three parts: perception, decision-making and control. Achieving accurate three-dimensional perception of vehicles and other objects in the road environment is an important prerequisite for decision-making and control. Compared with two-dimensional target detection, three-dimensional target detection is more complicated and requires the output of more parameters to determine the position and movement direction of the target. Single sensor methods such as cameras can obtain image information. Although there is rich semantic information and texture information, it is difficult to obtain accurate distance information. The lidar method, although it can obtain accurate coordinate information, has a lower resolution than the camera, and the long-distance point cloud is relatively sparse, making it difficult to obtain sufficient semantic information and texture information. Therefore, the multi-sensor fusion method that uses both cameras and lidars has become an effective method to solve the three-dimensional target detection in the perception of unmanned driving environments.

[0003] Multi-sensor fusion methods are mainly divided into three categories: data layer fusion, feature layer fusion and decision layer fusion. Although data layer fusion can obtain cross-information between multiple sensors and has the richest information, data matching is more difficult and the structure is complex, making it difficult to achieve real-time detection. Although decision layer fusion has a simple structure and is easy to implement, it only has data association during detection, which loses the advantage of data fusion to a certain extent. Compared with the above two methods, feature layer fusion, which is between the two, can not only effectively utilize the rich information obtained by multiple sensors, but also has relatively simple data processing and is easier to achieve real-time detection. Therefore, the feature layer data fusion method has become the most promising and widely used sensor fusion target detection algorithm.

[0004] Although various feature layer fusion methods have been proposed, the target detection effect is not satisfactory, and even in some scenarios the detection accuracy is lower than that of the LiDAR method. After a lot of research and discussion, the current feature layer fusion methods have the following shortcomings:

[0005] (1) Data fusion is too early. Data from different sources intersect with each other during the fusion process, which not only fails to increase the amount of information, but also destroys the semantic information of different sensors.

[0006] (2) In the process of data fusion, only simple splicing is used without considering the correlation and matching of data, resulting in information loss.

[0007] (3) In the multi-view fusion process, the information from different viewpoints is simply processed uniformly without considering the different amounts of information from different viewpoints in actual scenes.

[0008] (4) Only the fused information is used for target detection, which causes a certain degree of information loss and ignores the original point cloud and image information. Summary of the invention

[0009] In order to solve the defects of the prior art, the present invention aims to provide a new feature layer data fusion method, which uses a new network structure and data processing method to achieve more effective and faster 3D target detection.

[0010] In order to achieve the above object, the technical solution adopted by the present invention is: a new feature layer data fusion method, the method comprising the following steps:

[0011] Step 1: After projecting the original point cloud information collected by the laser radar onto the bird's-eye view and the foreground view, feature extraction is performed, and the final point cloud feature information is generated after feature point mapping;

[0012] Step 2: Extract features from the image information collected by the camera and generate the final image information through a multi-scale hierarchical fusion module;

[0013] Step 3: Superimpose and fuse the point cloud feature information and image feature information generated in the above steps;

[0014] Step 4: The original point cloud information is voxelized at multiple scales to serve as an additional source of point cloud information;

[0015] Step 5: The extracted image features are pooled through spatial pyramid as an additional source of image information;

[0016] Step 6: Concatenate and fuse the features generated in steps 3, 4, and 5 to generate the final fused features.

[0017] Step 7: Detect three-dimensional objects based on the generated final fusion features.

[0018] Furthermore, the original point cloud information in step 1 is projected to generate the height view and density view of the bird's-eye view, as well as the height view and distance view of the foreground view. The point cloud information generated above is extracted using layer-by-layer convolution, and then the point cloud features are further processed using the traditional point cloud feature extraction network structure. The point cloud feature mapping method based on the Kriging model is used to achieve voxel-to-pixel feature mapping.

[0019] Furthermore, in step 2, the information obtained from the left front and right front views is superimposed, and then the attention mechanism is used to obtain image information of appropriate proportion, and finally the image is spliced ​​and fused. The left rear and right rear are also processed according to the above method to obtain the feature information of the rear of the lanes on both sides. The image features of the front and rear views are spliced ​​and fused with the previously obtained fusion information to generate the final image feature information.

[0020] Furthermore, in step 4, it is necessary to select multiple resolutions of different sizes to divide the point cloud information into grids, obtain voxelized 3D models of different scales, and splice and fuse them to generate additional original point cloud information.

[0021] Furthermore, in step 5, three image feature scales of different sizes are selected to perform spatial pyramid pooling on the image features to generate multi-scale image feature information.

[0022] A new feature layer data fusion system for autonomous driving, including:

[0023] The final point cloud feature information generation module: the original point cloud information collected by the lidar is projected onto the bird's-eye view and the foreground view, and features are extracted. The final point cloud feature information is generated after feature point mapping.

[0024] The original point cloud information is projected onto the bird's-eye view to generate a height view and a density view of the bird's-eye view, and is projected onto the foreground view to generate a height view and a distance view of the foreground view.

[0025] The feature extraction is as follows: first, the sizes of the four views are adjusted to ensure that the four views are of the same size, and then the four views are spliced ​​in the depth direction to generate preliminary point cloud information; layer-by-layer convolution is used to extract features from the point cloud information generated above, and three layer-by-layer convolutions are used in total, and the size of the convolution kernel and the number of output channels are continuously increased to simplify subsequent feature extraction operations; and then the traditional point cloud feature extraction network is used to further process the point cloud features;

[0026] The feature point mapping is to realize the feature mapping from voxels to pixels by using the point cloud feature mapping based on the Kriging model. Specifically, the acquired point cloud feature information is subjected to drift Kriging interpolation layer by layer to generate dense point cloud feature information. The point cloud feature information is known to be Z(x). Considering that the point cloud feature will generate corresponding offset errors during mapping, it is assumed that it is composed of a deterministic drift amount m(x) and a residual part R(x). The specific formula is as follows:

[0027] z(x)=m(x)+R(x)

[0028] The drift is defined as the mathematical expectation of known point cloud feature information, that is, the drift describes the overall distribution characteristics of the point cloud feature information, where the linear function aL(x)+b can be obtained by least squares linear fitting. The specific formula is as follows:

[0029] E(z(x))=m(x)=aL(x)+b

[0030] Finally, the residual is expressed as follows:

[0031] R(x)=z(x)-m(x)=z(x)-aL(x)-b

[0032] z*(x 0 ) represents the interpolation point, i represents the point cloud feature information near the interpolation point, and λ represents the weight information corresponding to different nearby points. The specific formula is as follows:

[0033]

[0034] In the interpolation process, the variance of the interpolation point is solved. When the variance is the smallest, the optimal interpolation fitting model is solved. At this time, the weight matrix obtained is the required interpolation matrix. The specific variance expression is as follows:

[0035]

[0036] After completing the interpolation, by adjusting the interpolation resolution, you can get dense point cloud information with different densities. By adjusting the interpolation resolution to the pixel scale, you can get point cloud feature information that corresponds to the image pixel features one by one.

[0037] The final image feature information generation module: extracts features from the image information collected by the camera and generates the final image feature information through multi-scale hierarchical fusion; specifically:

[0038] Six cameras with different viewing angles placed on the roof of the vehicle are used to obtain images from six different viewing angles, namely, images from the left front view, right front view, left rear view, right rear view, front view, and rear view;

[0039] The left front view image and the right front view image are superimposed, and then the number of channels of the superimposed information is adjusted to meet the requirements through multiple 1×1 convolution kernels. Then, the proportion of the two views in the fusion process is obtained through the attention mechanism, and then multiplied with the original features to obtain image information with appropriate proportions. Finally, splicing and fusion are performed to obtain the image feature information of the front of the lanes on both sides;

[0040] The left rear view image and the right rear view image are also processed according to the above method to obtain image feature information of the rear of the lanes on both sides;

[0041] Superimposing the acquired image feature information of the front of the lanes on both sides with the image feature information of the rear of the lanes on both sides;

[0042] The front view image features and the rear view image features are directly concatenated and fused with the superimposed feature information obtained in S2.4 after the number of channels are adjusted by a 1×1 convolution kernel to generate the final image feature information;

[0043] The calculation formula for obtaining the proportion of these two perspectives in the fusion process through the attention mechanism is as follows:

[0044] F g.CV-FL =F CV-FL ×σ(Conv(F CV-FL ⊕F CV-FR ))

[0045] F g.CV-FR =F CV-FR ×σ(Conv(F CV-FL ⊕F CV-FR ))

[0046] F g.CV-BL =F CV-BL ×σ(Conv(F CV-BL ⊕F CV-BR ))

[0047] F g.CV-BR =F CV-BR ×σ(Conv(F CV-BL ⊕F CV-BR ));

[0048] Superimposed feature information generation module: superimposes and fuses the above-mentioned final point cloud feature information and the final image feature information to obtain superimposed feature information;

[0049] Additional point cloud feature information generation module: The original point cloud information is converted into multi-scale voxelized information as additional point cloud feature information;

[0050] The original point cloud information collected by the LiDAR is voxelized, and multiple resolutions of different sizes are selected to divide the point cloud information into grids, and voxelized 3D models of different scales are obtained, which are spliced ​​and fused to generate additional original point cloud information;

[0051] Additional image feature information generation module: The image features after feature extraction are pooled through spatial pyramid as additional image feature information;

[0052] Select three image feature scales of different sizes to perform spatial pyramid pooling on the image features of six different perspectives to generate multi-scale image feature information;

[0053] Fusion feature information generation module: splicing and fusing the superimposed feature information, additional point cloud feature information, and additional image feature information to generate the final fusion feature information.

[0054] The above modules can be integrated into the unmanned vehicle controller.

[0055] The beneficial effects of the present invention are:

[0056] 1. The layer-by-layer convolution data preprocessing method proposed in the present invention can avoid the crossover of data caused by premature fusion of data from different sources, and ensure the integrity of semantic information in the data fusion process.

[0057] 2. Before fusion, the point cloud features and image features were mapped to voxels and pixels point by point, which improved the fusion effect and accelerated the fusion rate.

[0058] 3. In the multi-perspective fusion process, the real vehicle scene is taken into consideration, and a hierarchical fusion method based on the attention mechanism is adopted to improve the efficiency of data processing.

[0059] 4. The additional use of multi-scale voxelized original point cloud data and spatial pyramid pooled image feature data effectively compensates for the information loss in the fusion process. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 This is a general principle diagram of a novel feature layer data fusion method of the present invention;

[0061] Figure 2 Schematic diagram of a point cloud feature mapping method based on a Kriging model of the present invention;

[0062] Figure 3 A schematic diagram of a multi-view hierarchical fusion method of image information based on an attention mechanism of the present invention;

[0063] Figure 4 Schematic diagram of the multi-scale voxel fusion method of the original point cloud information of the present invention; DETAILED DESCRIPTION

[0064] The present invention will be further described below in conjunction with the accompanying drawings.

[0065] In order to make the purpose and technical solution of the present invention more clearly explained, the specific implementation methods of the present invention are further described in detail below in conjunction with the accompanying drawings.

[0066] The present invention proposes a new feature layer data fusion method, such as Figure 1 As shown, the overall principle diagram of a new feature layer data fusion method provided for the implementation of this application mainly includes the following steps:

[0067] Step 1: After projecting the original point cloud information collected by the laser radar, feature extraction is performed, and the final point cloud feature information is generated after feature point mapping;

[0068] like Figure 1 As shown in the point cloud pipeline flow chart, the original 3D point cloud information is voxelized and projected to generate a bird's-eye view and a foreground view, which are further projected to generate a height view and a density view of the bird's-eye view, as well as a height view and a distance view of the foreground view. Next, the size of the four views is adjusted using a convolution kernel of appropriate size to ensure that the four views are of the same size, and then they are spliced ​​in the depth direction to generate preliminary point cloud information.

[0069] Use layer-by-layer convolution to extract features from the point cloud information generated above. A total of three layer-by-layer convolutions are used. The height and width of all convolution kernels are set to 3, and the number of output channels is set to 32, 64, and 128 respectively. High-dimensional features are gradually generated to simplify subsequent feature extraction operations.

[0070] The point cloud features are then further processed using the traditional point cloud feature extraction network structure. Here, three-dimensional sparse convolution is used to further extract point cloud features. The length, width, and height of the convolution kernel are all set to 3, and the number of convolution kernels is set to 128, 64, 32, and 32 respectively.

[0071] The point cloud feature mapping method based on the Kriging model is further used to realize the feature mapping from voxel to pixel, such as Figure 2 shown.

[0072] The point cloud feature information obtained in the above process is interpolated layer by layer by drift Kriging to generate dense point cloud feature information. The known point cloud feature information is z(x). Considering that the point cloud feature will produce corresponding offset errors when mapping, it is assumed that it consists of a deterministic drift m(x) and a residual part R(x). The specific formula is as follows:

[0073] z(x)=m(x)+R(x)

[0074] The drift is defined as the mathematical expectation E(z(x)) of the known point cloud feature information, that is, the drift describes the overall distribution characteristics of the point cloud feature information, where the linear function aL(x)+b can be obtained by least squares linear fitting. The specific formula is as follows:

[0075] E(z(x))=m(x)=aL(x)+b

[0076] Where a and b represent the scaling ratio and bias of the drift, respectively.

[0077] Finally, the residual can be expressed as the following structure:

[0078] R(x)=z(x)-m(x)=z(x)-aL(x)-b

[0079] Define z*(x 0 ) represents the interpolation point, n represents the n points near the interpolation point, i represents the i-th point near the interpolation point, and λ represents the weight information corresponding to different nearby points. The specific formula is as follows:

[0080]

[0081] In the interpolation process, the variance of the interpolation point is solved. When the variance is the smallest, the optimal interpolation fitting model is solved. The weight matrix obtained at this time is the interpolation matrix we need. The specific variance expression is as follows:

[0082]

[0083] After the interpolation is completed, by adjusting the interpolation resolution, you can get dense point cloud information of different densities. By adjusting the interpolation resolution to the pixel scale, you can get point cloud feature information that corresponds to the image pixel features one by one.

[0084] Step 2: Extract features from the image information collected by the camera and generate the final image information through a multi-scale hierarchical fusion module;

[0085] like Figure 3 As shown in the figure, the six cameras with different viewpoints placed on the roof acquire images from six different viewpoints, namely the left front viewpoint, the right front viewpoint, the left rear viewpoint, the right rear viewpoint, the front viewpoint and the rear viewpoint. Considering the driving scene of the driverless car, the cars in the left and right lanes of the driverless car have little impact on the driving of the car if they do not change lanes. Therefore, the proportion of the feature information of the left front, left rear, right front and right rear viewpoints in multi-scale fusion should be reduced.

[0086] First, the information obtained from the left and right front perspectives is superimposed, and then the number of channels of the superimposed information is adjusted through three 1×1 convolution kernels to meet the requirements. Then, the attention mechanism is used to obtain the proportion of the two perspectives in the fusion process, and then multiplied with the original features to obtain image information with a suitable proportion. Finally, splicing and fusion are performed to obtain the feature information of the front of the lanes on both sides. The image information of the front of the left and right lanes is not always equally important. For example, when there are more cars or obstacles in the left front and fewer in the right front, more attention should be paid to the dynamics of the cars or obstacles in the left front. Therefore, it is particularly important to obtain image information with a suitable proportion.

[0087] The left rear and right rear are also processed according to the above method to obtain the feature information of the rear of the lanes on both sides.

[0088] Furthermore, the proportion of image information is obtained through the attention mechanism, and its calculation formula is as follows:

[0089] F g.CV-FL =F CV-FL ×σ(Conv(F CV-FL ⊕F CV-FR ))

[0090] F g.CV-FR =F CV-FR ×σ(Conv(F CV-FL ⊕F CV-FR ))

[0091] F g.CV-BL =F CV-BL ×σ(Conv(F CV-BL ⊕F CV-BR ))

[0092] F g.CV-BR =F CV-BR ×σ(Conv(F CV-BL ⊕F CV-BR ))

[0093] In the formula, F CV-FL 、F CV-FR 、F CV-BL 、F CV-BR Respectively represent the original image information of the left front, right front, left rear, and right rear of the lane, while F g.CV-FL 、F g.CV-FR 、F g.CV-BL 、F g.CV-BR They represent the fused left front, right front, left back, and right back image information respectively, σ represents the activation function, and Conv represents the two-dimensional convolution operation.

[0094] Next, the acquired front information of the lanes on both sides is superimposed and fused with the rear information of the lanes on both sides;

[0095] Correspondingly, the driving status of the vehicles in front and behind the driverless car in the same lane has a greater impact on the vehicle. Therefore, the image features of the front and rear perspectives are directly spliced ​​and fused with the previously obtained fusion information after the number of channels is adjusted by the 1×1 convolution kernel to generate the final image feature information.

[0096] Step 3: Figure 1 As shown in the fusion module in, the point cloud feature information and image feature information generated in the above steps are superimposed and fused, that is, the point cloud feature information and the image feature information are spliced ​​in the depth direction using the Concat operation;

[0097] Step 4: The original point cloud information is voxelized at multiple scales to serve as an additional source of point cloud information;

[0098] like Figure 4 As shown, the original point cloud information collected by the lidar is voxelized, and four different resolutions of 0.2 meters, 0.4 meters, 0.6 meters and 0.8 meters are selected to divide the point cloud information into grids, and voxelized three-dimensional models of different scales are obtained. After splicing and fusion, additional original point cloud information is generated.

[0099] Step 5: The extracted image features are pooled through spatial pyramid as an additional source of image information;

[0100] like Figure 1 As shown in the figure, the image features of 6 different perspectives obtained by pooling kernels of sizes 3, 5, and 7 are selected for pooling operation, so as to obtain image features of three different scales, and then they are superimposed in the depth direction to realize spatial pyramid pooling and generate multi-scale image feature information.

[0101] The features generated in step 6, step 3, step 4, and step 5 are concatenated and fused using the Concat operation to generate the final fused features.

[0102] Step 7: Detect the three-dimensional target based on the final fusion features obtained in step 6. The specific detection method is as follows:

[0103] Here, the two-stage target detection algorithm is taken as an example. The fused features are input into the candidate box generation network to obtain the candidate box, and the feature map selected by the candidate box is pooled into the region of interest. Finally, bounding box regression and confidence prediction are performed based on the pooled features to achieve three-dimensional target detection.

[0104] The series of detailed descriptions listed above are only specific descriptions of feasible implementation methods of the present invention. They are not intended to limit the scope of protection of the present invention. All equivalent methods or changes that do not deviate from the technical creation of the present invention should be included in the scope of protection of the present invention.

Claims

1. A new feature layer data fusion method for autonomous driving. It is characterized in that The steps include: Step 1: Project the original point cloud information collected by the laser radar to the bird's-eye view and the foreground view respectively, extract features, and generate the final point cloud feature information after feature point mapping; Step 2: Extract features from the image information collected by the camera and generate the final image feature information through multi-scale hierarchical fusion; The specific implementation of step 2 includes: S2.1 Six cameras with different viewing angles placed on the roof of the vehicle acquire images with six different viewing angles, namely, images of the left front viewing angle, the right front viewing angle, the left rear viewing angle, the right rear viewing angle, the front viewing angle, and the rear viewing angle; S2.2 superimposes the left front view image and the right front view image, and then adjusts the number of channels of the superimposed information to meet the requirements through three 1×1 convolution kernels. Then, the attention mechanism is used to obtain the proportion of the two views in the fusion process, and then multiplies them with the original features to obtain image information with appropriate proportions. Finally, splicing and fusion are performed to obtain the image feature information of the front of the lanes on both sides. S2.3 The left rear view image and the right rear view image are also processed according to the method of S2.2 to obtain image feature information of the rear of the lanes on both sides; S2.4 superimposes the acquired image feature information of the front of the lanes on both sides with the image feature information of the rear of the lanes on both sides; S2.5 directly merges the front view image features and the rear view image features with the superimposed feature information obtained in S2.4 after adjusting the number of channels through a 1×1 convolution kernel to generate the final image feature information; Step 3: superimpose and fuse the point cloud feature information and image feature information generated in the above steps 1 and 2 to obtain superimposed feature information; Step 4: The original point cloud information is voxelized at multiple scales to serve as additional point cloud feature information; The implementation of step 4 includes: voxelizing the original point cloud information collected by the laser radar, selecting multiple resolutions of different sizes to divide the point cloud information into grids, obtaining voxelized three-dimensional models of different scales, splicing and fusing them, and generating additional original point cloud information; Step 5: The image features after feature extraction are pooled through spatial pyramid as additional image feature information; The implementation of step 5 includes: Select three image feature scales of different sizes to perform spatial pyramid pooling on the image features of six different perspectives to generate multi-scale image feature information; Step 6: Concatenate and fuse the three features generated in steps 3, 4, and 5 to generate the final fused feature information.

2. The novel feature layer data fusion method for unmanned driving according to claim 1, It is characterized in that In the step 1, the original point cloud information is projected onto the bird's-eye view to generate a height view and a density view of the bird's-eye view, and is projected onto the foreground view to generate a height view and a distance view of the foreground view.

3. The novel feature layer data fusion method for unmanned driving according to claim 1, It is characterized in that The feature extraction in step 1 first adjusts the size of the four views to ensure that the four views are of the same size, and then performs splicing in the depth direction to generate preliminary point cloud information; uses layer-by-layer convolution to extract features from the point cloud information generated above, and uses three layer-by-layer convolutions in total, continuously increasing the size of the convolution kernel and the number of output channels to simplify subsequent feature extraction operations; then uses three-dimensional sparse convolution to extract further point cloud features, the length, width and height of the convolution kernel are all set to 3, and the number of convolution kernels is set to 128, 64, 32, and 32 respectively.

4. The novel feature layer data fusion method for unmanned driving according to claim 1, It is characterized in that The feature point mapping in step 1 is to realize the feature mapping from voxels to pixels by using the point cloud feature mapping based on the Kriging model; the details are as follows: The acquired point cloud feature information is interpolated layer by layer by drift Kriging to generate dense point cloud feature information; the known point cloud feature information is z(x). Considering that the point cloud feature will produce corresponding offset errors when mapping, it is assumed that it consists of a deterministic drift m(x) and a residual part R(x). The specific formula is as follows: z(x)=m(x)+R(x) The drift is defined as the mathematical expectation E(z(x)) of the known point cloud feature information, that is, the drift describes the overall distribution characteristics of the point cloud feature information, where the linear function aL(x)+b can be obtained by least squares linear fitting. The specific formula is as follows: E(z(x))=m(x)=aL(x)+b Finally, the residual is expressed as follows: R(x)=z(x)-m(x)=z(x)-aL(x)-b z*(x 0 ) represents the interpolation point, i represents the point cloud feature information near the interpolation point, and λ represents the weight information corresponding to different nearby points. The specific formula is as follows: In the interpolation process, the variance of the interpolation point is solved. When the variance is the smallest, the optimal interpolation fitting model is solved. At this time, the weight matrix obtained is the required interpolation matrix. The specific variance expression is as follows: After completing the interpolation, by adjusting the interpolation resolution, dense point cloud information with different densities can be obtained. By adjusting the interpolation resolution to the pixel scale, point cloud feature information that corresponds one-to-one to the image pixel features can be obtained.

5. The novel feature layer data fusion method for unmanned driving according to claim 1, It is characterized in that The calculation method for obtaining the proportion of the two perspectives in the fusion process through the attention mechanism is as follows: In the formula, F CV-FL 、F CV-FR 、F CV-BL 、F CV-BR Respectively represent the original image information of the left front, right front, left rear, and right rear of the lane, F g.CV-FL 、F g.CV-FR 、F g.CV-BL 、F g.CV-BR They respectively represent the fused left front, right front, left back, and right back image information, σ represents the activation function, and Conv represents a two-dimensional convolution operation.

6. A new feature layer data fusion system for unmanned driving, It is characterized in that include: The final point cloud feature information generation module: the original point cloud information collected by the lidar is projected onto the bird's-eye view and the foreground view, and features are extracted. The final point cloud feature information is generated after feature point mapping. The original point cloud information is projected onto the bird's-eye view to generate a height view and a density view of the bird's-eye view, and is projected onto the foreground view to generate a height view and a distance view of the foreground view. The feature extraction is as follows: first, the sizes of the four views are adjusted to ensure that the four views are of the same size, and then the four views are spliced ​​in the depth direction to generate preliminary point cloud information; layer-by-layer convolution is used to extract features of the point cloud information generated above, and three layer-by-layer convolutions are used in total, and the size of the convolution kernel and the number of output channels are continuously increased to simplify subsequent feature extraction operations; then, three-dimensional sparse convolution is used to extract further point cloud features, and the length, width and height of the convolution kernel are all set to 3, and the number of convolution kernels is set to 128, 64, 32, and 32 respectively; The feature point mapping is to realize the feature mapping from voxels to pixels by using the point cloud feature mapping based on the Kriging model. Specifically, the acquired point cloud feature information is subjected to drift Kriging interpolation layer by layer to generate dense point cloud feature information. The point cloud feature information is known to be z(x). Considering that the point cloud feature will generate corresponding offset errors during mapping, it is assumed that it is composed of a deterministic drift amount m(x) and a residual part R(x). The specific formula is as follows: z(x)=m(x)+R(x) The drift is defined as the mathematical expectation E(z(x)) of the known point cloud feature information, that is, the drift describes the overall distribution characteristics of the point cloud feature information, where the linear function aL(x)+b can be obtained by least squares linear fitting. The specific formula is as follows: E(z(x))=m(x)=aL(x)+b Finally, the residual is expressed as follows: R(x)=z(x)-m(x)=z(x)-aL(x)-b z*(x 0 ) represents the interpolation point, i represents the point cloud feature information near the interpolation point, and λ represents the weight information corresponding to different nearby points. The specific formula is as follows: In the interpolation process, the variance of the interpolation point is solved. When the variance is the smallest, the optimal interpolation fitting model is solved. At this time, the weight matrix obtained is the required interpolation matrix. The specific variance expression is as follows: After completing the interpolation, by adjusting the interpolation resolution, you can get dense point cloud information with different densities. By adjusting the interpolation resolution to the pixel scale, you can get point cloud feature information that corresponds to the image pixel features one by one. The final image feature information generation module: extracts features from the image information collected by the camera and generates the final image feature information through multi-scale hierarchical fusion; Specifically: Six cameras with different viewing angles placed on the roof of the vehicle are used to obtain images from six different viewing angles, namely, images from the left front view, right front view, left rear view, right rear view, front view, and rear view; The left front view image and the right front view image are superimposed, and then the number of channels of the superimposed information is adjusted to meet the requirements through three 1×1 convolution kernels. Then, the proportion of the two views in the fusion process is obtained through the attention mechanism, and then multiplied with the original features to obtain image information with appropriate proportions. Finally, splicing and fusion are performed to obtain the image feature information of the front of the lanes on both sides; The left rear view image and the right rear view image are also processed according to the above-mentioned processing method of the left front view image and the right front view image to obtain image feature information of the rear of the lanes on both sides; Superimposing the acquired image feature information of the front of the lanes on both sides with the image feature information of the rear of the lanes on both sides; The front view image features and the rear view image features are directly concatenated and fused with the superimposed feature information obtained in S2.4 after the number of channels are adjusted by a 1×1 convolution kernel to generate the final image feature information; The calculation formula for obtaining the proportion of these two perspectives in the fusion process through the attention mechanism is as follows: In the formula, F CV-FL 、F CV-FR 、F CV-BL 、F CV-BR Respectively represent the original image information of the left front, right front, left rear, and right rear of the lane, F g.CV-FL 、F g.CV-FR 、F g.CV-BL 、F g.CV-BR They represent the fused left front, right front, left back, and right back image information respectively, σ represents the activation function, and Conv represents the two-dimensional convolution operation; Superimposed feature information generation module: superimposes and fuses the above-mentioned final point cloud feature information and the final image feature information to obtain superimposed feature information; Additional point cloud feature information generation module: The original point cloud information is converted into multi-scale voxelized information as additional point cloud feature information; The original point cloud information collected by the LiDAR is voxelized, and multiple resolutions of different sizes are selected to divide the point cloud information into grids, and voxelized 3D models of different scales are obtained, which are spliced ​​and fused to generate additional original point cloud information; Additional image feature information generation module: The image features after feature extraction are pooled through spatial pyramid as additional image feature information; Select three image feature scales of different sizes to perform spatial pyramid pooling on the image features of six different perspectives to generate multi-scale image feature information; Fusion feature information generation module: splicing and fusing the superimposed feature information, additional point cloud feature information, and additional image feature information to generate the final fusion feature information.

7. The target detection method based on the novel feature layer data fusion system for unmanned driving according to claim 6, It is characterized in that The fused features are input into the candidate box generation network to obtain the candidate box, and the feature map selected by the candidate box is pooled into the region of interest. Finally, bounding box regression and confidence prediction are performed based on the pooled features to achieve three-dimensional object detection.

Citation Information

Patent Citations

  • 3D target detection method based on multi-sensor data fusion

    CN111209840A