Multi-sensor feature fusion method, system and product

Through preprocessing and feature extraction and fusion of image data and point cloud data, the fusion problem of lidar and cameras in autonomous driving is solved, and the accuracy and efficiency of environmental perception are improved.

CN119942290APending Publication Date: 2025-05-06CHONGQING CHANGAN TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510102409.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the existing autonomous driving technology, lidar and cameras have their own limitations in environmental perception, and it is difficult to efficiently and accurately integrate multi-sensor features, affecting the environmental perception ability of the autonomous driving system.

Method used

After preprocessing the image data and point cloud data, the image feature extraction network and the sparse 3D convolutional network are used to extract features, and the feature fusion is combined with the attention mechanism and the hybrid attention mechanism to achieve the alignment and fusion of lidar and camera data from the BEV perspective.

Benefits of technology

The fusion accuracy and feature fusion efficiency of laser point cloud data and image data are improved, and the environmental perception accuracy and real-time processing efficiency of autonomous driving vehicles are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942290A_ABST
    Figure CN119942290A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-sensor feature fusion method, system and product, and the method comprises the steps: carrying out the preprocessing of an image and point cloud data, and obtaining target image data and target point cloud data; extracting features of the target image data to obtain a multi-scale image feature map; extracting features of the target point cloud data to obtain a point cloud feature map; enhancing key area features of the point cloud feature map through an attention mechanism to obtain a target point cloud feature map; projecting the multi-scale image feature map to a BEV visual angle based on calibrated camera internal and external parameters; converting the target point cloud feature map to a BEV visual angle based on calibration parameters between the camera and the laser radar; aligning the multi-scale image feature map and the target point cloud feature map under the BEV visual angle; and performing feature fusion on the aligned multi-scale image feature map and the target point cloud feature map through a mixed attention mechanism to obtain a fused feature map. The invention aims to improve the accuracy of multi-sensor feature fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of feature fusion technology, and in particular to a method, system and product for multi-sensor feature fusion. Background Art

[0002] With the rapid development of autonomous driving technology, multi-sensor fusion technology has become one of the key technologies to improve the environmental perception ability of autonomous driving systems. At present, autonomous driving vehicles mainly obtain information about the surrounding environment through sensors such as Lidar and cameras. Lidar can accurately measure the distance of objects, but it is expensive and greatly affected by weather; cameras have high recognition accuracy for the shape and category of objects, but it is difficult to accurately judge the distance and is easily affected by lighting conditions. Therefore, how to complement the advantages of multiple sensors and achieve efficient and accurate deep feature fusion is an urgent problem to be solved in the current field of autonomous driving. Summary of the invention

[0003] In view of this, the present application provides a method, system and product for multi-sensor feature fusion, aiming to improve the accuracy of multi-sensor feature fusion.

[0004] The first aspect of the present application provides a method for multi-sensor feature fusion, the method comprising:

[0005] Preprocess the collected image data and point cloud data respectively to obtain target image data and target point cloud data;

[0006] Performing feature extraction on the target image data through an image feature extraction network to obtain a corresponding multi-scale image feature map;

[0007] Performing spatial layout and geometric feature extraction on the target point cloud data through a sparse 3D convolutional network to obtain a corresponding point cloud feature map, wherein the point cloud feature map records the spatial layout and geometric features of the point cloud;

[0008] The features of the key areas in the point cloud feature map are enhanced by an attention mechanism to obtain a target point cloud feature map;

[0009] Based on the calibrated camera internal and external parameters, the multi-scale image feature map is projected from the 2D image space to the BEV perspective through perspective transformation;

[0010] Based on the calibration parameters between the camera and the laser radar, the target point cloud feature map is converted to the BEV perspective;

[0011] Align the multi-scale image feature map from the BEV perspective with the target point cloud feature map;

[0012] The aligned multi-scale image feature map and the target point cloud feature map are fused through the hybrid attention mechanism to obtain a fused feature map.

[0013] Optionally, the method further includes:

[0014] Inputting the fused feature map into a multi-task model for parallel multi-task processing;

[0015] Through a query with position encoding corresponding to the task, query the feature information of interest to the task at the corresponding position in the fused feature map;

[0016] The feature information of interest obtained from the query is input into the task header of each corresponding task for processing to obtain the output result of each task.

[0017] Optionally, feature fusion is performed on the aligned multi-scale image feature map and the target point cloud feature map through a hybrid attention mechanism to obtain a fused feature map, including:

[0018] Generate a set number of target queries according to the position information of the aligned multi-scale image feature map and the target point cloud feature map, wherein the target query includes a position code representing itself in space;

[0019] Calculate the attention weights between target queries through a multi-head self-attention mechanism;

[0020] Performing weighted summation on the target Query according to the attention weight to obtain a modified target Query;

[0021] Through the cross-attention mechanism, the corrected target Query is cross-attended with the aligned multi-scale image feature map and the target point cloud feature map to obtain a weighted fused fusion feature map.

[0022] Optionally, after performing cross attention calculation on the modified target Query, the aligned multi-scale image feature map and the target point cloud feature map through a cross attention mechanism to obtain a weighted fused fusion feature map, the method further includes:

[0023] Determine that the modified target Query after performing the cross attention calculation is a fused target Query, wherein the fused target Query includes a target Query that fuses information of a multi-scale image feature map and a target point cloud feature map;

[0024] Based on the fusion target Query obtained in each round, multiple rounds of cross-attention calculations are performed through the multi-head self-attention mechanism and the cross-attention mechanism to obtain the fusion feature map after multiple rounds of weighted fusion.

[0025] Optionally, preprocessing the collected image data to obtain target image data includes:

[0026] Performing denoising processing on the collected image data by Gaussian filtering and median filtering to obtain first image data;

[0027] Performing color adjustment on the first image data by using an automatic white balance algorithm to obtain second image data;

[0028] Performing consistency correction on the second image data by color space conversion and color histogram matching to obtain third image data;

[0029] Performing contrast enhancement processing on the third image data by histogram equalization to obtain fourth image data;

[0030] Mapping the grayscale range of the fourth image data to a target grayscale range by preset stretching parameters;

[0031] The fourth image data after grayscale range mapping is carefully optimized by contrast gain and brightness adjustment to obtain optimized image data;

[0032] The irrelevant background in the optimized image data is cropped at the original ratio of the optimized image data, and the cropped optimized image data is scaled to the size required for the model input to obtain the target image data, wherein the original ratio is proportional to the size required for the model input.

[0033] Optionally, the collected point cloud data is preprocessed to obtain target point cloud data, including:

[0034] According to a preset voxel size and a neighborhood point number threshold, voxel noise points are removed from the collected point cloud data to obtain first point cloud data;

[0035] According to a preset distance threshold and a point number threshold, statistical noise points are removed from the first point cloud data to obtain second point cloud data;

[0036] Determine a corresponding target voxel size according to the point cloud density and resolution of the second point cloud data;

[0037] voxelize the second point cloud data according to the target voxel size to obtain voxelized point cloud data, wherein each voxel grid records the number, average position, and density feature information of the point clouds in the voxel grid;

[0038] Based on the calibrated camera internal and external parameters and lidar parameters, the voxelized point cloud data is converted into the same coordinate system as the target image data to obtain the target point cloud data.

[0039] Optionally, when the output results of each task include the target detection result, the cross-attention calculation is performed on the modified target Query, the aligned multi-scale image feature map and the target point cloud feature map through the cross-attention mechanism to obtain the fused feature map after weighted fusion, including:

[0040] The target detection results of the historical frame are fused with the multi-scale image feature map and the target point cloud feature map after alignment processing of the current frame through the hybrid attention mechanism to obtain the fused multi-scale image feature map and the target point cloud feature map;

[0041] Motion compensation is performed on the target query of the historical frame through a preset algorithm to predict the position of each target in the target detection result in the current frame;

[0042] Align and fuse the target query of the motion compensated historical frame with the corrected target query of the current frame to obtain the aligned target query of the current frame;

[0043] Through the cross-attention mechanism, the obtained alignment target Query is cross-attended with the fused multi-scale image feature map and the target point cloud feature map to obtain the weighted fused fusion feature map.

[0044] Optionally, before aligning the multi-scale image feature map under the BEV perspective and the target point cloud feature map, the method further includes:

[0045] The target point cloud feature map under the BEV perspective is subjected to offset noise denoising by means of a qualified denoising convolutional network, and a denoised target point cloud feature map is obtained;

[0046] The aligning process of the multi-scale image feature map under the BEV perspective and the target point cloud feature map comprises: aligning the multi-scale image feature map under the BEV perspective and the denoised target point cloud feature map.

[0047] Optionally, train a qualified denoising convolutional network, including:

[0048] Constructing a training data set, wherein the training data set includes a multi-scale image feature map and a target point cloud feature map under a BEV perspective that have been aligned and have a deviation below a set threshold, and a standard fusion feature map that is a fusion of the multi-scale image feature map and the target point cloud feature map;

[0049] Generate a random offset noise feature, and perform a convolution operation on the random offset noise feature to obtain a target random offset noise feature;

[0050] The target point cloud feature map grid is sampled under the BEV perspective through the target random offset noise feature to obtain the noise point cloud feature map;

[0051] Extracting features of the noise point cloud feature map through a denoising convolutional network to obtain a point cloud feature map;

[0052] Performing feature fusion on the point cloud feature map and the multi-scale image feature map corresponding to the point cloud feature map in the training data set to obtain a fused feature map to be evaluated;

[0053] Evaluate the fused feature map to be evaluated and the corresponding standard fused feature map through a loss function to determine the denoising effect of the denoising convolutional network;

[0054] When the denoising effect satisfies the set conditions, determining that the denoising convolution network training is qualified;

[0055] When the denoising effect does not meet the set conditions, the parameters of the denoising convolutional network are updated and the model training is performed again until the denoising effect obtained meets the set conditions.

[0056] A second aspect of the present application provides a system for multi-sensor feature fusion, the system comprising:

[0057] A preprocessing module is used to preprocess the collected image data and point cloud data respectively to obtain target image data and target point cloud data;

[0058] A first feature extraction module, used to extract features from the target image data through an image feature extraction network to obtain a corresponding multi-scale image feature map;

[0059] A second feature extraction module is used to perform spatial layout and geometric feature extraction on the target point cloud data through a sparse 3D convolutional network to obtain a corresponding point cloud feature map, in which the spatial layout and geometric features of the point cloud are recorded;

[0060] A feature enhancement processing module, used for enhancing the features of the key areas in the point cloud feature map through an attention mechanism to obtain a target point cloud feature map;

[0061] A projection module, used for projecting the multi-scale image feature map from the 2D image space to the BEV perspective through perspective transformation based on the calibrated camera internal and external parameters;

[0062] A perspective conversion module, used for converting the target point cloud feature map to a BEV perspective based on calibration parameters between the camera and the laser radar;

[0063] An alignment processing module is used to align the multi-scale image feature map and the target point cloud feature map under the BEV perspective;

[0064] The feature fusion module is used to fuse the aligned multi-scale image feature map and the target point cloud feature map through a hybrid attention mechanism to obtain a fused feature map.

[0065] The third aspect of the present application provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and running on the processor, wherein when the computer program is executed by the processor, the steps in the method for multi-sensor feature fusion as described in the first aspect of the present application are implemented.

[0066] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for multi-sensor feature fusion as described in the first aspect of the present application are implemented.

[0067] The multi-sensor feature fusion method provided in this application has the following advantages:

[0068] The embodiment of the present application provides a method for multi-sensor feature fusion. First, the collected image data and point cloud data are pre-processed respectively to obtain target image data and target point cloud data; feature extraction is performed on the target image data through an image feature extraction network to obtain a corresponding multi-scale image feature map; spatial layout and geometric feature extraction are performed on the target point cloud data through a sparse 3D convolutional network to obtain a corresponding point cloud feature map, in which the spatial layout and geometric features of the point cloud are recorded; features of key areas in the point cloud feature map are enhanced through an attention mechanism to obtain a target point cloud feature map; based on calibrated camera internal and external parameters, the multi-scale image feature map is projected from the 2D image space to the BEV perspective through perspective transformation; based on the calibration parameters between the camera and the lidar, the target point cloud feature map is converted to the BEV perspective; the multi-scale image feature map and the target point cloud feature map under the BEV perspective are aligned; and feature fusion is performed on the aligned multi-scale image feature map and the target point cloud feature map through a hybrid attention mechanism to obtain a fused feature map. This application combines the self-attention mechanism and sparse perception to improve the fusion accuracy of laser point cloud data and image data while improving the efficiency of feature fusion, thereby improving the environmental perception accuracy and real-time processing efficiency of autonomous driving vehicles. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor.

[0070] Figure 1 A flowchart of a method for multi-sensor feature fusion is shown as an embodiment of the present application;

[0071] Figure 2 A schematic diagram of the arrangement position of a laser radar in a method for multi-sensor feature fusion according to an embodiment of the present application;

[0072] Figure 3 A schematic diagram of a multi-sensor feature fusion system is shown in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0073] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0074] refer to Figure 1 , Figure 1 This is a schematic diagram of a method for multi-sensor feature fusion according to an embodiment of the present application. Figure 1 As shown, the method includes:

[0075] Step S01: pre-process the collected image data and point cloud data respectively to obtain target image data and target point cloud data.

[0076] In this embodiment, the image data of the environment around the vehicle body is collected by the camera arranged around the vehicle body to obtain the panoramic image data of the current frame around the vehicle body, and the following description is based on the image data. The collected image data is subjected to preprocessing operations such as denoising, color correction, and image enhancement to improve the quality of the collected image data. After a series of preprocessing operations are performed on the collected image data, the corresponding target image data is obtained. The laser radar arranged around the vehicle body is used to collect laser point cloud data around the vehicle body. After the point cloud data of the current frame is collected, the point cloud data is subjected to preprocessing operations such as filtering, voxelization, and coordinate conversion, so as to facilitate the subsequent feature extraction of the point cloud data. After a series of preprocessing operations are performed on the collected point cloud data, the corresponding target point cloud data is obtained.

[0077] Step S02: extracting features from the target image data through an image feature extraction network to obtain a corresponding multi-scale image feature map.

[0078] In this embodiment, in order to improve the feature fusion effect, after preprocessing the collected image data and point cloud data of the current frame to obtain the corresponding target image data and target point cloud data through step S1, the obtained target image data is subjected to feature extraction and fusion at different scales through the image feature extraction network to obtain a multi-scale image feature map with rich semantic information, so that the feature map finally obtained by fusion can more accurately reflect the environmental characteristics of the current frame.

[0079] Step S03: Perform spatial layout and geometric feature extraction on the target point cloud data through a sparse 3D convolutional network to obtain a corresponding point cloud feature map, in which the spatial layout and geometric features of the point cloud are recorded.

[0080] In this embodiment, in order to improve the fusion efficiency of the entire multi-sensor feature fusion, the present application performs feature extraction on the target point cloud data obtained after preprocessing through a sparse 3D convolutional network (3D Convolutional Neural Network, 3DCNN, three-dimensional convolutional network), so that the subsequent feature fusion of the point cloud features extracted from the target point cloud data is more efficient. Specifically, the target point cloud data obtained in step S1 is spatially laid out and geometric features are extracted through a sparse 3D convolutional network to obtain a point cloud feature map corresponding to the target point cloud data, which records the spatial layout and geometric features of the point cloud.

[0081] Step S04: Enhance the features of the key areas in the point cloud feature map through the attention mechanism to obtain a target point cloud feature map.

[0082] In this embodiment, in order to improve the fusion effect of subsequent feature fusion, so that the feature map finally fused can more accurately reflect the environmental features of the current frame. The present application further enhances the point cloud feature map extracted by step S3 to improve the effectiveness of the feature representation of the point cloud feature map. Specifically, the present application enhances the features of the key areas in the point cloud feature map obtained by step S3 through the attention mechanism to obtain a target point cloud feature map corresponding to the point cloud feature map.

[0083] Step S05: Based on the calibrated camera internal and external parameters, the multi-scale image feature map is projected from the 2D image space to the BEV perspective through perspective transformation.

[0084] In this embodiment, after obtaining the extracted multi-scale image feature map, based on the camera internal and external parameters of the camera calibration, the obtained multi-scale image feature map is projected from the 2D image space to the BEV perspective (Bird's Eye View) through perspective transformation for feature fusion.

[0085] Step S06: Based on the calibration parameters between the camera and the laser radar, the target point cloud feature map is converted to the BEV perspective.

[0086] In this embodiment, after obtaining the extracted target point cloud feature map, based on the calibration parameters between the camera and the laser radar, the target point cloud feature map is also converted to the BEV perspective for feature fusion. Figure 2 As stated, Figure 2 The layout positions of the laser radar on the vehicle are shown, where 1 is the Falcon K laser radar, 2 is the AT 128 laser radar, 3 is the Pandar128 laser radar, 4 is the GI7683 combined inertial navigation, 5 is the Pandar XT32 laser radar, and 6 is the Pandar QT64 laser radar.

[0087] Step S07: Align the multi-scale image feature map under the BEV perspective and the target point cloud feature map.

[0088] In this embodiment, after obtaining the multi-scale image feature map and the target point cloud feature map under the same BEV perspective, spatial alignment processing is performed on the obtained multi-scale image feature map and the target point cloud feature map to ensure that the two are spatially consistent.

[0089] Step S08: Perform feature fusion on the aligned multi-scale image feature map and the target point cloud feature map through a hybrid attention mechanism to obtain a fused feature map.

[0090] In this embodiment, after the multi-scale image feature map and the target point cloud feature map of the current frame are spatially aligned through step S7, the aligned multi-scale image feature map and the target point cloud feature map are feature fused through a hybrid attention mechanism to obtain a fused feature map of the current frame.

[0091] The embodiment of the present application provides a method for multi-sensor feature fusion. First, the collected image data and point cloud data are pre-processed respectively to obtain target image data and target point cloud data; feature extraction is performed on the target image data through an image feature extraction network to obtain a corresponding multi-scale image feature map; spatial layout and geometric feature extraction are performed on the target point cloud data through a sparse 3D convolutional network to obtain a corresponding point cloud feature map, in which the spatial layout and geometric features of the point cloud are recorded; features of key areas in the point cloud feature map are enhanced through an attention mechanism to obtain a target point cloud feature map; based on calibrated camera internal and external parameters, the multi-scale image feature map is projected from a 2D image space (that is, a two-dimensional plane image, with only horizontal and vertical coordinates) to a BEV perspective through a perspective transformation; based on the calibration parameters between the camera and the laser radar, the target point cloud feature map is converted to the BEV perspective; the multi-scale image feature map and the target point cloud feature map under the BEV perspective are aligned; the aligned multi-scale image feature map and the target point cloud feature map are feature fused through a hybrid attention mechanism to obtain a fused feature map. This application combines the self-attention mechanism and sparse perception to improve the fusion accuracy of laser point cloud data and image data while improving the efficiency of feature fusion, thereby improving the environmental perception accuracy and real-time processing efficiency of autonomous driving vehicles.

[0092] In combination with the above embodiments, in one implementation, the present application embodiment further provides a method for multi-sensor feature fusion. In the multi-sensor feature fusion method, the method further includes steps S09 to S11:

[0093] Step S09: input the fused feature map into a multi-task model for parallel multi-task processing.

[0094] In this embodiment, the present application also involves a multi-task model, which performs parallel multi-task processing on the fused feature map obtained in step S08, and obtains the corresponding task output results for each task, such as lane line detection, semantic segmentation and other parallel tasks.

[0095] Step S10: query the feature information of interest to the task at the corresponding position in the fused feature map through a query with position coding corresponding to the task.

[0096] In this embodiment, for the multi-task model, each task has its own corresponding Query with position encoding. The Query corresponding to the task is used to query the feature information related to the task in the fusion feature map input into the multi-task model, that is, the feature information of interest to the task, and the position encoding carried by the Query is used to indicate that the Query is querying the corresponding one. Query is a query in the self-attention mechanism, representing a request for information, and is a signal used by the model to guide attention to a specific part when processing input data.

[0097] Step S11: Input each piece of feature information of interest obtained through the query into the task header of each corresponding task for processing, and obtain the output result of each task.

[0098] In this embodiment, after the position-encoded Query of each task queries the feature information of interest to its own task from the fused feature map, the queried feature information of interest to its own task is input into the task header of its own task for processing to obtain the respective task output results, such as the lane line detection results, semantic segmentation results, etc.

[0099] In combination with the above embodiments, in one implementation, the present application embodiment further provides a method for multi-sensor feature fusion. In the multi-sensor feature fusion method, step S08 may include steps S081 to S084:

[0100] Step S081: Generate a set number of target queries according to the position information of the aligned multi-scale image feature map and the target point cloud feature map, wherein the target query includes a position code representing itself in space.

[0101] In this embodiment, based on the position information of the aligned multi-scale image feature map and the target point cloud feature map, a set number of target queries are generated, each target query includes its own position code in space to locate the specific position in the feature map. The set number can be set according to the actual application scenario, and is not specifically limited here, such as 100, 200, etc.

[0102] Step S082: Calculate the attention weights between target queries through a multi-head self-attention mechanism.

[0103] In this embodiment, after obtaining a preset number of target queries generated through step S081, the attention weights between the target queries are calculated through a multi-head self-attention mechanism to achieve information integration and enhancement, so that each target query will fuse information from other target queries to obtain a richer context representation.

[0104] Specifically: Set the appropriate number of self-attention heads (such as 8, 16) and parameters (such as query, key, and value dimensions). In the multi-head self-attention mechanism, the original attention calculation is divided into multiple heads (such as 8 or 16), and each head performs calculations independently. In each head, the target query interacts with other target queries to calculate the attention weights between them. An optional implementation is to calculate the dot product similarity between the target queries or use other similarity metrics for calculation.

[0105] Step S083: Perform weighted summation on the target Query according to the attention weight to obtain a modified target Query.

[0106] In this embodiment, based on the calculated attention weights between the target query and other target queries, the feature representations of each target query are weighted and summed to obtain the corresponding new target query, which is the modified target query. Thus, the original representation of each target query will be updated to a new representation, namely the corresponding modified target query, which contains additional information fused from other target queries. This weighted summation process makes the representation of the target query more robust. This is because they not only have their own initial representations, but also fuse information from other target queries. The improvement in robustness enables the target query to better fuse the multi-scale image feature map and the target point cloud feature map in the subsequent cross-attention mechanism.

[0107] Step S084: Through the cross-attention mechanism, the corrected target Query is cross-attended with the aligned multi-scale image feature map and the target point cloud feature map to obtain a weighted fused fusion feature map.

[0108] In this embodiment, appropriate cross-attention weights and parameters are set. After each corrected target Query is obtained through step S083, the corrected target Query and the aligned multi-scale image feature map and the target point cloud feature map are cross-attention calculated through the cross-attention mechanism to obtain the corresponding weighted fused fusion feature map. Specifically: in the cross-attention mechanism, the corrected Query is used as the query, and the aligned multi-scale image feature map and the target point cloud feature map are used as the key (Key) and value (Value). By calculating the attention weights between the target corrected Query and these features, the aligned multi-scale image feature map and the target point cloud feature map are weightedly fused to obtain the corresponding fused feature map. Since the target corrected Query obtains a richer context representation through the multi-head self-attention mechanism, the target corrected Query can better guide the calculation of the cross-attention mechanism, thereby achieving more effective feature fusion.

[0109] In combination with the above embodiments, in one implementation, the present application embodiment further provides a method for multi-sensor feature fusion. In the multi-sensor feature fusion method, after step S084, the method further includes steps S085 to S086:

[0110] Step S085: Determine that the modified target Query after performing the cross-attention calculation is a fused target Query, wherein the fused target Query includes a target Query that fuses information of a multi-scale image feature map and a target point cloud feature map.

[0111] In this embodiment, after the modified target Query performs the cross-attention calculation in step S084, the modified target Query will be fused with the information of the multi-scale image feature map and the target point cloud feature map. At this time, the target Query1 that fuses the information of the multi-scale image feature map and the target point cloud feature map is determined as the fused target Query, which will be used to participate in the cyclic iterative fusion process.

[0112] Step S086: Based on the fusion target Query obtained in each round, multiple rounds of cross-attention calculations are performed through a multi-head self-attention mechanism and a cross-attention mechanism to obtain a fusion feature map after multiple rounds of weighted fusion.

[0113] In this embodiment, after the corrected target query after the cross-attention calculation is performed is used as the fused target query in step S085, the fused target query is used as the target query, and then step S082 is returned to perform a new round of feature weighted fusion. The execution process is the same as the execution process of the above steps S082 to step S085, except that the target query is replaced with the current fused target query. In each round of feature weighted fusion process, the fused target query obtained after the previous round of feature weighted fusion processing is used as the target query to return to step S082 for a new round of feature weighted fusion. The number of times the loop fusion is required is set in advance. After the set number of times the loop fusion is required is reached, the loop fusion is ended, and the fused feature map obtained by the last fusion is determined as the final fused feature map. The final fused feature map is subsequently input into the multi-task model for multi-task parallel processing.

[0114] In combination with the above embodiments, in one implementation, the embodiment of the present application also provides a method for multi-sensor feature fusion. In the method for multi-sensor feature fusion, the pre-processing of the image data in step S01 may include: performing denoising processing on the collected image data by Gaussian filtering and median filtering to obtain first image data; performing color adjustment on the first image data by an automatic white balance algorithm to obtain second image data; performing consistency correction on the second image data by color space conversion and color histogram matching to obtain third image data; performing contrast enhancement processing on the third image data by histogram equalization to obtain fourth image data; mapping the grayscale range of the fourth image data to the target grayscale range by preset stretching parameters; performing detailed optimization on the fourth image data after the grayscale range is mapped by contrast gain and brightness adjustment to obtain optimized image data; cropping the irrelevant background in the optimized image data at the original ratio of the optimized image data, and scaling the cropped optimized image data to the size required for the model input to obtain the target image data, wherein the original ratio is proportional to the size required for the model input.

[0115] In this embodiment, the image data is denoised by Gaussian filtering and median filtering, and the denoising evaluation algorithm is used to evaluate the denoising of the image data after the denoising, and the corresponding denoising evaluation result is obtained. When the denoising evaluation result meets the set conditions, the denoising process is terminated to obtain the corresponding first image data. When the denoising evaluation result does not meet the set conditions, the parameters of the denoising evaluation algorithm are updated to perform new denoising and denoising evaluation until the denoising evaluation result meets the set conditions. Among them, for Gaussian filtering, the filter size (such as 3x3, 5x5) and the standard deviation (such as 0.5, 1.0) are selected for smoothing; for median filtering, the filter size (such as 3x3) is selected for noise suppression; for the denoising evaluation algorithm, it can be a peak signal-to-noise ratio (PSNR) evaluation or a mean square error (MSE) evaluation algorithm.

[0116] In this embodiment, the color of the obtained first image data is adjusted by an automatic white balance algorithm (such as an algorithm based on the gray world assumption) to obtain corresponding second image data.

[0117] In this embodiment, the obtained second image data is then subjected to color consistency correction through color space conversion (such as RGB to Lab) and color histogram matching, and the second image data subjected to color consistency correction is subjected to color consistency evaluation through a color correction evaluation algorithm to obtain a color consistency evaluation effect. When the color consistency evaluation effect meets the corresponding set conditions, the color consistency correction is terminated, and the second image data subjected to color consistency correction is determined as the third image data. When the color consistency evaluation effect does not meet the corresponding set conditions, the parameters of the automatic white balance algorithm and / or the parameters in the color histogram matching process are adjusted to perform new color adjustments and / or color consistency corrections. The second image data subjected to color consistency correction is subjected to color consistency evaluation again through a color correction evaluation algorithm, and this cycle is repeated until the color consistency evaluation effect meets the corresponding set conditions. Among them, the automatic white balance algorithm is preferably an algorithm based on the grayscale world hypothesis; the color correction evaluation algorithm is preferably a visual inspection and color difference metric (such as ΔE) evaluation algorithm.

[0118] In this embodiment, contrast enhancement processing is performed on the obtained third image data by histogram equalization to obtain corresponding fourth image data, wherein a global or local equalization method is selected to perform the histogram equalization.

[0119] In this embodiment, the grayscale range of the fourth image data after contrast enhancement is mapped to the target grayscale range by setting stretching parameters (such as minimum and maximum grayscale values).

[0120] In this embodiment, the image data after grayscale range mapping is then carefully optimized by contrast gain and brightness adjustment, and the carefully optimized image data is evaluated by an image enhancement evaluation algorithm to obtain an enhanced evaluation effect. When the enhanced evaluation result meets the corresponding set conditions, the carefully optimized image data is determined to be the optimized image data; when the enhanced evaluation result does not meet the corresponding set conditions, the parameters in the image enhancement process are adjusted to be carefully optimized again; the carefully optimized image data is re-evaluated by the image enhancement evaluation algorithm until the enhanced evaluation effect meets the corresponding set conditions. Among them, the image enhancement evaluation algorithm selects visual inspection and contrast evaluation indicators (such as contrast ratio).

[0121] In this embodiment, the original scale of the image data is proportional to the size required by the network input. For example, if the size required by the network input is 10*20, the original scale of the image data is a multiple of 10*20. After obtaining the optimized image data, the irrelevant background in the optimized image data is cropped at the original scale of the optimized image data, and then the cropped optimized image data is scaled to the size required by the network input to obtain the target image data. The scaling method includes but is not limited to bilinear interpolation, bicubic interpolation, or nearest neighbor interpolation.

[0122] In combination with the above embodiments, in one implementation, the embodiment of the present application also provides a method for multi-sensor feature fusion. In the multi-sensor feature fusion method, the pre-processing of the point cloud data in step S01 may include: removing voxel noise points from the collected point cloud data according to a preset voxel size and a neighborhood point count threshold to obtain first point cloud data; removing statistical noise points from the first point cloud data according to a preset distance threshold and a point count threshold to obtain second point cloud data; determining the corresponding target voxel size according to the point cloud density and resolution of the second point cloud data; voxelizing the second point cloud data according to the target voxel size to obtain voxelized point cloud data, wherein each voxel grid records the number of point clouds, average position, and density feature information within the voxel grid; based on the calibrated camera internal and external parameters and lidar parameters, converting the voxelized point cloud data to the same coordinate system as the target image data to obtain target point cloud data.

[0123] In this embodiment, based on the preset voxel size (such as 0.1m, 0.2m, etc.) and the neighborhood point number threshold (such as 5, 10, etc.), the point cloud data is first subjected to voxel noise point removal to obtain the corresponding first point cloud data. Then, according to the preset distance threshold and point number threshold, the first point cloud data after the voxel noise point removal is subjected to statistical noise point removal, and the noise point removal effect of the first cloud data after the statistical noise point removal is determined by point cloud visualization inspection. If the noise point removal effect meets the corresponding set conditions, the first point cloud data is determined to be the second point cloud data; if the noise point removal effect does not meet the corresponding set conditions, the parameters of the voxel noise point removal and / or statistical noise point removal are adjusted, and the point cloud data is subjected to noise point removal again until the noise point removal effect meets the corresponding set conditions. After obtaining the second point cloud data, the target voxel size matching the point cloud density and resolution is determined according to the point cloud density and resolution of the second point cloud data, and the second point cloud data is divided into voxel grid data of the target voxel size, thereby obtaining voxelized point cloud data, that is, voxelized point cloud data, in which each voxel network in the voxelized point cloud data records the number of point clouds, average position, density and other characteristic information in the voxel grid. Based on the calibrated internal and external parameters of the camera and the laser radar parameters, the voxelized voxelized point cloud data is converted to the same coordinate system as the target image data, and then the alignment effect after the conversion is determined through visual inspection. If the alignment effect meets the corresponding set conditions, the voxelized target point cloud data after coordinate conversion is obtained; if the alignment effect does not meet the corresponding set conditions, the coordinate conversion is performed again until the alignment effect after the conversion meets the corresponding set conditions.

[0124] In combination with the above embodiments, in one implementation, the embodiment of the present application also provides a method for multi-sensor feature fusion. In the multi-sensor feature fusion method, when the output results of each task obtained include target detection results, step S084 may include: fusing the target detection results of the historical frame with the multi-scale image feature map and the target point cloud feature map after alignment processing of the current frame through a hybrid attention mechanism to obtain a fused multi-scale image feature map and a target point cloud feature map; performing motion compensation on the target query of the historical frame through a preset algorithm to predict the position of each target in the target detection result in the current frame; aligning and fusing the target query of the historical frame after motion compensation with the corrected target query of the current frame to obtain the aligned target query of the current frame; performing cross-attention calculation on the obtained aligned target query and the fused multi-scale image feature map and the target point cloud feature map through a cross-attention mechanism to obtain a fused feature map after weighted fusion.

[0125] In this embodiment, when the output results of each task include the target detection results, another implementation method is provided for step S084, specifically: after the target detection result has been obtained once, for the subsequent feature fusion, the target detection result of the historical frame is fused with the multi-size image feature map and the target point cloud feature map after the alignment processing of the current frame using the hybrid attention mechanism to obtain the fused multi-size image feature map and the target point cloud feature map. The target query of the historical frame is motion compensated by a preset algorithm (Kalman filtering, optical flow estimation) to predict the position of each target in the target detection result in the current frame, and the target query of the historical frame after motion compensation is aligned and fused with the corrected target query of the current frame to obtain the aligned target query of the current frame, such as using the cross attention mechanism for fusion. Then, the cross attention mechanism is used to perform weighted fusion on the aligned target query of the current frame obtained with the fused multi-size image feature map and the target point cloud feature map to obtain the weighted fused fused feature map, and then the fused feature map is used for parallel multi-task processing of the multi-task model.

[0126] In combination with the above embodiments, in one implementation, the embodiment of the present application also provides a method for multi-sensor feature fusion. In the multi-sensor feature fusion method, before step S07, the method also includes: performing offset noise denoising on the target point cloud feature map under the BEV perspective through a qualified denoising convolutional network to obtain a denoised target point cloud feature map; the multi-scale image feature map under the BEV perspective and the target point cloud feature map are aligned, including: aligning the multi-scale image feature map under the BEV perspective and the denoised target point cloud feature map.

[0127] In this embodiment, there may be errors in the calibration of the laser sensor in actual applications, which may cause the laser point cloud data and the visual image data to encounter a situation where the features cannot be aligned. To solve this problem, the present application pre-trains a qualified denoising convolutional network, and then uses the trained qualified denoising convolutional network to perform offset noise denoising on the target point cloud feature map obtained from the BEV perspective to obtain the denoised target point cloud data. In this implementation method, the alignment processing of step S07 is performed on the multi-scale image feature map from the BEV perspective and the denoised target point cloud feature map from the BEV perspective.

[0128] In combination with the above embodiments, in one implementation, the embodiments of the present application also provide a method for multi-sensor feature fusion. In the multi-sensor feature fusion method, training to obtain a qualified denoising convolutional network includes: constructing a training data set, the training data set includes a multi-scale image feature map and a target point cloud feature map under the BEV perspective that have been aligned and have a deviation below a set threshold, and a standard fusion feature map fused with the multi-scale image feature map and the target point cloud feature map; generating a random offset noise feature, and performing a convolution operation on the random offset noise feature to obtain a target random offset noise feature; sampling the target point cloud feature map grid under the BEV perspective through the target random offset noise feature to obtain a noise point cloud feature map; and sampling the target point cloud feature map grid through the denoising convolutional network. The noise point cloud feature map is extracted to obtain a point cloud feature map; the point cloud feature map is feature-fused with a multi-scale image feature map corresponding to the point cloud feature map in a training data set to obtain a fused feature map to be evaluated; the fused feature map to be evaluated and the corresponding standard fused feature map are evaluated by a loss function to determine the denoising effect of the denoising convolutional network; if the denoising effect meets the set conditions, it is determined that the denoising convolutional network is trained to be qualified; if the denoising effect does not meet the set conditions, the parameters of the denoising convolutional network are updated and the model is retrained until the denoising effect obtained meets the set conditions.

[0129] In this embodiment, for the training of the denoising convolutional network, the present application provides an optional implementation method, which is as follows:

[0130] First, a training data set required for training the network is constructed. Each training sample data in the training data set includes a multi-scale image feature map and a target point cloud feature map under the BEV perspective that have been aligned and whose deviations are lower than a set threshold, as well as a standard fusion feature map fused with the multi-scale image feature map and the target point cloud feature map. For each training sample data selected from the training data set, noise will be added to the target point cloud feature map under the BEV perspective in the training sample data. The noise addition method of each target point cloud feature map is the same. For ease of understanding, the noise addition of the target point cloud feature map under the BEV perspective in a training sample data is used as an example for explanation. For the selected training sample data, a corresponding random offset noise feature is randomly generated, and then a convolution operation is performed on the random offset noise feature to obtain a corresponding target random offset noise feature; the target point cloud feature map in the training sample data is grid sampled by the generated target random offset noise feature to obtain a corresponding noise point cloud feature map. Through the same implementation method, a corresponding noise point cloud feature map with noise added will be obtained for each selected training sample data. Then, the obtained noise point cloud feature map is extracted by the denoising convolution network to obtain a point cloud feature map. For the point cloud feature map extracted by the denoising convolution network, it is feature fused with the multi-scale image feature map in the corresponding training sample data to obtain a fused feature map to be evaluated. The obtained fused feature map to be evaluated and the standard fused feature map in the corresponding training sample data are evaluated by the loss function to determine the denoising effect of the denoising convolution network. If the denoising effect meets the corresponding set conditions, it is determined that the denoising convolution network is qualified for training; if the denoising effect does not meet the corresponding set conditions, the parameters of the denoising convolution network are updated and the model is retrained until the denoising effect meets the corresponding set conditions, and a qualified denoising convolution network is obtained. For example, noise is added to the target point cloud feature map a1 in the training sample data A to obtain the noise point cloud feature map a2, and the noise point cloud feature map a2 is extracted by the denoising convolutional network to obtain the point cloud feature map a3. The point cloud feature map a3 extracted by the denoising convolutional network is fused with the multi-scale image feature map b in the corresponding training sample data A to obtain the fused feature map a3b to be evaluated. The obtained fused feature map a3b to be evaluated and the standard fused feature map ab in the corresponding training sample data A are evaluated by the loss function to determine the denoising effect of the denoising convolutional network.

[0131] Based on the same inventive concept, an embodiment of the present application provides a system for multi-sensor feature fusion, such as Figure 3 As shown, the system 300 includes:

[0132] A preprocessing module 301 is used to preprocess the collected image data and point cloud data to obtain target image data and target point cloud data;

[0133] A first feature extraction module 302, configured to extract features from the target image data through an image feature extraction network to obtain a corresponding multi-scale image feature map;

[0134] The second feature extraction module 303 is used to perform spatial layout and geometric feature extraction on the target point cloud data through a sparse 3D convolutional network to obtain a corresponding point cloud feature map, in which the spatial layout and geometric features of the point cloud are recorded;

[0135] A feature enhancement processing module 304 is used to enhance the features of the key areas in the point cloud feature map through an attention mechanism to obtain a target point cloud feature map;

[0136] A projection module 305 is used to project the multi-scale image feature map from the 2D image space to the BEV perspective through perspective transformation based on the calibrated camera internal and external parameters;

[0137] A perspective conversion module 306 is used to convert the target point cloud feature map to a BEV perspective based on calibration parameters between the camera and the laser radar;

[0138] An alignment processing module 307 is used to align the multi-scale image feature map under the BEV perspective with the target point cloud feature map;

[0139] The feature fusion module 308 is used to perform feature fusion on the aligned multi-scale image feature map and the target point cloud feature map through a hybrid attention mechanism to obtain a fused feature map.

[0140] Optionally, the system 300 further includes:

[0141] A parallel processing module, used for inputting the fused feature map into a multi-task model to perform parallel multi-task processing;

[0142] A query module, used to query the feature information of interest to the task at a corresponding position in the fused feature map through a query with a position code corresponding to the task;

[0143] The execution module is used to input each feature information of interest obtained by the query into the task header of each corresponding task for processing, so as to obtain the output result of each task.

[0144] Optionally, the feature fusion module 308 includes:

[0145] A query generation module, used to generate a set number of target queries according to the position information of the aligned multi-scale image feature map and the target point cloud feature map, wherein the target query includes a position code representing itself in space;

[0146] The attention weight determination module is used to calculate the attention weights between target queries through a multi-head self-attention mechanism;

[0147] A query determination module, used for performing weighted summation on the target query according to the attention weight to obtain a modified target query;

[0148] The calculation module is used to perform cross-attention calculation on the corrected target Query, the aligned multi-scale image feature map and the target point cloud feature map through a cross-attention mechanism to obtain a fused feature map after weighted fusion.

[0149] Optionally, the system 300 further includes:

[0150] A first Query determination module is used to determine that the modified target Query after performing the cross-attention calculation is a fused target Query, wherein the fused target Query includes a target Query that fuses information of a multi-scale image feature map and a target point cloud feature map;

[0151] The fusion feature map determination module is used to perform multiple rounds of cross-attention calculations based on the fusion target Query obtained in each round through a multi-head self-attention mechanism and a cross-attention mechanism to obtain a fusion feature map after multiple rounds of weighted fusion.

[0152] Optionally, the preprocessing module 301 includes:

[0153] A denoising module, used to perform denoising processing on the collected image data through Gaussian filtering and median filtering to obtain first image data;

[0154] an adjustment module, configured to perform color adjustment on the first image data by using an automatic white balance algorithm to obtain second image data;

[0155] A correction module, used for performing consistency correction on the second image data through color space conversion and color histogram matching to obtain third image data;

[0156] an enhancement module, configured to perform contrast enhancement processing on the third image data through histogram equalization to obtain fourth image data;

[0157] A mapping module, used for mapping the grayscale range of the fourth image data to a target grayscale range by preset stretching parameters;

[0158] An optimization module, used for performing detailed optimization on the fourth image data after grayscale range mapping by contrast gain and brightness adjustment to obtain optimized image data;

[0159] The irrelevant background in the optimized image data is cropped at the original ratio of the optimized image data, and the cropped optimized image data is scaled to the size required for the model input to obtain the target image data, wherein the original ratio is proportional to the size required for the model input.

[0160] Optionally, the preprocessing module 301 includes:

[0161] A first removal module is used to remove voxel noise points from the collected point cloud data according to a preset voxel size and a neighborhood point number threshold to obtain first point cloud data;

[0162] A second removal module, configured to perform statistical noise point removal on the first point cloud data according to a preset distance threshold and a point number threshold, to obtain second point cloud data;

[0163] A voxel size determination module, used to determine a corresponding target voxel size according to the point cloud density and resolution of the second point cloud data;

[0164] A voxelization module, configured to voxelize the second point cloud data according to the target voxel size to obtain voxelized point cloud data, wherein each voxel grid records the number, average position, and density feature information of the point clouds in the voxel grid;

[0165] The conversion module is used to convert the voxelized point cloud data into the same coordinate system as the target image data based on the calibrated internal and external parameters of the camera and the laser radar parameters, so as to obtain the target point cloud data.

[0166] Optional, computing module, including:

[0167] A first fusion module is used to fuse the target detection results of the historical frame with the multi-scale image feature map and the target point cloud feature map after alignment processing of the current frame through a hybrid attention mechanism when the output results of each task include the target detection result, so as to obtain the fused multi-scale image feature map and the target point cloud feature map;

[0168] The compensation module is used to perform motion compensation on the target query of the historical frame through a preset algorithm to predict the position of each target in the target detection result in the current frame;

[0169] An alignment and fusion module is used to align and fuse the target query of the motion compensated historical frame with the corrected target query of the current frame to obtain the aligned target query of the current frame;

[0170] The calculation submodule is used to perform cross-attention calculation on the obtained alignment target Query and the fused multi-scale image feature map and target point cloud feature map through a cross-attention mechanism to obtain a weighted fused fusion feature map.

[0171] Optionally, the system 300 further includes:

[0172] An offset noise removal module is used to perform offset noise denoising on the target point cloud feature map under the BEV perspective through a trained qualified denoising convolutional network to obtain a denoised target point cloud feature map;

[0173] The alignment processing module 307 is used to align the multi-scale image feature map under the BEV perspective and the denoised target point cloud feature map.

[0174] Optionally, the system 300 further includes a model training module, which is used to train and obtain a qualified denoising convolutional network;

[0175] The model training module includes:

[0176] A data set construction module is used to construct a training data set, wherein the training data set includes a multi-scale image feature map and a target point cloud feature map under the BEV perspective that have been aligned and have a deviation below a set threshold, and a standard fusion feature map that is a fusion of the multi-scale image feature map and the target point cloud feature map;

[0177] A noise generation module, used for generating a random offset noise feature, and performing a convolution operation on the random offset noise feature to obtain a target random offset noise feature;

[0178] A sampling module is used to sample the target point cloud feature map grid under the BEV perspective through the target random offset noise feature to obtain the noise point cloud feature map;

[0179] A feature extraction module, used to extract features from the noise point cloud feature map through a denoising convolutional network to obtain a point cloud feature map;

[0180] A feature fusion module to be evaluated is used to perform feature fusion on the point cloud feature map and a multi-scale image feature map corresponding to the point cloud feature map in a training data set to obtain a fused feature map to be evaluated;

[0181] An evaluation module, used to evaluate the fused feature map to be evaluated and the corresponding standard fused feature map through a loss function to determine the denoising effect of the denoising convolutional network;

[0182] A first evaluation result determination module is used to determine that the denoising convolution network training is qualified when the denoising effect meets the set conditions;

[0183] The second evaluation result determination module is used to update the parameters of the denoising convolutional network and re-train the model when the denoising effect does not meet the set conditions until the denoising effect obtained meets the set conditions.

[0184] Based on the same inventive concept, an embodiment of the present application provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and running on the processor. When the computer program is executed by the processor, the steps in the method for multi-sensor feature fusion as described in the first aspect of the present application are implemented.

[0185] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for multi-sensor feature fusion as described in the first aspect of the present application are implemented.

[0186] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0187] It should be noted that, for the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.

[0188] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0189] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the embodiments of the present application may adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the embodiments of the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0190] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0191] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0192] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable terminal device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0193] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0194] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0195] The above is a detailed introduction to the method, system and product of multi-sensor feature fusion provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for multi-sensor feature fusion, characterized in that: The method comprises: Preprocess the collected image data and point cloud data respectively to obtain target image data and target point cloud data; Performing feature extraction on the target image data through an image feature extraction network to obtain a corresponding multi-scale image feature map; Performing spatial layout and geometric feature extraction on the target point cloud data through a sparse 3D convolutional network to obtain a corresponding point cloud feature map, wherein the point cloud feature map records the spatial layout and geometric features of the point cloud; The features of the key areas in the point cloud feature map are enhanced by an attention mechanism to obtain a target point cloud feature map; Based on the calibrated camera internal and external parameters, the multi-scale image feature map is projected from the 2D image space to the BEV perspective through perspective transformation; Based on the calibration parameters between the camera and the laser radar, the target point cloud feature map is converted to the BEV perspective; Align the multi-scale image feature map from the BEV perspective with the target point cloud feature map; The aligned multi-scale image feature map and the target point cloud feature map are fused through the hybrid attention mechanism to obtain a fused feature map.

2. The method for multi-sensor feature fusion according to claim 1, characterized in that: The method further comprises: Inputting the fused feature map into a multi-task model for parallel multi-task processing; Through a query with position encoding corresponding to the task, query the feature information of interest to the task at the corresponding position in the fused feature map; The feature information of interest obtained from the query is input into the task header of each corresponding task for processing to obtain the output result of each task.

3. The method for multi-sensor feature fusion according to claim 2, characterized in that: The feature fusion of the aligned multi-scale image feature map and the target point cloud feature map is performed through the hybrid attention mechanism to obtain a fused feature map, including: Generate a set number of target queries according to the position information of the aligned multi-scale image feature map and the target point cloud feature map, wherein the target query includes a position code representing itself in space; Calculate the attention weights between target queries through a multi-head self-attention mechanism; Performing weighted summation on the target Query according to the attention weight to obtain a modified target Query; Through the cross-attention mechanism, the corrected target Query is cross-attended with the aligned multi-scale image feature map and the target point cloud feature map to obtain a weighted fused fusion feature map.

4. The method for multi-sensor feature fusion according to claim 3, characterized in that: After performing cross-attention calculation on the modified target Query, the aligned multi-scale image feature map and the target point cloud feature map through a cross-attention mechanism to obtain a weighted fused fusion feature map, the method further includes: Determine that the modified target Query after performing the cross attention calculation is a fused target Query, wherein the fused target Query includes a target Query that fuses information of a multi-scale image feature map and a target point cloud feature map; Based on the fusion target Query obtained in each round, multiple rounds of cross-attention calculations are performed through the multi-head self-attention mechanism and the cross-attention mechanism to obtain the fusion feature map after multiple rounds of weighted fusion.

5. The method for multi-sensor feature fusion according to claim 1, characterized in that: Preprocess the collected image data to obtain target image data, including: Performing denoising processing on the collected image data by Gaussian filtering and median filtering to obtain first image data; Performing color adjustment on the first image data by using an automatic white balance algorithm to obtain second image data; Performing consistency correction on the second image data by color space conversion and color histogram matching to obtain third image data; Performing contrast enhancement processing on the third image data by histogram equalization to obtain fourth image data; Mapping the grayscale range of the fourth image data to a target grayscale range by preset stretching parameters; The fourth image data after grayscale range mapping is carefully optimized by contrast gain and brightness adjustment to obtain optimized image data; The irrelevant background in the optimized image data is cropped at the original ratio of the optimized image data, and the cropped optimized image data is scaled to the size required for the model input to obtain the target image data, wherein the original ratio is proportional to the size required for the model input.

6. The method for multi-sensor feature fusion according to claim 1, characterized in that: Preprocess the collected point cloud data to obtain target point cloud data, including: According to a preset voxel size and a neighborhood point number threshold, voxel noise points are removed from the collected point cloud data to obtain first point cloud data; According to a preset distance threshold and a point number threshold, statistical noise points are removed from the first point cloud data to obtain second point cloud data; Determine a corresponding target voxel size according to the point cloud density and resolution of the second point cloud data; voxelize the second point cloud data according to the target voxel size to obtain voxelized point cloud data, wherein each voxel grid records the number, average position, and density feature information of the point clouds in the voxel grid; Based on the calibrated camera internal and external parameters and lidar parameters, the voxelized point cloud data is converted into the same coordinate system as the target image data to obtain the target point cloud data.

7. The method for multi-sensor feature fusion according to claim 3, characterized in that: In the case where the output results of each task include the target detection result, the cross-attention mechanism is used to perform cross-attention calculation on the modified target Query, the aligned multi-scale image feature map and the target point cloud feature map to obtain a fused feature map after weighted fusion, including: The target detection results of the historical frame are fused with the multi-scale image feature map and the target point cloud feature map after alignment processing of the current frame through the hybrid attention mechanism to obtain the fused multi-scale image feature map and the target point cloud feature map; Motion compensation is performed on the target query of the historical frame through a preset algorithm to predict the position of each target in the target detection result in the current frame; Align and fuse the target query of the motion compensated historical frame with the corrected target query of the current frame to obtain the aligned target query of the current frame; Through the cross-attention mechanism, the obtained alignment target Query is cross-attended with the fused multi-scale image feature map and the target point cloud feature map to obtain the weighted fused fusion feature map.

8. The method for multi-sensor feature fusion according to claim 1, characterized in that: Before aligning the multi-scale image feature map under the BEV perspective and the target point cloud feature map, the method further includes: The target point cloud feature map under the BEV perspective is subjected to offset noise denoising by means of a qualified denoising convolutional network, and a denoised target point cloud feature map is obtained; The aligning process of the multi-scale image feature map under the BEV perspective and the target point cloud feature map comprises: aligning the multi-scale image feature map under the BEV perspective and the denoised target point cloud feature map.

9. The method for multi-sensor feature fusion according to claim 8, characterized in that: The training process obtains a qualified denoising convolutional network, including: Constructing a training data set, wherein the training data set includes a multi-scale image feature map and a target point cloud feature map under a BEV perspective that have been aligned and have a deviation below a set threshold, and a standard fusion feature map that is a fusion of the multi-scale image feature map and the target point cloud feature map; Generate a random offset noise feature, and perform a convolution operation on the random offset noise feature to obtain a target random offset noise feature; The target point cloud feature map grid is sampled under the BEV perspective through the target random offset noise feature to obtain the noise point cloud feature map; Extracting features of the noise point cloud feature map through a denoising convolutional network to obtain a point cloud feature map; Performing feature fusion on the point cloud feature map and the multi-scale image feature map corresponding to the point cloud feature map in the training data set to obtain a fused feature map to be evaluated; Evaluate the fused feature map to be evaluated and the corresponding standard fused feature map through a loss function to determine the denoising effect of the denoising convolutional network; When the denoising effect satisfies the set conditions, determining that the denoising convolution network training is qualified; When the denoising effect does not meet the set conditions, the parameters of the denoising convolutional network are updated and the model training is performed again until the denoising effect obtained meets the set conditions.

10. A multi-sensor feature fusion system, characterized in that: The system comprises: A preprocessing module is used to preprocess the collected image data and point cloud data respectively to obtain target image data and target point cloud data; A first feature extraction module, used to extract features from the target image data through an image feature extraction network to obtain a corresponding multi-scale image feature map; A second feature extraction module is used to perform spatial layout and geometric feature extraction on the target point cloud data through a sparse 3D convolutional network to obtain a corresponding point cloud feature map, in which the spatial layout and geometric features of the point cloud are recorded; A feature enhancement processing module, used for enhancing the features of the key areas in the point cloud feature map through an attention mechanism to obtain a target point cloud feature map; A projection module, used for projecting the multi-scale image feature map from the 2D image space to the BEV perspective through perspective transformation based on the calibrated camera internal and external parameters; A perspective conversion module, used for converting the target point cloud feature map to a BEV perspective based on calibration parameters between the camera and the laser radar; An alignment processing module is used to align the multi-scale image feature map and the target point cloud feature map under the BEV perspective; The feature fusion module is used to fuse the aligned multi-scale image feature map and the target point cloud feature map through a hybrid attention mechanism to obtain a fused feature map.

11. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and running on the processor, wherein when the computer program is executed by the processor, the steps in the method for multi-sensor feature fusion as described in claims 1 to 9 are implemented.

12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps in the method for multi-sensor feature fusion as claimed in claims 1 to 9 are implemented.

Citation Information

Cited By

  • Multi-view target detection method and device, multi-view target detection model training method and device, equipment, storage medium and program product

    CN120411485A

  • Traffic target tracking method and system, electronic equipment and storage medium

    CN121259048A

  • A fisheye image and point cloud multi-modal fusion method, system, device and medium

    CN122820462A