Efficient occupied grid prediction method based on multi-sensor fusion

Through multi-sensor fusion technology, the characteristics of cameras and lidar are dynamically integrated, and through optimized computing strategies, the limitations of a single sensor perception method in complex environments are solved, achieving efficient and robust autonomous driving perception capabilities.

CN119919905APending Publication Date: 2025-05-02BEIJING TIEMUNIU INTELLIGENT MACHINERY TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411998756.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing single sensor perception method has limitations when dealing with complex environments, cannot identify distinct obstacles, and is inefficient in computing, making it difficult to meet the needs of autonomous vehicles under various environmental conditions.

Method used

Using an efficient occupancy grid prediction method based on multi-sensor fusion, the multi-modal features of cameras and lidar are integrated to achieve effective fusion of features, and the computational efficiency is optimized through fast ray transformation and channel height conversion.

Benefits of technology

It improves the perceived accuracy and robustness of autonomous vehicles under various environmental conditions, enhances the operation efficiency and real-time nature of the algorithm, and is suitable for complex and changeable traffic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919905A_ABST
    Figure CN119919905A_ABST
Patent Text Reader

Abstract

The invention provides an efficient occupied grid prediction method based on multi-sensor fusion. The efficient occupied grid prediction method comprises the steps that vehicle surrounding image data and laser point cloud information data are collected; respectively extracting image features and point cloud features from the image data and the point cloud information data by using a convolutional network; performing view conversion by adopting fast ray transformation, and converting the image features into image BEV features; performing dynamic feature fusion on the image BEV features and the point cloud features to obtain a fusion feature map; and converting the fused feature map into an occupied grid feature map through channel height conversion, and performing semantic category prediction on single voxels in the space based on the occupied grid feature map. According to the efficient occupied grid prediction method based on multi-sensor fusion, multi-modal features from a camera and a laser radar can be integrated through dynamic fusion of a 3D / 2D module, and effective fusion of the features is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and in particular relates to an efficient occupancy grid prediction method based on multi-sensor fusion. Background Art

[0002] With the development of autonomous driving technology, the vehicle's ability to perceive the surrounding environment has become particularly important. Traditional 3D object detection methods mainly rely on single sensor data, such as cameras or lidar, which have limitations when dealing with complex environments and cannot identify irregular obstacles. For example, the performance of cameras degrades at night or in bad weather conditions, and although lidar can provide accurate distance measurements, it also faces challenges when dealing with dynamic objects and complex scenes. Therefore, the perception capabilities of a single sensor are limited and it is difficult to meet the needs of autonomous vehicles in various environmental conditions.

[0003] In order to overcome the limitations of a single sensor, multi-sensor fusion technology can improve the accuracy and robustness of perception by integrating data from different sensors. Occupancy grid prediction technology provides a more sophisticated environment representation method for autonomous vehicles by predicting the semantic category of each voxel in 3D space. However, existing occupancy grid prediction methods face many challenges when dealing with multi-sensor data fusion, such as feature fusion and computational efficiency. Direct fusion of raw data may lead to information redundancy and conflict. It is necessary to design an effective feature fusion strategy to extract complementary information and suppress noise. Multi-sensor fusion algorithms usually involve a lot of data processing and calculation, resulting in large computational delays and resource consumption. In real-time autonomous driving systems, computational efficiency is crucial, so the algorithm needs to be optimized to increase the running speed and reduce resource consumption. Summary of the invention

[0004] The purpose of the present invention is to solve the problems existing in the prior art and propose an efficient occupancy grid prediction method based on multi-sensor fusion, which can make full use of the advantages of different sensors and improve the perception accuracy and robustness of autonomous driving vehicles under various environmental conditions. By dynamically fusing 3D / 2D modules, the algorithm can integrate multimodal features from cameras and lidars to achieve effective fusion of features. In addition, the algorithm is designed with computational efficiency in mind, ensuring superior performance in different perception scenarios.

[0005] In order to achieve the above object, the present invention adopts the following technical scheme.

[0006] The efficient occupancy grid prediction method based on multi-sensor fusion,

[0007] Collect image data and point cloud information data around the vehicle;

[0008] Extracting image features and point cloud features from the image data and the point cloud information data respectively using a convolutional network;

[0009] Fast ray transformation is used for view transformation to convert image features into image BEV features;

[0010] Performing dynamic feature fusion on the image BEV feature and the point cloud feature to obtain a fused feature map;

[0011] The fused feature map is converted into an occupancy grid feature map through channel height conversion, and semantic category prediction is performed on a single voxel in the space based on the occupancy grid feature map.

[0012] Furthermore, the acquisition of vehicle surrounding image data and point cloud information data includes:

[0013] The vehicle surrounding image data includes image data of an all-round field of view around the vehicle;

[0014] The point cloud information data includes 3D point cloud data generated by collecting vehicle sensors, and the 3D point cloud data includes precise distance and shape information of the vehicle's surrounding environment.

[0015] Furthermore, the extracting image features and point cloud features from the image data and the point cloud information data respectively using a convolutional network comprises:

[0016] The captured image data is processed using a 2D convolutional network to extract color, texture, and shape information:

[0017] F 2D =CNN 2D ( I)

[0018] Among them, I is the input image data, F 2D is the extracted 2D image feature map;

[0019] Use 3D convolutional network to process point cloud information data and extract spatial position and geometric structure information:

[0020] F 3D =CNN 3D ( P)

[0021] Among them, P is the input environmental information data, F 3D is the extracted point cloud feature map.

[0022] Furthermore, the use of fast ray transformation to perform view conversion and convert image features into image BEV features includes:

[0023] F′ camera_BEV =T 2D→BEV( F 2D )

[0024] Among them, T 2D→BEV Represents a view transition operation.

[0025] Furthermore, it also includes:

[0026] After obtaining the intrinsic parameters of a single camera and the extrinsic parameters of the camera-to-vehicle coordinate system through calibration, the projection relationship between the image feature space and the BEV space is pre-calculated and the projection index is recorded;

[0027] When performing view transformation, for a single feature in the BEV space, a projection relationship is constructed between it and a single camera projection index and the corresponding camera image feature two-dimensional coordinate system (x, y), and all image features are projected into the BEV space.

[0028] Furthermore, the dynamically fusing the image BEV feature and the point cloud feature to obtain a fused feature map includes:

[0029] The feature fusion is performed by directly splicing the image BEV features and point cloud features:

[0030] F fusion =Concat ( F Lidar ,F Camera_BEV )

[0031] Where Concat represents the feature concatenation operation;

[0032] The image BEV features and point cloud features are used as the input of dynamic feature fusion, and the spliced ​​feature map is extracted through the convolutional network to obtain the fused feature map F. fusion :

[0033] F fusion =Conv2d(Concat(F Lidar ,F Camera_BEV )).

[0034] Furthermore, it also includes:

[0035] Perform global average pooling on the features of a single channel and use a fully connected layer to extract features to obtain attention weight information of different channels;

[0036] The fused feature map is adjusted according to the attention weight information to refine the feature representation:

[0037] F fusion =α·F fusion

[0038] Where α is the attention weight information.

[0039] Furthermore, converting the fused feature map into an occupied grid feature map by channel height conversion, and predicting a semantic category of a single voxel in the space based on the occupied grid feature map includes:

[0040] The fused feature map F obtained by fusion of dynamic features fusion , through channel height conversion, the fusion feature map F fusion Convert to an occupied grid feature map, and make semantic category predictions for individual voxels in 3D space based on the occupied grid feature map:

[0041] S=Segmentation ( F fusion )

[0042] Among them, S is the segmentation result, which represents the semantic category of a single voxel.

[0043] Calculate the occupancy probability of a voxel:

[0044] O = sigmoid(W·F fusion +b)

[0045] Among them, W and b are model parameters, and the sigmoid function is used to limit the output between 0 and 1 to represent the probability of occupancy.

[0046] In order to achieve the above-mentioned purpose, the present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a program running on the processor, and when the processor runs the program, the steps of the efficient occupancy grid prediction method based on multi-sensor fusion as described above are executed.

[0047] In order to achieve the above object, the present invention also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed, execute the steps of the efficient occupancy grid prediction method based on multi-sensor fusion as described above.

[0048] The present invention proposes an efficient occupancy grid prediction method based on multi-sensor fusion, which has the following beneficial effects:

[0049] The present invention proposes a dynamic multi-sensor fusion strategy, which improves the accuracy of 3D semantic occupancy prediction by improving the effect of feature fusion; multi-sensor input makes the algorithm more adaptable and robust to environmental changes (such as rainy and foggy weather and night lighting conditions); and the efficient occupancy grid prediction design through channel height conversion avoids complex 3D convolution calculations, thereby improving the operating efficiency and real-time performance of the algorithm.

[0050] Other features and advantages of the present invention will be set forth in the description which follows, and in part will be apparent from the description, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0052] Figure 1 The present invention is a flowchart of an efficient occupancy grid prediction method based on multi-sensor fusion. DETAILED DESCRIPTION

[0053] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0054] Example 1

[0055] Figure 1 The flowchart of the efficient occupancy grid prediction method based on multi-sensor fusion according to the present invention is as follows. Figure 1 , the efficient occupancy grid prediction method based on multi-sensor fusion of the present invention is described in detail.

[0056] In step 101, image data and environmental information around the vehicle are collected.

[0057] Optionally, image data is acquired from multiple cameras around the vehicle, which covers a 360° all-round view around the vehicle. At the same time, 3D point cloud data generated by the lidar sensor is collected, which provides accurate distance and shape information of the surrounding environment.

[0058] In step 102, 2D and 3D convolutional networks are used to extract image features and point cloud features from image and point cloud data respectively.

[0059] Optionally, a 2D convolutional network processes images captured by the camera, extracting color, texture, and shape information.

[0060] F 2D =CNN 2D ( I)

[0061] Where I is the input image, F 2D is the extracted 2D feature map.

[0062] Optionally, a 3D convolutional network processes point cloud data generated by lidar and radar to extract spatial position and geometric structure information.

[0063] F3D =CNN 3D ( P)

[0064] Among them, P is the input point cloud data, F 3D is the extracted feature map.

[0065] In step 103, fast ray transform is used to perform view conversion, and the 2D image features are converted into BEV (Bird's Eye View) features.

[0066] Optionally, convert 2D image features to BEV features:

[0067] F′ camera_BEV =T 2D→BEV ( F 2D )

[0068] Among them, T 2D→BEV Representing a view conversion operation, the present invention uses fast ray transformation to perform view conversion.

[0069] In this embodiment, the projection of image features from image space to BEV space will take up a lot of time. To reduce the computational complexity, this embodiment obtains the intrinsic parameters of each camera and the extrinsic parameters of the camera-to-vehicle coordinate system based on calibration, and then pre-calculates the projection relationship between the image feature space and the BEV space, and records it in a fixed projection index to avoid recalculating the index every time the inference is performed. When performing view transformation, for each feature in the three-dimensional BEV space, a projection relationship is directly constructed between it and each camera index and the corresponding camera image feature two-dimensional coordinate system (x, y), and all camera image features are projected into the same dense BEV space at one time, thus avoiding complex aggregation operations.

[0070] In step 104, the BEV features extracted from the camera and transformed by the view are dynamically fused with the point cloud features extracted from the point cloud to obtain a fused feature map.

[0071] Feature fusion is performed by directly splicing the camera image feature map and the lidar point cloud 3D features:

[0072] F fusion =Concat ( F Lidar ,F Camera_BEV )

[0073] Concat represents the feature concatenation operation.

[0074] In this embodiment, a dynamic feature fusion method is used to improve the feature fusion efficiency and feature expression ability. The aforementioned Lidar BEV image feature map and Camera BEV point cloud extraction feature map are used as the input of dynamic feature fusion and spliced. The spliced ​​feature map is sent to a 3x3 convolution for preliminary feature extraction to obtain a fused feature map F fusion :

[0075] F fusion =Conv2d(Concat(F Lidar ,F Camera_BEV ))

[0076] The features of each channel are globally averaged pooled, and the features are further extracted using a fully connected layer to obtain the attention weight information of different channels.

[0077] The fused feature map is adjusted according to the attention weight to refine the feature representation.

[0078] F fusion =α·F fusion

[0079] Where α is the attention weight information.

[0080] In step 105, the fused feature map is converted into an occupancy grid feature map by channel height conversion, and an occupancy grid prediction is performed for each voxel in the 3D space based on the occupancy grid feature map.

[0081] In this embodiment, in order to improve the computational efficiency, the occupancy grid prediction adopts an optimized computational strategy. By reducing redundant computations, the algorithm can reduce computational delays and resource consumption while maintaining high accuracy. This makes the algorithm suitable for real-time autonomous driving systems and meets real-time requirements.

[0082] Optionally, the feature map F obtained by fusion of the above dynamic features fusion , through channel height conversion, the 2D F fusion Convert to 3D fusion feature F 3D ,Based on the 3D fusion features, the algorithm of this embodiment predicts the semantic category of each voxel in the 3D space.

[0083] S=Segmentation ( F fusion )

[0084] Among them, S is the segmentation result, which represents the semantic category of each voxel.

[0085] Calculate the occupancy probability of a voxel:

[0086] O = sigmoid(W·F fusion+b)

[0087] Among them, W and b are model parameters, and the sigmoid function is used to limit the output between 0 and 1 to represent the probability of occupancy.

[0088] The present invention proposes an efficient occupancy grid prediction method based on multi-sensor fusion, which realizes feature extraction through image encoder and point cloud encoder respectively for multi-view camera images and lidar point cloud inputs, and proposes dynamic feature fusion to efficiently fuse the features obtained by image and point cloud encoders under the bird's-eye view, thereby improving the fusion accuracy. Through the efficient occupancy grid prediction design of channel height conversion, rapid detection and recognition of objects in different autonomous driving scenes can be achieved. It can adapt to complex and changeable traffic scenes, and enhance the robustness and stability of the autonomous driving system.

[0089] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a program running on the processor, and the processor executes the steps of the above-mentioned efficient occupancy grid prediction method based on multi-sensor fusion when running the program.

[0090] The present invention also provides a computer-readable storage medium having computer instructions stored thereon. When the computer instructions are executed, the steps of the above-mentioned efficient occupancy grid prediction method based on multi-sensor fusion are executed. The efficient occupancy grid prediction method based on multi-sensor fusion is described in the introduction of the previous part and will not be repeated here.

[0091] Those skilled in the art can understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions recorded in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An efficient occupancy grid prediction method based on multi-sensor fusion, characterized in that: include: Collect image data and point cloud information data around the vehicle; Extracting image features and point cloud features from the image data and the point cloud information data respectively using a convolutional network; Fast ray transformation is used for view transformation to convert image features into image BEV features; Performing dynamic feature fusion on the image BEV feature and the point cloud feature to obtain a fused feature map; The fused feature map is converted into an occupancy grid feature map through channel height conversion, and semantic category prediction is performed on a single voxel in the space based on the occupancy grid feature map.

2. The efficient occupancy grid prediction method based on multi-sensor fusion according to claim 1, characterized in that: The acquisition of vehicle surrounding image data and point cloud information data includes: The vehicle surrounding image data includes image data of an all-round field of view around the vehicle; The point cloud information data includes 3D point cloud data generated by collecting vehicle sensors, and the 3D point cloud data includes precise distance and shape information of the vehicle's surrounding environment.

3. The efficient occupancy grid prediction method based on multi-sensor fusion according to claim 1, characterized in that: The extracting of image features and point cloud features from the image data and the point cloud information data using a convolutional network respectively comprises: The captured image data is processed using a 2D convolutional network to extract color, texture, and shape information: F 2D =CNN 2D (I) Among them, I is the input image data, F 2D is the extracted 2D image feature map; Use 3D convolutional network to process point cloud information data and extract spatial position and geometric structure information: F 3D =CNN 3D (P) Among them, P is the input environmental information data, F 3D is the extracted point cloud feature map.

4. The efficient occupancy grid prediction method based on multi-sensor fusion according to claim 1, characterized in that: The method of using fast ray transformation to perform view conversion and converting image features into image BEV features includes: F′ camera_BEV =T 2D→BEV (F 2D ) Among them, T 2D→BEV Represents a view transition operation.

5. The efficient occupancy grid prediction method based on multi-sensor fusion according to claim 4, characterized in that: Also includes, After obtaining the intrinsic parameters of a single camera and the extrinsic parameters of the camera-to-vehicle coordinate system through calibration, the projection relationship between the image feature space and the BEV space is pre-calculated and the projection index is recorded; When performing view transformation, for a single feature in the BEV space, a projection relationship is constructed between it and a single camera projection index and the corresponding camera image feature two-dimensional coordinate system (x, y), and all image features are projected into the BEV space.

6. The efficient occupancy grid prediction method based on multi-sensor fusion according to claim 1, characterized in that: The dynamically fusing the image BEV feature and the point cloud feature to obtain a fused feature map comprises: The feature fusion is performed by directly splicing the image BEV features and point cloud features: F fusion =Concat(F Lidar ,F Camera_BEV ) Where Concat represents the feature concatenation operation; The image BEV features and point cloud features are used as the input of dynamic feature fusion, and the spliced ​​feature map is extracted through the convolutional network to obtain the fused feature map F. fusion : F fusion =Conv2d(Concat(F Lidar ,F Camera_BEV ))。 7. The efficient occupancy grid prediction method based on multi-sensor fusion according to claim 6, characterized in that: Also includes, Perform global average pooling on the features of a single channel and use a fully connected layer to extract features to obtain attention weight information of different channels; The fused feature map is adjusted according to the attention weight information to refine the feature representation: F fusion =α·F fusion Where α is the attention weight information.

8. The efficient occupancy grid prediction method based on multi-sensor fusion according to claim 1, characterized in that: The step of converting the fused feature map into an occupied grid feature map by channel height conversion, and predicting a semantic category of a single voxel in the space based on the occupied grid feature map comprises: The fused feature map F obtained by fusion of dynamic features fusion , through channel height conversion, the fusion feature map F fusion Convert to an occupied grid feature map, and make semantic category predictions for individual voxels in 3D space based on the occupied grid feature map: S=Segmentation(F fusion ) Among them, S is the segmentation result, which represents the semantic category of a single voxel. Calculate the occupancy probability of a voxel: O=sigmoid(W·F fusion +b) Among them, W and b are model parameters, and the sigmoid function is used to limit the output between 0 and 1 to represent the probability of occupancy.

9. An electronic device, characterized in that: It comprises a memory and a processor, wherein the memory stores a program running on the processor, and the processor executes an efficient occupancy grid prediction method based on multi-sensor fusion as described in any one of claims 1 to 8 when running the program.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed, an efficient occupancy grid prediction method based on multi-sensor fusion as described in any one of claims 1 to 8 is executed.

Citation Information

Cited By

  • Occupied grid map generation method fusing fish eye passable space detection

    CN120333476A

  • Dynamic self-adaptive grid occupying method for special scene

    CN121214401A