Vehicle sensing method and system based on multi-modal fusion

By fusing camera and lidar data in the vehicle perception system and extracting and fusion features using specific networks, the problems of low target object recognition accuracy and failure of field of view recognition in existing systems are solved, and target object recognition and environmental perception with higher accuracy are achieved.

CN120164185APending Publication Date: 2025-06-17CHERY COMMERCIAL VEHICLE (BOZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510196098.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When the existing vehicle perception system integrates camera sensors and lidar data, the detection target is low and susceptible to light and environmental factors, resulting in failure of obstacle identification in the blind spot of the field of view.

Method used

The image data is obtained by the camera and the point cloud data is obtained by the lidar, and the multimodal fusion is performed. The point cloud and image features in the BEV space are extracted by the SECOND network and the MoA-Transformer feature extraction network. After the fusion is performed, the 3D detection head is input to the target object recognition.

Benefits of technology

It improves the recognition accuracy of target objects in front of driving, enhances the vehicle's perception of the environment, especially target objects in blind spots in the field of view.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164185A_ABST
    Figure CN120164185A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle perception, and provides a vehicle perception method and system based on multi-modal fusion, and the method comprises the steps: (1) obtaining image data in front of a driving vehicle through a camera, and scanning point cloud data in front of the driving vehicle and point cloud data in sight blind areas at two sides of the driving vehicle through a laser radar; and (2) fusing the point cloud data in front of the vehicle with the image data, identifying a target object in a corresponding area based on the fused spatial features, and identifying the target object in the corresponding area based on the point cloud data of the sight blind areas on the two sides in front. According to the method, the image features of different sizes are extracted from the image data of the driving front, the image features of different sizes are spliced and then projected to the BEV space, the spliced image features are fused with the point cloud features in the BEV space, and the target object is identified based on the obtained fusion features in the BEV space, so that the identification precision of the target object of the driving front is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle perception, and provides a vehicle perception method and system based on multimodal fusion. Background Art

[0002] With the advancement of smart cities, the implementation of autonomous driving in various industries has accelerated. For example, unmanned sanitation vehicles will improve the operation methods of cities, reduce the safety risks of sanitation workers, and also reduce the operating costs and labor costs of sanitation companies, improve the operation efficiency of vehicles, and promote the further development of the industry.

[0003] Camera sensors can provide texture information and context information by obtaining image information around the vehicle, but they cannot provide sufficient geometric and physical information of the detection target, and cameras are easily affected by factors such as light and environment. LiDAR sensors are used to obtain the depth information of the target, but lack the semantic information of the target. Therefore, the sparsity of the point cloud will limit its performance, and due to the characteristics of the working principle of the sensor, heavy snow, fog, dust, etc. will be recognized as obstacles, resulting in detection failure.

[0004] The existing sensor solution directly converts the point cloud data and image data detected by LiDAR into the BEV space for fusion, and the integrated features obtained contain less feature information, thereby reducing the detection accuracy of target objects. Summary of the Invention

[0005] The present invention provides a vehicle perception method and system based on multimodal fusion, aiming to improve the above problems.

[0006] The present invention is implemented as follows. A vehicle perception method based on multimodal fusion is as follows:

[0007] (1) Obtain image data in front of the vehicle through a camera, and scan the point cloud data in front of the vehicle and the blind spots on both sides of the front line of sight of the vehicle through a LiDAR.

[0008] (2) Fuse the point cloud data and image data in front of the vehicle, identify the target objects in the corresponding area based on the fused spatial features, and identify the target objects in the corresponding area based on the point cloud data in the blind spots on both sides of the front line of sight.

[0009] Further, the process of identifying the target objects in front of the vehicle is as follows:

[0010] Extract the point cloud features from the point cloud data in front of the vehicle, extract the image features from the image data in front of the vehicle, fuse the extracted point cloud features and image features, and input the fused spatial features into a 3D detection head. The 3D detection head extracts the target objects based on the fused features.

[0011] Further, the fusion process of the point cloud features and the image features is as follows:

[0012] (21) Input the point cloud data in front of the vehicle into the SECOND network, and extract the point cloud features of the point cloud data in the BEV space through the SECOND network;

[0013] (22) Input the image data in front of the vehicle into the MoA-Transformer feature extraction network, extract the image features of different sizes in the image data, splice the image features of different sizes based on the ASFF algorithm, and project the spliced image features into the BEV space based on the EA-LSS algorithm to obtain the image features in the BEV space;

[0014] (23) Fuse the point cloud features in the BEV space and the image features in the BEV space to obtain the fused spatial features.

[0015] Further, the process of extracting the target objects in the blind spots of the front side vision is as follows:

[0016] (24) Input the point cloud data in front of the vehicle into the SECOND network, and extract the point cloud features of the point cloud data in the BEV space through the SECOND network;

[0017] (25) Input the point cloud features in the BEV space into the 3D detection head, and the 3D detection head extracts the target objects based on the point cloud features.

[0018] Further, the target objects include vehicles and pedestrians.

[0019] Further, after step (2), the following is also included:

[0020] (3) After detecting the target objects in front of the vehicle and in the blind spots of the front side vision based on the image data and point cloud data of the current frame, use the AB3DMOT target tracking algorithm to track the detected target objects.

[0021] The present invention is implemented as follows. A vehicle perception system based on multi-modal fusion, the system includes:

[0022] Three lidars, one camera and one processor provided on the vehicle, the lidars and the camera are connected to the processor; the three lidars are respectively used to scan the point cloud in front of the vehicle and in the blind spots of the front side vision, and send it to the processor; the camera is used to capture the image data in front of the vehicle and send it to the processor; the processor perceives the target objects in the surrounding environment based on the above vehicle perception method based on multi-modal fusion.

[0023] Further, the three lidars are one M2 solid-state lidar and two mechanical lidar sensors. The M2 solid-state lidar is installed on the front bumper of the vehicle, and the two mechanical lidar sensors are installed on both sides of the vehicle head.

[0024] Further, the M2 solid-state lidar is used to scan the point cloud data in front of the vehicle during driving, and the two mechanical lidars are used to scan the point cloud data in the blind spots on both sides of the front view of the vehicle during driving.

[0025] Extract image features of different sizes from the image data in the driving front. After splicing the image features of different sizes, project them onto the BEV space and fuse them with the point cloud features in the BEV space. Input the obtained fusion features in the BEV space into the 3D detection head for target object recognition. Use the fusion features of the image data and the point cloud data for target object recognition, which greatly improves the recognition accuracy of target objects in the driving front. In addition, for the front vision blind area, based on the point cloud data for target object recognition, further improves the vehicle's environmental perception ability. Description of the Drawings

[0026] Figure 1 It is a schematic structural diagram of a vehicle perception system based on multi-modal fusion provided by an embodiment of the present invention;

[0027] Figure 2 It is a flowchart of a vehicle perception method based on multi-modal fusion provided by an embodiment of the present invention. Detailed Embodiments

[0028] Next, with reference to the drawings, through the description of the optimal embodiments, the specific embodiments of the present invention will be further described in detail.

[0029] Lidar and camera are two different perception devices, and the data information and representation forms they capture are very different. Lidar mainly provides distance information and reflection intensity, generating a set of three-dimensional point cloud data; while the camera mainly provides color and texture information, generating two-dimensional image data. Each of these two types of data has its own advantages and disadvantages. For example, lidar data is less affected by lighting and weather conditions, but has a lower spatial resolution; camera data can provide rich color and texture information and has a higher spatial resolution, but is sensitive to lighting and weather conditions. Aligning and fusing the data of lidar and camera can combine the advantages of both and avoid their respective disadvantages, thus obtaining a more comprehensive and accurate environmental representation. Based on this, the present invention provides a vehicle perception system based on the fusion of camera and lidar. Figure 1 It is a schematic structural diagram of a vehicle perception system based on multi-modal fusion provided by an embodiment of the present invention. For the convenience of description, only the parts related to the embodiments of the present invention are shown. The vehicle perception system includes:

[0030] Three lidar sensors, one camera, and one processor are installed on a vehicle. The lidar sensors and the camera are connected to the processor. The three lidar sensors are respectively used to scan the point clouds in the front and the blind spots on both sides of the front of the vehicle during driving, and send them to the processor. The camera is used to capture the image data in the front of the vehicle during driving and send it to the processor. The processor perceives the target objects in the surrounding environment based on the following vehicle perception method based on multi-modal fusion. The target objects include vehicles and pedestrians, mainly used to perceive other vehicles or pedestrians in the front of the vehicle during driving, facilitating the auxiliary driving of the vehicle or realizing the side autonomous driving of the vehicle. In addition, perceiving other vehicles or pedestrians in the blind spot of the front view of the vehicle mainly aims to improve the safety during the vehicle driving process.

[0031] In the embodiment of the present invention, the three lidar sensors are respectively one M2 solid-state lidar sensor and two mechanical lidar sensors. Among them, the M2 solid-state lidar sensor is installed on the front bumper of the vehicle and is used to scan the point cloud data in the front of the vehicle during driving. The two mechanical lidar sensors are installed on both sides of the front of the vehicle head and are used to scan the point cloud data in the blind spots on both sides of the front of the vehicle during driving.

[0032] Figure 2 The following is a flowchart of the vehicle perception method based on multi-modal fusion provided by the embodiment of the present invention. The method is as follows:

[0033] (1) Obtain the image data in the front of the vehicle during driving through the camera, scan the point cloud data in the front and the blind spots on both sides of the front of the vehicle during driving through the lidar sensors, and perform noise reduction processing on the point cloud data. Since the noise reduction processing process is existing, the present invention will not elaborate on it here.

[0034] (2) Fuse the point cloud data in the front of the vehicle and the image data, perceive the target objects in the corresponding area based on the fused spatial features, and perceive the target objects in the corresponding area based on the point cloud data in the blind spots on both sides of the front of the vehicle during driving. The target objects are vehicles and pedestrians.

[0035] In the embodiment of the present invention, the point cloud features are extracted from the point cloud data in the front of the vehicle during driving, the image features are extracted from the image data in the front of the vehicle during driving, the extracted point cloud features and image features are fused, and the fused features are input into the 3D detection head. The 3D detection head extracts the target objects based on the fused features. In the embodiment of the present invention, the fusion method of the point cloud features in the point cloud data and the image features in the image data is as follows:

[0036] (21) Input the point cloud data in the front of the vehicle during driving into the SECOND network, and extract the point cloud features of the point cloud data in the BEV space through the SECOND network.

[0037] The SECOND network, as a feature extraction network for point cloud data, consists of a voxel feature extractor module and a sparse convolutional layer module. The voxel grid converted from the point cloud data is input into the voxel feature extractor module, and per-voxel max pooling is used to obtain the local aggregated features of each voxel. The tiling of the local features is concatenated with different points; the sparse convolutional layer module collects the features obtained in the previous step. By stacking multiple layers of sparse convolutional layers, the network can gradually extract features from local to global. Each layer of sparse convolution increases the abstraction level of the features and reduces the resolution of the features, and these features are used for subsequent feature fusion.

[0038] (22) Input the image data in front of the vehicle into the MoA-Transformer feature extraction network to extract image features of different sizes in the image data. Based on the ASFF algorithm, the image features of different sizes are concatenated to obtain a feature map with rich semantic information. Then, based on the EA-LSS algorithm, the concatenated image features are projected into the BEV space to obtain the image features in the BEV space;

[0039] The feature extraction network of MoA-Transformer adopts multi-resolution overlapping attention (MOA) for the image classification model of Transformer to achieve information interaction between nearby windows and all non-local windows. MoA-Transformer mainly includes the following aspects: multi-scale structure: MoA-Transformer adopts a multi-layer structure to extract the input image into feature maps of different scales, which is convenient for processing targets of different sizes. Attention mechanism: MoA-Transformer adopts multi-resolution overlapping attention to make full use of the global image information; Model training: MoA-Transformer is first pre-trained on a large-scale image dataset and then fine-tuned on special detection tasks to increase the generalization ability of the network.

[0040] The ASFF algorithm can learn the features of other scales, so as to retain useful semantic information for fusion. The key idea of ASFF is to adaptively learn the fusion weights of the feature maps of each scale, and then perform adaptive fusion on the feature maps of each layer, so that the discriminative features dominate and filter out unimportant image information, so as to achieve rich feature semantic information at each scale.

[0041] The EA-LSS algorithm converts 2D image information into point cloud information in the BEV vision and then uses it for feature fusion. The EA-LSS algorithm includes an edge-aware depth fusion module and a refined depth module. The edge-aware depth fusion module is used to process the original point cloud to obtain a multi-view edge map to alleviate the problem of rapid depth change. The refined depth module is used to calculate the difference between the predicted feature map and the real map to make the fine-grained perceived depth distribution and retain the original depth information.

[0042] (23)Fuse the point cloud features in the BEV space with the image features in the BEV space to obtain the fused features in the BEV space.

[0043] In the embodiments of the present invention, by assigning different weights to the point cloud features in the BEV space and the image features in the BEV space, the fusion of the point cloud features in the BEV space and the image features in the BEV space is achieved based on the weight values, forming the fused features in the BEV space. The fused features in the BEV space synthesize the image features and the point cloud features. Based on the fused features in the BEV space, the detection of target objects is carried out, greatly improving the detection accuracy of target objects.

[0044] In the embodiments of the present invention, the process of extracting target objects in the blind spots of the line of sight on both sides in front of the vehicle during driving is specifically as follows:

[0045] (24)Input the point cloud data in front of the driving into the SECOND network, and extract the point cloud features of the point cloud data in the BEV space through the SECOND network; (25) Input the point cloud features in the BEV space into the 3D detection head, and the 3D detection head extracts target objects based on the point cloud features. The target objects are mainly moving objects, including vehicles and pedestrians. The vehicle can be a sanitation vehicle.

[0046] In the embodiments of the present invention, after step (2), it further includes: (3) After detecting the target objects in front of the driving and in the blind spots of the line of sight on both sides in front based on the image data and point cloud data of the current frame, the AB3DMOT target tracking algorithm is used to track the detected target objects. By tracking the motion states of other vehicles and pedestrians around the vehicle, it is beneficial to predict the motion trajectories of other vehicles and pedestrians around.

[0047] Extract image features of different sizes from the image data in front of the driving. After splicing the image features of different sizes, project them onto the BEV space, fuse them with the point cloud features in the BEV space, input the obtained fused features in the BEV space into the 3D detection head for the recognition of target objects. Using the fused features of the image data and the point cloud data for the recognition of target objects greatly improves the recognition accuracy of target objects in front of the driving. In addition, for the blind area of the front vision, based on the point cloud data for the recognition of target objects, the environmental perception ability of the vehicle is further improved.

[0048] The present invention has been described exemplarily. Obviously, the specific implementation of the present invention is not limited by the above methods. As long as various non-substantive improvements are made by adopting the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.

Claims

1. A vehicle perception method based on multimodal fusion, characterized in that: The method is specifically as follows: (1) The camera obtains image data in front of the vehicle and the laser radar scans the point cloud data of the blind spots in front of the vehicle and on both sides of the front. (2) The point cloud data and image data in front of the vehicle are fused, and the target objects in the corresponding area are identified based on the fused spatial features, and the target objects in the corresponding area are identified based on the point cloud data of the blind spots on both sides of the front.

2. The vehicle perception method based on multimodal fusion as claimed in claim 1, characterized in that: The target object recognition process ahead of the vehicle is as follows: Point cloud features are extracted from the point cloud data ahead of the vehicle, and image features are extracted from the image data ahead of the vehicle. The extracted point cloud features are fused with the image features, and the fused spatial features are input into the 3D detection head. The 3D detection head extracts the target object based on the fused spatial features.

3. The vehicle perception method based on multimodal fusion as claimed in claim 2, characterized in that: The fusion process of point cloud features and image features is as follows: (21) Inputting the point cloud data in front of the vehicle into the SECOND network, and extracting the point cloud features of the point cloud data in the BEV space through the SECOND network; (22) Inputting the image data in front of the vehicle into the MoA-Transformer feature extraction network, extracting image features of different sizes from the image data, splicing the image features of different sizes based on the ASFF algorithm, and projecting the spliced ​​image features into the BEV space based on the EA-LSS algorithm to obtain the image features in the BEV space; (23) The point cloud features in the BEV space are fused with the image features in the BEV space to obtain the fusion features in the BEV space.

4. The vehicle perception method based on multimodal fusion as claimed in claim 1, characterized in that: The process of extracting target objects in the blind spots on both sides of the front is as follows: (24) Inputting the point cloud data in front of the vehicle into the SECOND network, and extracting the point cloud features of the point cloud data in the BEV space through the SECOND network; (25) The point cloud features in the BEV space are input into the 3D detection head, and the 3D detection head extracts the target object based on the point cloud features.

5. The vehicle perception method based on multimodal fusion according to any one of claims 1 to 4, characterized in that: Target objects include vehicles and pedestrians.

6. The vehicle perception method based on multimodal fusion as claimed in claim 1, characterized in that: After step (2), the method further includes: (3) After detecting the target objects in the blind spots in front and on both sides of the front based on the image data and point cloud data of the current frame, the AB3DMOT target tracking algorithm is used to track the detected target objects.

7. A vehicle perception system based on multimodal fusion, characterized in that: The system comprises: Three laser radars, a camera and a processor are arranged on a vehicle, and the laser radars and the camera are connected to the processor; the three laser radars are respectively used to scan the point clouds of the blind spots in front of the vehicle and on both sides of the front, and send them to the processor; the camera takes image data in front of the vehicle and sends it to the processor; the processor perceives target objects in the surrounding environment based on the vehicle perception method based on multimodal fusion as described in any one of claims 1 to 6.

8. The vehicle perception system based on multimodal fusion as claimed in claim 7, characterized in that: The three laser radars are one M2 solid-state laser radar and two mechanical laser radar sensors. The M2 solid-state laser radar is installed on the front bumper of the vehicle, and the two mechanical laser radar sensors are installed on both sides of the front of the vehicle.

9. The vehicle perception system based on multimodal fusion as claimed in claim 7, characterized in that: The M2 solid-state laser radar is used to scan the point cloud data in front of the vehicle, and the two mechanical laser radars are used to scan the point cloud data in the blind spots on both sides of the vehicle.