Road Perception Method and System for Multi-View Camera Image and LiDAR Fusion

Through the integration of multi-view camera images and lidar data, the deep learning network generates and aligns features, the accuracy problem of road perception in complex scenarios is solved, and high-precision road recognition and safe driving support in harsh environments are achieved.

CN120198881BActive Publication Date: 2025-07-22JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510689568.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-22
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The existing road perception technology has low detection accuracy in complex scenarios and harsh environments, especially when the field of view is limited or the environment is blurred, and the perception accuracy of multi-view camera images is insufficient.

Method used

By fusing multi-view camera images and lidar data, deep information features are generated using the Swin-Transformer network and VoxelNet model, feature interaction and alignment are performed, bird's-eye view enhancement features are generated in combination with the cross attention mechanism, and finally prediction is performed through the semantic segmentation head.

Benefits of technology

It improves the accuracy and robustness of road perception, ensures accurate identification of road conditions in unstructured roads and harsh environments, and enhances the vehicle's safe driving decision-making ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198881B_ABST
    Figure CN120198881B_ABST
Patent Text Reader

Abstract

The present invention proposes a road perception method and system for fusing multi-view camera images and lidar, and the method includes: obtaining the depth information features of the camera bird's-eye view and the depth information features of the lidar bird's-eye view based on the multi-view images of the vehicle-mounted camera and the lidar point cloud data; obtaining the camera bird's-eye view features with depth distribution based on the depth information features of the camera bird's-eye view and the depth information features of the lidar bird's-eye view; performing feature alignment on the camera bird's-eye view features with depth distribution and the camera bird's-eye view features without depth distribution to obtain the camera bird's-eye view image features containing depth information representation; obtaining the prediction result based on the camera bird's-eye view image features containing depth information representation and the depth information features of the lidar bird's-eye view. The present invention can, during vehicle driving, in real time infer and predict the road conditions of the current scene through the images of the vehicle-mounted camera, and can in real time perceive the current driving road conditions under unstructured road conditions with limited vision and harsh environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning and computer vision analysis, and particularly to a road perception method and system for fusing multi-view camera images and lidar. Background Art

[0002] The evolution of vision-centric autonomous driving technology largely depends on the accurate interpretation and real-time response to complex dynamic environments, which requires the perception system to be sufficiently robust. Among them, road perception, as one of the core perception tasks, can help autonomous vehicles reliably identify drivable areas, ensure safe navigation, and interact intelligently with the surrounding environment, thus laying a foundation for the wide application of autonomous driving technology in real traffic scenarios. Road perception technology is crucial in identifying road information, providing high-quality decision-making basis for vehicles, enabling them to perform operations such as cruise driving, following the vehicle in front, and changing lanes inside and outside the road, and ensuring coherent and safe intelligent driving.

[0003] To effectively prevent traffic accidents and improve vehicle safety performance, the research and development and popularization of intelligent safety systems have become particularly urgent. In the development process of autonomous driving technology, road perception is a crucial link, aiming to accurately identify and sense vehicle and road condition information on the road, helping the vehicle stay within the lane and make safe driving decisions.

[0004] Although current road perception technology has shown high detection accuracy in well-lit and simple road conditions, problems are still prominent in complex scenarios. For example, the driving area may be unclear due to being blocked or worn by other vehicles, or in adverse driving environments such as rain, night, and fog, the on-vehicle multi-view camera is affected by the environment, resulting in a blurred image, thereby reducing the accuracy of perception. Summary of the Invention

[0005] In view of the above situation, the main object of the present invention is to propose a road perception method and system for fusing multi-view camera images and lidar to solve the above technical problems.

[0006] The present invention proposes a road perception method for fusing multi-view camera images and lidar, and the method includes the following steps:

[0007] Step 1: Input the obtained multi-view images of on-vehicle cameras into the Swin-Transformer network, and obtain the depth information features of the camera bird's-eye view through multi-modal feature conversion;

[0008] Process the obtained lidar point cloud data through the VoxelNet model to obtain the depth information features of the lidar bird's-eye view;

[0009] Step 2: Based on the depth information features of the camera bird's-eye view and the depth information features of the lidar bird's-eye view, confirm to obtain camera information and radar information;

[0010] Based on the camera information and radar information, through bird's-eye view voxel pooling processing, obtain the camera bird's-eye view features without depth distribution and the image features without depth distribution;

[0011] Through the depth information features of the lidar bird's-eye view, insert the lidar depth features into the image features without depth distribution to obtain the camera bird's-eye view features with depth distribution;

[0012] Step 3: Align the camera bird's-eye view features with depth distribution and the camera bird's-eye view features without depth distribution to obtain the camera bird's-eye view image features containing depth information representation;

[0013] Step 4: Based on the camera bird's-eye view image features containing depth information representation and the depth information features of the lidar bird's-eye view, respectively confirm to obtain the corresponding query vector, key vector, and value vector;

[0014] Based on the corresponding query vector, key vector, and value vector, through cross-attention calculation, obtain the bird's-eye view enhanced features;

[0015] Step 5: Input the bird's-eye view enhanced features into the semantic segmentation head for prediction to generate the prediction result.

[0016] The present invention also proposes a road perception system for fusing multi-view camera images and lidar, and the system includes:

[0017] Content input module, used for:

[0018] Input the acquired multi-view images of the vehicle-mounted camera into the Swin-Transformer network, and obtain the depth information features of the camera bird's-eye view through multi-modal feature conversion;

[0019] Process the acquired lidar point cloud data through the VoxelNet model to obtain the depth information features of the lidar bird's-eye view;

[0020] Depth information interaction and fusion module, used for:

[0021] Based on the depth information features of the camera bird's-eye view and the depth information features of the lidar bird's-eye view, confirm to obtain camera information and radar information;

[0022] Based on the camera information and radar information, through bird's-eye view voxel pooling processing, obtain the camera bird's-eye view features without depth distribution and the image features without depth distribution;

[0023] Insert lidar depth features into the image features without depth distribution based on the depth information features of the lidar bird's-eye view to obtain camera bird's-eye view features with depth distribution;

[0024] Align the camera bird's-eye view features with depth distribution and the camera bird's-eye view features without depth distribution to obtain camera bird's-eye view image features containing depth information representation;

[0025] The bird's-eye view feature enhancement and fusion module is used for:

[0026] Based on the camera bird's-eye view image features containing depth information representation and the depth information features of the lidar bird's-eye view, respectively confirm the corresponding query vectors, key vectors, and value vectors;

[0027] Based on the corresponding query vectors, key vectors, and value vectors, obtain bird's-eye view enhanced features through cross-attention calculation;

[0028] The road perception prediction output module is used for:

[0029] Input the bird's-eye view enhanced features into the semantic segmentation head for prediction to generate a prediction result.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0031] 1. The present invention generates enhanced bird's-eye view features by interacting and fusing the depth information features of the point cloud bird's-eye view generated from lidar data with the camera bird's-eye view depth features, and aligns the radar point cloud features with the depth-enhanced bird's-eye view features, thereby making up for the lack of depth information in the camera bird's-eye view perspective to ensure the situation of unstructured road segments;

[0032] 2. The present invention deeply fuses the enhanced bird's-eye view features with the lidar bird's-eye view features through the cross-attention mechanism to obtain spatial perception multi-modal features, so as to obtain the final bird's-eye view level fusion features, thereby significantly improving the bird's-eye view perception ability and improving the accuracy and robustness of the model.

[0033] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the embodiments of the present invention. Brief Description of the Drawings

[0034] Figure 1 It is a flowchart of the road perception method for multi-view camera images and lidar fusion proposed by the present invention;

[0035] Figure 2 It is a schematic diagram of the overall framework of the road perception method for multi-view camera images and lidar fusion proposed by the present invention;

[0036] Figure 3Schematic diagram of the depth information interaction and fusion module of the road perception method for multi-view camera images and lidar fusion proposed by the present invention;

[0037] Figure 4 Schematic diagram of the bird's-eye view feature enhancement and fusion module of the road perception method for multi-view camera images and lidar fusion proposed by the present invention;

[0038] Figure 5 Schematic diagram of the overall framework of the road perception system for multi-view camera images and lidar fusion proposed by the present invention. Detailed implementation manners

[0039] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0040] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific implementation manners in the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0041] Please refer to Figure 1 and Figure 2 , the embodiments of the present invention propose a road perception method for multi-view camera images and lidar fusion, and the method includes the following steps:

[0042] Step 1: Input the acquired multi-view images of the vehicle-mounted camera into the Swin-Transformer network, and obtain the depth information features of the camera bird's-eye view through multi-modal feature conversion;

[0043] Process the acquired lidar point cloud data through the VoxelNet model to obtain the depth information features of the lidar bird's-eye view;

[0044] In step 1, when inputting the acquired multi-view images of the vehicle-mounted camera into the Swin-Transformer network and obtaining the depth information features of the camera bird's-eye view through multi-modal feature conversion, the relational expression existing in the corresponding process is:

[0045] ;

[0046] Among them, represents the input image, represents the depth information features of the camera bird's-eye view, Indicates the extraction operation performed from the th layer of the Swin-Transformer network, indicating the global average pooling operation;

[0047] The lidar point cloud data obtained is processed through the VoxelNet model to obtain the depth information features of the lidar bird's-eye view. The specific steps are as follows:

[0048] Map the lidar point cloud data to a 3D voxel grid. The relationship formula in the corresponding process is:

[0049] ;

[0050] where, represents the point cloud located at , represents the point cloud data, represents the voxel located at ;

[0051] Subsequently, use PointNet to extract the local geometric features of the points inside the voxel, and perform feature fusion through a multi-layer perceptron and a max pooling layer to obtain the voxel features. The relationship formula in the corresponding process is:

[0052] ;

[0053] where, represents the processing by the multi-layer perceptron for encoding the local features of the points, represents the max pooling operation, represents the voxel features;

[0054] Organize the voxel features into a sparse 3D feature map, input it into a 3D convolutional neural network for global feature extraction, and finally perform a bird's-eye view projection to obtain the depth information features of the lidar bird's-eye view.

[0055] Specifically, in this step, most of the Swin TransformerBlocks in the Swin-Transformer model are retained, and the intermediate layer close to the output layer but still containing semantic information is selected as the feature extraction layer. At the same time, the final fully connected classification layer is removed to ensure that the model focuses on extracting feature representations from the input image;

[0056] In this step, the nuScenes dataset is used as the training image input into the network, and all images are of a fixed size.

[0057] Step 2: Based on the depth information features of the camera bird's-eye view and the depth information features of the lidar bird's-eye view, confirm to obtain the camera information and the radar information;

[0058] Based on camera information and radar information, through bird's-eye view voxel pooling processing, a camera bird's-eye view feature without depth distribution and an image feature without depth distribution are obtained;

[0059] Through the depth information feature of the lidar bird's-eye view, the lidar depth feature is inserted into the image feature without depth distribution to obtain a camera bird's-eye view feature with depth distribution;

[0060] Please refer to Figure 3 , in step 2, based on camera information and radar information, through bird's-eye view voxel pooling processing, a camera bird's-eye view feature without depth distribution and an image feature without depth distribution are obtained, and the relationship existing in the corresponding process is:

[0061] ;

[0062] Among them, represents the camera bird's-eye view feature without depth distribution, represents the image feature without depth distribution, represents after bird's-eye view voxel pooling processing, represents the image size tensor, represents the camera information, represents the radar information;

[0063] Through the depth information feature of the lidar bird's-eye view, the lidar depth feature is inserted into the image feature without depth distribution to obtain a camera bird's-eye view feature with depth distribution, and the relationship existing in the corresponding process is:

[0064] ;

[0065] Among them, represents the camera bird's-eye view feature with depth distribution, represents after being processed by the sigmoid activation function, represents through the convolution function processing, represents after being processed by a single convolutional neural network layer with a parameter set of ; represents the lidar bird's-eye view information with depth distribution, represents the depth information feature of the lidar bird's-eye view.

[0066] Specifically, in this step, the camera features are first spatially transformed and projected into the bird's-eye view space to form a bird's-eye view feature map aligned with the lidar features. Then, the interaction between the two modal features is achieved by inserting the lidar depth feature: the camera features provide semantic guidance for the lidar features, while the lidar features supplement geometric constraints for the camera features. This two-way interaction ensures geometric consistency in the feature fusion process.

[0067] Step 3: Align the camera bird's-eye view features with depth distribution and the camera bird's-eye view features without depth distribution to obtain the camera bird's-eye view image features with depth information representation;

[0068] In Step 3, when aligning the camera bird's-eye view features with depth distribution and the camera bird's-eye view features without depth distribution to obtain the camera bird's-eye view image features with depth information representation, the relational expression in the corresponding process is:

[0069] ;

[0070] Among them, represents the camera bird's-eye view image features with depth information representation, represents being processed by the Feature Pyramid Network, represents performing self-calibration convolution processing.

[0071] Step 4: Based on the camera bird's-eye view image features with depth information representation and the depth information features of the lidar bird's-eye view, respectively confirm the corresponding query vector, key vector, and value vector;

[0072] Based on the corresponding query vector, key vector, and value vector, through cross-attention calculation, obtain the enhanced bird's-eye view features;

[0073] Please refer to Figure 4 , in Step 4, when obtaining the enhanced bird's-eye view features through cross-attention calculation based on the corresponding query vector, key vector, and value vector, the relational expression in the corresponding process is:

[0074] ;

[0075] Among them, represents the enhanced bird's-eye view features, , and respectively represent the query vector, key vector, and value vector of the camera bird's-eye view image features with depth information representation, , and respectively represent the query vector, key vector, and value vector of the depth information features of the lidar bird's-eye view, , and respectively represent the query vector, key vector, and value vector, represents the transpose symbol, represents the scaling factor, represents the dimension of the feature vector.

[0076] Specifically, in this step, the camera bird's-eye view image features containing depth information representation and the depth information features of the lidar bird's-eye view are respectively used as the keyword and query vector of the cross-attention mechanism. Among them, the camera bird's-eye view features containing depth information representation serve as the carrier of semantic information, providing rich environmental context information; while the depth information features of the lidar bird's-eye view serve as the carrier of geometric information, providing accurate spatial structure information. In the feature fusion process, the correlation weight between the camera features and the lidar features is calculated through the scale dot product attention mechanism.

[0077] Step 5: Input the enhanced bird's-eye view features into the semantic segmentation head for prediction to generate a prediction result;

[0078] In step 5, the enhanced bird's-eye view features are input into the semantic segmentation head for prediction to generate a prediction result. Among them, the mapping channels of the semantic segmentation head during the supervised training process are adopted by using the focal loss function, and the expression of the focal loss function is as follows:

[0079] ;

[0080] Among them, represents the focal loss function, represents the semantic category, represents all categories, represents the index of the category, represents the balance factor, represents the prediction probability of the model for the correct category, represents the adjustment factor, represents the adjustable parameter for solving the foreground-background imbalance, represents the number of categories.

[0081] Specifically, in this step, the enhanced bird's-eye view features first go through a series of upsampling operations and convolutional layers to gradually restore the spatial resolution of the feature map to be consistent with the resolution of the original input image or the bird's-eye view space. During the upsampling process, skip connections are used to fuse the low-level features and the high-level features to retain more detailed information. Finally, the category label of each pixel point is determined through the argmax operation to generate a complete road perception result. This result is presented in the form of a segmentation mask, clearly marking information such as the road area, drivable area, and obstacles in the scene.

[0082] Please refer to Figure 5 , the embodiment of the present invention also provides a road perception system for fusing multi-view camera images and lidar, and the system includes:

[0083] Content input module, used for:

[0084] Input the multi-view images of the in-vehicle camera obtained into the Swin-Transformer network, and obtain the depth information features of the camera bird's-eye view through multi-modal feature transformation;

[0085] Process the lidar point cloud data obtained through the VoxelNet model to obtain the depth information features of the lidar bird's-eye view;

[0086] The depth information interaction and fusion module is used for:

[0087] Based on the depth information features of the camera bird's-eye view and the depth information features of the lidar bird's-eye view, confirm the camera information and radar information;

[0088] Based on the camera information and radar information, through bird's-eye view voxel pooling processing, obtain the camera bird's-eye view features without depth distribution and the image features without depth distribution;

[0089] Through the depth information features of the lidar bird's-eye view, insert the lidar depth features into the image features without depth distribution to obtain the camera bird's-eye view features with depth distribution;

[0090] Align the camera bird's-eye view features with depth distribution and the camera bird's-eye view features without depth distribution to obtain the camera bird's-eye view image features containing depth information representation;

[0091] The bird's-eye view feature enhancement and fusion module is used for:

[0092] Based on the camera bird's-eye view image features containing depth information representation and the depth information features of the lidar bird's-eye view, respectively confirm the corresponding query vectors, key vectors, and value vectors;

[0093] Based on the corresponding query vectors, key vectors, and value vectors, through cross-attention calculation, obtain the bird's-eye view enhanced features;

[0094] The road perception prediction output module is used for:

[0095] Input the bird's-eye view enhanced features into the semantic segmentation head for prediction to generate the prediction results.

[0096] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0097] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0098] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. A road perception method for fusing multi-view camera images and lidar, characterized in that, The method includes the following steps: Step 1: Input the multi-view images of the vehicle-mounted camera obtained into the Swin-Transformer network, and through multi-modal feature transformation, obtain the depth information features of the camera bird's-eye view; Process the lidar point cloud data obtained through the VoxelNet model to obtain the depth information features of the lidar bird's-eye view; Step 2: Based on the depth information features of the camera bird's-eye view and the depth information features of the lidar bird's-eye view, confirm to obtain camera information and radar information; Based on the camera information and radar information, through bird's-eye view voxel pooling processing, obtain the camera bird's-eye view features without depth distribution and the image features without depth distribution; Through the depth information features of the lidar bird's-eye view, insert the lidar depth features into the image features without depth distribution to obtain the camera bird's-eye view features with depth distribution; Step 3: Align the camera bird's-eye view features with depth distribution and the camera bird's-eye view features without depth distribution to obtain the camera bird's-eye view image features containing depth information representation; Step 4: Based on the camera bird's-eye view image features containing depth information representation and the depth information features of the lidar bird's-eye view, respectively confirm to obtain the corresponding query vectors, key vectors, and value vectors; Based on the corresponding query vectors, key vectors, and value vectors, through cross-attention calculation, obtain the bird's-eye view enhanced features; Step 5: Input the bird's-eye view enhanced features into the semantic segmentation head for prediction to generate the prediction result.

2. The road perception method for fusing multi-view camera images and lidar according to claim 1, characterized in that In the above Step 1, when inputting the multi-view images of the vehicle-mounted camera obtained into the Swin-Transformer network and obtaining the depth information features of the camera bird's-eye view through multi-modal feature transformation, the relational formula existing in the corresponding process is: ; Among them, represents the input image, represents the depth information feature of the camera bird's-eye view, represents the extraction operation from the layer of the Swin-Transformer network, represents the operation of global average pooling.

3. The road perception method for fusing multi-view camera images and lidar according to claim 2, wherein In the above Step 1, when processing the lidar point cloud data obtained through the VoxelNet model to obtain the depth information features of the lidar bird's-eye view, the specific steps are as follows: Map the lidar point cloud data to a 3D voxel grid, and the relational formula existing in the corresponding process is: ; Among them, represents the point cloud located at ; represents the point cloud data, represents the voxel located at ; Use PointNet to extract the local geometric features of the points inside the voxels, and through a multi-layer perceptron and a max pooling layer for feature fusion to obtain voxel features, and the relational formula existing in the corresponding process is: ; Among them, represents the processing by the local feature multi-layer perceptron for encoding points, represents the max pooling operation, represents the body feature; Organize the voxel features into a sparse 3D feature map, input it into a 3D convolutional neural network for global feature extraction, and finally perform bird's-eye view projection to obtain the depth information features of the lidar bird's-eye view.

4. The road perception method for fusing multi-view camera images and lidar according to claim 3, characterized in that, In the above Step 2, based on the camera information and radar information, through bird's-eye view voxel pooling processing, obtain the camera bird's-eye view features without depth distribution and the image features without depth distribution, and the relational formula existing in the corresponding process is: ; Among them, represents the camera bird's-eye view feature without depth distribution, represents the image feature without depth distribution, represents being processed by bird's-eye view pixel pooling, represents the image size tensor, represents the camera information, represents the radar information.

5. The method for road perception by fusing multi-view camera images and lidar according to claim 4, characterized in that, In the above Step 2, through the depth information features of the lidar bird's-eye view, insert the lidar depth features into the image features without depth distribution to obtain the camera bird's-eye view features with depth distribution, and the relational formula existing in the corresponding process is: ; Among them, represents the camera bird's-eye view feature with depth distribution, represents being processed by the sigmoid activation function, represents being processed by the convolution function, represents being processed by a single convolutional neural network layer with a parameter set of ; represents the radar bird's-eye view information with depth distribution, represents the depth information feature of the lidar bird's-eye view.

6. The road perception method for fusing multi-view camera images and lidar according to claim 5, characterized in that, In the above Step 3, when aligning the camera bird's-eye view features with depth distribution and the camera bird's-eye view features without depth distribution to obtain the camera bird's-eye view image features containing depth information representation, the relational formula existing in the corresponding process is: ; Among them, represents the camera bird's-eye view image feature containing depth information representation, represents being processed by the feature pyramid network, represents performing self-calibration convolution processing.

7. The road perception method for fusing multi-view camera images and lidar according to claim 6, wherein In the step 4, based on the corresponding query vector, key vector, and value vector, through cross-attention calculation, an enhanced bird's-eye view feature is obtained. The relational expression existing in the corresponding process is as follows: ; Among them, represents the enhanced feature of the bird's-eye view, , and respectively represent the query vector, key vector, and value vector of the camera bird's-eye view image feature containing depth information representation, , and respectively represent the query vector, key vector, and value vector of the depth information feature of the lidar bird's-eye view, , and respectively represent the query vector, key vector, and value vector, represents the transpose symbol, represents the scaling factor, represents the dimension of the feature vector.

8. The road perception method for fusing multi-view camera images and lidar according to claim 7, characterized in that, In the step 5, the enhanced bird's-eye view feature is input into the semantic segmentation head for prediction to generate a prediction result. Among them, during the supervised training process of the semantic segmentation head by adopting the focal loss function, the expression of the focal loss function is as follows: ; Among them, represents the focal loss function, represents the semantic category, represents all categories, represents the index of the category, represents the balancing factor, represents the prediction probability of the model for the correct category, represents the adjustment factor, represents an adjustable parameter for solving the foreground-background imbalance, represents the number of categories.

9. A road perception system that fuses multi-view camera images and lidar, characterized in that, The system applies the road perception method for fusing multi-view camera images and lidar as described in any one of claims 1 to 8. The system includes: A content input module, which is used for: Inputting the acquired multi-view images of the vehicle-mounted camera into the Swin-Transformer network, and obtaining the depth information feature of the camera bird's-eye view through multi-modal feature conversion; Processing the acquired lidar point cloud data through the VoxelNet model to obtain the depth information feature of the lidar bird's-eye view; A depth information interaction and fusion module, which is used for: Based on the depth information feature of the camera bird's-eye view and the depth information feature of the lidar bird's-eye view, confirming the camera information and the radar information; Based on the camera information and the radar information, through bird's-eye view voxel pooling processing, obtaining the camera bird's-eye view feature without depth distribution and the image feature without depth distribution; Inserting the lidar depth feature into the image feature without depth distribution through the depth information feature of the lidar bird's-eye view to obtain the camera bird's-eye view feature with depth distribution; Aligning the camera bird's-eye view feature with depth distribution and the camera bird's-eye view feature without depth distribution to obtain the camera bird's-eye view image feature containing depth information representation; A bird's-eye view feature enhancement and fusion module, which is used for: Based on the camera bird's-eye view image feature containing depth information representation and the depth information feature of the lidar bird's-eye view, respectively confirming the corresponding query vector, key vector, and value vector; Based on the corresponding query vector, key vector, and value vector, through cross-attention calculation, obtaining the enhanced bird's-eye view feature; A road perception prediction output module, which is used for: Inputting the enhanced bird's-eye view feature into the semantic segmentation head for prediction to generate a prediction result.

Citation Information

Patent Citations

  • Point cloud reconstruction method and device based on multi-modal mask strategy

    CN117745934A

  • Three-dimensional sensing method based on millimeter wave radar and camera aerial view fusion

    CN118038396A