A Cross-Modal Multi-Task Environment Perception Method and System
By integrating multimodal sensor information under a unified framework to generate strong BEV features, the problems of low efficiency and poor robustness of the existing environmental perception system are solved, and efficient and robust environmental perception is achieved.
Patent Information
- Application Number
- CN202311203963.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-09-18
AI Technical Summary
The existing environment perception system performs online 3D detection and offline high-precision map generation under different frameworks, resulting in inefficiency and poor robustness in harsh environments, making it difficult to make full use of the observation information of multimodal sensors.
Under the unified framework, through multimodal information fusion, the information obtained by vehicle-mounted multi-view cameras and lidars is used to generate dense depth maps and BEV features, combined with attention mechanisms and feature extraction networks, and realize multi-task environment perception.
It improves the efficiency of the environment perception system and its robustness in harsh environments, and can efficiently sense dynamic and static information around the vehicle.
Smart Images

Figure CN117237895B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle autonomous driving, and particularly to a cross-modal multi-task environment perception method and system. Background Art
[0002] An autonomous driving vehicle includes a perception, decision-making, and planning and control module. Constructing a robust environment perception system that includes dynamic and static information around the vehicle under a unified framework helps to improve the performance of subsequent decision-making and planning tasks.
[0003] Existing environment perception systems take the observation information of multi-modal sensors as input. First, multi-modal information fusion is achieved through data-level fusion or feature-level fusion. Then, online 3D detection and offline high-precision map generation are respectively performed under different frameworks. Finally, the perception results under different frameworks are converted into a unified space to construct an environment perception system that includes dynamic and static information around the vehicle. The existing methods mainly have the following disadvantages:
[0004] 1) Existing methods need to respectively perform online 3D detection and offline high-precision map generation under different frameworks, and construct an environment perception system by converting the perception results under different frameworks into a unified space, which reduces the efficiency of environment perception.
[0005] 2) Existing methods need to construct an environment perception system based on an offline high-precision map, and the generation of an offline high-precision map is complex and expensive, and it is difficult to cover all road scenarios, which limits the application scope of autonomous driving vehicles.
[0006] 3) Existing environment perception methods based on data-level fusion or feature-level fusion cannot fully utilize the observation information of multi-modal sensors, which limits the robustness of the perception system in harsh environments, such as sensor misalignment and bad weather.
[0007] Therefore, fully fusing the observation information of in-vehicle multi-modal sensors and jointly performing 3D detection and local high-precision map generation under a unified framework is crucial for constructing an efficient and robust environment perception system. Summary of the Invention
[0008] The object of the present invention is to provide a cross-modal multi-task environment perception method and system, which can construct an efficient and robust environment perception system under a unified framework and realize the perception of dynamic and static information around the vehicle.
[0009] To achieve the above object, the present invention provides the following solutions:
[0010] A cross-modal multi-task environment perception method, comprising:
[0011] Obtain observation information; the observation information includes: image information obtained by using an in-vehicle multi-view camera and lidar point cloud information obtained by using a lidar;
[0012] Extract multi-scale features of the image by using a first feature extraction network, and construct a feature pyramid network;
[0013] Project the lidar point cloud information onto the image plane to obtain a sparse depth map, and use OpenCV morphological operations to complete the depth of the sparse depth map to obtain a dense depth map;
[0014] Use a fully convolutional network to fuse the dense depth map with the deepest feature map in the feature pyramid network to achieve multi-modal information data-level fusion. Predict the context vector and discrete depth probability of each pixel in the image according to the fused features, and project them onto the 3D space along the camera ray to generate image feature point clouds;
[0015] Use a bird's-eye view pooling operation to transform the image feature point cloud into the BEV space to generate camera BEV features;
[0016] Project the lidar points onto the image plane to capture the corresponding associated pixels, construct an associated region centered on the associated pixels, and use a max pooling operation to extract the associated vectors of the associated region;
[0017] Concatenate the lidar points with the corresponding associated vectors to achieve multi-modal information data-level fusion, and use a second feature extraction network to extract the feature information of the fused lidar point cloud to generate lidar BEV features;
[0018] Use an attention mechanism to fuse the camera BEV features and lidar BEV features in the shared BEV space to achieve multi-modal information BEV-level fusion and generate strong BEV features;
[0019] Jointly perform 3D detection and local high-precision map generation on the strong BEV features to construct an environment perception system.
[0020] Optionally, the first feature extraction network is a Swin-T network.
[0021] Optionally, the use of a fully convolutional network to fuse the dense depth map with the deepest feature map in the feature pyramid network to achieve multi-modal information data-level fusion, predicting the context vector and discrete depth probability of each pixel in the image according to the fused features, and projecting them onto the 3D space along the camera ray to generate image feature point clouds specifically includes the following:
[0022] p d =α d ×c;
[0023] Where p dis the feature information corresponding to pixel p in the image feature point cloud at a depth of d, α d is the discrete depth probability, and c is the context vector at pixel p.
[0024] Optionally, before using the bird's-eye view pooling operation to transform the image feature point cloud into the BEV space and generate the camera BEV feature, it further includes:
[0025] Optimizing the bird's-eye view pooling using the Precalculation method and the Interval Reduction method.
[0026] Optionally, the second feature extraction network is a VoxelNet network.
[0027] Optionally, using the attention mechanism to fuse the camera BEV feature and the radar BEV feature in the shared BEV space to achieve BEV-level fusion of multi-modal information and generate a strong BEV feature, specifically including the following formula:
[0028]
[0029] where F A is the fused feature map, Q is the query vector of the radar BEV feature, K and V are the key and value of the camera BEV feature respectively, Softmax is non-maximum suppression, and d k is the scaling factor of the channel dimension.
[0030] A cross-modal multi-task environment perception system, including:
[0031] An observation information acquisition module for acquiring observation information; the observation information includes: image information acquired by using an in-vehicle multi-view camera and radar point cloud information acquired by using a lidar;
[0032] An image feature extraction module for extracting multi-scale features of an image by using a first feature extraction network and constructing a feature pyramid network;
[0033] An image depth map generation module for projecting the radar point cloud information onto the image plane to obtain a sparse depth map and using OpenCV morphological operations to complete the depth of the sparse depth map to obtain a dense depth map;
[0034] An image feature point cloud generation module for fusing the dense depth map with the deepest feature map in the feature pyramid network by using a fully convolutional network to achieve data-level fusion of multi-modal information, predicting the context vector and discrete depth probability of each pixel in the image, and projecting along the camera ray into the 3D space to generate an image feature point cloud;
[0035] The camera BEV feature extraction module is used to convert the image feature point cloud to the BEV space by using the bird's-eye view pooling operation, and generate the camera BEV feature;
[0036] The correlation vector extraction module is used to project the radar points onto the image plane to capture the corresponding correlation pixels, construct a correlation region centered on the correlation pixels, and extract the correlation vector of the correlation region by using the max pooling operation;
[0037] The radar BEV feature extraction module is used to concatenate the radar points with the corresponding correlation vectors to achieve multi-modal information data-level fusion, and extract the feature information of the fused radar point cloud by using the second feature extraction network to generate the radar BEV feature;
[0038] The multi-modal feature adaptive fusion module is used to fuse the camera BEV feature and the radar BEV feature in the shared BEV space by using the attention mechanism to achieve multi-modal information BEV-level fusion and generate a strong BEV feature;
[0039] The multi-task head module is used to jointly perform 3D detection and local high-precision map generation on the strong BEV feature to construct an efficient and robust environment perception system.
[0040] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0041] A cross-modal multi-task environment perception method and system provided by the present invention generate a dense depth map according to sparse radar point cloud information; use a fully convolutional network to fuse the dense depth map with the feature map of the deepest layer in the feature pyramid network, predict the context vector and discrete depth probability of each pixel in the image according to the fused features, and project them onto the 3D space along the camera ray to generate image feature point clouds; use a bird's-eye view pooling operation to convert the image feature point clouds into the BEV space to generate camera BEV features; project the radar points onto the image plane to capture the corresponding associated pixels, construct an associated area centered on the associated pixels, and use a max pooling operation to extract the associated vector of the associated area; concatenate the radar points with the corresponding associated vectors, and use a radar feature extraction network to extract the feature information of the fused radar point clouds to generate radar BEV features; use an attention mechanism to fuse the camera BEV features and the radar BEV features to generate strong BEV features; jointly perform 3D detection and local high-precision map generation on the strong BEV features to construct an efficient and robust environment perception system; that is, by means of depth-guided camera view transformation, region-associated data-level fusion, and attention mechanism-based BEV-level fusion, the observation information of multi-modal sensors is fully utilized to generate strong BEV features and improve the robustness of the environment perception system in harsh environments; jointly perform 3D detection and local high-precision map generation under a unified framework to realize the perception of dynamic and static information around the vehicle and improve the efficiency of the environment perception system. The present invention provides a cross-modal multi-task environment perception method and system to solve the problems of low efficiency of the current environment perception system and poor robustness to harsh environments, and realize efficient and robust perception of dynamic and static information around the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0043] Figure 1 It is a schematic flowchart of a cross-modal multi-task environment perception method provided by the present invention;
[0044] Figure 2 It is a schematic diagram of the principle of a cross-modal multi-task environment perception system provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] The object of the present invention is to provide a cross-modal multi-task environment perception method and system, which can construct an efficient and robust environment perception system under a unified framework to realize the perception of dynamic and static information around the vehicle.
[0047] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0048] As Figure 1 shown, a cross-modal multi-task environment perception method provided by the present invention includes:
[0049] S101, obtaining observation information; the observation information includes: image information obtained by using an in-vehicle multi-view camera and radar point cloud information obtained by using a lidar.
[0050] S102, using a first feature extraction network to extract multi-scale features of the image and constructing a feature pyramid network.
[0051] As a specific embodiment, the first feature extraction network is a Swin-T network. The Swin-T network is used to extract image features with input sizes of 1 / 4, 1 / 8, and 1 / 16 and construct a feature pyramid network (referred to as L1, L2, and L3 respectively).
[0052] S103, projecting the radar point cloud information onto the image plane to obtain a sparse depth map, and using OpenCV morphological operations to complete the depth of the sparse depth map to obtain a dense depth map.
[0053] S104, using a fully convolutional network to fuse the dense depth map with the feature map of the L3 layer (the deepest layer) in the feature pyramid network to achieve multi-modal information data-level fusion, predicting the context vector and discrete depth probability of each pixel in the image according to the fused features, and projecting them along the camera ray into the 3D space to generate image feature point clouds.
[0054] First, predict the context vector c and |D| uniformly distributed discrete depth probabilities α at pixel p based on the fused features d ; then map pixel p along the camera ray to |D| discrete points, and by using the corresponding discrete depth probability α dScale the context vector c to generate the feature information for each point in the image feature point cloud. The calculation formula is as follows:
[0055] p d = α d × c.
[0056] Among them, p d is the feature information corresponding to the pixel p in the image feature point cloud at the depth d, α d is the discrete depth probability, and c is the context vector at the pixel p.
[0057] S105. Use the bird's-eye view (BEV) pooling operation to transform the image feature point cloud into the BEV space and generate the camera BEV feature.
[0058] Before S105, it also includes:
[0059] Optimize the bird's-eye view pooling using the Precalculation method and the Interval Reduction method.
[0060] S106. Project the radar points onto the image plane to capture the corresponding associated pixels, construct a 3×3 associated region centered on the associated pixels, and use the max pooling operation to extract the associated vector of the associated region.
[0061] S107. Concatenate the radar points with the corresponding associated vectors to achieve multi-modal information data-level fusion, and use the second feature extraction network to extract the feature information of the fused radar point cloud to generate the radar BEV feature; the second feature extraction network is the VoxelNet network.
[0062] S108. Use the attention mechanism to fuse the camera BEV feature and the radar BEV feature in the shared BEV space to achieve multi-modal information BEV-level fusion and generate a strong BEV feature; the strong BEV feature retains both the geometric information of the radar point cloud, the semantic information of the image, and the correlation information of the multi-modal sensors.
[0063] S108 specifically includes the following formula:
[0064]
[0065] Among them, F A is the fused feature map, Q represents the query vector of the radar BEV feature, K and V respectively represent the key and value of the camera BEV feature, Softmax represents non-maximum suppression, and d k is the scaling coefficient of the channel dimension.
[0066] S109. Jointly perform 3D detection and local high-precision map generation on strong BEV features to construct an efficient and robust environment perception system, and realize the perception of dynamic and static information around the vehicle.
[0067] As Figure 2 shown, a cross-modal multi-task environment perception system provided by the present invention includes:
[0068] An observation information acquisition module, configured to acquire observation information; the observation information includes: image information acquired by an in-vehicle multi-view camera and radar point cloud information acquired by a lidar.
[0069] An image feature extraction module, configured to extract multi-scale features of an image by using a first feature extraction network and construct a feature pyramid network.
[0070] An image depth map generation module, configured to project radar point cloud information onto an image plane to obtain a sparse depth map, and perform depth completion on the sparse depth map by using OpenCV morphological operations to obtain a dense depth map.
[0071] An image feature point cloud generation module, configured to fuse the dense depth map with the deepest feature map in the feature pyramid network by using a fully convolutional network to achieve multi-modal information data-level fusion, predict the context vector and discrete depth probability of each pixel in the image according to the fusion features, and project them onto the 3D space along the camera ray to generate an image feature point cloud.
[0072] A camera BEV feature extraction module, configured to convert the image feature point cloud into the BEV space by using a bird's-eye view pooling operation to generate a camera BEV feature;
[0073] An association vector extraction module, configured to project radar points onto the image plane to capture corresponding associated pixels, construct an association region centered on the associated pixels, and extract the association vector of the association region by using a max pooling operation.
[0074] A radar BEV feature extraction module, configured to concatenate the radar points with the corresponding association vectors to achieve multi-modal information data-level fusion, and extract the feature information of the fused radar point cloud by using a second feature extraction network to generate a radar BEV feature.
[0075] A multi-modal feature adaptive fusion module, configured to fuse the camera BEV feature and the radar BEV feature by using an attention mechanism in a shared BEV space to generate a strong BEV feature.
[0076] A multi-task head module, configured to jointly perform 3D detection and local high-precision map generation on the strong BEV feature to construct an efficient and robust environment perception system.
[0077] Among them, the image feature extraction module, the image depth map generation module, the image feature point cloud generation module, and the camera BEV feature extraction module constitute the camera encoder branch; the correlation vector extraction module and the radar BEV feature extraction module constitute the radar encoder branch.
[0078] The present invention discloses the following technical effects:
[0079] 1) The present invention takes the observation information of in-vehicle multi-modal sensors as input, and realizes the perception of dynamic and static information around the vehicle by jointly performing 3D detection and local high-precision map generation in a unified framework, improving the efficiency of the environmental perception system.
[0080] 2) The present invention makes full use of the observation information of multi-modal sensors through depth-guided camera view transformation, region-correlation-based data-level fusion, and attention mechanism-based BEV-level fusion to generate strong BEV features, improving the robustness of the environmental perception system in harsh environments.
[0081] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0082] In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.
Claims
1. A cross-modal multi-task environment perception method, characterized in that Including: Obtain observation information; The observation information includes: image information obtained by using an in-vehicle multi-view camera and radar point cloud information obtained by using a lidar; Use a first feature extraction network to extract multi-scale features of the image and construct a feature pyramid network; Project the radar point cloud information onto the image plane to obtain a sparse depth map, and use OpenCV morphological operations to complete the depth of the sparse depth map to obtain a dense depth map; Use a fully convolutional network to fuse the dense depth map with the deepest feature map in the feature pyramid network to achieve multi-modal information data-level fusion, predict the context vector and discrete depth probability of each pixel in the image according to the fused features, and project them onto the 3D space along the camera ray to generate an image feature point cloud; Use a bird's-eye view pooling operation to transform the image feature point cloud into the BEV space to generate a camera BEV feature; Project the radar points onto the image plane to capture the corresponding associated pixels, construct an associated region centered on the associated pixels, and use a max pooling operation to extract the associated vector of the associated region; Concatenate the radar points with the corresponding associated vectors to achieve multi-modal information data-level fusion, and use a second feature extraction network to extract the feature information of the fused radar point cloud to generate a radar BEV feature; Use an attention mechanism to fuse the camera BEV feature and the radar BEV feature in the shared BEV space to achieve multi-modal information BEV-level fusion and generate a strong BEV feature; Jointly perform 3D detection and local high-precision map generation on the strong BEV feature to construct an environmental perception system.
2. The cross-modal multi-task environment perception method according to claim 1, wherein The first feature extraction network is a Swin-T network.
3. A cross-modal multi-task environment perception method according to claim 1, characterized in that The use of the fully convolutional network to fuse the dense depth map with the deepest feature map in the feature pyramid network to achieve multi-modal information data-level fusion, predict the context vector and discrete depth probability of each pixel in the image according to the fused features, and project them onto the 3D space along the camera ray to generate an image feature point cloud specifically includes The following: p d = α d × c; where p d is the feature information corresponding to pixel p in the image feature point cloud at a depth of d, and α d is the discrete depth probability, and c is the context vector at pixel p.
4. A cross-modal multi-task environment perception method according to claim 1, characterized in that Before the use of the bird's-eye view pooling operation to transform the image feature point cloud into the BEV space to generate a camera BEV feature, it also includes: Optimize the bird's-eye view pooling using the Precalculation method and the Interval Reduction method.
5. A cross-modal multi-task environment perception method according to claim 1, characterized in that The second feature extraction network is a VoxelNet network.
6. A cross-modal multi-task environment perception method according to claim 1, characterized in that The use of the attention mechanism to fuse the camera BEV feature and the radar BEV feature in the shared BEV space to achieve multi-modal information BEV-level fusion and generate a strong BEV feature specifically includes the following formula: Among them, F A is the fused feature map, Q is the query vector of the radar BEV feature, K and V are the key and value of the camera BEV feature respectively, Softmax is non-maximum suppression, and d k is the scaling factor of the channel dimension.
7. A cross-modal multi-task environment perception system, characterized in that, Including: An observation information acquisition module for acquiring observation information; The observation information includes: image information obtained by using an in-vehicle multi-view camera and radar point cloud information obtained by using a lidar; An image feature extraction module for using a first feature extraction network to extract multi-scale features of the image and construct a feature pyramid network; An image depth map generation module for projecting the radar point cloud information onto the image plane to obtain a sparse depth map, and using OpenCV morphological operations to complete the depth of the sparse depth map to obtain a dense depth map; An image feature point cloud generation module, which is used to fuse the dense depth map with the feature map of the deepest layer in the Feature Pyramid Network using a fully convolutional network, achieve multi-modal information data-level fusion, predict the context vector and discrete depth probability of each pixel in the image according to the fused features, and project them onto the 3D space along the camera ray to generate an image feature point cloud; A camera BEV feature extraction module, which is used to convert the image feature point cloud to the BEV space using the Bird's Eye View Pooling operation to generate a camera BEV feature; An associated vector extraction module, which is used to project the radar points onto the image plane to capture the corresponding associated pixels, construct an associated region centered on the associated pixels, and use the max pooling operation to extract the associated vector of the associated region; A radar BEV feature extraction module, which is used to concatenate the radar points with the corresponding associated vectors to achieve multi-modal information data-level fusion, and use the second feature extraction network to extract the feature information of the fused radar point cloud to generate a radar BEV feature; A multi-modal feature adaptive fusion module, which is used to fuse the camera BEV feature and the radar BEV feature in the shared BEV space using the attention mechanism to achieve multi-modal information BEV-level fusion and generate a strong BEV feature; A multi-task head module, which is used to jointly perform 3D detection and local high-precision map generation on the strong BEV feature to construct an environmental perception system.
Citation Information
Patent Citations
Attention-based 4D millimeter wave radar and vision fusion method
CN116129234A
Multi-Task Multi-Sensor Fusion for Three-Dimensional Object Detection
US20200160559A1