Intelligent driving assistance method, device, equipment, storage medium and product
By generating image pseudo point clouds and laser point cloud feature maps and using the attention mechanism for feature fusion, the problem of the inability to effectively fuse camera images and laser point cloud data in existing technologies is solved, and intelligent driving path planning in complex scenarios is realized.
Patent Information
- Application Number
- CN202511053894.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing technologies are unable to effectively fuse camera images and laser point cloud data, making it difficult for intelligent driving systems to meet path planning requirements in complex scenarios.
By acquiring camera images and lidar point cloud data, image pseudo point cloud and laser point cloud feature map are generated, feature fusion is performed using the attention mechanism, and multi-scale sampling and feature splicing are performed to generate the final features to guide path planning.
The effective fusion of camera images and laser point cloud data meets the needs of intelligent driving in complex scenarios and achieves high-precision path planning and obstacle recognition.
Smart Images

Figure CN120552896B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent driving technology, and more specifically, to an intelligent driving assistance method, device, equipment, storage medium, and product. Background Art
[0002] With the booming development of intelligent driving technology, intelligent driving assistance systems in underground parking scenarios face many challenges. Underground parking spaces are usually narrow, with complex situations such as right-angle turns, steep slopes, and multi-story structures, which makes vehicle path planning more difficult. At present, intelligent driving assistance systems mainly rely on sensors such as lidar and cameras to perceive the surrounding environment. Although some existing technologies attempt to fuse lidar point cloud data with camera image data to assist vehicles in achieving intelligent driving, traditional fusion methods often simply splice the two data together or perform fusion processing at the decision-making level. This processing method cannot deeply explore the inherent relationship between the two types of data, resulting in difficulty in accurately supplementing the semantic information missing from the laser point cloud with image features, making it difficult to meet the actual driving needs of intelligent driving systems in underground parking scenarios. Summary of the Invention
[0003] Based on this, the present invention provides an intelligent driving assistance method, device, equipment, storage medium and product to solve the defect of the existing technology that it is difficult to meet the intelligent driving needs in complex scenarios due to the inability to effectively fuse camera images and laser point cloud data.
[0004] To achieve the above objectives, an embodiment of the present invention provides an intelligent driving assistance method, comprising:
[0005] Acquire a camera image captured by a camera and laser point cloud data collected by a laser radar; wherein the camera and the laser radar are installed on a vehicle;
[0006] generating image pseudo point cloud data based on the camera image;
[0007] Converting the image pseudo point cloud data and the laser point cloud data into an image feature map and a laser point cloud feature map respectively;
[0008] Generate a query vector required for the attention mechanism based on the image feature map; generate a key vector and a value vector required for the attention mechanism based on the laser point cloud feature map; based on the attention mechanism, integrate the image feature map into the laser point cloud feature map according to the query vector, the key vector and the value vector to generate a fusion feature;
[0009] Performing multi-scale sampling and feature splicing processing on the fusion features in sequence to generate final features;
[0010] A map is constructed based on the final features, a path is planned for the vehicle based on the map to generate a target path, and the vehicle is controlled to perform intelligent driving assistance operations based on the target path.
[0011] To achieve the above objectives, an embodiment of the present invention further provides an intelligent driving assistance device, comprising:
[0012] A data acquisition module, configured to acquire camera images captured by a camera and laser point cloud data captured by a laser radar; wherein the camera and the laser radar are provided on the vehicle;
[0013] A data generation module, configured to generate image pseudo point cloud data based on the camera image;
[0014] A feature map generation module, used to convert the image pseudo point cloud data and the laser point cloud data into an image feature map and a laser point cloud feature map respectively;
[0015] a fusion module, configured to generate a query vector required for an attention mechanism based on the image feature map; generate a key vector and a value vector required for the attention mechanism based on the laser point cloud feature map; and, based on the attention mechanism, fuse the image feature map into the laser point cloud feature map according to the query vector, the key vector, and the value vector to generate a fused feature;
[0016] A feature conversion module is used to perform multi-scale sampling and feature splicing processing on the fusion features in sequence to generate final features;
[0017] A driving assistance module is used to construct a map based on the final features, plan a path for the vehicle based on the map, generate a target path, and control the vehicle to perform intelligent driving assistance operations based on the target path.
[0018] To achieve the above-mentioned objectives, an embodiment of the present invention further provides an intelligent driving assistance device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the intelligent driving assistance method as described in any of the above-mentioned embodiments is implemented.
[0019] To achieve the above-mentioned purpose, an embodiment of the present invention further provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the intelligent driving assistance method as described in any of the above embodiments.
[0020] To achieve the above objectives, an embodiment of the present invention further provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the intelligent driving assistance method as described in any of the above embodiments.
[0021] Compared with the prior art, the intelligent driving assistance method, device, equipment, storage medium and product disclosed in the embodiments of the present invention first obtain a camera image captured by a camera and laser point cloud data captured by a laser radar, wherein the camera and the laser radar are installed on a vehicle; then, image pseudo point cloud data is generated based on the camera image, and the image pseudo point cloud data and the laser point cloud data are converted into an image feature map and a laser point cloud feature map respectively; then, a query vector required for the attention mechanism is generated based on the image feature map, and a key vector and a value vector required for the attention mechanism are generated based on the laser point cloud feature map, and then the image feature map is integrated into the laser point cloud feature map according to the query vector, key vector and value vector using the attention mechanism to generate a fused feature, and the fused feature is sequentially subjected to multi-scale sampling and feature splicing processing to generate a final feature; finally, a map is constructed based on the final feature, a path is planned for the vehicle based on the map, a target path is generated, and the vehicle is controlled to perform an intelligent driving assistance operation based on the target path. It can be seen from this that the embodiment of the present invention obtains the camera image captured by the camera and the laser point cloud data collected by the lidar, generates an image pseudo point cloud from the camera image and converts it into an image feature map, converts the laser point cloud data into a laser point cloud feature map, and uses the attention mechanism to fuse the image feature map and the laser point cloud feature map. After multi-scale sampling and feature splicing processing, the final feature is obtained, which effectively fuses the camera image and the laser point cloud data. Finally, the final feature is used to guide the vehicle's path planning and guide the vehicle to perform intelligent driving assistance operations, meeting the intelligent driving needs in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for use in the implementation. Obviously, the drawings described below are only some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 This is a flow chart of an intelligent driving assistance method provided by one embodiment of the present invention;
[0024] Figure 2 1 is a schematic structural diagram of an MSFDNet module provided by one embodiment of the present invention;
[0025] Figure 3is a schematic structural diagram of an intelligent driving assistance device provided by one embodiment of the present invention;
[0026] Figure 4 It is a structural diagram of an intelligent driving assistance device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0028] See also Figure 1 , is a flow chart of an intelligent driving assistance method provided by one embodiment of the present invention. Specifically, the intelligent driving assistance method includes steps S1 to S6:
[0029] S1. Acquire a camera image captured by a camera and laser point cloud data collected by a laser radar; wherein the camera and the laser radar are installed on a vehicle;
[0030] S2. generating image pseudo point cloud data based on the camera image;
[0031] S3, converting the image pseudo point cloud data and the laser point cloud data into an image feature map and a laser point cloud feature map respectively;
[0032] S4. Generate a query vector required for the attention mechanism based on the image feature map; generate a key vector and a value vector required for the attention mechanism based on the laser point cloud feature map; based on the attention mechanism, integrate the image feature map into the laser point cloud feature map according to the query vector, the key vector, and the value vector to generate a fusion feature;
[0033] S5, performing multi-scale sampling and feature splicing processing on the fusion features in sequence to generate final features;
[0034] S6. Construct a map based on the final features, perform path planning for the vehicle based on the map, generate a target path, and control the vehicle to perform intelligent driving assistance operations based on the target path.
[0035] It is worth noting that the vehicle can be a new energy vehicle, a hybrid vehicle, a gasoline vehicle, or a robot (such as a vehicle-moving robot). The specific type of vehicle is not limited here. The intelligent driving assistance method can be executed by the vehicle or by a cloud server, without limitation.
[0036] For example, a vehicle is equipped with multiple sensors, including two lidars and four cameras: one lidar at each front and rear, and one camera at each corner. The front lidar is designated as the primary lidar. By jointly calibrating the cameras and lidars, we obtain the extrinsic parameter matrix for all sensors converted to the primary sensor (i.e., the primary lidar), as well as the intrinsic parameter matrix for the cameras.
[0037] Assume that the coordinate systems of the two lidars are 、 , the camera coordinate system is , the intrinsic parameter matrix of the camera is:
[0038] ;
[0039] in, is the number of pixels in the horizontal direction, is the number of pixels in the vertical direction; is the pixel coordinate of the horizontal principal point of the image, is the vertical pixel coordinate of the principal point of the image, For the The intrinsic parameter matrix of each camera, For the A camera coordinate system;
[0040] The external parameter matrix is:
[0041] ;
[0042] in, Represents data from the camera coordinate system Or from the lidar coordinate system Transform to the main lidar coordinate system The external parameter matrix of represents the rotation matrix, Represents the translation vector.
[0043] In step S2, after feature extraction is performed on the camera image, the feature data is converted based on the camera's extrinsic parameter matrix and intrinsic parameter matrix to obtain three-dimensional (3D) spatial feature data (i.e., image pseudo point cloud data).
[0044] Before step S3, the laser point cloud data of the laser radar is first converted from the coordinate system Convert to coordinate system , to achieve the splicing of the front and rear lidar point cloud data. The coordinate system conversion formula of the laser point cloud data is as follows:
[0045] ;
[0046] in, Indicates that the data is from the lidar coordinate system Transform to the main lidar coordinate system The external parameter matrix of Represents the coordinate system The laser point cloud data under express Data after coordinate conversion.
[0047] After executing steps S1 to S3, execute steps S4 to S6 to convert the image pseudo point cloud data into an image feature map and the laser point cloud data into a laser point cloud feature map. Then, perform a linear transformation on the image feature map and the laser point cloud feature map to generate the query vector, key vector, and value vector required by the attention mechanism. The image feature map is integrated into the laser point cloud feature map using the attention mechanism to obtain a fused feature. The fused feature is then subjected to multi-scale sampling and feature splicing processing to generate the final feature. The final feature is then used to construct a high-precision map to identify the vehicle's environment. Based on the constructed map, a path is planned for the vehicle according to the obtained current and destination locations of the vehicle, a target path is generated, and the vehicle is controlled to travel along the target path. Furthermore, during vehicle driving, the intelligent driving assistance method can be used to achieve obstacle avoidance.
[0048] Compared with the existing technology, the embodiment of the present invention obtains camera images captured by a camera and laser point cloud data collected by a lidar, generates an image pseudo point cloud from the camera image and converts it into an image feature map, converts the laser point cloud data into a laser point cloud feature map, and uses the attention mechanism to fuse the image feature map and the laser point cloud feature map. After multi-scale sampling and feature splicing processing, the final feature is obtained, which effectively fuses the camera image and the laser point cloud data. Finally, the final feature is used to guide the vehicle's path planning and guide the vehicle to perform intelligent driving assistance operations, meeting the intelligent driving needs in complex scenarios.
[0049] In a preferred embodiment, based on steps S1 to S6, generating image pseudo point cloud data based on the camera image includes:
[0050] Performing multi-scale feature extraction on the camera image to obtain initial image features;
[0051] generating depth feature data based on the initial features of the image;
[0052] The depth feature data is converted into image pseudo point cloud data in the coordinate system of the laser radar according to the pre-acquired intrinsic parameter matrix of the camera and the extrinsic parameter matrix converted from the camera to the laser radar.
[0053] Specifically, a Multi-scale Feature Depth Network (MSFDNet) module is designed. Through the MSFDNet module with shared parameters, the multi-scale features of the camera images captured by four cameras are extracted as the initial image features. Then, depth feature data is generated based on the initial image features. Finally, based on the camera's intrinsic parameter matrix and the extrinsic parameter matrix converted from the camera to the lidar, the depth feature data is converted into three-dimensional space to obtain image pseudo point cloud data.
[0054] Exemplarily, the process of generating image pseudo point cloud data is as follows:
[0055] The camera images acquired by the four cameras are represented as:
[0056] ;
[0057] in, Indicates the camera images from each camera, The dimension is The set of real numbers;
[0058] MSFDNet consists of a multi-scale fusion network MSFNet and a depth analysis network DepNet. MSFNet is used to lightweight extract multi-scale features of images. Multi-scale features are input into DepNet as initial features of the image. DepNet generates deep feature data based on the initial features of the image:
[0059] ;
[0060] ;
[0061] in, express The multi-scale features of express The corresponding image pseudo point cloud data, are the height, width and number of channels respectively, Indicates the dimension The set of real numbers, Indicates the dimension The set of real numbers.
[0062] See also Figure 2 The structural diagram of the MSFDNet module is shown below. The following is a more detailed introduction to the generation process of image pseudo point cloud data:
[0063] MSFNet uses a lightweight sliding window Transformer module (SwinTransformerBlock) as the backbone network, and performs multi-scale feature extraction and fusion processing on camera images through a hierarchical feature extraction mechanism to obtain the initial image features. Divided into 16×16 sub-pixel images, the embedding feature dimension of each sub-pixel image is set to =96, so the resolution is × × Subgraph of , subgraph Local and global features are interacted through SwinTransformerBlock (based on window multi-head self-attention W-MSA and shifted window multi-head self-attention SW-MSA), and the output resolution is × × Feature map . For feature maps The 4×4 blocks are reorganized and the channel projection is expanded by 2 times through the linear neural network after splicing, and then down-sampled through the pooling layer; the down-sampled feature map is processed by SwinTransformerBlock and converted to a resolution of × × Feature map . Feature map Continue feature sampling, and then process the transformation through SwinTransformerBlock to obtain a resolution of × × Feature map . The feature continues to be sampled, and then processed by SwinTransformerBlock to obtain a resolution of × × Feature map Finally, the multi-scale features are obtained using the following formula (Image initial features):
[0064] ;
[0065] ;
[0066] ;
[0067] ;
[0068] in, is upsampling; It is a point convolution module, which includes Conv2d module, BN2d module and SiLU module. Conv2d is a two-dimensional convolution, BN2d is the normalization of the two-dimensional convolution output, and SiLU is the Swish activation function. Indicates the dimension The set of real numbers; Indicates the dimension The set of real numbers; Indicates the dimension The set of real numbers.
[0069] Will Enter DepNet and first generate the discrete depth probability of each pixel:
[0070] ;
[0071] in, Pixel The discrete depth probability of the pixel points is indivual, The dimension is ; The dimension of the discrete features for each pixel; The final dimension is .
[0072] Combining Discrete Deep Probability with Deep Feature Networks Multiply to get discrete depth features :
[0073] ;
[0074] in, is a trainable parameter feature with dimension ;
[0075] Generated by the following formula The corresponding features containing depth information (Deep feature data):
[0076] ;
[0077] Finally, based on the camera's intrinsic parameter matrix and the extrinsic parameter matrix converted from the camera to the lidar, the depth feature data is converted into three-dimensional spatial feature data (image pseudo point cloud data) through spatial transformation:
[0078] ;
[0079] in, Indicates the Pseudo point cloud data of the camera images, represents three-dimensional Euclidean space, Indicates the The intrinsic parameter matrix of each camera, Represents data from the camera coordinate system Convert the external parameter matrix to the laser radar coordinate system. It is worth noting that if there are multiple laser radars, the data needs to be converted from the camera coordinate system to the external parameter matrix of the laser radar coordinate system. Transform to the main lidar coordinate system.
[0080] Preferably, the models, intrinsic parameter matrices, and resolutions of all cameras are the same, and all camera images are processed using the same MSFDNet module to obtain pseudo point cloud data corresponding to each camera.
[0081] In this embodiment, MSFDNet achieves efficient multi-scale extraction of image features through a lightweight architecture with parameter sharing, and completes cross-dimensional conversion from two-dimensional image perspective to three-dimensional space coordinates.
[0082] In a preferred embodiment, based on any of the above embodiments, before converting the image pseudo point cloud data and the laser point cloud data into an image feature map and a laser point cloud feature map respectively, the method further includes:
[0083] Clustering the image pseudo point cloud data and the laser point cloud data respectively to obtain a plurality of image pseudo point cloud sets and a plurality of laser point cloud sets;
[0084] Performing matching processing on the image pseudo point cloud set and the laser point cloud set to obtain a matching relationship between the image pseudo point cloud set and the laser point cloud set;
[0085] Based on the matching relationship, perspective correction is performed on the matched image pseudo point cloud set according to the laser point cloud set, and the image pseudo point cloud data is updated.
[0086] Specifically, the image pseudo point cloud data generated based on the camera image is generated through probabilistic calculations and is therefore not absolute spatial information. To improve accuracy and better achieve alignment and fusion of camera images and laser point cloud data, a perspective correction module is designed. First, clustering is performed on the image pseudo point cloud data and the laser point cloud data in parallel. Then, a matching laser point cloud set is found for each image pseudo point cloud set. The distance information between each pair of matching image pseudo point cloud sets and laser point cloud sets is calculated in space, and the image pseudo point cloud data is corrected using this distance information. It can be seen that in this embodiment, by designing a perspective correction module, the spatial features of the camera image are corrected based on the spatial distance of the laser point cloud data. The similarity calculation of multimodal data is achieved through spatial clustering and distance difference calculation, correcting the spatial deviation of the image pseudo point cloud data generated based on the camera image.
[0087] Furthermore, the matching processing of the image pseudo point cloud set with the laser point cloud set to obtain a matching relationship between the image pseudo point cloud set and the laser point cloud set includes:
[0088] Calculating a first Hausdorff distance between a first image pseudo point cloud set and a first laser point cloud set; wherein the first image pseudo point cloud set is any one of all the image pseudo point cloud sets, and the first laser point cloud set is any one of all the laser point cloud sets;
[0089] If the first Hausdorff distance is less than or equal to a preset distance threshold, the first image pseudo point cloud set matches the first laser point cloud set.
[0090] Exemplarily, the matching relationship between the image pseudo point cloud set and the laser point cloud set is determined in the following manner:
[0091] Laser point cloud data and image pseudo point cloud data Perform clustering to obtain the respective target instance sets:
[0092] ;
[0093] ;
[0094] in, represents the set of all laser point clouds, represents the set of pseudo point clouds of all images, Indicates the A collection of laser point clouds, Indicates the A set of pseudo point clouds of images, Indicates the total number of laser point cloud sets, Indicates the total number of sets of image pseudo point cloud sets;
[0095] For each pair , calculate the Hausdorff distance to measure the spatial distribution difference:
[0096] ;
[0097] ;
[0098] in, express and The Hausdorff distance between Indicates from arrive The one-way Hausdorff distance of Indicates from arrive The one-way Hausdorff distance of express The point in express The point in
[0099] like ≤ , then it is believed that and For the same goal, and Match each other and establish a matching relationship between the two; among them, It is the preset distance threshold, and its specific value is set according to the actual situation.
[0100] Furthermore, based on the matching relationship, performing perspective correction on the matched image pseudo point cloud set according to the laser point cloud set to update the image pseudo point cloud data includes:
[0101] Calculating the centroid of the second image pseudo point cloud set and the centroid of the second laser point cloud set; wherein the second image pseudo point cloud set is any one of all the image pseudo point cloud sets, and the second laser point cloud set is a laser point cloud set that matches the second image pseudo point cloud set;
[0102] Calculating a correction offset of the second image pseudo point cloud set according to the centroid of the second image pseudo point cloud set and the centroid of the second laser point cloud set;
[0103] A translation transformation process is performed on each point in the second image pseudo point cloud set according to the correction offset of the second image pseudo point cloud set to update the image pseudo point cloud data in the second image pseudo point cloud set.
[0104] Exemplarily, the method of using the laser point cloud set to perform perspective correction on the image pseudo point cloud set is as follows:
[0105] For each matching target pair , respectively calculated and The center of mass:
[0106] ;
[0107] ;
[0108] in, 、 and Indicates a point The coordinates of 、 and Indicates a point coordinates of express The center of mass, express The center of mass;
[0109] Calculate the correction offset: ;
[0110] right Each point , apply a translation transformation:
[0111] ;
[0112] The corrected image pseudo point cloud data is .
[0113] In a preferred embodiment, based on any of the above embodiments, converting the image pseudo point cloud data and the laser point cloud data into an image feature map and a laser point cloud feature map respectively includes:
[0114] Performing voxel grid division on the image pseudo point cloud data and the laser point cloud data respectively to obtain pseudo point cloud voxel data and laser point cloud voxel data;
[0115] Based on a preset voxel feature encoder, feature encoding is performed on the pseudo point cloud voxel data and the laser point cloud voxel data respectively to generate pseudo point cloud voxel features and laser point cloud voxel features;
[0116] Three-dimensional sparse convolution processing is performed on the pseudo point cloud voxel features and the laser point cloud voxel features respectively, and the processing results of the three-dimensional sparse convolution processing are converted into a bird's-eye view representation to obtain an image feature map and a laser point cloud feature map.
[0117] For example, we first use laser point cloud data Pseudo point cloud data for images Perform perspective correction to obtain the corrected image pseudo point cloud data , and then for the rectified image pseudo point cloud data and laser point cloud data The specific conversion principle is as follows:
[0118] First, define the point cloud data in a general form The variant of the sparsely embedded convolutional detection network (SECOND) is used to process the point cloud data in general form, setting the voxel size to (0.1m×0.1m×0.2m), the general form of point cloud data is divided into 512×512×16 voxel grids. For each non-empty voxel, the voxel feature encoder (VFE) in the SECOND network is used to encode the voxel feature, and then the feature is transformed by the three-dimensional spatial sparse convolution module (3DSparseConv) and the bird's-eye view reshaping module (Reshape toBEV) in sequence to obtain a feature map represented by a bird's-eye view. The above principle is used to correct the pseudo point cloud data of the image and laser point cloud data Convert and get the image feature map and laser point cloud feature map , the size of the two feature maps is 512×512×16.
[0119] In a preferred embodiment, based on any of the above embodiments, the fusion feature is calculated in the following manner:
[0120] ;
[0121] ;
[0122] in, is the query vector, is the key vector, is the value vector, is the laser point cloud feature map, is the preset scaling factor, is the normalization function, For attention output; is the fusion feature, For splicing / fusion operations, is layer normalization;
[0123] Specifically, the image feature map and the laser point cloud feature map are linearly transformed to obtain the query vector required by the attention mechanism , key vector Sum value vector :
[0124] ;
[0125] ;
[0126] ;
[0127] in, is the image feature map, is the laser point cloud feature map, 、 and is the learnable weight matrix, Indicates the dimension The set of real numbers, represents the dimension of the key vector;
[0128] The information of the image feature map is integrated into the laser point cloud feature map through the following formula:
[0129] ;
[0130] .
[0131] In a preferred embodiment, based on any of the above embodiments, the fusion features are sequentially subjected to multi-scale sampling and feature splicing processing to generate final features, including:
[0132] The fusion features are subjected to multi-scale sampling and feature splicing processing using the following formula:
[0133] ;
[0134] ;
[0135] ;
[0136] ;
[0137] in, is the first scale feature, is the second scale feature, is the third scale feature, is the fusion feature, is the three-dimensional convolution function, is the upsampling operation, is the feature splicing operation, The final feature.
[0138] For example, for the rectified image pseudo point cloud data and laser point cloud data A cross-alignment and fusion module was designed to achieve alignment and fusion in three-dimensional space, outputting high-precision results and assisting in mapping and perception in complex environments such as underground parking lots. The specific methods for data alignment and fusion are as follows:
[0139] 1. Use the attention mechanism to integrate the information of the image feature map into the laser point cloud feature map to obtain the fusion feature ;
[0140] 2. Fusion features Perform multi-scale sampling:
[0141] ;
[0142] ;
[0143] ;
[0144] 3. Output final features:
[0145] ;
[0146] in, Indicates the dimension The set of real numbers, Indicates the dimension The set of real numbers, Indicates the dimension The set of real numbers, It is the feature obtained by multi-scale sampling (such as the first scale feature , second scale features and third-scale features ), Indicates the dimension The set of real numbers, The final feature feature dimension; it can be understood that the final feature is the data that effectively integrates the image semantic features and point cloud spatial features.
[0147] After obtaining the final features, perform the following steps:
[0148] ;
[0149] in, Represents a feedforward neural network (Feed Forward Network), which is the final feature Mapping to the output space of the detection task, that is, converting abstract features into specific target detection results; Represents the detection results, including the size, location information, and semantic category information of vehicles, people, pillars, walls, obstacles, etc. It can provide more accurate spatial features containing semantic information for vehicles in underground parking lots (such as car-moving robots) during the mapping process, and provide lightweight and high-precision perception results during dynamic perception.
[0150] Compared with the prior art, the intelligent driving assistance method disclosed in an embodiment of the present invention first obtains a camera image captured by a camera and laser point cloud data captured by a laser radar, wherein the camera and the laser radar are installed on a vehicle; then, image pseudo point cloud data is generated based on the camera image, and the image pseudo point cloud data and the laser point cloud data are converted into an image feature map and a laser point cloud feature map respectively; then, a query vector required for an attention mechanism is generated based on the image feature map, and a key vector and a value vector required for the attention mechanism are generated based on the laser point cloud feature map, and then the image feature map is integrated into the laser point cloud feature map according to the query vector, the key vector and the value vector using the attention mechanism to generate a fused feature, and the fused feature is sequentially subjected to multi-scale sampling and feature splicing processing to generate a final feature; finally, a map is constructed based on the final feature, a path is planned for the vehicle based on the map, a target path is generated, and the vehicle is controlled to perform an intelligent driving assistance operation based on the target path. It can be seen from this that the embodiment of the present invention obtains the camera image captured by the camera and the laser point cloud data collected by the lidar, generates an image pseudo point cloud from the camera image and converts it into an image feature map, converts the laser point cloud data into a laser point cloud feature map, and uses the attention mechanism to fuse the image feature map and the laser point cloud feature map. After multi-scale sampling and feature splicing processing, the final feature is obtained, which effectively fuses the camera image and the laser point cloud data. Finally, the final feature is used to guide the vehicle's path planning and guide the vehicle to perform intelligent driving assistance operations, meeting the intelligent driving needs in complex scenarios.
[0151] See also Figure 3 , an embodiment of the present invention further provides an intelligent driving assistance device, comprising:
[0152] A data acquisition module 21 is used to acquire camera images captured by a camera and laser point cloud data collected by a laser radar; wherein the camera and the laser radar are installed on the vehicle;
[0153] A data generating module 22 is configured to generate image pseudo point cloud data based on the camera image;
[0154] A feature map generating module 23 is used to convert the image pseudo point cloud data and the laser point cloud data into an image feature map and a laser point cloud feature map respectively;
[0155] A fusion module 24 is configured to generate a query vector required for an attention mechanism based on the image feature map; generate a key vector and a value vector required for the attention mechanism based on the laser point cloud feature map; and, based on the attention mechanism, fuse the image feature map into the laser point cloud feature map according to the query vector, the key vector, and the value vector to generate a fused feature.
[0156] A feature conversion module 25 is used to perform multi-scale sampling and feature splicing processing on the fusion features in sequence to generate final features;
[0157] The driving assistance module 26 is configured to construct a map based on the final features, perform path planning for the vehicle based on the map, generate a target path, and control the vehicle to perform intelligent driving assistance operations based on the target path.
[0158] It is worth noting that the specific working process of the intelligent driving assistance device can refer to the working process of the intelligent driving assistance method described in the above embodiment, which will not be repeated here.
[0159] See also Figure 4 The embodiment of the present invention further provides an intelligent driving assistance device, comprising a processor 31, a memory 32, and a computer program stored in the memory 32 and configured to be executed by the processor 31. When the processor 31 executes the computer program, the steps in the above-mentioned intelligent driving assistance method embodiment are implemented, for example Figure 1 or, the processor 31 implements the functions of the modules in the above-mentioned device embodiments when executing the computer program.
[0160] Exemplarily, the computer program can be divided into one or more modules, and the one or more modules are stored in the memory 32 and executed by the processor 31 to complete the present invention. The one or more modules can be a series of computer program instruction segments that can perform specific functions, and the instruction segments are used to describe the execution process of the computer program in the intelligent driving assistance device. For example, the computer program can be divided into multiple modules. The specific working process of each module can refer to the working process of the intelligent driving assistance model described in the above embodiment, and will not be repeated here.
[0161] The intelligent driving assistance device can be a computing device such as a desktop computer, laptop, PDA, or cloud server. The intelligent driving assistance device may include, but is not limited to, a processor 31 and a memory 32. Those skilled in the art will appreciate that the intelligent driving assistance device may also include input and output devices, network access devices, buses, and the like.
[0162] The processor 31 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor 31 is the control center of the intelligent driving assistance device, connecting various components of the entire intelligent driving assistance device using various interfaces and lines.
[0163] The memory 32 can be used to store the computer programs and / or modules. The processor 31 implements the various functions of the intelligent driving assistance device by running or executing the computer programs and / or modules stored in the memory 32 and accessing the data stored in the memory 32. The memory 32 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as image playback), while the data storage area may store data generated based on the use of the mobile phone. Furthermore, the memory 32 may include high-speed random access memory (RAM) and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0164] If the integrated module of the intelligent driving assistance device is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can also implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by the processor 31, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal, and a software distribution medium.
[0165] An embodiment of the present invention further provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the intelligent driving assistance method as described in any of the above embodiments.
[0166] Compared with the prior art, the intelligent driving assistance method, apparatus, equipment storage medium and product disclosed in the embodiments of the present invention obtain camera images captured by a camera and laser point cloud data collected by a lidar, generate an image pseudo point cloud from the camera image and convert it into an image feature map, convert the laser point cloud data into a laser point cloud feature map, utilize the attention mechanism to perform feature fusion on the image feature map and the laser point cloud feature map, and then obtain the final features through multi-scale sampling and feature splicing processing, effectively fusing the camera image and laser point cloud data. Finally, based on the final features, the vehicle's path planning is guided, and the vehicle is guided to perform intelligent driving assistance operations, thereby meeting the needs of intelligent driving in complex scenarios.
[0167] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. An intelligent driving assistance method, characterized in that: include: Acquire a camera image captured by a camera and laser point cloud data collected by a laser radar; wherein the camera and the laser radar are installed on a vehicle; generating image pseudo point cloud data based on the camera image; Converting the image pseudo point cloud data and the laser point cloud data into an image feature map and a laser point cloud feature map respectively; Generate a query vector required for the attention mechanism based on the image feature map; generate a key vector and a value vector required for the attention mechanism based on the laser point cloud feature map; based on the attention mechanism, integrate the image feature map into the laser point cloud feature map according to the query vector, the key vector and the value vector to generate a fusion feature; Performing multi-scale sampling and feature splicing processing on the fusion features in sequence to generate final features; constructing a map based on the final features, performing path planning for the vehicle based on the map to generate a target path, and controlling the vehicle to perform intelligent driving assistance operations based on the target path; The generating of image pseudo point cloud data based on the camera image comprises: Performing multi-scale feature extraction on the camera image to obtain initial image features; generating depth feature data based on the initial features of the image; The depth feature data is converted into image pseudo point cloud data in the coordinate system of the laser radar according to the pre-acquired intrinsic parameter matrix of the camera and the extrinsic parameter matrix converted from the camera to the laser radar; The converting the image pseudo point cloud data and the laser point cloud data into an image feature map and a laser point cloud feature map respectively comprises: Performing voxel grid division on the image pseudo point cloud data and the laser point cloud data respectively to obtain pseudo point cloud voxel data and laser point cloud voxel data; Based on a preset voxel feature encoder, feature encoding is performed on the pseudo point cloud voxel data and the laser point cloud voxel data respectively to generate pseudo point cloud voxel features and laser point cloud voxel features; Three-dimensional sparse convolution processing is performed on the pseudo point cloud voxel features and the laser point cloud voxel features respectively, and the processing results of the three-dimensional sparse convolution processing are converted into a bird's-eye view representation to obtain an image feature map and a laser point cloud feature map.
2. The intelligent driving assistance method according to claim 1, wherein: Before converting the image pseudo point cloud data and the laser point cloud data into an image feature map and a laser point cloud feature map respectively, the method further includes: Clustering the image pseudo point cloud data and the laser point cloud data respectively to obtain a plurality of image pseudo point cloud sets and a plurality of laser point cloud sets; Performing matching processing on the image pseudo point cloud set and the laser point cloud set to obtain a matching relationship between the image pseudo point cloud set and the laser point cloud set; Based on the matching relationship, perspective correction is performed on the matched image pseudo point cloud set according to the laser point cloud set, and the image pseudo point cloud data is updated.
3. The intelligent driving assistance method according to claim 2, wherein: The matching process of the image pseudo point cloud set and the laser point cloud set to obtain a matching relationship between the image pseudo point cloud set and the laser point cloud set includes: Calculating a first Hausdorff distance between a first image pseudo point cloud set and a first laser point cloud set; wherein the first image pseudo point cloud set is any one of all the image pseudo point cloud sets, and the first laser point cloud set is any one of all the laser point cloud sets; If the first Hausdorff distance is less than or equal to a preset distance threshold, the first image pseudo point cloud set matches the first laser point cloud set; The step of performing perspective correction on the matched image pseudo point cloud set according to the laser point cloud set based on the matching relationship and updating the image pseudo point cloud data includes: Calculating the centroid of the second image pseudo point cloud set and the centroid of the second laser point cloud set; wherein the second image pseudo point cloud set is any one of all the image pseudo point cloud sets, and the second laser point cloud set is a laser point cloud set that matches the second image pseudo point cloud set; Calculating a correction offset of the second image pseudo point cloud set according to the centroid of the second image pseudo point cloud set and the centroid of the second laser point cloud set; A translation transformation process is performed on each point in the second image pseudo point cloud set according to the correction offset of the second image pseudo point cloud set to update the image pseudo point cloud data in the second image pseudo point cloud set.
4. The intelligent driving assistance method according to claim 1, wherein: The fusion feature is calculated in the following way: ; ; in, is the query vector, is the key vector, is the value vector, is the laser point cloud feature map, is the preset scaling factor, is the normalization function, For attention output; is the fusion feature, For splicing / fusion operations, is layer normalization; The step of sequentially performing multi-scale sampling and feature splicing processing on the fusion features to generate final features includes: The fusion features are subjected to multi-scale sampling and feature splicing processing using the following formula: ; ; ; ; in, is the first scale feature, is the second scale feature, is the third scale feature, is the three-dimensional convolution function, is the upsampling operation, is the feature splicing operation, The final feature.
5. An intelligent driving assistance device, characterized in that: include: A data acquisition module, configured to acquire camera images captured by a camera and laser point cloud data captured by a laser radar; wherein the camera and the laser radar are provided on the vehicle; A data generation module, configured to generate image pseudo point cloud data based on the camera image; A feature map generation module, used to convert the image pseudo point cloud data and the laser point cloud data into an image feature map and a laser point cloud feature map respectively; a fusion module, configured to generate a query vector required for an attention mechanism based on the image feature map; generate a key vector and a value vector required for the attention mechanism based on the laser point cloud feature map; and, based on the attention mechanism, fuse the image feature map into the laser point cloud feature map according to the query vector, the key vector, and the value vector to generate a fused feature; A feature conversion module is used to perform multi-scale sampling and feature splicing processing on the fusion features in sequence to generate final features; a driving assistance module, configured to construct a map based on the final features, perform path planning for the vehicle based on the map, generate a target path, and control the vehicle to perform intelligent driving assistance operations based on the target path; The data generation module is specifically used to: Performing multi-scale feature extraction on the camera image to obtain initial image features; generating depth feature data based on the initial features of the image; The depth feature data is converted into image pseudo point cloud data in the coordinate system of the laser radar according to the pre-acquired intrinsic parameter matrix of the camera and the extrinsic parameter matrix converted from the camera to the laser radar; The feature map generation module is specifically used to: Performing voxel grid division on the image pseudo point cloud data and the laser point cloud data respectively to obtain pseudo point cloud voxel data and laser point cloud voxel data; Based on a preset voxel feature encoder, feature encoding is performed on the pseudo point cloud voxel data and the laser point cloud voxel data respectively to generate pseudo point cloud voxel features and laser point cloud voxel features; Three-dimensional sparse convolution processing is performed on the pseudo point cloud voxel features and the laser point cloud voxel features respectively, and the processing results of the three-dimensional sparse convolution processing are converted into a bird's-eye view representation to obtain an image feature map and a laser point cloud feature map.
6. An intelligent driving assistance device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the intelligent driving assistance method according to any one of claims 1 to 4 when executing the computer program.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the intelligent driving assistance method according to any one of claims 1 to 4.
8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the intelligent driving assistance method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Cognitive map construction method fusing image and laser point cloud
CN116310176A
Object detection using radar and lidar fusion
US20230109909A1