Vehicle-road collaborative sensing method and system based on refined fusion vehicle-road characteristics
Through the multi-convolution kernel feature pyramid network and attention mechanism fusion module, the problem of rough feature fusion in vehicle-road collaborative perception is solved, and a higher precision three-dimensional object detection is achieved.
Patent Information
- Application Number
- CN202510449130.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
The existing vehicle-road collaborative perception algorithm is relatively rough in feature fusion, making it difficult to effectively integrate information from vehicles and infrastructure, resulting in insufficient detection accuracy.
The multi-convolution kernel feature pyramid network and the fusion module based on attention mechanism are adopted to alleviate the semantic gap between vehicle-end and road-end features through multi-scale information aggregation and feature interaction, and achieve better detection effects.
It improves the fusion accuracy of vehicle-road collaborative perception and improves the accuracy and consistency of three-dimensional object detection.
Smart Images

Figure CN120298673A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision-based three-dimensional object detection, and particularly to a vehicle-road collaborative perception method and system based on refined fusion of vehicle-road features. Background Art
[0002] In recent years, autonomous driving technology has attracted wide attention from all walks of life due to its outstanding advantages of reducing driver burden and enhancing driving safety. A complete modern autonomous driving vehicle system is usually divided into three major modules: perception, planning, and control. The perception module plays the role of "ears and eyes" in the entire system, mainly responsible for perceiving the driving environment and accurately positioning the vehicle. The perception data it outputs provides reliable data support for the orderly operation of the downstream planning and control modules. Among the numerous vision tasks involved in the perception module, object detection and segmentation, lane line detection, and semantic and instance segmentation are the most common. Among them, three-dimensional object detection occupies a core position in the perception system and is indispensable. This technology focuses on identifying key objects in the driving environment, determining their positions, sizes, and categories, and finally presenting the results in the form of three-dimensional object detection frames.
[0003] In the evolution process of autonomous driving technology, the traditional self-vehicle perception mode has long relied on on-vehicle sensors (such as lidar, cameras, millimeter-wave radars, etc.) to achieve environmental detection, but its physical limitations have become increasingly prominent: on the one hand, due to the installation position limitations of on-vehicle sensors, there are fixed blind spots, making it difficult to handle scenarios such as curves and large vehicle occlusions; on the other hand, extreme environments such as rain, fog, and strong light will seriously weaken the sensor performance, resulting in a decline in object detection accuracy. With the increasing complexity of traffic scenarios and the improvement of high-level autonomous driving requirements, the perception bottleneck of single-vehicle intelligence has become more significant. At the same time, the rapid development of vehicle networking technology, technologies such as 5G / 6G communication and edge computing provide support for data interaction between vehicles, roadside units, and other traffic participants. At the policy level (such as the planning of intelligent transportation systems in various countries), the "single-vehicle intelligence + vehicle-road collaboration" integration path is continuously promoted, prompting the autonomous driving perception mode to upgrade from closed self-vehicle perception to open collaborative perception.
[0004] At present, cooperative perception can be divided into three categories: early cooperation, mid-term cooperation, and late cooperation. Early cooperation utilizes the raw data fusion at the network input. Since the raw data contains the richest information, early cooperation can fundamentally solve the occlusion and long-distance perception problems that plague ego vehicle perception, thereby maximizing performance. In contrast, late cooperation fuses the prediction results of the model. Compared with early cooperation, it has a smaller transmission volume and lower complexity. However, since the detection results of a single agent may be noisy and incomplete, late cooperation usually produces the worst perception performance. Considering the large transmission volume of early cooperation and the poor perception performance of late cooperation, some studies have proposed mid-term cooperation methods. In mid-term cooperation, the infrastructure usually transmits deep semantic features to vehicles. This makes a trade-off between perception performance and transmission volume, and thus becomes the most popular choice for cooperative perception. Previous mid-term cooperation research has mainly focused on how to extract more representative features or how to efficiently transmit information, but has ignored the fact that effectively fusing vehicle-side features and road-side features is actually an extremely important part of cooperative perception. For example, FFNet uses a simple feature concatenation and convolution method to fuse vehicle-road features (Flow-Based Feature Fusion for Vehicle-Infrastructure Cooperative 3DObject Detection, Haibao Yu1 Yingjuan Tang Enze Xie Jilei Mao Ping Luo Zaiqing Nie The University of Hong Kong Institute for AI Industry Research (AIR), Tsinghua University Beijing Institute of Technology 4Shanghai AI Laboratory). This process lacks in-depth exploration of the intrinsic correlation between vehicle-road features, and the fusion method is relatively rough, making it difficult to achieve efficient integration and collaborative use of vehicle-road information. In addition, due to differences in cost and sensor preferences between vehicle manufacturers and roadside infrastructure manufacturers, vehicles and infrastructure may be equipped with sensors of different brands and specifications. Therefore, there are differences in the semantic information of features extracted from vehicles and infrastructure. In order to improve the perception capability of the entire system, how to effectively fuse the features of vehicles and infrastructure is an important issue that needs to be solved urgently. Summary of the invention
[0005] Since the current vehicle-road collaborative perception algorithm is relatively rough in feature fusion, the present invention proposes a vehicle-road collaborative perception method and system based on refined fusion of vehicle-road features, which provides a richer receptive field through a multi-convolution kernel feature pyramid network and aggregates multi-scale information. The fusion module based on the attention mechanism effectively fuses vehicle-road features, alleviates the semantic gap between vehicle-end features and road-end features to a certain extent, and achieves a better detection effect.
[0006] The present invention is implemented by at least one of the following technical solutions.
[0007] A vehicle-road collaborative perception method based on refined fusion of vehicle-road features, comprising the following steps:
[0008] Step 1: Obtain the original point cloud of the lidar and use a voxel encoding network to perform feature encoding on the vehicle-end point cloud and the road-end point cloud respectively to obtain the voxel features of the vehicle end and the road end;
[0009] Step 2: Feed the voxel features of the vehicle end and the road end into a 3D object detection backbone network to obtain BEV features of different scales;
[0010] Step 3: Feed the BEV features of different scales into a multi-convolution kernel feature pyramid network to aggregate features of different scales and obtain BEV features containing multi-scale information;
[0011] Step 4: Feed the BEV features containing multi-scale information into a fusion module based on the attention mechanism to obtain fusion features that aggregate vehicle-road information;
[0012] Step 5: Generate 3D object detection results according to the vehicle-road fusion feature map to obtain the specific positions and categories of the targets.
[0013] Further, in step 1, using a voxel encoding network to perform feature encoding on the vehicle-end point cloud and the road-end point cloud respectively includes: discretizing the 3D point cloud space into equal-size voxel units, mapping the point cloud to the corresponding voxel units according to the spatial coordinates; based on the sparse characteristics of the point cloud data, for non-empty voxels, using a sparse voxel feature extractor to perform voxel feature extraction to achieve efficient learning of features.
[0014] Further, in step 2, the 3D object detection backbone network uses multiple sparse 3D convolutional layers to downsample the voxel features of the vehicle end and the road end to obtain 3D feature maps of different scales, perform feature summation on the 3D feature maps of different scales in the height dimension, and convert the multiple 3D feature maps of different scales into multiple BEV feature maps.
[0015] Further, in step 3, the BEV features of different scales are fed into a multi-convolution kernel feature pyramid network to aggregate features of different scales, and the BEV features containing multi-scale information are obtained, which specifically includes:
[0016] The BEV features of each scale are divided into N groups in the channel dimension, where N represents the number of convolution kernels of different sizes. Then, depthwise separable convolution and transposed convolution operations are performed on each group of features using convolution kernels of different sizes. Finally, the BEV features of different scales are concatenated in the channel dimension to obtain features that aggregate multi-scale information.
[0017] Further, in step 3, the formula of the multi-convolution kernel feature pyramid network is as follows:
[0018] F m = Concat(Deconv(DWConv(F i=1,2,3,… )));
[0019] F m is the feature that fuses multi-scale information, F i is the feature of the i-th scale extracted by the 3D object detection backbone network, DWConv(.) is the depthwise separable convolution operation, Deconv(.) is the transposed convolution operation, and Concat(.) is the feature concatenation operation; by performing depthwise separable convolution and transposed convolution operations on the feature F i of each scale, the enhanced multi-scale feature F' i is obtained, and then the F' i=1,2,3,… is aggregated through the feature concatenation operation to obtain a fusion feature containing rich multi-scale information.
[0020] Further, in step 4, the fusion module based on the attention mechanism includes the fusion module based on the attention mechanism including channel attention, spatial attention, and cross attention;
[0021] Among them, the calculation formula of the channel attention is as follows:
[0022] F c = Mlp(MaxPool(F)) + Mlp(AvgPool(F));
[0023] F c is the channel feature extracted by the channel attention, F is the original point cloud feature, MaxPool(.) represents the max pooling function, AvgPool(.) represents the average pooling function, and Mlp(.) represents the multi-layer perceptron; the representation of the feature in the channel dimension is extracted through the channel attention;
[0024] The calculation formula of the cross attention is as follows:
[0025]
[0026] Query is a query, which is generated by vehicle-side channel features or roadside channel features; Key is a key, and Value is a value. Key and Value are generated by roadside channel features or vehicle-side channel features. Here, two cross-attention calculations are required, that is, one is to use vehicle feature F vec_c as the query Query, and roadside feature F inf_c as the key Key and value Value to generate the vehicle-road cross-attention Vec-Inf Cross Attention. The other is to use roadside feature F inf_c as the query Query, and vehicle feature F vec_c as the key Key and value Value to generate the road-vehicle cross-attention Inf-Vec Cross Attention; softmax is a normalization function; the information interaction between vehicle-side features and roadside features is realized through cross-attention;
[0027] The calculation formula of the spatial attention is as follows:
[0028] F s = CNN(MaxPool(F)) + Mlp(AvgPool(F));
[0029] F s is the channel feature extracted by spatial attention. F is the original point cloud feature. MaxPool(.) represents the max pooling function, and AvgPool(.) represents the average pooling function. The two pooling functions act on the spatial dimension of the feature; CNN(.) represents the convolutional neural network; the representation of the feature in the spatial dimension is extracted through spatial attention.
[0030] Furthermore, in step 4, the calculation formula of the fusion module based on the attention mechanism is as follows:
[0031] F′ vec = CrossAttn(CA(F vec ), CA(F inf )) ⊙ F vec + F vec ;
[0032] F′ inf = CrossAttn(CA(F inf ), CA(F vec )) ⊙ F inf + F inf ;
[0033] F Fusion = CNN(Concat(SA(F′ vec ) ⊙ F′vec +F′ vec , SA(F′ inf )⊙F′ inf +F′ inf ));
[0034] Among them, F Fusion is the fused feature obtained after being processed by the fusion module based on the attention mechanism, F′ vec and F′ inf are respectively the vehicle-end feature and the road-end feature enhanced by channel attention and cross attention, F vec and F inf are respectively the original vehicle-end feature and road-end feature. CA(.) represents channel attention, CrossAttn(.) represents cross attention, SA(.) represents spatial attention, and ⊙ represents the Hadamard product.
[0035] Furthermore, in step 5, according to the fused feature, a three-dimensional object detection result is generated to obtain the specific position and category of the object, specifically including:
[0036] Using three Anchor3DHead detection heads to respectively obtain detection results of three major categories from the fused feature; among them, the detection result of each major category includes a heat map representing the position and type of the object center point, the horizontal offset of the object center point, the height of the object center point, the size and orientation of the object.
[0037] To implement the system of a vehicle-road collaborative perception method based on refined fusion of vehicle-road features, including:
[0038] A data acquisition module for obtaining the point cloud voxel features of the vehicle end and the road end;
[0039] A three-dimensional object detection backbone network module, where the point cloud voxel features of the vehicle end and the road end are downsampled multiple times through the three-dimensional object detection backbone network to obtain BEV features of multiple different scales;
[0040] A multi-convolution kernel feature pyramid network module, using the multi-convolution kernel feature pyramid network to aggregate BEV features of different scales into features of different scales, obtaining BEV features containing multi-scale information, for providing a richer receptive field and effectively aggregating multi-scale information;
[0041] A fusion module based on the attention mechanism, including a channel attention module, a spatial attention module and a cross attention module, fusing the vehicle-road features through multiple attention modules to obtain features aggregating vehicle-road information;
[0042] A detection module, including multiple detection heads, and generating a three-dimensional object detection result from the features aggregating vehicle-road information through the multiple detection heads.
[0043] A computer device of the present invention includes: a memory, a processor, and a computer program stored on the memory. When the computer program is executed on the processor, the described vehicle-road collaborative perception method based on refined fusion of vehicle-road features is implemented.
[0044] Compared with the existing technologies, the beneficial effects of the present invention are as follows:
[0045] A vehicle-road collaborative perception method based on refined fusion of vehicle-road features provided by the present invention sends the point cloud voxel features into a three-dimensional object detection backbone network to obtain BEV features of different scales; uses a multi-convolution kernel feature pyramid network with rich receptive fields to aggregate features of different scales to obtain BEV features containing multi-scale information; sends the BEV features containing multi-scale information into a fusion module based on an attention mechanism to obtain fusion features aggregating vehicle-road information; and sends the fusion features into a detection head to generate three-dimensional object detection results, obtaining the specific positions and categories of the targets. The multi-convolution kernel feature pyramid network provides a richer receptive field and aggregates multi-scale information, while the fusion module based on the attention mechanism effectively fuses vehicle-road features, alleviating to a certain extent the semantic gap between vehicle-end features and road-end features and achieving a better detection effect.
[0046] The present invention realizes a more refined fusion of vehicle-road features. Aiming at the problem of differences in semantic information of features extracted from vehicles and infrastructure, an attention mechanism-based fusion module is designed to adaptively fuse vehicle-road features, and a multi-convolution kernel feature pyramid network provides a richer receptive field, effectively aggregating multi-scale information, ultimately achieving a higher fusion accuracy and a better detection effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the following-described drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0048] Figure 1 It is a flowchart of a vehicle-road collaborative perception method based on refined fusion of vehicle-road features of the present invention.
[0049] Figure 2 It is a schematic diagram of the overall framework of a vehicle-road collaborative perception method based on refined fusion of vehicle-road features of the present invention.
[0050] Figure 3 It is a schematic diagram of the process of the multi-convolution kernel feature pyramid network provided by the present invention.
[0051] Figure 4 Schematic diagram of the fusion module based on the attention mechanism provided by the present invention.
[0052] Figure 5 Schematic diagram of various attention mechanisms provided by the present invention. Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] The purpose of the present invention is to propose a vehicle-road collaborative perception method based on refined fusion of vehicle-road features to solve the problems that the receptive field scale of the existing method is relatively single and the vehicle-road feature fusion method is too rough, and to achieve better detection effects.
[0055] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0056] Embodiment 1
[0057] See Figures 1 - 5 , a vehicle-road collaborative perception method based on refined fusion of vehicle-road features in this embodiment includes the following steps:
[0058] Step 1: Obtain the original point cloud of the lidar and use a voxel encoding network to perform feature encoding on the vehicle endpoint cloud and the road endpoint cloud respectively to obtain the voxel features of the vehicle end and the road end:
[0059] Discretize the three-dimensional point cloud space into equal-size voxel units, and map the point cloud to the corresponding voxel units according to the spatial coordinates; based on the sparse characteristics of the point cloud data, the vast majority of voxel units are empty, and only a small number of voxels contain valid point cloud data. For non-empty voxels, a sparse voxel feature extractor (Voxel Feature Extractor) is used to extract voxel features to achieve efficient feature learning.
[0060] Step 2: Feed the voxel features of the point cloud into a three-dimensional object detection backbone network, and obtain BEV features of different scales through multiple downsampling methods.
[0061] Step 3: Feed the BEV features of different scales into a multi-convolution kernel feature pyramid network to aggregate features of different scales, and obtain BEV features containing multi-scale information;
[0062] The formula of the multi-convolution kernel feature pyramid network is as follows:
[0063] F m = Concat(Deconv(DWConv(F i=1,2,3,… )));
[0064] F m is the feature that aggregates multi-scale information, and F i is the feature of the i-th scale extracted by the 3D object detection backbone network. DWConv(.) is the depthwise separable convolution operation, Deconv(.) is the deconvolution operation, and Concat(.) is the feature concatenation operation. By performing the depthwise separable convolution and deconvolution operations on the feature F i of each scale, the enhanced multi-scale feature F' i is obtained, and then the feature F' i=1,2,3,… is aggregated through the feature concatenation operation to obtain the aggregated feature containing rich multi-scale information.
[0065] As an embodiment, the BEV features of each scale are divided into N groups in the channel dimension, where N represents the number of convolution kernels of different sizes. For example, a combination of convolution kernel sizes (3×3, 5×5) is selected, and at this time N = 2. Then, depthwise separable convolution and deconvolution operations are performed on each group of features using convolution kernels of different sizes, and finally, the BEV features of different scales are concatenated in the channel dimension to obtain the feature aggregating multi-scale information.
[0066] Step 4: Send the BEV feature containing multi-scale information into the fusion module based on the attention mechanism to obtain the fusion feature aggregating vehicle-road information.
[0067] The fusion module based on the attention mechanism is designed based on multiple attention mechanisms (channel attention, spatial attention, and cross attention). The fusion module based on the attention mechanism includes a channel attention module, a cross attention module, and a spatial attention module. First, the channel attention module respectively performs feature screening on the vehicle-end feature and the road-end feature in the channel dimension to generate the representations of the vehicle-end and road-end features in the channel dimension. Then, the cross attention module is used to construct a bidirectional interaction mechanism, calculate the attention matrices from the vehicle-end to the road-end and from the road-end to the vehicle-end, and through the Hadamard element-wise product operation, fuse the attention matrices with the corresponding original features to obtain the vehicle-end and road-end features after interaction enhancement. Subsequently, the vehicle-end and road-end features after interaction enhancement are input into the spatial attention module, and through the key region focusing processing in the spatial dimension, the features are further enhanced. Finally, the vehicle-end and road-end features after enhancement are concatenated in the channel dimension, and the feature fusion is completed through the convolution operation to form the final feature representation aggregating vehicle-road collaborative information.
[0068] The addition of channel attention and spatial attention enables the network to pay more attention to features with important information, while cross-attention is used for the interaction between vehicle-side features and road-side features. This module effectively alleviates the semantic gap between features and further aggregates vehicle-road features.
[0069] Among them, the calculation formula of the channel attention is as follows:
[0070] F c = Mlp(MaxPool(F)) + Mlp(AvgPool(F))
[0071] F c is the channel feature extracted by channel attention, F is the original point cloud feature, MvgPool(.) represents the max pooling function, AvgPool(.) represents the average pooling function, and here these two pooling functions act on the channel dimension of the feature. Mlp(.) represents the multi-layer perceptron. The representation of the feature in the channel dimension is extracted through channel attention.
[0072] The calculation formula of the cross-attention is as follows:
[0073]
[0074] Query is the query, which is generated from the vehicle-side channel feature or the road-side channel feature, Key is the key, and Value is the value, both of which are generated from the road-side channel feature or the vehicle-side channel feature. Here, two cross-attentions need to be calculated, that is: one is to use the vehicle feature F vec_c as the query (Query), and the roadside feature F inf_c as the key (Key) and value (Value) to generate the vehicle-road cross-attention (Vec-Inf Cross Attention), and the other is to use the roadside feature F inf_c as the query (Query), and the vehicle feature F vec_c as the key (Key) and value (Value) to generate the road-vehicle cross-attention (Inf-Vec CrossAttention). softmax is the normalization function. The information interaction between vehicle-side features and road-side features is realized through cross-attention.
[0075] The calculation formula of the spatial attention is as follows:
[0076] F s = CNN(MaxPool(F)) + Mlp(AvgPool(F))
[0077] F sis the channel feature obtained by spatial attention extraction. F is the original point cloud feature. MaxPool(.) represents the max pooling function, and AvgPool(.) represents the average pooling function. Here, these two pooling functions act on the spatial dimension of the feature. CNN(.) represents the convolutional neural network. The feature representation in the spatial dimension is extracted through spatial attention.
[0078] The overall calculation formula of the fusion module based on the attention mechanism is as follows:
[0079] F′ vec = CrossAttn(CA(F vec ), CA(F inf )) ⊙ F vec + F vec
[0080] F′ inf = CrossAttn(CA(F inf ), CA(F vec )) ⊙ F inf + F inf
[0081] F Fusio = CNN(Concat(SA(F′ vec ) ⊙ F′ vec + F′ vec , SA(F′ inf ) ⊙ F′ inf + F′ inf ))
[0082] F Fusion is the fusion feature obtained after being processed by this module. F′ vec and F′ inf are the vehicle-end feature and road-end feature enhanced by channel attention and cross attention respectively. F vec and F inf are the original vehicle-end feature and road-end feature respectively. CA(.) represents channel attention, CrossAttn(.) represents cross attention, SA represents spatial attention, and ⊙ represents the Hadamard product. Thus, the fusion feature that fully aggregates vehicle-end information and road-end information is obtained.
[0083] Step 5: Send the fusion feature into the detection head to generate 3D object detection results, obtaining the specific position and category of the object, specifically including: using three Anchor3DHead detection heads to obtain detection results of three major categories from the fusion feature respectively; where the detection result of each major category includes a heat map representing the position and type of the object center point, the horizontal offset of the object center point, the height of the object center point, the size and orientation of the object.
[0084] Example 2
[0085] A vehicle-road collaborative perception method based on refined integration of vehicle and road features in this embodiment includes the following steps:
[0086] Step 1: Obtain the original point cloud map of the lidar and use a voxel encoding network to perform feature encoding on the vehicle endpoint cloud and the road endpoint cloud respectively to obtain the voxel features of the vehicle end and the road end;
[0087] The lidar, as a commonly used sensor in the autonomous driving perception system, plays a key role in the 3D object detection algorithm. In the autonomous driving perception system, the lidar generates high-density point cloud data containing three-dimensional coordinates (x, y, z) and reflection intensity r through multi-beam laser high-frequency scanning of the environment. Its centimeter-level spatial positioning accuracy and anti-environmental interference characteristics provide accurate object spatial position, size profile, and motion trajectory information for the 3D object detection algorithm, and it is the core sensor for realizing vehicle safety decision-making. Step 1 divides the point cloud map into multiple voxels, performs feature encoding and feature extraction on the voxels to obtain the features of each voxel point cloud, specifically including:
[0088] Step 1.1: Perform voxel division on the entire three-dimensional space occupied by the original point cloud map, and perform spatial classification according to the three-dimensional coordinates of the point cloud, and classify the point cloud into the corresponding voxels.
[0089] First, filter the point cloud data from the specified spatial range, and perform voxel division on the three-dimensional space at a resolution of D×W×H. Where D represents the depth dimension, W represents the width dimension, and H represents the height dimension. Then, perform voxel classification according to the three-dimensional coordinates of the point cloud, allocate the point cloud to the corresponding voxels, and only retain the non-empty voxels with the number of internal point clouds greater than zero.
[0090] Step 1.2: For the voxels with the number of point clouds in the voxel exceeding the threshold T, use the random downsampling method to retain T points to avoid information redundancy; for the voxels with less than T points, perform filling through zero-padding technology to ensure that all voxels have a unified structured data format, thereby eliminating the sparsity difference of the original point cloud distribution and providing a standardized input for subsequent convolutional neural network processing.
[0091] Step 1.3: Aggregate all the point clouds in the voxel through the coordinate mean pooling operation, use the three-dimensional coordinate mean as the initial voxel feature vector, and then input it into the three-dimensional sparse convolutional layer for feature learning, use sparse calculation to efficiently extract the voxel features, and finally output the point cloud voxel feature F V 。
[0092] Step 2: Send the point cloud voxel feature into the three-dimensional object detection backbone network to obtain BEV features of different scales, specifically including:
[0093] Step 2.1: Use multiple sparse 3D convolutional layers to downsample the voxel features of the point cloud to obtain 3D feature maps of different scales Among them, the sparse 3D convolutional layer consists of a 3D sparse convolution, a regularization function, and an activation function. C represents the number of channels of the feature map, represents the scale of the feature map
[0094] Step 2.2: For the 3D feature maps of different scales, perform feature summation in the height dimension D to convert the 3D feature maps of multiple different scales into multiple BEV feature maps of different scales
[0095] Step 3: Feed the BEV features of different scales into a multi-convolution kernel feature pyramid network to aggregate features of different scales and obtain BEV features containing multi-scale information. This step is as follows Figure 3 shown, specifically including:
[0096] Step 3.1: For the BEV feature maps of different scales, divide the features into N groups in the channel dimension C, and perform depthwise separable convolution on each group of features using convolution kernels of different sizes to obtain corresponding features
[0097] Step 3.2: Perform upsampling operations on the features obtained by the above depthwise separable convolution respectively to obtain enhanced features
[0098] Step 3.3: Concatenate the enhanced features of different scales in the channel dimension to obtain features containing multi-scale information
[0099] Step 4: Feed the BEV features containing multi-scale information into a fusion module based on an attention mechanism to obtain fusion features that aggregate vehicle-road information. This step is as follows Figure 4 and Figure 5 shown, specifically including:
[0100] Denote the feature F ms extracted from the vehicle endpoint cloud as F vec and denote the feature F ms extracted from the road endpoint cloud as F inf .
[0101] Use the channel attention module to extract the representation of the vehicle endpoint feature and the road endpoint feature in the channel dimension respectively
[0102] Subsequently, use the cross-attention module to calculate the vehicle-road attention matrix Vehicle-Road Attention Matrix
[0103] Immediately afterwards, the Vehicle-Road Attention Matrix and the Road-Vehicle Attention Matrix are respectively subjected to Hadamard product with the corresponding original features to obtain enhanced vehicle-end features and road-end features
[0104] Then, the enhanced vehicle-end features and road-end features are fed into the spatial attention module for further feature enhancement to obtain features
[0105] Finally, the features obtained through feature enhancement are concatenated in the channel dimension, and the features are fused by means of convolution to obtain the final features aggregating vehicle-road information
[0106] Step 5: Feed the final features aggregating vehicle-road information F fusion into the detection head to generate 3D object detection results, obtaining the specific positions and categories of the objects
[0107] Use a fusion module based on the attention mechanism to fuse vehicle-road features to obtain the features aggregating vehicle-road information F fusion After that, use the AnchorHead detection head to obtain detection results for three major categories from the fused features respectively; among them, the detection results for each major category include heatmaps representing the positions and types of object centers, horizontal offsets of object centers, heights of object centers, sizes and orientations of objects
[0108] In Step 5, feeding the fused features into the detection head to generate 3D object detection results, obtaining the specific positions and categories of the objects, specifically including the following steps
[0109] Step 5.1: Use three detection heads to detect the Car, Pedestrian, and Cyclist classes in the DAIR-V2X dataset respectively, so that each detection head can focus on specific classification targets, and the specific classification is shown in Table 1
[0110] Table 1 Target Detection Head Classification Table
[0111] Detection head type Detection target Head_0 Car Head_1 Pedestrian Head_2 Cyclist
[0112] Step 5.2: Use three anchor-based detection heads, AnchorHead, to obtain the detection results of three types of targets from the vehicle-road fusion feature map. The detection results of each type of target include the heat map of the target center point position and category, the size of the target, the offset of the target center point, the size and orientation of the target. Add the detection results of each type to the detection result list to form the final detection result.
[0113] Step 5.3: Integrate the category and position parameters of the target bounding box based on the detection results of the three types of targets as the 3D target detection results.
[0114] Finally, integrate the category and position parameters (x, y, z, w, l, h, θ) of the target bounding box from the detection results of the above three types of targets. Among them, (x, y, z) are the 3D coordinates of the target center point, (w, l, h) are the length, width, and height of the target, and θ is the yaw angle (heading angle) of the target.
[0115] A vehicle-road collaborative perception method based on refined vehicle-road feature fusion proposed by the present invention can be regarded as a vehicle-road collaborative 3D target detection model based on refined vehicle-road feature fusion as a whole. Its loss function is defined as the sum of classification loss, orientation loss, and regression loss:
[0116] Loss = λ1L reg + λ2L dir + λ3L cls (1)
[0117] Where L cls is the classification loss, L dir is the orientation loss, L reg is the regression loss, and λ1, λ2, λ3 represent the balance coefficients of the three types of losses. The classification loss L cls acts on the heat map of the predicted output. To address the problem of unbalanced positive and negative samples, the focal loss function Focal Loss is used.
[0118] Next, experiments are conducted to verify the technical effects of the method of the present invention.
[0119] As an example, the publicly available dataset DAIR-V2X is used for experiments. This dataset is a large-scale, real-world vehicle-road collaborative perception dataset, containing images and point clouds from more than 100 scenarios and 18,000 pairs of data points. These data are collected by the perception units of autonomous vehicles and roadside infrastructure respectively. When an autonomous vehicle crosses an intersection, data can be captured simultaneously from the perspectives of both infrastructure and vehicle sensors. Collaborative 3D annotations of the vehicle-infrastructure view are provided for 9311 pairs of data. The dataset is divided into a training set, a validation set, and a test set in a ratio of 5:2:3, and the model evaluation is carried out on the validation set.
[0120] The model training of the present invention uses the adam optimizer, with the initial learning rate set to 0.01. It is trained using 8 RTX3090s, with the batch size of each card set to 2, and a total of 40 epochs are trained, and the model converges well.
[0121] In terms of model evaluation, the three-dimensional mean average precision mAP@3D and the bird's-eye view mean average precision mAP@BEV metrics proposed by the DAIR-V2X dataset are used for evaluation. To fully verify the excellent performance of the method of the present invention, it is compared with the baseline algorithm and other representative detection models (V2VNet, V2X-ViT, Where2comm, FFNet). The experimental comparison results are shown in Table 2.
[0122] Table 2 Performance evaluation table of the method of the present invention and existing models in the field
[0123] Method / model name mAP@3D mAP@BEV V2VNet 52.02 60.78 V2X-ViT 53.57 59.36 Where2comm 49.36 57.34 FFNet 55.48 63.14 The method (model) of the present invention 57.47 67.24
[0124] Table 2 gives the experimental result comparison between the method of the present invention and other representative models in the field. It can be seen from the data in Table 2 that the mAP@3D and mAP@BEV metrics of the method of the present invention are higher than those of other representative models in the field, which is feasible and effective.
[0125] The method proposed by the present invention introduces a multi-convolution kernel feature pyramid network in the process of vehicle-road feature fusion. The multi-convolution kernel feature pyramid network provides a richer receptive field and aggregates multi-scale information, while the fusion module based on the attention mechanism effectively fuses the vehicle-road features, alleviating the semantic gap between the vehicle-side features and the road-side features to a certain extent and achieving a better detection effect.
[0126] In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A vehicle-road collaborative perception method based on refined integration of vehicle and road features, characterized in that It includes the following steps: Step 1: Obtain the original point cloud of the lidar and use a voxel encoding network to perform feature encoding on the vehicle endpoint cloud and the road endpoint cloud respectively to obtain the voxel features of the vehicle end and the road end; Step 2: Feed the voxel features of the vehicle end and the road end into a 3D object detection backbone network to obtain BEV features of different scales; Step 3: Feed the BEV features of different scales into a multi-convolution kernel feature pyramid network to aggregate features of different scales and obtain a BEV feature containing multi-scale information; Step 4: Feed the BEV feature containing multi-scale information into a fusion module based on the attention mechanism to obtain a fusion feature that aggregates vehicle-road information; Step 5: Generate a 3D object detection result based on the vehicle-road fusion feature map to obtain the specific position and category of the target.
2. The vehicle-road collaborative perception method based on refined integration of vehicle and road features according to claim 1, characterized in that In Step 1, using the voxel encoding network to perform feature encoding on the vehicle endpoint cloud and the road endpoint cloud respectively includes: discretizing the 3D point cloud space into equal-size voxel units, mapping the point cloud to the corresponding voxel units according to the spatial coordinates; based on the sparse characteristics of the point cloud data, for non-empty voxels, using a sparse voxel feature extractor to perform voxel feature extraction to achieve efficient learning of features.
3. A vehicle-road collaborative perception method based on refined fusion of vehicle and road features according to claim 1, characterized in that, In Step 2, the 3D object detection backbone network uses multiple sparse 3D convolutional layers to downsample the voxel features of the vehicle end and the road end to obtain 3D feature maps of different scales, perform feature summation on the 3D feature maps of different scales in the height dimension, and convert multiple 3D feature maps of different scales into multiple BEV feature maps of different scales.
4. The vehicle-road collaborative perception method based on refined integration of vehicle and road features according to claim 1, wherein, In Step 3, feeding the BEV features of different scales into a multi-convolution kernel feature pyramid network to aggregate features of different scales and obtain a BEV feature containing multi-scale information specifically includes: Dividing the BEV features of each scale into N groups in the channel dimension, where N represents the number of convolution kernels of different sizes, then performing depthwise separable convolution and transposed convolution operations on each group of features using convolution kernels of different sizes, and finally splicing the BEV features of different scales in the channel dimension to obtain a feature that aggregates multi-scale information.
5. The vehicle-road collaborative perception method based on refined integration of vehicle and road features according to claim 1, characterized in that In Step 3, the formula of the multi-convolution kernel feature pyramid network is as follows: F m = Concat(Deconv(DWConv(F i=1,2,3,… ))); F m is a feature that fuses multi-scale information, and F i is the feature of the i-th scale extracted by the 3D object detection backbone network. DWConv(.) is the depthwise separable convolution operation, Deconv(.) is the transposed convolution operation, and Concat(.) is the feature concatenation operation. By performing the depthwise separable convolution and the transposed convolution operations on the feature F i of each scale, the enhanced multi-scale feature F i ′ is obtained. Then, by aggregating F i ′ =1,2,3,… through the feature concatenation operation, a fused feature containing rich multi-scale information is obtained.
6. The vehicle-road collaborative perception method based on refined integration of vehicle and road features according to claim 1, characterized in that, In Step 4, the fusion module based on the attention mechanism includes channel attention, spatial attention, and cross attention; Among them, the calculation formula of the channel attention is as follows: Fc = Mlp(MaxPool(F)) + Mlp(AvgPool(F)); Fc is the channel feature extracted by channel attention, F is the original point cloud feature, MaxPool(.) represents the max pooling function, AvgPool(.) represents the average pooling function, and Mlp(.) represents the multi-layer perceptron; the representation of the feature in the channel dimension is extracted through channel attention; The calculation formula of the cross attention is as follows: Query is a query, generated from vehicle-side channel features or roadside channel features; Key is a key, and Value is a value. Key and Value are generated from roadside channel features or vehicle-side channel features. Here, two cross-attention calculations are required, that is, one is to use the vehicle feature F vec_c as the query Query, and the roadside feature F inf_c as the key Key and value Value to generate the vehicle-road cross-attention Vec-InfCrossAttention. The other is to use the roadside feature F inf_c as the query Query, and the vehicle feature F vec_c as the key Key and value Value to generate the road-vehicle cross-attention Inf-VecCrossAttention; softmax is a normalization function; information interaction between vehicle-side features and roadside features is achieved through cross-attention; The calculation formula of the spatial attention is as follows: F s = CNN(MaxPool(F)) + Mlp(AvgPool(F)); F s is the channel feature obtained by spatial attention extraction. F is the original point cloud feature. AvgPool(.) represents the max pooling function, and AvgPool(.) represents the average pooling function. The two pooling functions act on the spatial dimension of the feature. CNN(.) represents the convolutional neural network. The feature extracted by spatial attention represents the feature in the spatial dimension.
7. A vehicle-road collaborative perception method based on refined integration of vehicle and road features according to claim 1, characterized in that, In Step 4, the calculation formula of the fusion module based on the attention mechanism is as follows: F v ′ ec = CrossAttn(CA(F vec ), CA(F inf )) ⊙ F vec + F vec ; F i ′ nf = CrossAttn(CA(F inf ), CA(F vec )) ⊙ F inf + F inf ; F Fusion = CNN(Concat(SA(F v ′ ec )) ⊙ F v ′ ec + F v ′ ec , SA(F i ′ nf )) ⊙ F i ′ nf + F i ′ nf )); Among them, F Fusi is the fused feature obtained after being processed by the attention-based fusion module, F v ′ ec and F i ′ nf are the vehicle-end feature and the road-end feature enhanced by channel attention and cross attention respectively, F vec and F inf are the original vehicle-end feature and road-end feature respectively. CA(.) represents channel attention, CrossAttn(.) represents cross attention, SA(.) represents spatial attention, and ⊙ represents Hadamard product.
8. A vehicle-road collaborative perception method based on refined integration of vehicle and road features according to any one of claims 1 to 7, characterized in that, In Step 5, generating a 3D object detection result based on the fusion feature to obtain the specific position and category of the target specifically includes: Three Anchor3DHead detection heads are used to obtain detection results of three major categories from the fused features respectively; among them, the detection results of each major category include a heat map representing the position and category of the target center point, the horizontal offset of the target center point, the height of the target center point, the size and orientation of the target.
9. A system for implementing the vehicle-road collaborative perception method based on refined integration of vehicle-road features according to claim 8, characterized in that, It includes: A data acquisition module for acquiring the point cloud voxel features of the vehicle end and the road end; A three-dimensional object detection backbone network module. The point cloud voxel features of the vehicle end and the road end are downsampled multiple times through the three-dimensional object detection backbone network to obtain BEV features of multiple different scales; A multi-convolution kernel feature pyramid network module. The multi-convolution kernel feature pyramid network is used to aggregate the BEV features of different scales into features of different scales to obtain BEV features containing multi-scale information, which is used to provide a richer receptive field and effectively aggregate multi-scale information; A fusion module based on the attention mechanism, including a channel attention module, a spatial attention module and a cross-attention module. The vehicle-road features are fused through multiple attention modules to obtain features that aggregate vehicle-road information; A detection module, including multiple detection heads. The features that aggregate vehicle-road information generate three-dimensional object detection results through multiple detection heads.
10. A computer device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory. When the computer program is executed on the processor, it implements a vehicle-road collaborative perception method based on refined fusion of vehicle-road features as described in any one of claims 1 to 8.
Citation Information
Cited By
Target detection method, electronic equipment, storage medium and vehicle
CN121095537A
Vehicle-road cooperative auxiliary driving method and system for agglomerate fog scene
CN121122051A
Inter-vehicle communication blind compensation method and system based on laser radar point cloud target detection and semantic joint source channel coding, and storage medium
CN122457628A