Vehicle-road multi-modal data instance-level feature transmission cooperative perception system and method, and storage medium

The multimodal vehicle-road cooperative perception system, which incorporates unified vehicle-road location modeling, dynamic channel fusion, and high-confidence target screening modules, solves the problems of limited information transmission bandwidth and insufficient multimodal data fusion accuracy in vehicle-road cooperative perception systems, and achieves efficient and accurate transmission of vehicle-road cooperative perception information.

CN118741460BActive Publication Date: 2025-11-07JIANGSU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410715790.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-04
Publication Date
2025-11-07
Estimated Expiration
2044-06-04

AI Technical Summary

Technical Problem

Existing vehicle-road cooperative perception systems suffer from bandwidth limitations, insufficient accuracy in multimodal data fusion, and redundancy control issues during information transmission. In particular, they cannot effectively transmit vehicle-road cooperative perception information in high-concurrency scenarios, and traditional algorithms lack the ability to filter and identify the validity of transmitted content.

Method used

A multimodal vehicle-road cooperative perception system based on instance-level feature transmission is adopted. Through unified vehicle-road position modeling, dynamic channel fusion, and high-confidence target screening modules, the system achieves effective extraction and propagation of features from both the vehicle and roadside. This system unifies vehicle and road sensor information into the same coordinate system and employs multimodal data fusion and high-confidence target screening mechanisms to select multimodal instance-level transmission features, thereby enhancing transmission efficiency and accuracy.

Benefits of technology

It achieves efficient information transmission for vehicle-road cooperative perception under bandwidth-constrained conditions, reduces transmission bandwidth requirements, reduces false positive results, and improves the accuracy and transmission density of multimodal data fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118741460B_ABST
    Figure CN118741460B_ABST
Patent Text Reader

Abstract

The application discloses a vehicle-road multi-modal data instance-level feature transmission cooperative perception system and method and a storage medium. A vehicle-road unified position modeling module is established to place vehicle-road multi-source sensor coordinates in a unified vehicle-end coordinate system, thereby improving information matching accuracy. A dynamic channel fusion module is designed to realize effective fusion of vehicle single-side multi-modal data through dynamic adjustment of multi-modal data representation. Finally, a high-confidence target feature screening algorithm is designed to distinguish target features and non-target features, thereby reducing transmission bandwidth and reducing false positive results of perception results.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of transportation, and relates to a vehicle-road multi-modal data instance-level feature transmission cooperative perception system and method and a storage medium. BACKGROUND

[0002] With the continuous improvement and rapid development of the automobile industry, cars are becoming an indispensable means of transportation for daily travel and production and life. The continuous popularization of autonomous vehicles can effectively reduce the workload of drivers and avoid traffic accidents, and is expected to realize the beautiful vision of smart travel. The environmental perception capability of autonomous vehicles is an important basis for the driving system, providing core guidance information for path planning and vehicle dynamic control. However, the existing single-vehicle sensor intelligent driving system often faces problems such as blind spots caused by close-range occlusion and sparse long-range perception. In response to this, the vehicle-road cooperative perception technology centered on intelligent networking technology has emerged. Vehicle-road cooperative perception fuses the perception information of different sensors at the vehicle end and the roadside to generate overall perception of the surrounding environment. Compared with single-vehicle intelligence, the addition of roadside sensors forms a multi-agent cooperative perception that can reduce the visual blind area caused by occlusion in congested road conditions. In addition, due to the relatively high position of roadside devices compared to vehicles, intelligent networked vehicles can obtain global perspective and long-range perception information of the surrounding environment, which can significantly improve the environmental perception capability of vehicles.

[0003] However, the emerging field of vehicle-road multi-agent cooperative perception still has many problems. First, the transmission of vehicle-road sensor perception information requires the support of 5G and other wireless communication networks, but existing 5G communication networks cannot support network communication demand in high-concurrency scenarios. Existing cooperative perception strategies are divided into three categories according to the transmission content: pre-fusion (data-level fusion), intermediate fusion (feature-level fusion), and post-fusion (target-level fusion). The pre-fusion strategy is to fuse vehicle-road sensor information before data input, and there is basically no information loss, so the detection accuracy is high; the post-fusion strategy is to directly fuse the target perception results of multiple sensors. Since the data is processed by target detection, a large amount of effective features are lost, and the cooperative perception has limited improvement compared to single perception. At the same time, when a single sensor makes a false detection, there is no effective error correction mechanism. Intermediate feature fusion is a compromise between the above two modes, which reduces the transmission bandwidth demand and ensures the ideal cooperative perception effect. It is the current mainstream information transmission strategy. Although existing algorithms based on intermediate feature transmission consider feature fusion under different conditions, they only consider the overall compression of transmission features under effective transmission bandwidth, but lack research on transmission content screening and effectiveness identification. The transmission content contains a lot of environmental information unrelated to the perception target, and there is a major defect in the transmission information density.

[0004] Secondly, the multi-modal data feature fusion problem. The existing vehicle-road cooperative perception cooperation often only uses the transmission of laser radar data for data perception due to the need for alignment of vehicle and road sensors and features. Although laser radar has excellent spatial construction and all-weather anti-interference capability, making the industry and academia generally recognize the role of laser radar in advanced automatic driving technology. However, it is undeniable that single laser radar perception has inherent problems such as difficulty in effective information screening and low semantic information density. Visual sensors have relatively rich texture details and strong object color perception details, but the fusion accuracy of visual sensor data between vehicles and roads is reduced. The existing multi-modal data fusion of a single vehicle mainly goes through the following steps: first, 3D detection and segmentation are performed using laser radar data to obtain 3D spatial elements; then, 2D detection is performed on the camera data to obtain 2D spatial information; finally, multi-modal data fusion algorithm is used for sensor data fusion to output the final result. The existing L2 level driving assistance system tends to use the above multi-modal data post-fusion mode. However, this mode has many problems for vehicle-road cooperative systems. First, the fusion result of the multi-sensor post-fusion mode has a large loss of data for each sensor, and the accuracy of the fusion result is limited. Second, the data transmission bandwidth of the vehicle-road cooperative system is limited, and in the transmission link centered on laser radar, there is a complementary information redundancy control problem between visual sensor data and laser radar data, that is, both laser radar and visual sensor data can realize object detection, and transmitting two sets of data will not improve the detection accuracy but will occupy valuable communication resources. SUMMARY

[0005] In order to solve the above technical problems, the application provides a multi-modal vehicle-road cooperative perception system and method based on instance-level feature transmission and a storage medium. Through in-depth research on vehicle-road cooperative information transmission under bandwidth limited conditions, the optimal feature extraction texture is explored, and based on the instance-level feature extraction framework, the traditional algorithm of transmitting all perception data or compressing the feature dimension is abandoned. By establishing a unified position coding for vehicle and road features, vehicle and road feature position relationship modeling is realized, and a high confidence target discrimination mechanism is used to select multi-modal instance-level transmission features, so as to realize effective extraction and transmission of roadside features. Finally, the vehicle and road multi-source feature extraction and transmission of the proposed framework are realized.

[0006] The multi-modal vehicle-road cooperative perception system based on instance-level feature transmission provided by the application has three main contents: a vehicle-road unified position modeling module, a dynamic channel fusion module and a high confidence target screening module. The following will introduce the above three modules.

[0007] (1) For the vehicle-road unified position modeling module

[0008] The main way is to unify the vehicle-road sensor information perception coordinates in the cooperative perception space to the same coordinate system. In the traditional non-cooperative perception framework, the road side and the vehicle end sensors use different coordinate systems, and the feature transfer between the two cannot be accurately matched. Position encoding is more important for instance-level fusion, as the environment description coordinates of the multi-source sensor system need to be consistent. As for the selection of the vehicle-road center coordinate, since the ultimate goal of the vehicle-road cooperative architecture is to supplement and enhance the vehicle's ability to perceive the surrounding scene, it is appropriate to use the vehicle coordinate system as the unified coordinate system.

[0009] The module needs to adapt to the information transmission task in the multi-modal scene, so the bird's eye view BEV perspective is selected. First, the relative position relationship between the vehicle end and the road side in the BEV perspective needs to be determined. Due to the lack of height information in the BEV perspective, the position transformation relationship is simplified to three degrees of freedom (tx, ty, θy) from six degrees of freedom (tx, ty, tz, θp, θr, θy), where tx, ty, tz represent the translation along the x, y, z axes, and θp, θr, θy represent the roll angle along the x axis, the pitch angle along the y axis, and the yaw angle along the z axis. Then use the transformation matrix to transform the road side independent position encoding to the vehicle side BEV space, as shown in the following formula:

[0010]

[0011] Where x', y' represent the transformed coordinates. For the numerical precision problem caused by coordinate transformation, the nearest neighbor method is used to align the converted position encoding to the vehicle side. The specific formula is as follows:

[0012]

[0013] Where round() represents the rounding operator, xout represents the x-axis component after precision alignment, and yout represents the y-axis component after precision alignment. The above steps generate unified position encoding. The original data space with position information can be aligned before information transmission, thereby reducing the feature misalignment problem during feature fusion.

[0014] (2) Dynamic channel fusion module

[0015] After completing the vehicle-road unified position encoding, the original perception data or feature extracted data between vehicle-road sensors can be aligned. Since the traditional method uses the point cloud data transmission method, it ignores the effective acquisition of multi-modal data features, so the multi-modal data transmission method can be used to enhance the target detection performance. Effective fusion of vehicle and road multi-modal fusion data requires the help of a dynamic channel fusion module to complete multi-modal data fusion.

[0016] Firstly, the multi-modal data feature extraction step needs to be performed: in order to obtain effective features from the features of different modalities, a backbone network of a specific modality can be adopted to extract the data stream features of the vision and lidar. The input data is a set of vision-lidar related scene data:

[0017] Data = {(L1, C1), (L2, C2),..., (Li, Ci)}, i e {1,..., N}

[0018] wherein Li represents the lidar data, Ci represents the vision data stream, and i represents different time points.

[0019] For the vision data stream, a backbone network based on ResNet is adopted to extract effective features, and the specific process is as follows:

[0020]

[0021] wherein ResNet(·) contains three 3x3 convolutional networks with batch normalization (Batch Normalization, BN) and ReLU activation function. H, W and Ccam in the multi-scale feature map of the vision branch represent the height, width and feature map channel number at different resolutions.

[0022] For the lidar point cloud data processing branch, a point cloud feature extractor based on the PointPillar backbone network is adopted, and the original point cloud is recorded as L i = {p1, p2,..., p c}(p c = (x c , y c , z c , r)) wherein xc, yc, zc, r and c represent the spatial coordinates, reflectivity and the number of point clouds in the scene respectively. The main function of the 3D point cloud pillar feature network (pillar feature net, PFN) is to convert the 3D point cloud into a stacked Pillar tensor. The specific process is that the 3D point cloud is divided into fixed size grids in the X-Y plane in the overhead perspective, and then a Pillar is formed in the Z axis direction. The Pillar features are extracted using the PointNet architecture. Then the tensor is projected into a 2D pseudo map with a size of HxWxC, wherein H and W represent the height and width of the corresponding pseudo map canvas respectively, and C represents the channel number of the pseudo map. The pseudo map, like a picture, has three dimensions, and the feature extraction is relatively convenient. The features are scattered back to the X-Y plane to generate a BEV pseudo map, and then a 2D convolution network PointPillar is used to generate dense lidar BEV features from the BEV pseudo map, and the specific process is as follows:

[0023]

[0024] where H, W and Clidar obey the high, wide and the number of channels of the laser radar feature map at different resolutions.

[0025] After obtaining the corresponding BEV view data of vision and laser radar, it is necessary to use a dynamic channel fusion framework to fuse the feature data.

[0026] The input data of the dynamic channel fusion module is the laser radar feature and the visual feature The traditional fusion method is to directly splice or sum the feature data of the two, to generate a multi-modal data feature set for enhancing multi-modal features. However, due to the inherent heterogeneity of vision and laser radar, and due to the lack of sufficient object semantic supervision of sensor features, direct splicing or summing operation often leads to spatial misplacement and ultimately leads to rough information fusion granularity. However, the present application uses a dynamic channel fusion algorithm to fuse the context information of the vision image and the laser radar in the form of channels. Specifically, a 3x3 convolution is used to explore valuable semantic and geometric feature information to generate reconstructed features Fconv. In order to highlight the distinguishability of the target, a global average pooling operator GAP(·) is applied to the channel features, and then a multilayer perceptron with a sigmoid activation function δ(·) is used to generate the activation probability of the channel feature reweighting, and finally it is multiplied by the convolution feature Fconv to generate a single multi-modal feature Fsingle for inter-agent feature transmission. The overall process is as follows:

[0027]

[0028] Fsingle = δ(MLP(GAP(Fconv)))·Fconv

[0029] (3) High confidence target screening module

[0030] After the dynamic channel fusion module, the fusion features Fsingle are generated from the vision sensor and the laser radar data. Since the vehicle end and the road end both have the fusion features, in order to facilitate description, they are denoted as Fveh and Finf respectively. Traditional algorithms will directly transmit the vehicle end and road side features, or use compression, dimensionality reduction and other methods to reduce the transmission bandwidth required. However, the above method contains a large amount of information unrelated to the perception target in the transmission feature, including irrelevant perception information such as roads, pedestrians, trees and other irrelevant information in the environment. The transmission of such information will greatly reduce the transmission density of effective information and occupy valuable transmission bandwidth.

[0031] To this the application adopts screening mechanism, first the scene is divided into several grid, then the confidence evaluation is carried out to all grid object features in the grid scene, then the low confidence level grid features are filtered out using the filtering module, and the filtered vehicle end and road end feature information is transmitted. The above algorithm has two obvious advantages compared with the traditional algorithm, first, the transmission density of effective perception information can be further enhanced, thereby reducing the demand of transmission bandwidth; second, the transmission features after screening, due to the fact that the grid blocks in the features almost do not contain non-road information, the features after vehicle-road scene fusion will not generate vehicle features at non-road, finally the generation of false positive results in the perception result can be reduced.

[0032] The main content of the high confidence target screening module is as follows, first, the fused features are input into the confidence feature discrimination module, and the screened vehicle end and roadside end features are output. The specific process is as follows

[0033] Aconf=Prob(Fveh)

[0034] Wherein Prob(·) represents a feature confidence generation network, which is composed of two layers of 3x3 CNN convolutional neural network and ReLU activation function, Aconf represents a feature confidence matrix, which is in the form as follows:

[0035]

[0036] m×n represents that the perception space with the size of HxW is divided into feature grid blocks with the length of m and the width of n, since the feature grid is derived from the convolution of the perception space value, the values of m and n are determined by the size of the scene HxW. The element aij in the matrix represents the confidence value of the element in the confidence matrix, the numerical value represents the probability of the grid block having vehicle features, and the larger the numerical value, the higher the possibility of the grid block having vehicles.

[0037] Then the feature confidence matrix is sent to the screening module, and the specific process is as follows:

[0038] Aconf'=Filter(Aconf)

[0039] Wherein Filter(·) function represents a screening function, which mainly sets the value in the confidence interval to 0 or 1, and the specific process is as follows:

[0040]

[0041] Wherein ε represents the confidence value, which is represented by a constant. The generated Aconf' represents the position matrix of the low confidence feature grid blocks filtered out in the scene, and the effective transmission content of the vehicle end and the road end is obtained by the dot product of the position matrix Aconf' and the feature matrix Fveh, and the specific process is as follows.

[0042] Ftrans=Aconf'·Fveh

[0043] Wherein Ftrans represents the vehicle effective feature data in the corresponding scene between the vehicle end and the road end, that is, the perception content for information transmission between the vehicle and the road.

[0044] The above is a multi-modal vehicle-road cooperative perception system based on instance-level feature transmission designed by the application. The core is to explore the vehicle-road cooperative information transmission mode under the condition of limited bandwidth and to explore the optimal feature texture extraction texture. The unified position modeling mechanism, the dynamic channel fusion module and the high confidence target screening module are mainly described.

[0045] Based on the above perception system, the application also proposes a vehicle-road multi-modal data instance-level feature transmission cooperative perception method, comprising the following steps:

[0046] S1, unified position modeling of vehicle and road, which realizes the content of the unified position modeling module of vehicle and road;

[0047] S2, dynamic channel fusion, which realizes the content of the dynamic channel fusion module;

[0048] S3, high confidence target screening, which realizes the content of the high confidence target screening module.

[0049] The application also proposes a storage medium, which contains the program code of the perception method.

[0050] The application has the following advantages:

[0051] 1. Unified position modeling of vehicle and road is realized. Since the vehicle end and the road end perception data use different coordinate systems, the vehicle-road feature cooperative transmission cannot be matched, so the coordinates of the vehicle-road multi-source sensor cooperation need to be kept consistent. At the same time, since the ultimate goal of vehicle-road cooperation is to supplement and enhance the surrounding scene perception ability, the vehicle coordinate system is selected as the vehicle-road reference coordinate system. And to solve the problem of accuracy difference caused by vehicle-road coordinate change, the nearest neighbor method is used to align the position code to the vehicle end side, reducing the accuracy loss.

[0052] 2. Dynamic channel fusion is realized. Traditional vehicle-road cooperative perception cooperation considers that data fusion is difficult, so the transmission information content is limited to point cloud data, ignoring that multi-modal data such as vision-laser radar can enhance target detection performance. Therefore, the application designs a data fusion module based on vision and laser radar multi-modal data, which realizes the effective fusion of single-sided multi-modal data of the vehicle end or the road side through the dynamic adjustment between the feature channels of multi-modal sensors.

[0053] 3High confidence target feature screening is realized. Since the existing vehicle-road cooperation feature level transmission mode usually adopts feature dimension reduction and compression to reduce the required bandwidth of transmission, a large amount of scene information irrelevant to the perception target is contained, mainly including irrelevant perception information such as road points, trees and pedestrians in the scene. In view of this, the high confidence perception feature screening mechanism is designed, only the scene grid containing the vehicle feature in the scene is retained. The required bandwidth of data transmission is greatly reduced, and the false positive results of the perception result can be reduced. BRIEF DESCRIPTION OF DRAWINGS

[0054] Figure 1 The vehicle-road unified position modeling schematic diagram proposed by the application;

[0055] Figure 2 The multi-modal data dynamic channel fusion framework schematic diagram proposed by the application;

[0056] Figure 3 The high confidence target feature screening mechanism schematic diagram proposed by the application. DETAILED DESCRIPTION

[0057] The application will be further described below in combination with the drawings and embodiments.

[0058] The application embodiment: the experimental setting is specifically as follows, the real scene data set DAIR-V2X is used as a verification environment. In terms of hardware, in terms of data acquisition system, a 1-inch global exposure CMOS camera is used at the roadside end, the sampling frame rate is 25Hz, the image format is RGB format, and the resolution is compressed to 1920*1080 and saved as a JPEG format; the laser radar adopts a 300-line solid-state laser radar. The sampling frame rate is 10Hz, and the maximum detection range is 280m. The vehicle end is equipped with a camera with a sampling frame rate of 20Hz, and the image format is a JPEG format image with a resolution of 1920*1080; the laser radar adopts a 40-line 360-surrounding laser radar, the sampling frame rate is 10Hz, and the maximum detection range is 200m.

[0059] In terms of cooperative distance, in the cooperative scenario, the communication distance between the intelligent agents composed of vehicles and roads is set to 70 m. If the distance exceeds the communication distance, it is considered unnecessary to enter the cooperative range, and the intelligent agent will be ignored and not broadcast the perception data. The information transmission part can be compressed, and the compression ratio can be set to 8, 16, and 32. At the same time, the intelligent agent laser radar environment perception measurement range needs to be set, which is [-140.8, 140.8] and [-38.4, 38.4] in the x-axis and y-axis, respectively, and the range of the z-axis is [-1, 3]. Since the point cloud beyond the above laser radar perception range has a large sparsity, it cannot provide effective information for vehicle target perception, so the points beyond the range are deleted. Since the laser radar space needs to be voxelized, the same settings are used for vehicle-end and road-side voxelization, and the length, width, and height of a single space voxel are set to 0.4 m, 0.4 m, and 4 m, respectively.

[0060] (1) The design of unified position modeling of vehicle and road is shown in FIG. 1. In order to achieve spatial synchronization between different sensors, vehicle-road cooperation needs to use sensor parameter information for coordinate system conversion. Since both the vehicle end and the road side have cameras and laser radars, the first step is to convert the spatial position relationship of the camera-laser radar two homologous sensors. Figure 1

[0061] The laser radar coordinate system is the origin point with the laser radar sensor as the geometric center, the x-axis is horizontal forward, the y-axis is horizontal right, and the z-axis is vertical upward, which conforms to the right-hand rule. Since the road end device has a ground with a pitch angle, in order to facilitate research, the road-side laser radar extrinsic data is uniformly converted into a virtual laser radar coordinate system; the camera coordinate system is with the camera optical center as the origin, the x-axis and y-axis are parallel to the image plane coordinate system x-axis and y-axis, and the z-axis is parallel to the camera optical axis forward and perpendicular to the image plane. Through the extrinsic parameter coordinate system matrix of the laser radar to the camera, the virtual laser radar coordinate system can be converted to the camera coordinate system. Using the camera intrinsic parameter, the projection from the camera coordinate to the image coordinate can be realized. The image coordinate is the position encoding of the vehicle-end or road-side sensor.

[0062] Regarding the position modeling of vehicle-road cooperation, in the traditional single vehicle perception framework, it is reasonable for the road side and the vehicle end to use different coordinates for cooperation. However, for the cooperative system, using different coordinate systems will cause the problem of not matching the scene features. At the same time, since the framework of vehicle-road cooperation is to enhance the surrounding scene perception ability, the vehicle coordinate system needs to be used as the unified coordinate system.

[0063] ​Since the module needs to complete the information transmission task in the multi-modal scene, it is necessary to complete the fusion of point cloud and image first. Therefore, the bird's eye view (BEV) perspective based collaborative method is selected for the vehicle side and roadside side devices. First, the relative position relationship between the vehicle side and roadside side in the BEV perspective needs to be determined. The perception data of the roadside side is converted to the vehicle side BEV perspective through independent position encoding using a matrix transformation matrix, as shown in the following formula

[0064]

[0065] For the precision problem caused by coordinate transformation, the round() function is used to completely convert the numerical precision to the position encoding of the vehicle side.

[0066] According to the above steps, the unified encoding of homologous sensors is first realized, and then the unified position encoding of the vehicle side and roadside side is realized. The multi-modal data can be aligned in the original data space with position information before transmission, thereby reducing the feature layering problem during vehicle-road feature fusion.

[0067] (2) The design of the multi-module data dynamic channel fusion framework is shown in Figure 2 . The main purpose is to enhance the target detection performance of the traditional single point cloud data transmission method by transmitting multi-modal data. Multi-modal data effective fusion needs to use the dynamic channel fusion module to complete multi-modal data fusion.

[0068] The dynamic channel fusion framework inputs the coordinate corresponding scene data of the vehicle side or roadside side:

[0069] Data={(L1,C1),(L2,C2),...,(Li,Ci)},i∈{1,...,N}

[0070] Where Li represents the laser radar data stream, Ci represents the visual data stream, and i represents different time points. The visual data stream Ci uses the input based on the ResNet backbone network to extract effective scene BEV features The specific process is as follows:

[0071]

[0072] Where the visual feature extraction network ResNet(·) contains multiple 3x3 convolution kernels with batch normalization layer BN and ReLU activation function. The feature of the visual branch is a multi-scale feature map containing H, W and Ccam, which represent height, width and feature map channel number respectively.

[0073] The laser radar data stream Li is input into a laser radar point cloud processing branch, and the branch uses PointPillar as a point cloud feature extractor of a backbone network.

[0074] The 3D point cloud is converted into a stacked Pillar tensor through a 3D point cloud Point Feature Net module, and the specific process is that the 3D point cloud is divided into a fixed-size grid in the X-Y plane in the bird's eye view, and then a Pillar is formed in the Z-axis direction. Then, the fixed-size Pillar grid is input into a simple feature extractor based on the PointNet architecture to extract features of the Pillar. Then, the tensor is projected into a 2D pseudo graph with a size of HxWxC, where H and W represent the height and width of the corresponding pseudo graph canvas respectively, and C represents the channel number of the pseudo graph. The pseudo graph and the picture have three dimensions, and the feature extraction is convenient. The features are scattered back to the X-Y plane to generate a BEV pseudo graph, and then a 2D convolution network PointPillar is used to generate dense laser radar BEV features from multiple BEV pseudo graphs, and the specific process is as follows:

[0075]

[0076] Where H, W and Clidar represent the height, width and laser radar feature map channel number at different resolutions respectively.

[0077] After obtaining the corresponding BEV view data of the vision and the laser radar, the feature data is input into a dynamic channel fusion framework to dynamically fuse the feature data. The input data of the dynamic channel fusion framework is laser radar features and vision features The traditional fusion method is to directly splice or sum the feature data of the two, but the direct splicing or summing operation often causes spatial misalignment and finally leads to rough information fusion granularity.

[0078] In order to solve the above problems, the dynamic channel fusion algorithm is adopted to fuse the image and laser radar context information in the form of channels. Specifically, a feature representation module composed of multiple 3x3 convolution kernels is used to enhance the feature depth, explore valuable semantic and geometric feature information, and generate reconstructed features Fconv. The specific process is as follows

[0079]

[0080] Meanwhile, a global average pooling operator GAP(·) is applied to the channel features of the vehicle and the road to detect the distinguishability of the target. Then, a multi-layer perception with a sigmoid activation function δ(·) is used to generate the activation probability value of the channel feature re-weighting, which is multiplied by the convolution feature Fconv to generate the single multi-modal feature Fsingle for inter-agent feature transmission. The overall process is shown as follows:

[0081] Fsingle = δ(MLP(GAP(Fconv))) · Fconv

[0082] The generated Fsingle is a single-sided feature with multi-modal information at the vehicle end or the road end, which contains all scene perception information, including vehicle target information and non-target information. Therefore, it is necessary to waste a large amount of communication resources, and further filtering of the transmission information is needed.

[0083] (3) The design of the high-confidence target feature filtering mechanism is shown in Figure 3 The core of the high-confidence target feature discrimination mechanism is to score the feature data in advance, and to distinguish the target information and non-target information of the feature, so as to enhance the effective information transmission density. The main content of the structure is as follows: first, the fused features are input into the confidence feature discrimination module, and the filtered vehicle end and roadside end features are output. The two processes are the same, and the following takes the vehicle end feature Fveh filtering as an example. The specific process is shown as follows:

[0084] Aconf = Prob(Fveh)

[0085] wherein Prob(·) represents a feature confidence generation network composed of two layers of 3x3 CNN convolutional neural network and ReLU activation function, and Aconf represents the generated feature confidence matrix, which has the following matrix form structure:

[0086]

[0087] wherein i and j represent the grid block length and width serial numbers, and m x n represents that the original spatial size H x W perception space is divided into a feature grid block with a length of m and a width of n. Since the feature grid is derived from the convolution operation on the perception space value, the value range of m and n is determined by the size of the scene H x W. The element aij in the matrix represents the confidence value of the element in the confidence matrix, and the numerical size represents the probability size of the vehicle feature in the grid block. Then the feature confidence matrix is sent to the filtering function, and the specific process is shown as follows:

[0088] Aconf' = Filter(Aconf)

[0089] wherein the Filter(·) function represents a filter function, and the main function is to set the value in the confidence interval to 0 or 1. The specific result is shown as follows:

[0090]

[0091] wherein ε represents a confidence value, and a constant is used to represent it. The generated Aconf' represents a position matrix of filtering out low-confidence feature grid blocks in the scene, and the valid transmission content Ftrans between the vehicle end and the road end is obtained by the dot product of the position matrix Aconf' and the feature matrix Fveh. The specific process is shown as follows.

[0092] Ftrans=Aconf'·Fveh

[0093] wherein Ftrans represents the valid feature data of the vehicle in the corresponding scene between the vehicle end and the road end, that is, the perception content for information transmission between the vehicle and the road. Then, it is sent to the feature fusion module to complete the effective fusion between the vehicle and the road data.

[0094] The series of detailed descriptions listed above are only specific descriptions of the feasible implementation manners of the present application, and are not used to limit the protection scope of the present application. Any equivalent manners or changes without departing from the present application should be included in the protection scope of the present application.

Claims

1. A vehicle-road multi-modal data instance-level feature transmission cooperative perception system, characterized in that, Comprise: A vehicle-road unified position modeling module, a dynamic channel fusion module, and a high-confidence target screening module; The vehicle-road unified position modeling module unifies vehicle-road sensor information perception coordinates in a cooperative perception space to a vehicle coordinate system, adopts a bird's eye view (BEV) perspective, and adapts to information transmission tasks in a multi-modal scene; The dynamic channel fusion module is configured to effectively fuse multi-modal data at a vehicle end and a road end, and generate fused features of visual and laser radar point clouds; The high-confidence target screening module uses a screening mechanism to divide a scene into a plurality of grids, evaluate the confidence of all grid object features in the grid scene, and then filter out low-confidence grid features to obtain filtered fused features.

2. The vehicle-road multi-modal data instance-level feature transmission cooperative perception system according to claim 1, characterized in that, The vehicle-road unified position modeling module comprises: Firstly, the relative position relationship between the vehicle end and the roadside end in the BEV perspective needs to be determined. Due to the lack of height information in the BEV perspective, the position transformation relationship is simplified from six degrees of freedom (t x ,t y ,t z ,θ p ,θ r ,θ y ) to three degrees of freedom (t x ,t y ,θ y ); then the independent position code of the roadside end is transformed into the BEV space on the vehicle side using a transformation matrix, as shown in the following formula: To solve the numerical precision problem caused by coordinate transformation, the converted position code is aligned to the vehicle side using the nearest neighbor method, and the specific formula is as follows: Wherein round() represents a rounding operator.

3. The vehicle-road multi-modal data instance-level feature transmission cooperative perception system according to claim 1, characterized in that, The dynamic channel fusion module, input data is a laser radar feature and visual features The dynamic channel fusion algorithm is used to fuse visual features and laser radar features in the form of channels, specifically as follows: A 3x3 convolution is used to explore valuable semantic and geometric feature information, generating reconstruction features F conv To highlight the distinguishability of the target, a global average pooling operator GAP(·) is applied on the channel features, and then a multi-layer perceptron with a sigmoid activation function δ(·) is used to generate the activation probability of the channel feature reweighting. Finally, it is multiplied by the convolutional features F conv to generate a single multi-modal feature F single for inter-agent feature transmission.

4. The vehicle-road multi-modal data instance-level feature transmission cooperative perception system according to claim 3, characterized in that, The formula corresponding to the dynamic channel fusion algorithm is as follows: F single = δ(MLP(GAP(F conv ))) · F conv .

5. The vehicle-road multi-modal data instance-level feature transmission cooperative perception system according to claim 3, characterized in that, The visual features Through ResNet-based backbone network extraction, specifically as follows: where ResNet(·) contains three 3x3 convolutional networks with batch normalization layer BN and ReLU activation function, visual feature is a multi-scale feature map containing H, W and C cam H, W and C cam represent height, width and feature map channel number, respectively.

6. The vehicle-road multi-modal data instance-level feature transmission cooperative perception system according to claim 3, characterized in that, The lidar features Point cloud features are extracted using a PointPillar backbone network-based point cloud feature extractor, in particular as follows: Let the original 3D point cloud of the LiDAR be denoted as P = {p1, p2, ..., p...} c }, p c =(x c ,y c ,z c ,r) where x c y c , z c Represents spatial coordinates, where r and c represent reflectivity and the number of point clouds in the scene, respectively; 3D point cloud is converted into a stacked Pillar tensor, and the specific process is as follows: 3D point cloud is divided into fixed-size grids in the X-Y plane under the bird's eye view, and a column, i.e., Pillar, is formed in the Z-axis direction; PointNet architecture is used to extract features of the columnar features; Then, the Pillar tensor is projected into a 2D pseudo graph with a size of HxWxC, where H and W represent the height and width of the corresponding pseudo graph canvas, respectively, and C represents the number of channels of the pseudo graph; the features are scattered back to the X-Y plane to generate a BEV pseudo graph, and then a 2D convolution network PointPillar is used to generate dense laser radar BEV features for the BEV pseudo graph, and the specific process is as follows: where H, W and C lidar represent the height, width and number of channels of the LiDAR feature map at different resolutions.

7. The vehicle-road multi-modal data instance-level feature transmission cooperative perception system according to claim 1, characterized in that, The high-confidence target screening module comprises a confidence feature discrimination module and a screening module; The confidence feature discrimination module inputs the fused features and outputs the screened vehicle end and road side features, and the specific formula is as follows: A conf = Prob(F veh ) where F veh denotes the fusion feature of the vehicle end, Prob(·) denotes the feature confidence generation network, which is composed of two layers of 3 × 3 CNN convolutional neural network and ReLU activation function, A conf denotes the feature confidence matrix, which has the following form: mxn represents dividing the HxW perceptual space into m-length and n-width feature grid blocks, since the feature grid is derived from the convolution of the perceptual space values, the values of m and n are determined by the scene HxW size, the element a ij represents the confidence value of the element in the confidence matrix, the numerical size represents the probability size of the vehicle feature in the grid block, the larger the numerical value, the higher the possibility of the vehicle in the grid block; The screening module inputs the feature confidence matrix, filters out the low-confidence grid features, and obtains the filtered fused features; the specific formula is as follows: A conf ' = Filter(A conf ) where A conf Filter(·) function represents a screening function, which function is to set the value in the confidence interval to 0 or 1, and the specific process is as follows: Wherein ε represents a confidence value, and is a constant.

8. The vehicle-road multi-modal data instance-level feature transmission cooperative perception system according to claim 7, characterized in that, The high-confidence target screening module further comprises the generated A conf and the feature matrix F veh The point product is obtained to obtain the effective transmission content of the vehicle end and the road end, and the specific process is as follows: F trans = A conf '·F veh wherein F trans represents the vehicle effective feature data in the corresponding scene between the vehicle end and the road end, i.e. the perception content for information transmission between the vehicle and the road.

9. A vehicle-road multi-modal data instance-level feature transmission collaborative perception method, characterized in that, The method comprises the following steps: S1, vehicle-road unified position modeling, which realizes the content of the vehicle-road unified position modeling module in any one of claims 1-8; S2, dynamic channel fusion, which realizes the content of the dynamic channel fusion module in any one of claims 1-8; S3, high-confidence target screening, which realizes the content of the high-confidence target screening module in any one of claims 1-8.

10. A storage medium, characterized by The storage medium contains the program code of the perception method in claim 9.

Citation Information

Patent Citations

  • Multi-mode-based automatic driving perception method and device, equipment and medium

    CN115879060A

  • Vehicle detection method based on multi-modal data fusion

    CN116863461A