Unmanned automobile cloud cooperative wide-area sensing method and system based on sky and ground feature fusion

Through the vehicle-cloud collaborative perception method, lightweight processing on the vehicle side and high computing power fusion technology on the cloud side are used to solve the problem of limited perception range of bicycles and low cloud integration efficiency, achieving global perception and real-time improvement, reducing the risk of autonomous driving.

CN120467359APending Publication Date: 2025-08-12JIANGSU UNIV

Patent Information

Application Number
CN202510618089.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The limited perception range of existing bicycles and low cloud integration efficiency lead to delays or misjudgment of autonomous vehicles in complex scenarios, lacking a global perspective, and increasing the risk of accidents.

Method used

The method of lightweight preprocessing on the vehicle side and high computing power fusion in the cloud is adopted. PointPillars body-columnization combined with FPN-ResNet to process the lidar point cloud to generate BEV feature maps, and the AB3DMOT multi-objective tracking algorithm is used to generate trajectory collections, and self-attention fusion and satellite data enhancement are carried out in the cloud, combining trajectory prediction and feature distortion technology to achieve coordinated perception of the vehicle and cloud.

Benefits of technology

Effectively expand the scope of perception, enhance the semantic understanding of global road structures and traffic rules, improve perception robustness and real-timeness, reduce the pressure on the computing power of the vehicle, and ensure the accuracy of autonomous driving decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120467359A_ABST
    Figure CN120467359A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned vehicle cloud cooperative wide area sensing method and system based on sky-ground feature fusion, and the method comprises the steps: vehicle end data processing: carrying out the processing of a laser radar point cloud through PointPill volume cylindricity in combination with an FPN-ResNet backbone network, generating a BEV feature map, outputting a target detection result containing the information of a three-dimensional coordinate, a course angle and the like through a detection head, and carrying out the detection of the target detection result; then, an AB3DMOT multi-target tracking algorithm is utilized to associate detection results into a track set, and then the BEV feature map, the track set and vehicle pose information are packaged, compressed and uploaded to a cloud end; cloud data processing: after receiving the data, the cloud decompresses the data to obtain a BEV feature map, a track set and vehicle pose information, and then performs feature fusion and track fusion; and the vehicle end fuses the received information fused by the cloud end with the characteristics of the vehicle at the current moment, eliminates the delay, and finally obtains a vehicle-cloud cooperative sensing result. The real-time performance of the sensing result can be guaranteed, and timely and accurate environment information support is provided for automatic driving decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of environmental perception technology for intelligent self-driving cars, and specifically designs a vehicle-cloud collaborative wide-area perception method and system for unmanned vehicles based on the fusion of sky and earth features. Background Art

[0002] In recent years, with the advancement of intelligent driving technology, the single-vehicle perception performance of autonomous vehicles has rapidly improved. However, due to the physical characteristics of sensors, single-vehicle perception often cannot cover the entire area surrounding the vehicle, resulting in blind spots. Furthermore, due to the limitations of the data processing capabilities of on-board equipment, single-vehicle perception may not be able to process all perception data in a timely and accurate manner. This can lead to delays or misjudgments in complex or high-speed driving scenarios, increasing the risk of accidents. Finally, single-vehicle perception focuses primarily on changes in the local environment surrounding the vehicle and lacks a global perspective. This means that autonomous vehicles may not fully understand traffic conditions and potential risks across the entire road network.

[0003] Collaborative perception technology combines all collaborative autonomous vehicles (CAVs) and roadside intelligent infrastructure in the area to obtain more comprehensive and accurate road and environmental information. When a vehicle's own sensors have blind spots, other collaborative participants can monitor the conditions in these blind spots and send this information to the vehicle in real time, allowing better decision-making. Roadside infrastructure also plays an important role in this task, leveraging its higher installation position and various sensing devices to provide vehicles with accurate road condition information in adverse conditions (such as dense fog, heavy rain, etc.) and complex dynamic traffic scenarios (such as complex urban intersections). By supplementing perception information from other sources, collaborative perception technology can significantly improve the accuracy and effective range of perception tasks. As the foundation of autonomous driving technology, it provides effective support for other perception tasks.

[0004] However, collaborative sensing technology also faces challenges. Current collaborative sensing research focuses solely on the cloud as a communication mechanism. Specifically, a collaborative task involves n participants, and each participant must send data once and receive data n-1 times during a single round of communication, fusing data from n sources. Considering the large number of participants involved in this process, this process undoubtedly poses significant challenges to both communication and computing power.

[0005] How to effectively utilize the integrated vehicle-cloud communication and computing platform, add the cloud as a link in perception, make full use of the advantages of high computing power and low latency of the cloud platform, and further upgrade collaborative perception to vehicle-cloud collaborative perception is an important technical development direction of autonomous vehicle environmental perception technology, which is more conducive to comprehensively promoting the upgrade and transformation of the comprehensive performance of intelligent transportation systems. Summary of the Invention

[0006] In response to the shortcomings of the existing technology, the present invention provides a vehicle-cloud collaborative wide-area perception method and system for unmanned vehicles based on the fusion of earth-sky features, aiming to solve the problems of limited perception range of existing single vehicles and low cloud fusion efficiency, reduce the communication and computing power pressure in vehicle-road collaborative tasks, and improve the range and robustness of intelligent vehicle perception.

[0007] The vehicle-cloud collaborative wide-area perception method for unmanned vehicles based on the fusion of earth-ground features in the present invention includes vehicle-side processing and cloud-side processing; the details are as follows:

[0008] First, the vehicle side processes the lidar point cloud through PointPillars columnarization combined with the FPN-ResNet backbone network to generate a BEV feature map. The detection head outputs the target detection results containing three-dimensional coordinates, heading angles and other information. The AB3DMOT multi-target tracking algorithm is then used to associate the detection results into a trajectory collection. The BEV feature map, trajectory and vehicle posture information are then packaged and compressed and uploaded to the cloud to achieve lightweight preprocessing on the vehicle side.

[0009] Secondly, after receiving the data, the cloud first unifies the BEV feature map into the world coordinate system according to the position of each vehicle. The stacked aggregate tensor is compressed into cloud-based fusion features through the self-attention mechanism to capture the global spatial correlation from the perspective of multiple vehicles.

[0010] The extended perception area union is then calculated based on the multiple vehicle poses and the lidar detection radius. High-resolution images and road vector data of the area are requested from the satellite. Semantic segmentation of the satellite imagery is performed using the DeepLabv3+ model to generate a road mask. Prior features such as road curvature and speed limit are extracted from the road vector data. After bilinear upsampling, the data is concatenated with the road mask in the channel dimension. The data is then mapped to a feature space compatible with the lidar features using a linear transformation matrix, forming satellite-enhanced features that include global road semantics and traffic rules.

[0011] Subsequently, a "vehicle-side local features-satellite global semantics" fusion module was constructed, using cloud-side fusion features as queries and satellite-enhanced features as keys and values. Combined with satellite prior feature encoding, the weights were dynamically adjusted through the attention mechanism to achieve deep fusion of vehicle-side features and satellite semantics, suppressing noise characteristics in non-road areas.

[0012] In the trajectory processing stage, the cloud uses the Hungarian matching algorithm to associate multiple vehicle trajectories, retaining the longest trajectory of the same instance, and fitting trajectory trends through the Transformer encoder in trajectory prediction, completing short trajectories based on motion similarity.

[0013] After the ego vehicle receives the fused information from the cloud, it addresses the spatiotemporal misalignment problem caused by data processing delays by separating repeated instances through Hungarian matching of historical target detection results and fused predicted trajectories. It then generates a feature stream based on the actual trajectory obtained by the ego vehicle's multi-target tracking, distorts the cloud-based fused features to the current moment, avoids feature reconstruction noise, and finally outputs the vehicle-cloud collaborative perception results through the detection head.

[0014] The above specific processing process will be described in detail in the following specific implementation method section.

[0015] Beneficial effects of the present invention:

[0016] (1) The present invention proposes a vehicle-cloud collaborative wide-area perception method and system for unmanned vehicles based on the fusion of earth-sky features, constructs a collaborative architecture of lightweight vehicle-side acquisition and cloud-side centralized intelligence, and migrates tasks such as multi-source fusion and trajectory prediction that require high computing power to the cloud. While reducing the computing power pressure on the vehicle side, by integrating multi-vehicle perception information with satellite remote sensing data, it breaks through the local perception limitations of a single vehicle, effectively expands the perception range, and enhances the semantic understanding ability of the global road structure and traffic rules.

[0017] (2) Trajectory prediction is used to construct feature streams. To address the time delay problem in vehicle-cloud data processing, trajectory matching and feature distortion technology are used to dynamically compensate for spatiotemporal misalignment, avoid feature reconstruction noise, and significantly improve perception robustness in complex scenarios (such as high-speed driving and sensor occlusion).

[0018] (3) Through the cloud-based self-attention mechanism and the ground-ground integration module, the dynamic fusion of local detail features on the vehicle side and global semantic features on the satellite is achieved, the recognition of drivable areas and traffic rules constraints are strengthened, and the feature representation capability is improved; at the same time, the efficient cloud-based computing power allocation controls the processing delay at a low level, ensures the real-time nature of the perception results, and provides timely and accurate environmental information support for autonomous driving decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Flowchart for the implementation of the present invention;

[0020] Figure 2 This is the overall network architecture diagram of the vehicle-side preprocessing of the present invention;

[0021] Figure 3 This is a structural diagram of the cloud feature fusion of the present invention;

[0022] Figure 4 This is a structural diagram of the cloud trajectory fusion of the present invention;

[0023] Figure 5 This is a structural diagram of the characteristic flow compensation of the present invention. DETAILED DESCRIPTION

[0024] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings. The contents mentioned in the embodiment are not intended to limit the present invention.

[0025] The implementation process of the vehicle-cloud collaborative wide-area perception method for unmanned vehicles based on the fusion of earth and sky features is as follows: Figure 1 As shown, it specifically includes the following steps:

[0026] Step 1: Build the vehicle-side platform to pre-process the vehicle-side data and compress and upload it. The details are as follows:

[0027] As a crucial component of vehicle-cloud collaboration, the vehicle-side platform is responsible for initial data collection. It can also preprocess data from various sensors. For example, it can perform noise reduction and contrast enhancement on camera images, and filter radar data to remove interference. This example uses lidar as an example to illustrate vehicle-side data preprocessing.

[0028] Assume that a single cooperative autonomous driving vehicle at time t is CAV i ∈{CAV}, i=1,2,…,m. Where m represents the total number of all cooperative autonomous vehicles. The lidar point cloud input by Pointpillars (encoder for point cloud target detection) is composed of N p points, each of which is represented by a four-dimensional vector p = (x, y, z, r), where x, y, z are the coordinates of the point on the X, Y, and Z axes in the lidar coordinate system, respectively, and r is the reflection intensity. First, the original lidar point cloud is columnarized using the Pointpillars encoding method. The backbone of Pointpillars adopts a multi-scale feature pyramid FPN strategy and a residual learning (ResNet) strategy. FPN is a multi-scale feature pyramid structure designed to generate multi-scale feature maps with rich semantic information. Residual blocks are introduced to replace the convolution blocks in the original FPN. Each residual block contains 3 convolution layers and skip connections. The 3 convolution layers are 1×1 convolutions, which are used to reduce the number of channels; 3×3 convolutions are the main convolution layers of the residual blocks; 1×1 convolutions are used to restore the number of channels; the cross-layer connection mechanism can solve the problems of gradient disappearance and gradient explosion during training. The columnarized lidar point cloud passes through the backbone part to obtain the BEV feature map at time t. Its spatial dimensions are:

[0029]

[0030] in, represents the real number space, It is a three-dimensional real tensor, H is the height of the BEV feature map, W is the width of the BEV feature map, and C is the number of channels of the BEV feature map.

[0031] Then the detection head is used to detect the BEV feature map Perform fine classification and regression to obtain a collection of target detection results for collaborative autonomous driving vehicles n represents the total number of detected targets, and each element It consists of a tuple [x′,y′,z′,h,w,l,θ,s], which contains the 3D coordinates of the instance bounding box center [x′,y′,z′], the 3D size of the instance bounding box [h,w,l], the heading angle θ and the confidence score s.

[0032] The detected instance bounding box is associated with the trajectory segment using a multi-target tracking algorithm. The present invention adopts AB3DMOT, which is an online multi-target tracking algorithm. First, a 3D Kalman filter is used to predict the trajectory from the previous frame to the current frame; then, a data association module is used to match the predicted trajectory of the 3D Kalman filter with the instance bounding box detected in the current frame; the 3D Kalman filter updates the instance trajectory based on the matching result. This process creates a trajectory for the new instance and deletes the trajectory of the disappeared instance. After the above steps, a trajectory collection is obtained. See the process Figure 2 .

[0033] Vehicle-side data preprocessing can be expressed by the following formula:

[0034] F=Pillar(PC) (2)

[0035]

[0036] Among them, PC represents the input lidar point cloud, Pillar represents the volume pillarization operation, FPN_ResNet represents the multi-scale feature pyramid and residual learning strategy, Head represents the detection head, and AB3DMOT represents the AB3DMOT multi-target tracking algorithm.

[0037] After completing all the above tasks, the vehicle side packages and compresses And the pose information P i =(x i ,y i ,θ i ) and sent to the cloud. Among them, (x i ,y i ) represents the center coordinate of the vehicle, θ i Indicates the heading angle.

[0038] Step 2: Build a cloud platform to integrate the perception information sent by all collaborative autonomous vehicles. The specific implementation is as follows:

[0039] Cloud platforms boast a vast server cluster, providing computing power far exceeding that of a single device. In this method, the cloud platform now handles the multi-source fusion task, previously performed by the vehicle itself and remote sensing satellites. This approach effectively reduces the computing pressure and hardware requirements on the vehicle-side devices. Leveraging the cloud platform's powerful computing power, the accuracy and real-time performance of the entire method can also be improved.

[0040] When the cloud receives the perception data from each collaborative autonomous driving vehicle, it first decompresses it to obtain the BEV feature map collection Track Collection Group and pose information P i =(x i ,y i ,θ i ). The features and trajectories are then fused separately.

[0041] In the feature fusion task, the present invention innovatively embeds satellite remote sensing data in order to provide autonomous driving with prior information such as global road structure and traffic flow, break through the local perception blind spot of a single vehicle, expand the perception range and enhance the semantic understanding ability in complex scenes. The process is shown in Figure 3 .

[0042] Specifically, self-attention is first used to fuse BEV feature maps from different vehicle sources. In computer vision, the advantage of self-attention is that its focus is global. It can obtain the global spatial information of the feature map through simple query and assignment. Specifically, the BEV feature map is first unified to the world coordinate system according to the posture of each vehicle; the present invention stacks all BEV feature maps in the world coordinate system to obtain the aggregate tensor Where N is the total number of BEV feature maps received by the cloud. Then, the self-attention mechanism is used to reduce the aggregate tensor from N dimensions to 1 dimension to obtain the cloud fusion feature The advantage of this is that for different feature maps of the same local area, self-attention will operate on multiple feature maps to infer interactions and thus capture the most representative features.

[0043] Then, the cloud calculates the perception expansion area A of each vehicle based on the detection radius R and posture information of the lidar. i (P i The area union of all vehicles is A. total =∪Α i This is the satellite data request range. Send A to the satellite total The coordinate range of the area is used to obtain high-resolution remote sensing satellite imagery and road vector data for the area. Satellite imagery may have brightness distortion due to factors such as sensor response differences and atmospheric scattering, which is corrected using histogram matching.

[0044] After correction, the DeepLabv3+ model is used to perform semantic segmentation on the satellite image and output the satellite road mask. The mask value 1 indicates a drivable road, and 0 indicates an obstacle. At the same time, the road curvature and speed limit are extracted from the road vector data, encoded and concatenated, and bilinearly upsampled to the size of the lidar BEV feature map H×W to obtain the satellite prior feature.

[0045] Satellite road mask Satellite prior characteristics To synthesize. First, directly in the channel dimension and Perform splicing to obtain the spliced features right Apply a linear transformation matrix (C sat is the satellite feature target dimension) is mapped to a feature space compatible with the lidar feature to form a satellite enhanced feature containing global road semantics and traffic rules, that is, Optimize W through training sat ,make Can maintain road semantics, and can integrate The prior rules provide high-quality input for subsequent fusion with lidar features.

[0046] Get the lidar cloud fusion features and satellite augmentation features Finally, a ground-to-ground fusion module of "vehicle-side local features-satellite global semantics" is constructed, and the idea of multi-head attention mechanism is used to fusion features of lidar cloud. and satellite augmentation features Perform dynamic fusion, adjust the importance of different features, and solve the problems of unclear feature relationships and information redundancy in simple splicing.

[0047]

[0048] Among them, W pri represents the satellite prior feature coding matrix, represents the linear transformation matrix, Q represents the query vector of the lidar cloud fusion feature, Q sat Represents satellite prior features After encoding matrix W pri The query vector obtained after is K, which represents the key vector and V, which represents the value vector. Represents the final fusion feature, MultiAttention represents the multi-head attention mechanism, softmax represents the normalized activation function, d k represents the dimension of the key vector K, and γ represents the weight coefficient. This formula dynamically adjusts the influence of satellite prior features on fusion through γ. For example, in a complex intersection scene, increasing γ makes the satellite road mask and prior features Enhance the guidance of fusion features, suppress the noise characteristics of lidar in non-road areas, and ensure It includes the local details of LiDAR while also conforming to the global road rules and structures provided by satellites.

[0049] In the trajectory fusion task, the present invention first uses the Hungarian matching algorithm to determine the trajectories from different sources that belong to the same instance, and then retains the longest trajectory of the same instance, which is beneficial for the subsequent trajectory prediction task. The specific implementation is shown in the following pseudo code:

[0050]

[0051] The present invention uses a Transformer encoder to perform trajectory prediction tasks. Since the trajectory lengths of various instances are inconsistent, the trajectories are used as inputs to the Transformer encoder in reverse chronological order, and equal-length vectors are generated through an embedding algorithm, wherein empty trajectories are masked by a mask. After the equal-length vectors pass through the encoding block, they act as keys and values in the attention mechanism. The present invention obtains the trajectory trend by fitting the trajectory, and selects different fitting methods according to the road in which they are located: for instances on straight roads, linear fitting is used, and for instances at intersections, nonlinear fitting is used. For shorter trajectories (less than 3 consecutive frames), based on motion similarity, shorter trajectories use trajectories of instances with similar motion states. Specifically, if the postures of two moving instances meet certain similarity conditions, then their motion changes in a short period of time will also be similar. The certain similarity conditions are: limiting the center distance between the two instances to less than 10m and the interpolation of the heading angle θ to less than 20 degrees. After the above steps, the possible low-confidence trajectory of each instance in the future is obtained, and these trajectories are used as queries in the attention mechanism. Finally, through the decoding block, a high-confidence fusion prediction trajectory is obtained. See the process Figure 4 .

[0052] Step 3: The vehicle integrates the received cloud information with the current characteristics of the vehicle, and finally obtains the vehicle-cloud collaborative perception result, such as Figure 5 As shown, the details are as follows:

[0053] By the time data is finally received by the vehicle, it has already passed a certain period of time since the data was collected. The actual location of instances in the traffic scene at the current moment may be somewhat different from the time of collection. This results in a lack of real-time perception results and an inability to provide a reliable basis for subsequent decision-making and control tasks. To address this issue, the present invention proposes a composite feature flow, combining predicted and actual trajectories to generate a feature flow that moves features to appropriate locations. This method has the advantage of not regenerating features, thus avoiding the generation of new noise.

[0054] The specific implementation of this method is as follows: At time t+2, the vehicle receives the final fusion feature map from the cloud. and fusion prediction trajectory According to the timestamp, the vehicle collects the instance bounding boxes of the target detection at time t in the historical information Fusion prediction trajectory at time t Do Hungarian matching, and move duplicate instances from Separation; Repeat the example using the real trajectory at time t+2 obtained by the multi-target tracking task of the vehicle Fusion prediction trajectory The remaining examples and the actual trajectory at time t obtained by the self-vehicle multi-target tracking task Add together to get the possible positions of all instances at the current moment Collection by possible location Get the flow n , the flow calculation formula is as follows:

[0055]

[0056] in, Represents the fusion prediction trajectory The coordinate values of the possible positions of all instances at time t obtained after processing, Represents the fusion prediction trajectory The coordinate values of the possible positions of all instances obtained after processing at time t+2;

[0057] After completing the flow construction, the complete flow set {flow n}, the flow graph can be obtained from the flow set Specifically: the flow graph is equal to flow in the ROI area n , zero padding is used outside the ROI area. The process of feature distortion is:

[0058]

[0059] in, It represents the final fusion feature map of the cloud at time t, h represents the flow map height, and w represents the flow map width. Represents the distorted feature map;

[0060] Subsequently, the BEV feature of the vehicle at time t+2 is integrated by fusion of the feature of the lidar BEV uploaded to the cloud by the vehicle in step 2. and The fusion feature of the vehicle is obtained, and then the detection head is used for fine classification and regression to obtain the target detection result collection of the vehicle. That is, the vehicle-cloud collaborative perception result of the vehicle itself.

[0061] Step 4: Training and loss function settings are as follows:

[0062] In the present invention, the target detection part and trajectory prediction are trained separately.

[0063] For the training of target detection tasks, the collaborative perception tasks in the prior art do not consider the positive or negative heading angles. However, in step 2 of the present invention, it is necessary to judge the motion similarity based on the heading angles. Therefore, the loss function of the target detection in the present invention is det Including smooth-L1 loss loss for positioning loc , cross entropy loss loss for object classification cls And the Softmax loss of positive and negative angles dir , the calculation formula is as follows:

[0064] loss det =loss cls +β loc loss loc +β dir loss dir (10)

[0065] Among them, the coefficient β loc =2, coefficient β dir =0.125, both are determined by empirical values.

[0066] For the training of trajectory prediction, the loss function loss pred Use root mean square error loss mse Calculate the offset between the predicted trajectory and the actual trajectory. The calculation formula is as follows:

[0067] loss pred =loss mse (11)

[0068] An embodiment of the present invention also proposes a vehicle-cloud collaborative wide-area perception system for unmanned vehicles based on the fusion of heaven and earth features, including a vehicle side and a cloud side. The vehicle side can execute the contents of steps one and three above; the cloud side can execute the contents of step two above.

[0069] The series of detailed descriptions listed above are only specific descriptions of feasible implementation methods of the present invention. They are not intended to limit the scope of protection of the present invention. Any equivalent methods or changes that do not deviate from the technology of the present invention should be included in the scope of protection of the present invention.

Claims

1. A vehicle-cloud collaborative wide-area perception method for unmanned vehicles based on the fusion of earth-ground features, characterized by: include: S1. Vehicle-side data processing: The lidar point cloud is processed through PointPillars columnarization combined with the FPN-ResNet backbone network to generate a BEV feature map. The detection head outputs target detection results containing information such as 3D coordinates and heading angles. The AB3DMOT multi-target tracking algorithm is then used to associate the detection results into a trajectory collection. The BEV feature map, trajectory collection, and vehicle posture information are then packaged and compressed and uploaded to the cloud. S2. Cloud data processing: After receiving the data, it first decompresses it to obtain the BEV feature map, trajectory collection, and vehicle posture information, and then performs feature fusion and trajectory fusion; S3. The vehicle integrates the received cloud information with the current characteristics of the vehicle itself, eliminates delays, and obtains the vehicle-cloud collaborative perception results.

2. The vehicle-cloud collaborative wide-area perception method for unmanned vehicles based on the fusion of earth-ground features according to claim 1 is characterized in that: The implementation of S1 is as follows: S1.1 Assume that a single cooperative autonomous driving vehicle at time t is a CAV i ∈{CAV}, i=1,2,…,m, Pointpillars is used as the encoder for point cloud target detection, and its input lidar point cloud is composed of N p points, each point is represented by a four-dimensional vector p = (x, y, z, r), where x, y, z are the coordinates of the point on the X, Y, and Z axes in the lidar coordinate system, and r is the reflection intensity; Pointpillars encoding is first performed to columnarize the original lidar point cloud. The backbone of Pointpillars adopts the multi-scale feature pyramid FPN strategy and the residual learning (ResNet) strategy, and introduces residual blocks to replace the convolution blocks in the original FPN. Each residual block contains 3 convolution layers and jump connections. The 3 convolution layers are: 1×1 convolution, used to reduce the number of channels; 3×3 convolution, as the main convolution layer of the residual block; 1×1 convolution, used to restore the number of channels; through the cross-layer connection mechanism, the gradient disappearance and gradient explosion problems in the training process are solved; the columnarized lidar point cloud passes through the backbone part to obtain the BEV feature map at time t Its spatial dimensions are: Among them, H is the height of the BEV feature map, W is the width of the BEV feature map, and C is the number of channels of the BEV feature map; S1.2 Using detection head to detect BEV feature map Perform fine classification and regression to obtain a collection of target detection results for collaborative autonomous driving vehicles n represents the total number of detected targets, and each element It consists of a tuple [x′, y′, z′, h, w, l, θ, s], which contains the 3D coordinates of the instance bounding box center [x′, y′, z′], the 3D size of the instance bounding box [h, w, l], the heading angle θ and the confidence score s; S1.3 uses the multi-target tracking algorithm AB3DMOT to associate the detected instance bounding box with the trajectory segment; first, use the 3D Kalman filter to predict the trajectory from the previous frame to the current frame; then, match the predicted trajectory of the 3D Kalman filter with the instance bounding box detected in the current frame; the 3D Kalman filter updates the instance trajectory based on the matching result, which creates a trajectory for the new instance and deletes the trajectory of the disappeared instance; after the above steps, the trajectory collection is obtained 3. The vehicle-cloud collaborative wide-area perception method for unmanned vehicles based on the fusion of earth-sky features according to claim 2 is characterized in that: The processing of S1 can be expressed by the following formula: F=Pillar(PC) (2) Among them, PC represents the input lidar point cloud, Pillar represents the volume pillarization operation, FPN_ResNet represents the multi-scale feature pyramid and residual learning strategy, Head represents the detection head, and AB3DMOT represents the AB3DMOT multi-target tracking algorithm.

4. The vehicle-cloud collaborative wide-area perception method for unmanned vehicles based on the fusion of earth-sky features according to claim 2 is characterized in that: The implementation method of the S2 feature fusion is as follows: S2.1 uses the self-attention mechanism to fuse BEV feature maps from different vehicle sources to obtain cloud-based fusion features S2.2 According to the detection radius R and posture P of the laser radar i , calculate the perception expansion area Α of each vehicle i , the regional union of all vehicles Α total =∪Α i That is the satellite data request range, send A to the satellite total The coordinate range of the area is used to obtain high-resolution remote sensing satellite images and road vector data of the area; S2.3 Use the DeepLabv3+ model to perform semantic segmentation on satellite images and output satellite road masks The mask value 1 indicates a drivable road and 0 indicates an obstacle. At the same time, the road curvature and speed limit are extracted from the road vector data, encoded and spliced, and bilinearly upsampled to the size of the lidar BEV feature map H×W to obtain the satellite prior feature. S2.4 Satellite road mask Satellite prior characteristics To synthesize; first directly in the channel dimension and Perform splicing to obtain the spliced features right Apply a linear transformation matrix Mapped to a feature space compatible with lidar features, C sat For the satellite feature target dimension, a satellite enhanced feature containing global road semantics and traffic rules is formed Optimize W through training sat ,make Can maintain road semantics, and can integrate A priori rules; S2.5 Get the LiDAR cloud fusion features and satellite augmentation features Finally, a space-ground fusion module of "vehicle-side local features-satellite global semantics" is constructed, and the idea of multi-head attention mechanism is used to integrate the features of lidar cloud. and satellite augmentation features Perform dynamic fusion.

5. The method for wide-area perception of unmanned vehicles based on vehicle-cloud collaboration based on fusion of earth-ground features according to claim 4 is characterized in that: The implementation of S2.1 is as follows: First, according to the posture of each vehicle, the BEV feature map is unified into the world coordinate system, and all BEV feature maps in the world coordinate system are stacked to obtain the aggregate tensor Where N is the total number of BEV feature maps received by the cloud; then the self-attention mechanism is used to reduce the aggregate tensor from N dimensions to 1 dimension to obtain the cloud fusion feature 6. The vehicle-cloud collaborative wide-area perception method for unmanned vehicles based on the fusion of earth-sky features according to claim 4 is characterized in that: In S2.2, the perception expansion area A of each vehicle i The composition is as follows: P i The center is a square with a side length of 2(R+ΔR), where ΔR is the extended redundancy.

7. The method for wide-area perception of unmanned vehicles based on vehicle-cloud collaboration based on fusion of earth-ground features according to claim 4 is characterized in that: The implementation of S2.5 is shown in the following formula: Among them, W pri represents the satellite prior feature coding matrix, represents the linear transformation matrix, Q represents the query vector of the lidar cloud fusion feature, Q sat Represents satellite prior features After encoding matrix W pri The query vector obtained after is K, which represents the key vector and V, which represents the value vector. Represents the final fusion feature, MultiAttention represents the multi-head attention mechanism, softmax represents the normalized activation function, d k Represents the dimension of the key vector K, γ represents the weight coefficient, and this formula dynamically adjusts the influence of satellite prior features on fusion through γ.

8. The vehicle-cloud collaborative wide-area perception method for unmanned vehicles based on the fusion of earth-sky features according to claim 2 is characterized in that: The implementation method of trajectory fusion in S2 is as follows: First, use the Hungarian matching algorithm to determine the trajectories from different sources that belong to the same instance, and then retain the longest trajectory of the same instance; Then use the Transformer encoder to predict the trajectory; since the trajectory lengths of each instance are inconsistent, the trajectories are used as the input of the Transformer encoder in reverse chronological order, and equal-length vectors are generated through the embedding algorithm. Among them, the empty trajectories are masked by the mask. After the equal-length vectors pass through the encoding block, they act as keys and values in the attention mechanism. The trajectory trend is obtained by fitting the trajectory, and different fitting methods are selected according to the road: for instances on straight roads, linear fitting is used, and for instances at intersections, nonlinear fitting is used; for shorter trajectories, according to motion similarity, shorter trajectories use trajectories of instances with similar motion states. Specifically, if the postures of two motion instances meet similar conditions, then their motion changes in a short period of time will also be similar; after the above steps, the possible low-confidence trajectory of each instance in the future is obtained, and these trajectories are used as queries in the attention mechanism; finally, the decoding block is used to obtain a high-confidence fusion prediction trajectory 9. The vehicle-cloud collaborative wide-area perception method for unmanned vehicles based on the fusion of earth-ground features according to claim 8 is characterized in that: The implementation of S3 is as follows: Using the composite feature flow idea, the predicted trajectory and the real trajectory are combined to generate a feature flow, and the features are moved to the appropriate position, as follows: Assume that the vehicle-side data upload is time t, the cloud fusion is time t+1, and the vehicle download is time t+2; At time t+2, the vehicle receives the final fusion feature map from the cloud. and fusion prediction trajectory According to the timestamp, the vehicle collects the instance bounding boxes of the target detection at time t in the historical information Fusion prediction trajectory at time t Do Hungarian matching, and move duplicate instances from Separation; Repeat the example using the real trajectory at time t+2 obtained by the multi-target tracking task of the vehicle Fusion prediction trajectory The remaining examples and the actual trajectory at time t obtained by the self-vehicle multi-target tracking task Add together to get the possible positions of all instances at the current moment Collection by possible location Get the flow n , the flow calculation formula is as follows: in, Represents the fusion prediction trajectory The coordinate values of the possible positions of all instances at time t obtained after processing, Represents the fusion prediction trajectory The coordinate values of the possible positions of all instances obtained after processing at time t+2; After completing the flow construction, the complete flow set {flow n }, the flow graph can be obtained from the flow set Specifically: the flow graph is equal to flow in the ROI area n , zero padding is used outside the ROI area, and the feature distortion process is: in, It represents the final fusion feature map of the cloud at time t, h represents the flow map height, and w represents the flow map width. Represents the distorted feature map; Subsequently, the BEV feature of the vehicle at time t+2 is integrated by fusion of the feature of the lidar BEV uploaded to the cloud by the vehicle in step S2. and The fusion feature of the vehicle is obtained, and then the detection head is used for fine classification and regression to obtain the target detection result collection of the vehicle. That is, the vehicle-cloud collaborative perception result of the vehicle itself.

10. A vehicle-cloud collaborative wide-area perception system for unmanned vehicles based on the fusion of earth-sky features, characterized by: It includes a vehicle side and a cloud side, wherein the vehicle side can execute the contents of S1 and S3 described in any one of claims 1-9; the cloud side can execute the contents of S2 described in any one of claims 1-9.

Citation Information

Patent Citations

  • Automatic driving track prediction method based on space-time pyramid

    CN115049130A

  • Feature-result level fusion vehicle infrastructure cooperative sensing method, medium and electronic equipment

    CN116958763A

  • Autonomous vehicle road cloud fusion sensing method

    CN117111085A

  • Feature-level multi-source multi-mode cloud sensing fusion system and method

    CN118097600A

  • Multi-source heterogeneous sensing information fusion target detection and tracking network model and method

    CN118823308A

Cited By

  • Lightweight parameter consensus aggregation method and system for intelligent driving scene

    CN121505851A

  • Multi-vehicle cooperative sensing method based on risk prediction

    CN122067434A