An air-ground collaborative perception method based on attention-enhanced hierarchical reinforcement learning
Patent Information
- Application Number
- CN202610991708.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-06
AI Technical Summary
但是,完整特征图传输存在数据冗余量大、带宽资源消耗高的问题;并且,计算与通信模型相互独立,难以结合切换时延完成端到端整体时延的统一量化;同时,线性加权优化方式难以实现感知精度与时延的均衡最优配置,单层固定决策模式无法适配空地节点动态变化,调度灵活性不足
通过本申请提供的一种基于注意力增强分层强化学习的空地协同感知方法,能够通过对异构智能体原始点云数据进行鸟瞰图特征映射与空间置信度掩码处理,生成轻量化第一稀疏特征图,实现感知特征高效压缩与低冗余传输;并且,结合异构智能体计算能力、任务角色以及空地动态信道特性,精准构建计算开销与异构通信传输模型,合成包含计算、通信、切换时延的端到端总时延,依托切比雪夫距离建立感知精度与时延联合的多目标优化问题,实现两项矛盾目标的均衡权衡;同时,采用高层与低层智能体分层决策机制分别确定协同更新间隔和合作伙伴选择策略,经多源特征压缩传输、空间对齐与加权融合输出精准目标检测结果,并通过轨迹数据采集、广义优势估计与近端策略优化算法联合迭代更新网络参数,持续适配空地环境动态变化,显著提升空地协同感知的决策合理性、任务实时性与目标检测精度。
Smart Images

Figure CN122533686B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular to an air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning. Background Technology
[0002] With the rapid development of intelligent connected autonomous driving, edge computing, and air-to-ground wireless communication technologies, the demands for environmental perception in complex urban road conditions are becoming increasingly stringent. A heterogeneous intelligent agent cluster consisting of intelligent vehicles, roadside units, and drones, utilizing LiDAR point cloud data for air-to-ground collaborative perception, has become a crucial technological approach for expanding the perception field of view, avoiding blind spots, and enhancing the environmental perception capabilities of autonomous driving.
[0003] Existing technologies commonly employ encoding the raw point clouds of heterogeneous agents into complete bird's-eye view feature maps for direct transmission and interaction. Separate modeling is used to construct agent computational overhead models and air-to-ground communication link models. A linear weighted approach is used to optimize perception accuracy and system latency. Simultaneously, a single-layer decision architecture or a fixed-time-period mechanism is employed to select collaborating partners and update policies. Traditional reinforcement learning methods are used to iterate model parameters, enabling multi-agent data sharing and joint perception detection. However, transmitting complete feature maps suffers from large data redundancy and high bandwidth consumption. Furthermore, the independent computation and communication models make it difficult to unify and quantify end-to-end overall latency by incorporating switching latency. Additionally, the linear weighted optimization method struggles to achieve a balanced optimal configuration between perception accuracy and latency, and the single-layer fixed decision model cannot adapt to dynamic changes in air-to-ground nodes, resulting in insufficient scheduling flexibility.
[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] To address the aforementioned issues, this application provides an air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning, which can balance perception accuracy and transmission latency, optimize agent cooperative decision-making, and significantly improve the real-time performance and detection accuracy of air-ground cooperative perception.
[0006] To achieve the objectives of this application, the following technical solution is provided:
[0007] This application provides an air-to-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning, including: The original point cloud data of the heterogeneous intelligent agent is acquired, and the point cloud data is mapped into a bird's-eye view feature map through a feature encoder. A first binary mask is generated based on spatial confidence, and the bird's-eye view feature map is masked to obtain a first sparse feature map for transmission. A computational overhead model is constructed based on the computational capabilities and roles of the heterogeneous intelligent agents. A communication transmission model for the heterogeneous links is constructed based on the dynamic channel characteristics. The end-to-end total latency is synthesized based on a parallel processing mechanism. The end-to-end total latency includes computational latency, communication latency, and communication switching latency. With sensing accuracy and end-to-end total latency as optimization objectives, we define sensing accuracy normalized distance and latency normalized distance respectively, and construct a multi-objective optimization problem based on Chebyshev distance that minimizes the weighted maximum deviation. Solving the multi-objective optimization problem involves a high-level agent making decisions on the collaborative update interval and a low-level agent making decisions on partner selection to obtain the optimal collaborative perception strategy; the optimal collaborative perception strategy includes a binary vector. According to the collaborative perception strategy, a collaborative task is performed to compress and transmit multi-source features, perform spatial alignment and weighted fusion, and obtain the final target detection result. Collect trajectory data of decision execution, calculate generalized advantage estimates, and use the proximal policy optimization algorithm to jointly update the network parameters of high-level and low-level agents.
[0008] In one possible implementation, the steps of acquiring the raw point cloud data of the heterogeneous intelligent agent, mapping the point cloud data into a bird's-eye view feature map through a feature encoder, generating a first binary mask based on spatial confidence, and performing masking processing on the bird's-eye view feature map to obtain a first sparse feature map for transmission include: Acquire the raw point cloud data collected by each heterogeneous intelligent agent; wherein, the raw point cloud data is a point cloud set containing multiple points, and each point contains three-dimensional coordinates and reflection intensity information; The preprocessed point cloud is mapped into bird's-eye view features by a feature encoder, wherein the bird's-eye view features are three-dimensional tensors; A spatial confidence map is generated based on the confidence level output by the detection head, a first binary mask is obtained, and the bird's-eye view features are multiplied element-wise with the first binary mask to obtain the first sparse feature map used for transmission.
[0009] In one possible implementation, the steps of constructing a computational overhead model based on the computational capabilities and roles of the heterogeneous intelligent agents, constructing a communication transmission model for the heterogeneous links based on dynamic channel characteristics, and synthesizing the end-to-end total latency based on a parallel processing mechanism include: Construct a computational overhead model, define the computational load according to the role of the intelligent agent, and calculate the processing latency in combination with the hardware computing capabilities; Construct a general communication link model and calculate the data transmission rate, communication delay, and communication handover delay based on transmission power, path loss, and dynamic bandwidth allocation. A heterogeneous channel state model is constructed, and the probability of occurrence of line-of-sight and non-line-of-sight links and the corresponding path loss are calculated for UAV-to-vehicle, roadside unit-to-vehicle, and vehicle-to-vehicle links respectively. By using a parallel processing mechanism, the overall time cost of system decision-making is determined based on computation latency, communication latency, and communication switching latency.
[0010] In one possible implementation, the step of defining the normalized distance for sensing accuracy and the normalized distance for delay as optimization objectives, and constructing a multi-objective optimization problem based on Chebyshev distance to minimize the weighted maximum deviation, includes: The normalized distance of perception accuracy is obtained by subtracting the actual average accuracy under the current state and the selected strategy from the preset maximum average accuracy, and then dividing by the range of accuracy variation. The normalized distance of the delay is obtained by subtracting the preset minimum possible delay from the total end-to-end delay and then dividing by the range of delay variation. Based on the normalized distance of the perception accuracy and the normalized distance of the time delay, a multi-objective optimization problem is constructed to minimize the weighted maximum deviation.
[0011] In one possible implementation, the multi-objective optimization problem is:
[0012] in, The binary vector chosen for the partner For the integer variable of the collaborative update interval, The weighting coefficients for perception accuracy. Normalized distance for sensing accuracy. The weighting coefficients for end-to-end delay. The normalized distance represents the time delay.
[0013] In one possible implementation, the step of solving the multi-objective optimization problem, through high-level agents deciding on the collaborative update interval and low-level agents deciding on partner selection, to obtain the optimal collaborative perception strategy includes: Construct a high-level Markov decision process model and use high-level intelligent agents to make decisions on collaborative update intervals; Construct a low-level Markov decision process model, use low-level agents to make partner selection decisions, and obtain the optimal collaborative perception strategy; A proximal policy optimization algorithm is used to jointly update the network parameters of high-level agents and low-level agents.
[0014] In one possible implementation, the step of performing a cooperative task according to the cooperative perception strategy, performing compressed transmission, spatial alignment, and weighted fusion of multi-source features to obtain the final target detection result includes: Based on the binary vector, the selected cooperative node generates a spatial confidence map through the confidence output by the detection head, generates a second binary mask, and performs masking processing on the local bird's-eye view feature map to obtain a second sparse feature map and transmits it to the vehicle. The vehicle receives all the second sparse feature maps and performs spatial transformation according to the pose transformation matrix of each cooperative node to align the second sparse feature maps to the vehicle coordinate system. By using an attention-based feature fusion operator, the local bird's-eye view feature map of the vehicle is aggregated with all aligned second sparse feature maps to obtain a fused feature map, and the final target detection result is obtained by prediction through the target detection head.
[0015] In one possible implementation, the formula for calculating the fused feature map is:
[0016] in, The fused feature map For attention-based feature fusion functions, Features of a bird's-eye view For spatial transformation functions, This is the second sparse feature map. Here is the pose transformation matrix. This is the set of collaborative nodes selected at the current moment.
[0017] In one possible implementation, the steps of collecting trajectory data of decision execution, calculating generalized advantage estimates, and jointly updating the network parameters of high-level and low-level agents using a proximal policy optimization algorithm include: Collect experience trajectory data and store the state, action, reward, and next state at each moment in the experience replay buffer; The advantage estimate for each time step is calculated using generalized advantage estimation. Construct a total loss function and update the policy network parameters and value network parameters of high-level and low-level agents using stochastic gradient descent.
[0018] In one possible implementation, the advantage estimate for each time step is:
[0019] in, for The advantage estimate, As a discount factor, For smoothing parameters, For at any time Time difference residuals; The total loss function is:
[0020] in, For the total loss function, A shearing loss function optimized for near-end policies. The mean squared error loss of the value network. For the entropy of the strategy, This is the weighting coefficient for value loss. is the weighting coefficient of the entropy regularization term.
[0021] The technical solution provided in this application may include the following beneficial effects: This application provides an air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning. This method generates a lightweight first sparse feature map by performing bird's-eye view feature mapping and spatial confidence masking on the raw point cloud data of heterogeneous agents, achieving efficient compression and low-redundancy transmission of perception features. Furthermore, by combining the computing power, task roles, and dynamic air-ground channel characteristics of the heterogeneous agents, a precise computational overhead and heterogeneous communication transmission model is constructed, synthesizing the end-to-end total latency including computation, communication, and handover delays. A multi-objective optimization problem jointly addressing perception accuracy and latency is established based on Chebyshev distance, achieving a balance between these two conflicting objectives. Simultaneously, a hierarchical decision-making mechanism is employed between high-level and low-level agents to determine the cooperative update interval and partner selection strategy. Accurate target detection results are output through multi-source feature compression transmission, spatial alignment, and weighted fusion. Network parameters are iteratively updated through trajectory data acquisition, generalized advantage estimation, and near-end policy optimization algorithms, continuously adapting to dynamic changes in the air-ground environment, significantly improving the decision-making rationality, task real-time performance, and target detection accuracy of air-ground cooperative perception.
[0022] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0024] Figure 1A flowchart illustrating an air-ground collaborative perception method based on attention-enhanced hierarchical reinforcement learning, provided for an embodiment of this application; Figure 2 A flowchart illustrating step S100 of an air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning provided in an embodiment of this application; Figure 3 A flowchart illustrating step S200 of an air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning provided in an embodiment of this application; Figure 4 A flowchart illustrating step S300 of an air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning provided in an embodiment of this application; Figure 5 A flowchart illustrating step S400 of an air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning provided in an embodiment of this application; Figure 6 A flowchart illustrating step S500 of an air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning provided in an embodiment of this application; Figure 7 This is a flowchart illustrating step S600 of an air-ground collaborative perception method based on attention-enhanced hierarchical reinforcement learning, provided in an embodiment of this application. Detailed Implementation
[0025] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0026] This example implementation first provides an air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning. (Reference) Figure 1 As shown, the air-to-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning may include the following steps: Step S100: Obtain the original point cloud data of the heterogeneous intelligent agent, map the point cloud data into a bird's-eye view feature map through a feature encoder, generate a first binary mask based on spatial confidence, perform masking processing on the bird's-eye view feature map, and obtain a first sparse feature map for transmission.
[0027] Step S200: Construct a computational overhead model based on the computational capabilities and roles of the heterogeneous intelligent agents, construct a communication transmission model for the heterogeneous links based on the dynamic channel characteristics, and synthesize the end-to-end total latency based on a parallel processing mechanism; the end-to-end total latency includes computational latency, communication latency, and communication switching latency.
[0028] Step S300: Taking the sensing accuracy and the end-to-end total delay as optimization objectives, define the normalized distance of sensing accuracy and the normalized distance of delay respectively, and construct a multi-objective optimization problem based on Chebyshev distance to minimize the weighted maximum deviation.
[0029] Step S400: Solve the multi-objective optimization problem by making high-level agents decide on the collaborative update interval and low-level agents decide on partner selection to obtain the optimal collaborative perception strategy; the optimal collaborative perception strategy includes a binary vector.
[0030] Step S500: Execute a collaborative task according to the collaborative perception strategy, perform compressed transmission, spatial alignment and weighted fusion of multi-source features, and obtain the final target detection result.
[0031] Step S600: Collect trajectory data of decision execution, calculate the generalized advantage estimate, and use the proximal policy optimization algorithm to jointly update the network parameters of high-level and low-level agents.
[0032] Below, we will refer to Figures 2 to 7 The steps of the air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning described in this example implementation will be explained in more detail.
[0033] In step S100, the original point cloud data of the heterogeneous intelligent agent is obtained, the point cloud data is mapped into a bird's-eye view feature map through a feature encoder, and a first binary mask is generated based on spatial confidence. The bird's-eye view feature map is then masked to obtain a first sparse feature map for transmission.
[0034] It should be noted that the heterogeneous intelligent agents include, but are not limited to, connected autonomous vehicles, roadside units, and drones. Different types of intelligent agents have different sensor models, installation locations, and acquisition perspectives, resulting in variations in the density, range, and noise level of the raw point cloud data. Mapping this data to a bird's-eye view feature space using a unified feature encoder can eliminate representational differences between heterogeneous data sources, providing a unified coordinate system and semantic space for subsequent cross-node feature fusion. Employing a binary mask based on spatial confidence for sparsity processing can significantly reduce the amount of data in the feature map to be transmitted while preserving the feature information of key target areas, effectively alleviating the problem of limited communication bandwidth in air-to-ground collaborative scenarios.
[0035] In one possible implementation, step S100 may further include the following sub-steps: In step S110, each heterogeneous intelligent agent is obtained. Raw point cloud data collected Wherein, the original point cloud data contains A point cloud collection of points, each point Includes three-dimensional coordinates and reflection intensity information .
[0036] It should be noted that the raw point cloud data is typically acquired by 3D sensors such as LiDAR. Each point contains 3D coordinates and reflection intensity information. The coordinates reflect the spatial position of the object's surface relative to the sensor, while the reflection intensity reflects the material properties of the object's surface, which can be used to help distinguish different types of targets.
[0037] In step S120, the preprocessed point cloud is mapped to bird's-eye view features by a feature encoder. The bird's-eye view feature is a three-dimensional tensor; the three-dimensional tensor is... ,in, For feature map height, The width of the feature map. This represents the number of feature channels.
[0038] It should be noted that the bird's-eye view feature encoder can adopt classic architectures such as PointPillars. This involves dividing the point cloud vertically into columnar sections, aggregating features from points within each column, and then processing this data through a two-dimensional convolutional network to generate a bird's-eye view feature map of dimension H×W×C. Here, H and W correspond to the rasterization resolution in the physical space along the length and width directions, respectively, and C corresponds to the number of feature channels at each grid location, encoding local geometric and semantic information such as height distribution, point density, and reflection intensity at that location. This feature map will serve as the foundational representation for all subsequent collaborative perception by intelligent agents.
[0039] In step S130, a spatial confidence map is generated based on the confidence level output by the detection head. The first binary mask is obtained. and the bird's-eye view features With the first binary mask Element-wise multiplication is performed to obtain the first sparse feature map used for transmission. .
[0040] It should be noted that the detection head is a lightweight convolutional network jointly trained with the feature encoder, and its output spatial confidence map represents the probability of a target existing at each spatial grid location. By setting a preset confidence threshold, a binary mask can be generated, where a value of 1 corresponds to a high-confidence foreground target region, and a value of 0 corresponds to a low-confidence background region. After element-wise multiplying the original bird's-eye view feature map with the binary mask, only the features of the foreground region are retained, while the features of the background region are set to zero, thereby achieving feature sparsity.
[0041] In step S200, a computational overhead model is constructed based on the computational capabilities and roles of the heterogeneous intelligent agents, a communication transmission model of the heterogeneous links is constructed based on the dynamic channel characteristics, and the end-to-end total latency is synthesized based on a parallel processing mechanism; the end-to-end total latency includes computational latency, communication latency, and communication switching latency.
[0042] It should be noted that in air-ground cooperative sensing systems, improvements in sensing performance come at the cost of additional computational load and communication overhead. Different types of intelligent agents differ significantly in computing power, communication conditions, and spatial location, and these factors directly determine the real-time performance and effectiveness of cooperative decision-making. By establishing a joint computational model of sensing accuracy and end-to-end latency, precise quantitative basis is provided for subsequent multi-objective optimization decisions, enabling the system to make an optimal trade-off between sensing gains and time costs.
[0043] In one possible implementation, step S200 may further include the following sub-steps: In step S210, a computational overhead model is constructed, the computational load is defined according to the role of the intelligent agent, and the processing latency is calculated in combination with the hardware computing capabilities.
[0044] It's important to note that the core of the computational overhead model lies in differentiating the computational load of different agents within the same collaborative task. For the collaborative sender, its computational tasks include generating the full feature map for its own perception, as well as additional mask generation and inference performed for communication sparsity, resulting in a heavy computational load. For the vehicle receiver, its computational tasks mainly involve lightweight preprocessing of its own point cloud and aligning and fusing multiple received collaborative features; the computational load is positively correlated with the number of selected collaborative nodes. Accurately modeling the computational load of each agent is a prerequisite for accurately estimating processing latency.
[0045] Furthermore, step S210 also includes: In step S211, the calculation delay is defined. The calculation formula is as follows:
[0046] in, The number of floating-point operations (FLOPs) required for the task. The computing power of the intelligent agent's hardware; In step S212, the computational load of the cooperating sender is determined. For each collaborative node Its workload includes full feature extraction workload. and communication mask to generate payload ,Right now ; In step S213, the calculation of the vehicle receiver is determined. Its workload includes lightweight preprocessing workload. and integration The fusion load required for each collaborative feature .
[0047] It should be noted that steps S211 to S213 together construct a computational overhead model for air-to-ground cooperative sensing. The core design idea of this model is to combine the abstract definition of computational latency with the specific load composition of two different roles in the cooperative sensing task, thereby accurately quantifying the computational time cost of each cooperative decision for each participating node. By splitting the load of the cooperating sender into full feature extraction and communication mask generation, and the load of the vehicle receiver into lightweight preprocessing and multi-source feature fusion, this model not only accurately reflects the asymmetric computational burden of each node in the cooperative sensing task, but also provides a computational latency input that can be independently calculated for each node in the subsequent step S240, which is based on a parallel processing mechanism for end-to-end latency synthesis. This fine-grained modeling approach, differentiated by role, allows subsequent partner selection decisions to fully consider the available computing resources of each candidate node, avoiding the excessive allocation of computational tasks to nodes with limited computing power, thus ensuring the overall real-time performance of the cooperative sensing system.
[0048] In step S220, a general communication link model is constructed, and the data transmission rate and communication delay are calculated based on the transmission power, path loss, and dynamic bandwidth allocation.
[0049] It should be noted that in air-ground collaborative sensing scenarios, the quality of communication links is dynamically affected by various factors such as distance between nodes, obstacle obstruction, channel fading, and available bandwidth. The purpose of constructing a general communication link model is to estimate the achievable data transmission rate of each candidate link in real time based on the current network state and physical environment parameters, thereby accurately predicting the time cost required to transmit a specific size of characteristic load, and providing key communication quality basis for partner selection decisions.
[0050] Furthermore, step S220 also includes: In step S221, the desired received power is calculated. Its formula is:
[0051] in, For transmission power, and These are the transmit and receive antenna gains, respectively. For miscellaneous losses, For channel The probability of occurrence It is a linear decay factor that includes path loss and shadow fading.
[0052] In step S222, the expected achievable rate is calculated based on Shannon's formula. The Shannon formula is as follows:
[0053] in, The frequency band bandwidth allocated to this link, This is the signal-to-interference-to-noise ratio.
[0054] In step S223, the communication delay is calculated. For data size The characteristic payload has a transmission delay of .
[0055] It should be noted that steps S221 to S223 above collectively construct a general communication link model for air-to-ground cooperative sensing. First, by comprehensively considering multiple factors such as transmit power, antenna gain, channel probability, path loss, and shadow fading, the expected received power in a dynamic air-to-ground environment is calculated, providing a physical layer basis for link quality assessment. Then, the expected received power is converted into a signal-to-interference-plus-noise ratio, and substituted into the Shannon formula to obtain the upper bound of the expected achievable rate of the link, thus mapping the physical channel conditions to a data transmission capability indicator understandable by the network layer. Finally, using the data size of the first sparse feature map to be transmitted as input, combined with the expected achievable rate, the communication delay of the link is calculated, completing the full-link quantization from channel characteristics to time cost. This provides the communication delay input of each candidate link for the subsequent step S240 to synthesize the end-to-end total delay, enabling subsequent partner selection decisions to fully consider the instantaneous channel quality and available bandwidth of each link, prioritizing nodes with low communication costs and high transmission reliability for cooperation, effectively controlling system communication overhead while ensuring sensing gain.
[0056] In step S230, a heterogeneous channel state model is constructed, and the probability of occurrence of line-of-sight and non-line-of-sight links and the corresponding path loss are calculated for UAV-to-vehicle, roadside unit-to-vehicle, and vehicle-to-vehicle links, respectively.
[0057] It should be noted that there are three typical communication link types in the air-to-ground cooperative sensing system: UAV-to-vehicle link, roadside unit-to-vehicle link, and vehicle-to-vehicle link. These three links differ fundamentally in node altitude, relative speed, obstruction distribution, and channel statistical characteristics, and cannot be described using a unified channel model. This step establishes a dedicated line-of-sight probability calculation method and path loss model for each link type to accurately capture the propagation characteristics of heterogeneous air-to-ground channels, providing a reliable basis for subsequent link quality assessment and latency prediction.
[0058] Furthermore, step S230 also includes: In step S231, for the UAV-CAV link, the elevation angle is calculated based on the height difference and horizontal distance between the UAV and the vehicle. And the Al-Hourani model was used to calculate the line-of-sight probability. .
[0059] In step S232, for the roadside unit-to-vehicle link (RSU-CAV), based on the communication distance... Calculate the probability of sight distance The path loss was calculated for both line-of-sight and non-line-of-sight conditions.
[0060] In step S233, for vehicle-to-vehicle links (CAV-CAV), the line-of-sight (LoS), vehicle-occluded (NLoSv), or non-line-of-sight (NLoS) states are determined based on the geometric occlusion relationship between vehicles, and the corresponding path loss and shadow fading are calculated.
[0061] It should be noted that steps S231 to S233 above construct differentiated channel state models for three typical heterogeneous communication links in the air-ground cooperative sensing system. The communication environment in air-ground cooperative scenarios is extremely complex. The three types of nodes—UAVs, roadside units, and vehicles—have fundamental differences in deployment altitude, mobility characteristics, and the distribution of surrounding obstructions, making it impossible to accurately describe their propagation characteristics using a unified channel model. Step S231, for the UAV-to-vehicle link, uses the elevation angle as a key geometric parameter and calculates the line-of-sight probability using the Al-Hourani model, fully capturing the channel advantages brought by the UAV's high-altitude perspective. Step S232, for the roadside unit-to-vehicle link, classifies the probabilities of line-of-sight and non-line-of-sight states based on communication distance, reflecting the characteristics of roadside units being fixedly deployed and located at a high position. Step S233, for the vehicle-to-vehicle link, introduces the intermediate channel state obstructed by vehicles, accurately characterizing the additional signal attenuation caused by dynamic vehicle occlusion between ground vehicles.
[0062] In step S240, the overall time cost of system decision-making is determined based on computation delay, communication delay, and communication switching delay through a parallel processing mechanism.
[0063] It should be noted that the end-to-end latency of the air-to-ground cooperative sensing task is not a simple sum of the individual sub-latencies, because the computation and communication processes of multiple cooperating nodes can be executed in parallel. A parallel processing mechanism is used to synthesize the latency, meaning that the overall time cost of the system depends on the sum of the slowest cooperating node completing the task and the vehicle's post-processing time.
[0064] Furthermore, step S240 also includes: In step S241, the data readiness delay is determined. This delay is taken as the maximum of the vehicle's local computation completion time and the task completion times of all collaborating partners, where the collaborating partners... The task completion time is used to calculate the delay. With communication delay sum.
[0065] In step S242, the total end-to-end delay is calculated. The data ready delay With vehicle after-processing delay and communication switching latency Add them together to get the result.
[0066] It should be noted that the communication handover latency To establish a constant value, step S241 compares the vehicle's local computation completion time with the task completion times of all selected collaborating nodes and takes the maximum value. The task completion time of each collaborating node is the sum of its own computation latency and communication latency, following a sequential dependency of local computation followed by cross-node transmission. Step S242 adds the data readiness latency, the vehicle's post-processing latency, and the switching latency to comprehensively depict the time consumption of the entire link from the initiation of the collaborative task to the output of fused features. The vehicle's post-processing latency includes the computation time required for spatial alignment and attention fusion, while the switching latency considers the signaling overhead that may be introduced when the vehicle moves between different network coverage areas. This total end-to-end latency will serve as the input variable for the latency-normalized distance in step S300.
[0067] In step S300, with sensing accuracy and the total end-to-end delay as optimization objectives, the normalized distance for sensing accuracy and the normalized distance for delay are defined respectively, and a multi-objective optimization problem based on Chebyshev distance is constructed to minimize the weighted maximum deviation.
[0068] It's important to note that in air-to-ground cooperative sensing, sensing accuracy and end-to-end latency are contradictory optimization objectives. Transmitting more cooperative feature maps can improve sensing accuracy, but it also increases communication and computation latency. Traditional methods typically use simple weighted summation to transform the multi-objective problem into a single-objective problem, but this approach struggles to find the optimal solution on a non-convex Pareto front. By introducing a multi-objective optimization modeling method based on Chebyshev distance, and minimizing the weighted maximum deviation, a Pareto optimal solution can be guaranteed under arbitrary weight settings, thus achieving precise control over the trade-off between sensing accuracy and latency.
[0069] In one possible implementation, step S300 may further include the following sub-steps: In step S310, the preset maximum average accuracy Subtract the current state and selection strategy Actual average accuracy Divide by the range of accuracy variation to obtain the normalized distance of perception accuracy. It is used to map the perceptual precision of different dimensions to a unified metric space.
[0070] It should be noted that the maximum average accuracy is the upper bound of the perception accuracy achievable under ideal conditions, i.e., all candidate nodes participate in cooperation and there are no communication and computation delay constraints. The actual average accuracy is the average accuracy value actually output by the detection head after multi-source feature fusion under the current state and the selected strategy. The difference between the two reflects the perception performance loss of the current decision relative to the ideal state. Dividing by the range of accuracy variation normalizes the perception accuracy loss to an interval, eliminating its dimensions and allowing it to be compared with the normalized latency index on the same scale.
[0071] In step S320, the current end-to-end total delay Subtract the preset minimum possible delay Divide by the range of time delay variation to obtain the normalized distance of time delay. It is used to map time delay data of different dimensions to a unified metric space.
[0072] It should be noted that the minimum possible latency is the lower bound of the end-to-end latency achievable under ideal conditions, i.e., using only the vehicle's local perception without any collaboration. The current total end-to-end latency is the actual total time cost calculated in step S200. The difference between the two reflects the additional time overhead due to the introduction of collaborative perception. Dividing by the latency variation range, the latency cost can be normalized to an interval, allowing it to be weighed against the perception accuracy loss within the same metric space.
[0073] In step S330, a multi-objective optimization problem is constructed to minimize the weighted maximum deviation based on the normalized distance of the sensing accuracy and the normalized distance of the time delay.
[0074] It should be noted that the objective function for minimizing the weighted maximum deviation takes the form of min-max. Its physical meaning is: given the perception accuracy weight and the latency weight, it considers the deviation between the two objectives simultaneously, with the optimization guided by controlling the performance of the worst-case objective. When the perception accuracy weight is greater than the latency weight, the system prioritizes maintaining perception accuracy, tolerating greater latency to ensure detection performance; conversely, when the latency weight is greater than the perception accuracy weight, the system prioritizes real-time performance, prioritizing collaborative schemes with lower communication and computational costs. The negative value of this objective function will serve as the instantaneous reward signal at each time step in the subsequent hierarchical reinforcement learning algorithm, guiding the agent to learn the optimal decision-making strategy.
[0075] Furthermore, the multi-objective optimization problem is as follows:
[0076] This represents the maximum weighted deviation value minimized over the entire task cycle, where... The binary vector chosen for the partner For the integer variable of the collaborative update interval, The weighting coefficients for perception accuracy. Normalized distance for sensing accuracy. The weighting coefficients for end-to-end delay. The normalized distance represents the time delay. and Used to adjust the relative importance of the two optimization objectives in the total cost.
[0077] In step S400, the multi-objective optimization problem is solved by making high-level agents decide on the collaborative update interval and low-level agents decide on the partner selection to obtain the optimal collaborative perception strategy; the optimal collaborative perception strategy includes a binary vector.
[0078] It's important to note that decision-making in air-to-ground cooperative sensing inherently possesses a hierarchical structure. The cooperative update interval, a long-term macro-level decision, determines the update frequency of lower-level strategies and needs adjustment based on the overall dynamic trends of the network. Partner selection, a short-term micro-level decision, requires refined screening based on the instantaneous states of each candidate node at each cooperative moment. Jointly optimizing both within the same action space results in a massive number of combinations, leading to training difficulties. By constructing a hierarchical reinforcement learning framework, the high-level and low-level decisions are decoupled. This reduces the dimensionality of their respective action spaces and allows the macro-level planning of the higher levels to constrain and guide the micro-level decisions of the lower levels, achieving efficient hierarchical optimization.
[0079] In one possible implementation, step S400 may further include the following sub-steps: In step S410, a high-level Markov decision process model is constructed, and a high-level intelligent agent is used to make decisions on the collaborative update interval.
[0080] It's important to note that the high-level Markov decision process model is responsible for making sequential decisions regarding the macroscopic variable of the collaborative update interval. Since the output of the high-level policy determines whether the low-level policy remains frozen for multiple subsequent time steps, the high-level agent needs to make decisions based on a global environmental summary rather than local details to capture the long-term trends of the entire system. Modeling high-level decisions as independent Markov processes allows the decision-making cycle of the high-level process to be much longer than that of the low-level process, adapting to the dynamic characteristics of the environment at different time scales.
[0081] Furthermore, step S410 also includes: In step S411, high-level status information is obtained. This state consists of a global environment summary and historical decision information; the global environment summary includes the average spatial feature matrix of all candidate nodes. Average confidence matrix and average resource index matrix Historical decision-making information includes the update interval of the previous decision-making cycle. And the accumulated reward value.
[0082] In step S412, the high-level agent determines the state based on... Output through policy network The action From a predefined discrete set Select an integer value This value determines the policy update frequency of the low-level agent, that is, the frequency of the low-level policy update in the next... It remains unchanged within each time step.
[0083] In step S413, the high-level reward value is calculated. It is defined as the interval between updates. The sum of instantaneous rewards for all lower-level time steps during the duration, wherein the instantaneous rewards consist of negative values of the Chebyshev distance as defined in step S300.
[0084] It should be noted that the above steps construct a complete closed loop of the high-level Markov decision process, realizing sequential decision-making on the macroscopic variable of the collaborative update interval. The high-level state information defined in step S411, by aggregating and averaging the individual states of all candidate nodes, shields the dimensional uncertainty caused by the dynamic changes in the number of nodes, providing the high-level agent with a compact input representation with fixed dimensions. Simultaneously, the introduction of historical decision information allows the high-level agent to perceive its own historical behavioral trajectory, thereby learning a smoother and more continuous interval adjustment strategy. In step S412, the high-level action is designed as an integer value selected from a predefined discrete set. This discretization design reduces the learning difficulty of the policy network while making the physical meaning of the update interval clear and explicit. In step S413, the high-level reward is defined as the cumulative value of the instantaneous rewards of all lower-level time steps during the interval's duration. This delayed feedback mechanism allows the high-level agent to accurately evaluate the comprehensive performance of the selected update interval over a longer period, thereby learning to make the optimal trade-off between decision response speed and computational cost.
[0085] In step S420, a low-level Markov decision process model is constructed, and the low-level agent is used to make decisions on partner selection to obtain the optimal cooperative perception strategy.
[0086] It should be noted that the low-level Markov decision process model is responsible for refining the selection of partners at each collaborative time point within the collaborative update interval specified by the high-level agents. Unlike the high-level model, which focuses on global trends, the low-level model needs to deeply understand the specific state of each candidate node and the interrelationships between them. Since the number and identity of candidate nodes change dynamically over time, the dimension of the low-level state space is variable, which places demands on the design of the policy network to handle variable-length inputs.
[0087] Furthermore, step S420 also includes: In step S421, low-level state information is obtained. The state contains three matrices: the candidate space matrix. Candidate confidence matrix and candidate resource matrix ; wherein, the space matrix each line Including the The one-hot encoded type vector, three-dimensional coordinates, velocity scalar, and Euler angle direction of each candidate node; the resource matrix Each row contains information about the node's computing power, line-of-sight probability, non-line-of-sight probability, and available bandwidth.
[0088] In step S422, inter-node dependency features are extracted based on the attention mechanism. The constructed node features are then input into a multi-head attention network, and the node embeddings are updated using the attention calculation formula:
[0089] in, For the first The output feature matrix of each attention layer; For the first The input feature matrix of the layer; This is a multi-head attention operation function used to capture the relationships between nodes; This is the layer normalization operation function, used to stabilize network training.
[0090] In step S423, the partner selection decision is output. The policy network of the low-level agent takes the enhanced node features as input and outputs a binary vector. , where the binary vector Dimensions and number of candidate nodes Consistent, its first If the nth element is 1, then the nth element is selected. If a node is selected as a partner, and the number is 0, then no node is selected.
[0091] It should be noted that the above steps construct a low-level partner selection model based on an attention mechanism, enabling refined dynamic screening of candidate nodes within the collaborative update interval specified by the high-level agent. Step S421, by organizing the state information of each candidate node into a structured representation with three dimensions—a spatial matrix, a confidence matrix, and a resource matrix—fully characterizes the comprehensive features of the node in terms of spatial situation, perception quality, and communication and computing resources. The introduction of one-hot encoded type vectors allows the model to distinguish the differentiated value of three types of heterogeneous nodes—connected autonomous vehicles, roadside units, and drones—in collaborative perception. Step S422 automatically learns the interdependencies between candidate nodes through a multi-head attention network, enabling the model to dynamically evaluate the contribution of each node to the vehicle's perception gain based on the current task context, without relying on manually designed heuristic screening rules. Layer normalization ensures the numerical stability of deep network training. Step S423 outputs a binary vector that transforms the node relationships extracted by the attention mechanism into explicit discrete choices. The vector dimension adaptively changes with the number of candidate nodes, ensuring the robustness of the model in scenarios where nodes dynamically join or leave.
[0092] In step S430, the network parameters of the high-level agent and the low-level agent are jointly updated using the near-end policy optimization algorithm.
[0093] It should be noted that, since the collaborative update interval of the high-level agent's decision-making directly determines how many time steps the low-level agent keeps its policy frozen, the decision-making effects of the two layers are highly coupled. A policy deviation in one layer will affect the quality of the other's learning signal. If the two networks are trained independently, the high-level agent will find it difficult to accurately assess the long-term impact of its chosen interval on the low-level partner selection effect, and the low-level agent will also find it difficult to adapt to the non-stationary environment brought about by changes in the high-level policy. By adopting a joint update mechanism, the high-level and low-level networks share the multi-objective optimization reward signal defined in step S300, and synchronize parameter updates through a unified proximal policy optimization objective. This allows the two policies to adapt to each other and co-evolve during training, ultimately converging to the globally optimal hierarchical collaborative perception policy.
[0094] In step S500, a collaborative task is performed according to the collaborative perception strategy, including compressed transmission, spatial alignment and weighted fusion of multi-source features, to obtain the final target detection result.
[0095] It should be noted that the selected collaborative node performs feature sparsity processing and cross-node transmission, while the autonomous vehicle is responsible for receiving, aligning, and fusing multi-source features, and finally outputting the target detection result.
[0096] In one possible implementation, step S500 may further include the following sub-steps: In step S510, based on the binary vector The selected collaborating node generates a spatial confidence map based on the confidence level output by the detection head, and then generates a second binary mask. The local bird's-eye view feature map is then masked to obtain the second sparse feature map. And transmit it to the vehicle.
[0097] It should be noted that the second binary mask is independently generated by the selected collaborating node after receiving the collaboration request. Due to differences in observation perspective, sensor performance, and environment among different collaborating nodes, the spatial confidence maps generated by each node are also different, resulting in node-specific binary masks and second sparse feature maps. The second sparse feature map generated locally by each node retains only the features of the foreground target region with high confidence from its own perspective, maximizing the compression of the amount of data that needs to be transmitted while preserving the most critical perceptual information.
[0098] Furthermore, the first and second binary masks will be explained in more detail: The first binary mask is used for global state evaluation before decision-making. It is generated by all candidate agents before the algorithm makes a decision. Its purpose is to extract lightweight initial features to be input into the reinforcement learning model to calculate the cost and solve the optimal policy. The second binary mask is used for the actual task execution after the decision. It is generated only by the selected cooperative nodes after the algorithm makes the decision. Its purpose is to accurately compress local features for the final decision result and to perform wireless transmission and fusion across nodes.
[0099] In step S520, the vehicle receives all the second sparse feature maps. And based on the pose transformation matrix of each cooperating node Perform spatial transformation Align the second sparse feature map to the vehicle coordinate system.
[0100] It should be noted that because the coordinate systems of each collaborating node and the vehicle are different, the generated bird's-eye view feature maps cannot be directly aligned in space. The spatial transformation operation maps each spatial position in the collaborative feature map from the collaborating node's coordinate system to the vehicle's coordinate system using the pose transformation matrix of the collaborating node relative to the vehicle. This transformation matrix is typically calculated jointly by the node's own localization system and the vehicle's localization system. The transformation process ensures that all received collaborative features are strictly aligned in space with the vehicle's local features.
[0101] In step S530, a feature fusion operator based on an attention mechanism is used. The local bird's-eye view feature map of the vehicle is aligned with all the second sparse feature maps. Aggregation is performed to obtain a fused feature map. The final target detection result is obtained by predicting the target using the target detection head.
[0102] It should be noted that the attention-based feature fusion operator differs from simple element-wise addition or concatenation. It adaptively assigns weights to features from different sources during the fusion process. When a feature provided by a collaborating node has higher confidence at a certain spatial location or stronger complementarity with the vehicle's features, the attention mechanism automatically increases its fusion weight at that location; conversely, for feature regions that may contain noise or redundancy, its weight is correspondingly reduced. This soft-weighted fusion method effectively improves the overall quality of the fused feature map and the accuracy of subsequent object detection. The object detection head is a convolutional network jointly trained with the feature encoder, directly predicting the object's category and bounding box on the fused feature map. The fused feature map serves as input to the detection network, where the detection head performs classification and regression tasks, thereby outputting information such as the object's location, category, size, and orientation.
[0103] Furthermore, the calculation formula for the fused feature map is as follows:
[0104] in, The fused feature map For attention-based feature fusion functions, Features of a bird's-eye view For spatial transformation functions, This is the second sparse feature map. Here is the pose transformation matrix. This is the set of collaborative nodes selected at the current moment.
[0105] In step S600, trajectory data of decision execution is collected, generalized advantage estimate is calculated, and network parameters of high-level and low-level agents are jointly updated using the proximal policy optimization algorithm.
[0106] It should be noted that after completing one or more collaborative perception tasks, the policy network and value network of the hierarchical reinforcement learning need to be updated using the collected experience data to continuously optimize decision quality. Step S400 is responsible for policy inference and action sampling, step S500 is responsible for policy execution and reward signal acquisition, and step S600 is responsible for experience data processing and network parameter updates. Through iteratively completing this closed-loop process, the model gradually learns the optimal collaborative update interval and partner selection strategy in a dynamic air-ground environment.
[0107] In one possible implementation, step S600 may further include the following sub-steps: In step S610, experience trajectory data is collected, and the state, action, reward, and state at each moment are stored in the experience playback buffer.
[0108] It should be noted that the empirical trajectory data contains a complete interaction record for each time step. Here, the state is the environmental information observed by the high-level or low-level agent at that moment; the action is the cooperative update interval or partner binary vector selected by the agent based on the current policy network; the reward is the instantaneous feedback signal calculated from the negative Chebyshev distance defined in the multi-objective optimization problem of step S300; and the next state is the new observation state that the environment transitions to after the agent performs the action. Storing this trajectory data in the empirical replay buffer can break the temporal correlation between samples in subsequent parameter updates, improving training stability and sample utilization.
[0109] In step S620, the advantage estimate for each time step is calculated using generalized advantage estimation (GAE). .
[0110] It should be noted that Generalized Advantage Estimation (GAE) is a commonly used advantage function estimation method in near-end policy optimization algorithms. By introducing a smoothing parameter, GAE can achieve a flexible trade-off between bias and variance. When the variance is close to 0, the advantage estimate has a small variance but may introduce a large bias; when the variance is close to 1, the advantage estimate has a small bias but a large variance. The time-difference residual measures the difference between the actual reward of the current state and the value network prediction. By weighting and summing this residual according to the discount factor and the smoothing parameter, a comprehensive advantage estimate considering multi-step reward signals can be obtained.
[0111] Furthermore, the advantage estimate for each time step is:
[0112] in, for The advantage estimate, This is a discount factor used to balance current rewards with future rewards; This is a smoothing parameter used to balance bias and variance; For at any time Time difference residuals.
[0113] In step S630, the total loss function is constructed. The policy network parameters and value network parameters of high-level and low-level agents are updated using stochastic gradient descent.
[0114] It should be noted that the total loss function consists of three terms: the shearing loss function for near-end policy optimization, which constrains the probability ratio between the old and new policies to prevent policy collapse caused by excessively large single update steps; the mean squared error loss of the value network, which makes the value estimation closer to the true reward and provides an accurate baseline for advantage estimation; and the policy entropy regularization term, which encourages the agent to maintain a certain degree of exploration randomness and prevents premature convergence to local optima. By jointly optimizing these three losses through stochastic gradient descent, the decision-making level of both high-level and low-level agents can be gradually improved while maintaining update stability. The weight coefficients control the relative importance of the value loss and entropy regularization term in the overall optimization objective, respectively.
[0115] The total loss function is:
[0116] in, For the total loss function, A shearing loss function optimized for near-end policies is used to limit the magnitude of policy updates. The mean squared error loss of the value network. The entropy of the strategy is used to encourage exploration. This is the weighting coefficient for value loss. is the weighting coefficient of the entropy regularization term.
[0117] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
[0118] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. This application is not limited to the exact structures described above and illustrated in the accompanying drawings, and it should not be considered that the specific implementation of this application is limited to these descriptions. For those skilled in the art, various changes and modifications made without departing from the concept of this application should be considered to fall within the protection scope of this application.
Claims
1. An air-ground collaborative perception method based on attention-enhanced hierarchical reinforcement learning, characterized in that, include: The original point cloud data of the heterogeneous intelligent agent is acquired, and the point cloud data is mapped into a bird's-eye view feature map through a feature encoder. A first binary mask is generated based on spatial confidence, and the bird's-eye view feature map is masked to obtain a first sparse feature map for transmission. A computational overhead model is constructed based on the computational capabilities and roles of the heterogeneous intelligent agents. A communication transmission model for the heterogeneous links is constructed based on the dynamic channel characteristics. The end-to-end total latency is synthesized based on a parallel processing mechanism, including: A computational overhead model is constructed, defining the computational load according to the role of the intelligent agent and calculating the processing latency based on hardware computing capabilities; a general communication link model is constructed, calculating the data transmission rate, communication latency, and communication handover latency based on transmit power, path loss, and dynamic bandwidth allocation; a heterogeneous channel state model is constructed, calculating the probability of line-of-sight and non-line-of-sight occurrences and the corresponding path loss for UAV-to-vehicle, roadside unit-to-vehicle, and vehicle-to-vehicle links respectively; and the overall time cost of system decision-making is determined based on computational latency, communication latency, and communication handover latency through a parallel processing mechanism; the end-to-end total latency includes computational latency, communication latency, and communication handover latency; wherein, the step of constructing the computational overhead model, defining the computational load according to the role of the intelligent agent, and calculating the processing latency based on hardware computing capabilities includes: defining the computational latency. The calculation formula is as follows: ,in, The number of floating-point operations (FLOPs) required for the task. The computational capabilities of the intelligent agent's hardware; determining the computational load of the cooperating sender. For each collaborative node Its workload includes full feature extraction. and communication mask to generate payload ,Right now ; Calculation to determine the recipient of the vehicle Its workload includes lightweight preprocessing workload. and integration The fusion load required for each collaborative feature ; The steps for constructing the heterogeneous channel state model, specifically for UAV-to-vehicle, roadside unit-to-vehicle, and vehicle-to-vehicle links, to calculate the probability of line-of-sight and non-line-of-sight occurrences and the corresponding path loss, include: for UAV-to-vehicle links, calculating the elevation angle based on the altitude difference and horizontal distance between the UAV and the vehicle. And the Al-Hourani model was used to calculate the line-of-sight probability. For roadside unit-to-vehicle links, based on communication distance Calculate the probability of sight distance The path loss is calculated for line-of-sight and non-line-of-sight states respectively. For vehicle-to-vehicle links, the line-of-sight, vehicle-occluded, or non-line-of-sight states are determined based on the geometric occlusion relationship between vehicles, and the corresponding path loss and shadow fading are calculated. Using sensing accuracy and the total end-to-end delay as optimization objectives, we define the normalized distance for sensing accuracy and the normalized distance for delay, respectively, and construct a multi-objective optimization problem based on Chebyshev distance that minimizes the weighted maximum deviation, including: The normalized distance of the perception accuracy is obtained by subtracting the actual average accuracy under the current state and selected strategy from the preset maximum average accuracy, and then dividing by the range of accuracy variation. The normalized distance of the delay is obtained by subtracting the preset minimum possible delay from the current end-to-end total delay, and then dividing by the range of delay variation. Based on the normalized distance of the perception accuracy and the normalized distance of the delay, a multi-objective optimization problem is constructed to minimize the weighted maximum deviation. The multi-objective optimization problem is as follows: ,in, The binary vector chosen for the partner For the integer variable of the collaborative update interval, The weighting coefficients for perception accuracy. Normalized distance for sensing accuracy. The weighting coefficients for end-to-end delay. Normalized distance for time delay; Solving the multi-objective optimization problem involves a high-level agent making decisions on the collaborative update interval and a low-level agent making decisions on partner selection to obtain the optimal collaborative perception strategy; the optimal collaborative perception strategy includes a binary vector. According to the collaborative perception strategy, a collaborative task is performed to compress and transmit multi-source features, perform spatial alignment and weighted fusion, and obtain the final target detection result. Collect trajectory data of decision execution, calculate generalized advantage estimates, and use the proximal policy optimization algorithm to jointly update the network parameters of high-level and low-level agents.
2. The method of claim 1, wherein, The steps of acquiring the raw point cloud data of the heterogeneous intelligent agent, mapping the point cloud data into a bird's-eye view feature map through a feature encoder, generating a first binary mask based on spatial confidence, and performing masking processing on the bird's-eye view feature map to obtain a first sparse feature map for transmission include: Acquire the raw point cloud data collected by each heterogeneous intelligent agent; wherein, the raw point cloud data is a point cloud set containing multiple points, and each point contains three-dimensional coordinates and reflection intensity information; The preprocessed point cloud is mapped into bird's-eye view features by a feature encoder, wherein the bird's-eye view features are three-dimensional tensors; A spatial confidence map is generated based on the confidence level output by the detection head, a first binary mask is obtained, and the bird's-eye view features are multiplied element-wise with the first binary mask to obtain the first sparse feature map used for transmission.
3. The method of claim 1, wherein, The steps for solving the multi-objective optimization problem, which involve high-level agents deciding on the collaborative update interval and low-level agents deciding on partner selection to obtain the optimal collaborative perception strategy, include: Construct a high-level Markov decision process model and use high-level intelligent agents to make decisions on collaborative update intervals; Construct a low-level Markov decision process model, use low-level agents to make partner selection decisions, and obtain the optimal collaborative perception strategy; A proximal policy optimization algorithm is used to jointly update the network parameters of high-level agents and low-level agents.
4. The air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning according to claim 1, characterized in that, The step of executing a collaborative task according to the collaborative perception strategy, performing compressed transmission, spatial alignment, and weighted fusion of multi-source features to obtain the final target detection result includes: Based on the binary vector, the selected cooperative node generates a spatial confidence map through the confidence output by the detection head, generates a second binary mask, and performs masking processing on the local bird's-eye view feature map to obtain a second sparse feature map and transmits it to the vehicle. The vehicle receives all the second sparse feature maps and performs spatial transformation according to the pose transformation matrix of each cooperative node to align the second sparse feature maps to the vehicle coordinate system. By using an attention-based feature fusion operator, the local bird's-eye view feature map of the vehicle is aggregated with all aligned second sparse feature maps to obtain a fused feature map, and the final target detection result is obtained by prediction through the target detection head.
5. The air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning according to claim 4, characterized in that, The formula for calculating the fused feature map is: in, The fused feature map For attention-based feature fusion functions, Features of a bird's-eye view For spatial transformation functions, This is the second sparse feature map. Here is the pose transformation matrix. This is the set of collaborative nodes selected at the current moment.
6. The air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning according to claim 1, characterized in that, The steps of collecting trajectory data of decision execution, calculating generalized advantage estimates, and jointly updating the network parameters of high-level and low-level agents using a proximal policy optimization algorithm include: Collect experience trajectory data and store the state, action, reward, and next state at each moment in the experience replay buffer; The advantage estimate for each time step is calculated using generalized advantage estimation. Construct a total loss function and update the policy network parameters and value network parameters of high-level and low-level agents using stochastic gradient descent.
7. The air-ground cooperative perception method based on attention-enhanced hierarchical reinforcement learning according to claim 6, characterized in that, The advantage estimate for each time step is: in, for The advantage estimate, As a discount factor, For smoothing parameters, For at any time Time difference residuals; The total loss function is: in, For the total loss function, A shearing loss function optimized for near-end policies. The mean squared error loss of the value network. For the entropy of the strategy, The weighting coefficient for value loss. is the weighting coefficient of the entropy regularization term.