A Multi-Agent Air-Ground Network Resource Allocation Method Based on Federated Learning

By employing a multi-agent resource allocation method based on federated learning and deep reinforcement learning, the dynamic and security issues of resource allocation in air-to-ground networks are addressed, achieving low-latency and highly reliable air-to-ground network communication and meeting the data privacy and security requirements of intelligent transportation systems.

CN116546462BActive Publication Date: 2026-05-26NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2023-04-26
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In intelligent transportation systems, the dynamic changes in the motion states of vehicles and drones make it difficult to accurately model the air-to-ground network topology, hindering network resource allocation and posing data sharing security threats. Existing centralized machine learning methods cannot meet the requirements for high latency and data privacy protection.

Method used

We adopt a multi-agent resource allocation method based on federated learning. By constructing a deep reinforcement learning model, we optimize the spectrum resource allocation of V2I and V2U links. We combine hybrid spectrum access technology and orthogonal frequency division multiplexing, and use the Fed-D3QN algorithm to optimize the resource allocation strategy to protect user privacy and data security.

Benefits of technology

It achieves reduced total latency of V2I links, improved reliability and transmission rate of V2U links in highly dynamic air-to-ground networks, meets the service quality requirements of different links, and has good stability and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116546462B_ABST
    Figure CN116546462B_ABST
Patent Text Reader

Abstract

This invention discloses a multi-agent air-to-ground network resource allocation method based on federated learning. In the air-to-ground network, the ground network constitutes high-data-rate V2I links; the air network constitutes V2U links for direct communication with ground vehicles; the V2U links share the spectrum resources of the V2I links and use hybrid spectrum access technology for transmission; a network resource allocation system model consisting of M pairs of V2I links and K pairs of V2U links is constructed; a multi-agent resource allocation method is adopted, and a deep reinforcement learning model is constructed with the goal of minimizing the total transmission delay of the V2I link channels; federated learning is used to optimize the deep reinforcement learning model; during the execution phase, the V2U links obtain their current state based on observations, and the optimal resource allocation strategy is obtained using the trained model. This invention exhibits good stability in highly dynamic air-to-ground networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle networking technology, and particularly relates to a resource allocation method for vehicle networking, especially a multi-agent air-ground network resource allocation method based on federated learning. Background Technology

[0002] In intelligent transportation systems, vehicle-to-everything (V2X) networks relying solely on ground infrastructure have limited coverage and complex propagation environments, making it difficult to meet the application requirements of vehicular networks in autonomous driving, dynamic intelligent traffic management, and harsh environments such as emergency response and wilderness use. On the other hand, space-based platforms built using high-orbit satellites often introduce significant latency, failing to meet the needs of most latency-sensitive services in V2X. Therefore, air-to-ground networks are more suitable for the requirements of V2X. Low-altitude unmanned aerial vehicles (UAVs), with their high response speed, high bandwidth, highly reliable line-of-sight transmission, and flexible maneuverability, can constitute an aerial platform for V2X, serving as an important supplement to ground-based V2X networks. Through effective cooperation between vehicle-to-UAV (V2U) links, sensor data and control information can be transmitted between ground and air subnetworks, further enhancing the computing power of ground-based V2X networks and providing computing resources for vehicles to assist in achieving low-latency, highly reliable communication.

[0003] However, because both vehicles and drones are in motion, the air-to-ground network topology changes rapidly, and the spatial dimensions of vehicle and drone states and actions are constantly increasing, making it difficult to obtain accurate information about the global environment. While precise network description and modeling are challenging, the increase in terminal devices, the massive growth in data, and the demand for quality of service (QoS) make the allocation of scarce network resources extremely difficult. Furthermore, terminal entities face more severe security risks than ordinary ground-based terminal entities. Due to untrusted network environments, unreliable tracking of improper behavior, and low-quality shared data, data sharing between vehicles and drones presents potential security threats.

[0004] Machine learning (ML), particularly deep reinforcement learning (DRL), is an emerging algorithm for processing big data and data analysis. Leveraging the learning and predictive capabilities of deep learning, it can effectively support resource management in vehicular networks. Most proposed DRL models are centralized, failing to consider the unreliable communication connections in highly dynamic air-to-ground networks. However, if data exists in silos to protect user privacy, centralized machine learning becomes impractical, failing to meet the demands of network intelligence for data labels and feature dimensions. Furthermore, stringent latency requirements and limited local training data pose significant challenges to training DRL models.

[0005] Therefore, the above problems urgently need to be solved. Summary of the Invention

[0006] Purpose of the invention: The purpose of this invention is to provide a multi-agent air-to-ground network resource allocation method based on federated learning. This method addresses the QoS requirements of different links in the air-to-ground network, with the optimization objective being to minimize the total transmission delay of the vehicle-to-infrastructure (V2I) link channel. It ensures the reliability of the V2U link by constraining power delay and exhibits good stability in highly dynamic air-to-ground networks.

[0007] Technical Solution: To achieve the above objectives, this invention discloses a multi-agent air-to-ground network resource allocation method based on federated learning, comprising the following steps:

[0008] (1) In the air-ground network, the ground network consists of infrastructure and vehicle user equipment, forming a V2I link for high data rate service; the air network consists of unmanned aerial vehicles, forming a V2U link for direct communication with ground vehicles; the V2U link is used to collect important information related to driving safety.

[0009] (2) The V2U link shares the spectrum resources of the V2I link and uses hybrid spectrum access technology for transmission;

[0010] (3) Construct a network resource allocation system model consisting of M pairs of V2I links and K pairs of V2U links;

[0011] (4) Using a multi-agent resource allocation method, a deep reinforcement learning model is constructed with the goal of minimizing the total transmission delay of the V2I link channel, taking into account the reliability and latency of the V2U link.

[0012] (5) In order to improve the performance of multi-agent deep reinforcement learning models while protecting user privacy and data security, federated learning is used to optimize deep reinforcement learning models.

[0013] (6) During the execution phase, the V2U link obtains the current state based on observation and uses the trained model to obtain the optimal resource allocation strategy.

[0014] Step (3) includes the following specific steps:

[0015] (3.1) The network resource allocation system model includes M pairs of V2I links and K pairs of V2U links, which are represented by sets M = {1, 2, ..., M} and K = {1, 2, ..., K}, respectively; V2I and V2U communication adopts orthogonal frequency division multiplexing technology, which divides the channel into M flat fading orthogonal sub-channels with bandwidth W, and the m-th VUE user occupies the m-th sub-channel in advance for communication;

[0016] (3.2) The channel power gain of the k-th V2U link in the m-th V2I link subband is defined as:

[0017] g k [m]=η k h k [m]

[0018] Where, η k For large-scale fading, including path loss and shadowing fading, it is assumed to be frequency-independent; h k [m] represents small-scale fading, which follows Rayleigh fading patterns within the subband and at uncorrelated times;

[0019] The SINR of the m-th V2I link can be expressed as:

[0020]

[0021] The channel capacity of the m-th V2I link, calculated using Shannon's formula, can be expressed as:

[0022]

[0023] in, and Let σ represent the transmit power of the m-th VEU and the k-th UAV, respectively. 2 G represents noise power. m [m] represents the power gain of the m-th V2I channel. ρ represents the interference power gain from the k-th V2U link to the m-th V2I link; k [m] represents the binary subband allocation indicator for the spectrum multiplexing flag, ρ k [m] = 1 indicates that the k-th UAV reuses the spectrum of the m-th VEU; otherwise, ρ k [m] = 0;

[0024] (3.3) For the k-th V2U link, its subband selection information is:

[0025] ρ k ={ρ k [1], ρ k [2],…,ρ k [m],…,ρ k [M]}

[0026] The rule stipulates that each link can only select one resource block for transmission at any given time.

[0027] (3.4) The SINR of the k-th V2U link in the m-th subband can be expressed as:

[0028]

[0029] The channel capacity of the k-th V2U link in the m-th subband can be expressed as:

[0030]

[0031] in,

[0032]

[0033]

[0034] These represent interference from V2I links using the same spectrum and interference from other V2U pairs, respectively. It is the interference gain of the k′-th V2U link on the k-th V2U link;

[0035] (3.5) The V2U link is mainly responsible for reliably transmitting safety-critical information, which is generated periodically according to the mobility of the vehicle. Considering the decentralized resource allocation at the UAV end, only the transmission delay is considered as the delay of the V2U link, and the constraint of the V2U link on the delay is determined.

[0036] (3.6) The reliability constraints of V2U communication can be achieved by controlling the probability of interruption events, which can be described as received events. Below the predetermined threshold Determine reliability requirements;

[0037] (3.7) In air-to-ground communication networks, the design objective is to minimize the total latency of V2I link transmission, defined as...

[0038]

[0039] (3.8) Taking into account the service quality requirements of different links, establish the objective function and optimization conditions.

[0040] Furthermore, the V2U link latency constraint in step (3.5) can be written as follows:

[0041]

[0042] Among them, B k For the remaining payload that the UAV needs to transmit, T k ≤T max To reduce the maximum tolerable delay T max The remaining delay is calculated at the beginning.

[0043] Preferably, the reliability requirement in step (3.6) is expressed as follows:

[0044]

[0045] Where Pr{·} represents the probability of the input. The minimum SINR required to establish a reliable link for UAV, where p0 is the minimum interruptibility probability of the V2U link; under Rayleigh fading conditions, the reliability constraint is further transformed into:

[0046]

[0047] Where, γ th It is the SINR threshold of the UAV receiver on the k-th V2U link.

[0048] Furthermore, the objective function and optimization conditions established in step (3.8) are as follows:

[0049]

[0050]

[0051]

[0052]

[0053]

[0054]

[0055] The objective function is to minimize the total latency of the V2I link. Constraints C1 and C2 are the reliability and latency constraints of the V2U link. Constraint C3 states that the total power transmitted by the UAV on all subbands cannot exceed the maximum rated transmission power. Constraints C4 and C5 mean that each V2U link can only be assigned to one subband, but multiple V2U links can access the same subband.

[0056] Furthermore, step (4) includes the following specific steps:

[0057] (4.1) Define the state space to include all local observation information related to resource allocation.

[0058] (4.2) Define the action space of each agent as the selection of the spectrum sub-bands and the control of the transmission power of the V2U link, expressed as follows:

[0059]

[0060] in, Let C be the transmit power of the k-th V2U link user.k ∈{1, 2, ..., M} indicates that the k-th V2U link user has accessed the m-th subchannel;

[0061] (4.3) Define a reward function to reflect the optimization objective of the constraint problem. The optimization objective of minimizing the instantaneous total latency of all V2I links and the constraint of V2U link service quality are reflected in the reward at each time t, expressed as follows:

[0062]

[0063] in,

[0064]

[0065] Among them, D m [m, t] represents the latency of each V2I link; the smaller the total latency, the greater the reward. k (t) represents the effective transmission rate of the V2U link, reflecting the success rate of V2U link transmission; λ v and λ u It is a positive weight used to balance the contribution of V2I and V2U goals to the reward function, and needs to be adjusted based on experience;

[0066] (4.4) Introducing deep learning, a deep neural network is used to replace the Q-table to fit the state-action value Q, resulting in DQN;

[0067] (4.5) Introducing the dual-depth Q-network method, two neural networks with the same structure but different parameters are constructed as the training network Q. E (s, a; θ) and the target network Q T (s, a; θ), decoupling action selection from Q-value estimation, is expressed as:

[0068]

[0069] The training network selects actions by continuously updating parameters θ, and the target network Q... T Used to estimate the Q value, parameter θ - The parameters are kept fixed and replaced with the latest estimated network parameters θ at regular intervals.

[0070] (4.6) Introduce a competitive network, where the neural network consists of a state-value network V(s; θ). V ) and action advantage function network A(s, a; θ AThe algorithm is composed of two parts: the value function V(s) of the training network and the value function A(s,a) of the target network, which together balance the influence of actions on the Q-value in the training network and the target network, remove reward bias, and finally output the Q-value by adding the value function V(s) of the state and the advantage function A(s,a) of each action, and subtracting the average of the advantage functions of all actions in a certain state to ensure that for a given Q-value, the range of Q-values ​​is narrowed, and there are uniquely determined V(s) and A(s,a), removing redundant degrees of freedom and improving the stability of the algorithm. This can be expressed as:

[0071]

[0072] (4.7) The action advantage function represents the relative merit of a particular action compared to other actions in a given state. It is a measure of the relative merit of different actions in the current state. Unlike DQN, which directly learns all Q values, D3QN can distinguish whether the current reward is caused by the state itself or by the chosen action. Combined with the DDQN estimation method, the D3QN estimate is expressed as:

[0073]

[0074] (4.8) After each agent estimates the Q-value, it uses gradient descent to minimize the loss function and update the neural network parameters. The target network parameters are copied from the training network parameters at fixed intervals to complete the update of the target network, denoted as:

[0075] .

[0076] Furthermore, the state space is defined in step (4.1) as follows:

[0077] s t (k)={{G k [m]} m∈M ,{I k [m]} m∈M B k ,T k ,e,ε}

[0078] in,

[0079] G k [m]={g m [m],g m,k [m],g k [m],{g k',k [m]} k'≠k}

[0080] It is the set of local instantaneous channel information on the uplink of subchannel m, g m [m] represents the channel power gain of the m-th V2I link. G represents the interference power gain from the k-th V2U link to the m-th V2I link.k [m] is the channel power gain of the k-th V2U link. It is the interference power gain of the k'-th V2U link on the k-th V2U link;

[0081] in,

[0082]

[0083] They are V2I links on the same spectrum. Interference with other V2U pairs The sum;

[0084] Among them, B k and T k These represent the remaining load and remaining latency that the V2U user needs to transmit.

[0085] Preferably, in step (4.4), DQN is specifically represented as follows:

[0086]

[0087] Where θ is the neural network parameter and γ is the discount factor.

[0088] Furthermore, step (5) includes the following specific steps:

[0089] (5.1) The V2U link client uploads its local model to the server to execute the aggregation algorithm and obtain global network parameters. The aggregation algorithm performs a weighted average of all client models participating in federated learning according to their contribution to utilize global experience for training and maximize the aggregation effect. The specific formula is as follows:

[0090]

[0091] Where, θ t These are the parameters of the server's neural network at time t. N represents the neural network parameters of the k-th client at time t; k N and N are the training batch sizes of the k-th client and all clients, respectively. The ratio of these batch sizes is used to measure the contribution of the k-th client and serves as the weight value for aggregation.

[0092] (5.2) After the central server aggregates and averages, the resulting global network feedback is downloaded to the corresponding V2U link client. The client's training network and target network are updated to the received global model. They each perform a certain number of training rounds using local experience. If the number of training rounds is less than the preset value, then proceed to step (5.1). Training ends after the aggregation interval is reached.

[0093] Furthermore, step (6) includes the following steps:

[0094] (6.1) The deep reinforcement learning model trained using the Fed-D3QN algorithm is input with the state information s at a certain time. t (k);

[0095] (6.2) Output the optimal action strategy To obtain the optimal V2I user transmit power and allocation channel C k .

[0096] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: This invention employs a deep reinforcement learning algorithm to jointly optimize channel selection and power control, introduces federated learning to ensure user privacy and data security, can meet the service quality requirements of different links, reduces the total latency of V2I link channel transmission, and simultaneously improves the effective payload transmission rate of V2U links. This invention uses the Fed-D3QN algorithm to rationally and efficiently utilize limited spectrum resources to achieve resource sharing, exhibiting excellent stability in highly dynamic air-to-ground networks. While ensuring reasonable resource allocation and meeting the service quality requirements of both V2I and V2U links, the multi-agent air-to-ground network resource allocation method based on federated deep reinforcement learning proposed in this invention is feasible and superior for resource allocation problems in highly dynamic environments. Attached Figure Description

[0097] Figure 1 This is a schematic diagram of the multi-agent air-to-ground network resource allocation method in this invention;

[0098] Figure 2 The figure shows the simulation results of the average accumulated return and the number of iterations in this invention;

[0099] Figure 3 The figure shows the simulation results of the relationship between the total delay and load of the V2I link in this invention.

[0100] Figure 4 This is a simulation result diagram showing the relationship between the transmission success rate and load of the V2U link in this invention. Detailed Implementation

[0101] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0102] The core idea of ​​this invention is to propose a constrained optimization problem based on link requirements, define the state space, action space, and reward function for reinforcement learning, train the neural network parameters using D3QN, and introduce federated learning to upload the locally learned reinforcement learning data to the base station for aggregated average training. Based on the Fed-D3QN model, the optimal resource allocation strategy for V2I and V2U links in the air-to-ground network is obtained.

[0103] like Figure 1 As shown, the present invention discloses a multi-agent air-to-ground network resource allocation method based on federated learning, characterized by comprising the following steps:

[0104] (1) In the air-ground network, the ground network consists of infrastructure and vehicle user equipment, forming a V2I link with high data rate service; the air network consists of unmanned aerial vehicles, forming a V2U link that communicates directly with ground vehicles; the V2U link has a wide communication coverage and is used to collect important information related to driving safety.

[0105] (2) The V2U link shares the spectrum resources of the V2I link and uses hybrid spectrum access technology for transmission;

[0106] (3) Construct a network resource allocation system model consisting of M-to-V2I links and K-to-V2U links; specifically including the following steps:

[0107] (3.1) The network resource allocation system model includes M pairs of V2I links and K pairs of V2U links, which are represented by sets M = {1, 2, ..., M} and K = {1, 2, ..., K}, respectively; V2I and V2U communication adopts orthogonal frequency division multiplexing technology, which divides the channel into M flat fading orthogonal sub-channels with bandwidth W. To simplify the analysis, the m-th VUE user occupies the m-th sub-channel in advance for communication;

[0108] (3.2) The channel power gain of the k-th V2U link in the m-th V2I link subband is defined as:

[0109] g k [m]=η k h k [m]

[0110] Where, η k For large-scale fading, including path loss and shadowing fading, it is assumed to be frequency-independent; h k [m] represents small-scale fading, which follows Rayleigh fading patterns within the subband and at uncorrelated times;

[0111] The SINR of the m-th V2I link can be expressed as:

[0112]

[0113] The channel capacity of the m-th V2I link, calculated using Shannon's formula, can be expressed as:

[0114]

[0115] in, and Let σ represent the transmit power of the m-th VEU and the k-th UAV, respectively. 2 G represents noise power. m [m] represents the power gain of the m-th V2I channel. ρ represents the interference power gain from the k-th V2U link to the m-th V2I link; k [m] represents the binary subband allocation indicator for the spectrum multiplexing flag, ρ k [m] = 1 indicates that the k-th UAV reuses the spectrum of the m-th VEU; otherwise, ρ k [m] = 0;

[0116] (3.3) For the k-th V2U link, its subband selection information is:

[0117] ρ k ={ρ k [1],ρ k [2],…,ρ k [m],…,ρ k [M]}

[0118] The rule stipulates that each link can only select one resource block for transmission at any given time.

[0119] (3.4) The SINR of the k-th V2U link in the m-th subband can be expressed as:

[0120]

[0121] The channel capacity of the k-th V2U link in the m-th subband can be expressed as:

[0122]

[0123] in,

[0124]

[0125]

[0126] These represent interference from V2I links using the same spectrum and interference from other V2U pairs, respectively. It is the interference gain of the k′-th V2U link on the k-th V2U link;

[0127] (3.5) The V2U link is primarily responsible for reliably transmitting safety-critical information. This information is generated periodically based on vehicle mobility, placing high demands on communication reliability and low latency. Considering the decentralized resource allocation at the UAV end, only transmission latency is considered as the latency of the V2U link. The latency constraint of the V2U link can be written as...

[0128]

[0129] Among them, B k For the remaining payload that the UAV needs to transmit, T k ≤T max To reduce the maximum tolerable delay T max The remaining delay to begin calculation;

[0130] (3.6) The reliability constraints of V2U communication can be achieved by controlling the probability of interruption events, which can be described as received events. Below the predetermined threshold Reliability requirements are expressed as follows:

[0131]

[0132] Where Pr{·} represents the probability of the input. The minimum SINR required to establish a reliable link for UAV, where p0 is the minimum interruptibility probability of the V2U link; under Rayleigh fading conditions, the reliability constraint is further transformed into:

[0133]

[0134] Where, γ th It is the SINR threshold of the UAV receiver on the k-th V2U link;

[0135] (3.7) In air-to-ground communication networks, the design objective is to minimize the total latency of V2I link transmission, defined as...

[0136]

[0137] (3.8) Taking into account the service quality requirements of different links, the objective function and optimization conditions are established as follows:

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144] The objective function is to minimize the total latency of the V2I link. Constraints C1 and C2 are the reliability and latency constraints of the V2U link. Constraint C3 states that the total power transmitted by the UAV on all subbands cannot exceed the maximum rated transmission power. Constraints C4 and C5 mean that each V2U link can only be assigned to one subband, but multiple V2U links can access the same subband.

[0145] (4) Using a multi-agent resource allocation method, a deep reinforcement learning model is constructed with the goal of minimizing the total transmission delay of the V2I link channel, taking into account the reliability and latency of the V2U link.

[0146] Furthermore, step (4) includes the following specific steps:

[0147] (4.1) Define the state space as including all local observation information related to resource allocation, expressed as:

[0148] s t (k)={{G k [m]} m∈M ,{I k [m]} m∈M B k ,T k ,e,ε}

[0149] in,

[0150] G k [m]={g m [m],g m,k [m],g k [m],{g k',k [m]} k'≠k}

[0151] It is the set of local instantaneous channel information on the uplink of subchannel m, g m [m] represents the channel power gain of the m-th V2I link. G represents the interference power gain from the k-th V2U link to the m-th V2I link. k [m] is the channel power gain of the k-th V2U link. It is the interference power gain of the k'-th V2U link on the k-th V2U link;

[0152] in,

[0153]

[0154] They are V2I links on the same spectrum. Interference with other V2U pairs The sum;

[0155] Among them, B k and T k These represent the remaining load and remaining latency that the V2U user needs to transmit, respectively.

[0156] (4.2) Define the action space of each agent as the selection of the spectrum sub-bands and the control of the transmission power of the V2U link, expressed as follows:

[0157]

[0158] in, Let C be the transmit power of the k-th V2U link user. k ∈{1,2,...,M} indicates that the k-th V2U link user has accessed the m-th subchannel;

[0159] (4.3) Define a reward function to reflect the optimization objective of the constraint problem. The optimization objective of minimizing the instantaneous total latency of all V2I links and the constraint of V2U link service quality are reflected in the reward at each time t, expressed as follows:

[0160]

[0161] in,

[0162]

[0163] Among them, D m [m,t] represents the latency of each V2I link; the smaller the total latency, the greater the reward. k (t) represents the effective transmission rate of the V2U link, reflecting the success rate of V2U link transmission; λ v and λ u It is a positive weight used to balance the contribution of V2I and V2U goals to the reward function, and needs to be adjusted based on experience;

[0164] (4.4) Introducing deep learning, a deep neural network is used to replace the Q-table for fitting the state-action value Q. DQN is specifically represented as:

[0165]

[0166] Where θ is the neural network parameter and γ is the discount factor;

[0167] (4.5) Introducing the dual-depth Q-network method, two neural networks with the same structure but different parameters are constructed as the training network Q. E (s,a;θ) and the target network Q T (s,a;θ), decoupling action selection from Q-value estimation, is expressed as:

[0168]

[0169] The training network selects actions by continuously updating parameters θ, and the target network Q... T Used to estimate the Q value, parameter θ - The parameters are kept fixed and replaced with the latest estimated network parameters θ at regular intervals.

[0170] (4.6) Introduce a competitive network, where the neural network consists of a state-value network V(s; θ). V ) and action advantage function network A(s,a;θ) A The algorithm is composed of two parts: the value function V(s) of the training network and the value function A(s,a) of the target network, which together balance the influence of actions on the Q-value in the training network and the target network, remove reward bias, and finally output the Q-value by adding the value function V(s) of the state and the advantage function A(s,a) of each action, and subtracting the average of the advantage functions of all actions in a certain state to ensure that for a given Q-value, the range of Q-values ​​is narrowed, and there are uniquely determined V(s) and A(s,a), removing redundant degrees of freedom and improving the stability of the algorithm. This can be expressed as:

[0171]

[0172] (4.7) The action advantage function represents the relative merit of a particular action compared to other actions in a given state. It is a measure of the relative merit of different actions in the current state. Unlike DQN, which directly learns all Q values, D3QN can distinguish whether the current reward is caused by the state itself or by the chosen action. Combined with the DDQN estimation method, the D3QN estimate is expressed as:

[0173]

[0174] (4.8) After each agent estimates the Q-value, it uses gradient descent to minimize the loss function and update the neural network parameters. The target network parameters are copied from the training network parameters at fixed intervals to complete the update of the target network, denoted as:

[0175]

[0176] (5) To improve the performance of multi-agent deep reinforcement learning models while protecting user privacy and data security, federated learning is used to optimize deep reinforcement learning models, including the following steps:

[0177] (5.1) The V2U link client uploads its local model to the server to execute the aggregation algorithm and obtain global network parameters. The aggregation algorithm performs a weighted average of all client models participating in federated learning according to their contribution to utilize global experience for training and maximize the aggregation effect. The specific formula is as follows:

[0178]

[0179] Where, θ t These are the parameters of the server's neural network at time t. N represents the neural network parameters of the k-th client at time t; k N and N are the training batch sizes of the k-th client and all clients, respectively. The ratio of these batch sizes is used to measure the contribution of the k-th client and serves as the weight value for aggregation.

[0180] (5.2) After the central server aggregates and averages, the resulting global network feedback is downloaded to the corresponding V2U link client. The client's training network and target network are updated to the received global model. They each perform a certain number of training rounds using local experience. If the number of training rounds is less than the preset value, then proceed to step (5.1). The training ends after the aggregation interval is reached.

[0181] (6) During the execution phase, the V2U link obtains the current state s based on observations. t (k) The optimal resource allocation strategy is obtained using the trained model, including the following steps:

[0182] (6.1) The deep reinforcement learning model trained using the Fed-D3QN algorithm is input with the state information s at a certain time. t (k);

[0183] (6.2) Output the optimal action strategy To obtain the optimal V2I user transmit power and allocation channel C k .

[0184] like Figure 1 As shown, the structure of a multi-agent air-to-ground network resource allocation method based on Fed-D3QN is described. It uses deep reinforcement learning to seek the optimal solution, generates training data with priority experience and state, and introduces federated learning to enhance user privacy protection and data security.

[0185] like Figure 2 As shown, the simulation results of the average accumulated reward and the number of iterations under the Fed-D3QN algorithm are described. It can be seen that as the number of iterations increases, the average accumulated reward increases and eventually tends to stabilize, indicating effective convergence.

[0186] like Figure 3 As shown, the simulation results describe the relationship between the total delay of the V2I link and the load under the Fed-D3QN algorithm. Under different V2U link load conditions, the total delay of the V2I link under the Fed-D3QN algorithm can be reduced by about 6% compared with the D3QN algorithm, and by 12% compared with the random algorithm.

[0187] like Figure 4As shown, the simulation results describe the relationship between the transmission success rate and load of V2U links under the Fed-D3QN algorithm. Under different V2U link load conditions, the transmission success rate of V2U links under the Fed-D3QN algorithm is higher than that of the D3QN algorithm and the random algorithm, and the performance is better and more stable.

[0188] Based on the description of the present invention, those skilled in the art should readily recognize that the present invention can improve network performance and protect user privacy.

Claims

1. A method for multi-agent air-ground network resource allocation based on federated learning, characterized in that, Includes the following steps: (1) In the air-ground network, the ground network consists of infrastructure and vehicle user equipment, forming a V2I link for high data rate service; the air network consists of unmanned aerial vehicles, forming a V2U link for direct communication with ground vehicles. V2U links are used to collect important information related to driving safety; (2) The V2U link shares the spectrum resources of the V2I link and uses hybrid spectrum access technology for transmission; (3) Construct a network resource allocation system model consisting of M pairs of V2I links and K pairs of V2U links; Step (3) includes the following specific steps: (3.1) The network resource allocation system model includes M pairs of V2I links and K pairs of V2U links, which are represented by sets M={1,2,...,M} and K={1,2,...,K}, respectively; V2I and V2U communication adopts orthogonal frequency division multiplexing technology, which divides the channel into M flat fading orthogonal sub-channels with bandwidth W, and the m-th VUE user occupies the m-th sub-channel in advance for communication; (3.2) The channel power gain of the k-th V2U link in the m-th V2I link subband is defined as: , wherein is the large scale fading, including path loss and shadow fading, assumed to be frequency independent; is the small scale fading, varying according to Rayleigh fading within a subband and over uncorrelated time. The SINR of the m-th V2I link can be expressed as: , The channel capacity of the m-th V2I link, calculated using Shannon's formula, can be expressed as: , wherein, and Pm(k) and Pk(m) represent the transmit power of the mth VEU and the kth UAV, respectively, Pn represents the noise power, Gm(m) represents the power gain of the mth V2I channel, Gk(m) represents the interference power gain of the kth V2U link to the mth V2I link; Bm(k) represents the binary sub-band allocation indicator of the spectrum reuse flag, Bm(k) represents the spectrum reused by the kth UAV for the mth VEU, otherwise ; (3.3) For the k-th V2U link, its subband selection information is: , The rule stipulates that each link can only select one resource block for transmission at any given time. ; (3.4) The SINR of the k-th V2U link on the m-th subband can be expressed as: , The channel capacity of the k-th V2U link in the m-th subband can be expressed as: , in, , , These represent interference from V2I links using the same spectrum and interference from other V2U pairs, respectively. It is the first The interference gain of one V2U link on the k-th V2U link; (3.5) The V2U link is mainly responsible for reliably transmitting safety-critical information, which is generated periodically according to the mobility of the vehicle; considering the decentralized resource allocation at the UAV end, only the transmission delay is considered as the delay of the V2U link, and the constraint of the V2U link on the delay is determined. (3.6) The reliability constraint of V2U communication can be achieved by controlling the probability of interruption events, which can be described as the received SINR. Below the predetermined threshold Determine reliability requirements; (3.7) In air-to-ground communication networks, the design objective is to minimize the total latency of V2I link transmission, defined as: , (3.8) Taking into account the service quality requirements of different links, establish the objective function and optimization conditions; (4) Using a multi-agent resource allocation method, and considering the reliability and latency of the V2U link, a deep reinforcement learning model is constructed with the goal of minimizing the total transmission latency of the V2I link channel; (5) In order to improve the performance of multi-agent deep reinforcement learning models while protecting user privacy and data security, federated learning is used to optimize deep reinforcement learning models; (6) During the execution phase, the V2U link obtains the current state based on observation and uses the trained model to obtain the optimal resource allocation strategy.

2. The multi-agent air-to-ground network resource allocation method based on federated learning according to claim 1, characterized in that: The latency constraint of the V2U link in step (3.5) can be written as: , in, The remaining payload that the UAV needs to transmit. To minimize tolerable latency The remaining delay is calculated at the beginning.

3. The multi-agent air-to-ground network resource allocation method based on federated learning according to claim 2, characterized in that: The reliability requirement in step (3.6) is expressed as follows: , in, The probability of input. The minimum SINR required to establish a reliable link for a UAV The minimum interruptibility probability of the V2U link; under Rayleigh fading conditions, the reliability constraint is further transformed into: , in, It is the SINR threshold of the UAV receiver on the k-th V2U link.

4. The multi-agent air-to-ground network resource allocation method based on federated learning according to claim 3, characterized in that: The objective function and optimization conditions established in step (3.8) are as follows: , , , , , , The objective function is to minimize the total latency of the V2I link. Constraints C1 and C2 are the reliability and latency constraints of the V2U link. Constraint C3 states that the total power transmitted by the UAV on all subbands cannot exceed the maximum rated transmission power. Constraints C4 and C5 mean that each V2U link can only be assigned to one subband, but multiple V2U links can access the same subband.

5. The multi-agent air-to-ground network resource allocation method based on federated learning according to claim 4, characterized in that: Step (4) includes the following specific steps: (4.1) Define the state space to include all local observation information related to resource allocation. (4.2) Define the action space of each agent as the selection of the spectrum sub-bands and the control of the transmission power of the V2U link, expressed as follows: , in, Let k be the transmit power of the k-th V2U link user. This indicates that the k-th V2U link user has accessed the m-th subchannel; (4.3) Define a reward function to reflect the optimization objective of the constraint problem. The optimization objective of minimizing the instantaneous total latency of all V2I links and the constraint of V2U link service quality are reflected in the reward at each time t, expressed as follows: , in, , in, The reward is based on the latency of each V2I link; the lower the total latency, the greater the reward. The effective transmission rate of a V2U link reflects its transmission success rate. and It is a positive weight used to balance the contribution of V2I and V2U goals to the reward function, and needs to be adjusted based on experience; (4.4) Introducing deep learning, a deep neural network is used to replace the Q-table to fit the state-action value Q, resulting in DQN; (4.5) Introduce the dual-depth Q-network method algorithm to construct two neural networks with the same structure but different parameters, which are used as training networks respectively. With the target network Decoupling action selection from Q-value estimation is expressed as: , Training the network involves continuously updating its parameters. Select action, target network Used to estimate the Q value, parameters The parameters remain fixed and are periodically replaced with the latest estimated network parameters. ; (4.6) Introduce a competitive network, where the neural network is a state-value network. Action Advantage Function Network Together, they form a system that balances the impact of actions on the Q-value in both the training and target networks, removes reward bias, and finally outputs a Q-value that is a product of the state value function V(s) and the advantage function A(s, ...). Adding these together, and subtracting the average of all action advantage functions in a given state, ensures that, given a Q value, the range of Q values ​​is narrowed, resulting in uniquely determined V(s) and A(s). To remove redundant degrees of freedom and improve algorithm stability, this can be represented as: , (4.7) The action advantage function represents the relative merit of a certain action compared to other actions in the current state. It is a measure of the relative merit of different behaviors in the current state. Unlike DQN, which directly learns all Q values, D3QN can distinguish whether the current reward is caused by the state itself or by the chosen action. Combined with the DDQN estimation method, the D3QN estimate is expressed as: , (4.8) After each agent estimates the Q-value, it uses gradient descent to minimize the loss function and update the neural network parameters. The target network parameters are copied from the training network parameters at fixed intervals to complete the update of the target network, as shown below. 。 6. The multi-agent air-to-ground network resource allocation method based on federated learning according to claim 5, characterized in that: In step (4.1), the state space is defined as follows: , in, , It is the set of local instantaneous channel information on the uplink of subchannel m. This represents the channel power gain of the m-th V2I link. This represents the interference power gain from the k-th V2U link to the m-th V2I link. It is the channel power gain of the k-th V2U link. It is the first The interference power gain of one V2U link to the k-th V2U link; in, , It is interference from V2I links on the same spectrum. Interference from other V2U links The sum; in, and These represent the remaining load and remaining latency that the V2U user needs to transmit.

7. A multi-agent air-to-ground network resource allocation method based on federated learning according to claim 6, characterized in that: In step (4.4), DQN is specifically represented as follows: , in, These are neural network parameters. It is a discount factor.

8. A multi-agent air-to-ground network resource allocation method based on federated learning according to claim 7, characterized in that: Step (5) includes the following specific steps: (5.1) The V2U link client uploads its local model to the server to execute the aggregation algorithm and obtain global network parameters. The aggregation algorithm performs a weighted average of all client models participating in federated learning according to their contribution to utilize global experience for training and maximize the aggregation effect. The specific formula is as follows: , in, These are the parameters of the server's neural network at time t. Let be the neural network parameters of the k-th client at time t; and These are the training batch sizes of the k-th client and all clients, respectively. The ratio between these batch sizes is used to measure the contribution of the k-th client and serves as the weight value for aggregation. (5.2) After the central server aggregates and averages, the resulting global network feedback is downloaded to the corresponding V2U link client. The client's training network and target network are updated to the received global model. They each perform a certain number of training rounds using local experience. If the number of training rounds is less than the preset value, then proceed to step (5.1). Training ends after the aggregation interval is reached.

9. A multi-agent air-to-ground network resource allocation method based on federated learning according to claim 8, characterized in that: Step (6) includes the following steps: (6.1) The deep reinforcement learning model trained using the Fed-D3QN algorithm is input with the state information at a certain time. ; (6.2) Output the optimal action strategy To obtain the optimal V2I user transmit power and channel allocation .