A performance gap driven air-ground integrated load balancing method
Patent Information
- Application Number
- CN202611096152.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-08-18
AI Technical Summary
[0002]空地一体化网络作为6G的关键使能技术,正推动通信架构从传统地面集中式向立体协同框架演进,但也因卫星、无人机及地面基站等节点在空间位置、通信能力等方面的高度异构性,以及网络拓扑的时变性与业务分布的不均衡性,带来了严峻的负载均衡挑战
[0031] (1) Based on the consideration of differentiated service needs, this invention introduces the performance dissatisfaction rate as an evaluation index, and prioritizes unloading users with high dissatisfaction to the airspace layer. This can avoid the problem of users whose needs are met being forcibly unloaded, resulting in poor performance experience, while improving the performance of each service and ensuring relative fairness for users.
Smart Images

Figure CN122601064A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of air-to-ground communication and artificial intelligence. More specifically, it relates to a performance gap-driven air-to-ground integrated load balancing method, which aims to reduce ground load, optimize resource efficiency, and ensure fairness among users. Background Technology
[0002] As a key enabling technology for 6G, integrated air-ground networks are driving the evolution of communication architecture from traditional terrestrial centralized systems to a three-dimensional collaborative framework. However, the high heterogeneity of nodes such as satellites, drones, and ground base stations in terms of spatial location and communication capabilities, as well as the time-varying nature of network topology and the uneven distribution of services, bring severe load balancing challenges. Existing load balancing methods do not fully consider customized needs; different services have different performance requirements (such as coverage probability and latency), which may lead to uncontrollable and unpredictable user experience. A very small number of priority-based load balancing methods consider multiple performance indicators, but their offloading schemes do not consider whether the offloading users meet their performance requirements, and cannot ensure user fairness. In addition, traditional resource optimization faces problems such as optimization difficulties or high complexity, while multi-agent deep reinforcement learning algorithms suffer from policy interference problems due to memory sharing and possible policy interactions. Independent learning algorithms decouple the decision-making process of each agent, do not share memory or policy information, and fundamentally avoid interference problems caused by policy interactions in multi-agent deep reinforcement learning. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention provides a performance gap-driven integrated air-ground load balancing method, comprising the following technical solutions:
[0004] S1. Based on the characteristics of air-to-ground and ground-to-ground wireless channels, establish a channel model, analyze the performance indicators of high throughput, low latency and wide coverage services, including total throughput, average latency and average signal-to-interference-plus-noise ratio, and establish a performance indicator analysis model.
[0005] Downlink channel fading between user and base station Following the Rayleigh channel model, in time slot The channel coefficients of user k associated with subcarrier n and base station m The calculation is as follows:
[0006]
[0007] in, It is the distance between the base station and the user, and α is the path loss exponent;
[0008] Downlink channel fading between user and drone Characterized using the Ricean channel model, in time slots The channel coefficient between user k and associated subcarrier n and UAV v The calculation is as follows:
[0009]
[0010] in, h0 is the distance between the drone and the user; h0 is the average channel power gain when the reference distance is 1 meter; R is the fading factor. Represents the line-of-sight component, and ; Indicates the non-line-of-sight component; Indicates the modulus.
[0011] The performance index analysis model is as follows:
[0012]
[0013]
[0014]
[0015] In the formula, , , These are total throughput, average latency, and average signal-to-interference-plus-noise ratio, respectively. and These represent the data rate and latency of user k in time slot t, respectively. This is an indicator function that outputs 1 if the input is true, and 0 otherwise. Assign joint factors to node associations and subcarriers; K1 represents the signal-to-interference-plus-noise ratio (SNR) of user k associated with node i and subcarrier n; K2 and K3 represent the number of users requesting low-latency service and wide-coverage service, respectively; K1 represents the set of users requesting high-throughput service; K2 represents the set of users requesting low-latency service; K3 represents the set of users requesting wide-coverage service; I represents the set of base station and UAV nodes; and N represents the set of subcarriers.
[0016] S2. Based on the performance index analysis model, further quantify the user performance gap, define the performance dissatisfaction rate, and design a cross-layer offloading strategy based on the performance dissatisfaction rate.
[0017] The cross-layer offloading strategy is as follows: users are given priority access to the ground layer; when the ground layer is overloaded, all users are offloaded to the ground layer. k (t) Sort in ascending order and unload excess users to the airspace layer in descending order.
[0018] S3. Establish a scalar-based utility function based on the performance indicators, and propose a joint optimization problem for the performance indicators based on the utility function;
[0019] Based on the different performance metrics of the three types of services, a scalar-based utility function is established:
[0020]
[0021] in, These represent the weights of total throughput, average latency, and average signal-to-interference-plus-noise ratio, respectively. This represents the initial maximum benefit of latency;
[0022] Based on the above utility function, the joint optimization problem of performance indicators is proposed as follows:
[0023]
[0024] Where C1 represents a binary variable. The range of values for C2; C2 indicates the time slot. Each user is allocated at most one subcarrier; C3 indicates that in the time slot Each user can be associated with at most one network component; furthermore, C4 restricts the non-negativity of transmit power, while C5 indicates that the number of users associated with the same base station cannot exceed the maximum capacity. ; This represents the joint correlation factor between the subcarrier and the base station; This represents the joint correlation factor between subcarriers and UAVs; This indicates the power allocation from the base station to the user; This represents the power allocation from the drone to the user; M is the set of base stations; K is the set of all users in the system; and V is the set of drone nodes.
[0025] S4. Design of a condition-based independent multi-agent optimization algorithm; Based on the cross-layer offloading strategy, a condition-triggered learning algorithm is proposed: The DDPG agent in the ground layer is used to optimize the association between users and base stations and resource allocation. If there is base station overload in the ground layer, the DDPG agent in the airspace layer is used to optimize the offloading of the association between users and UAVs and resource allocation. Otherwise, the agent in the airspace layer will skip the current iteration.
[0026] Ground-level agents based on the current state The optimization actions for the current time slot are given—user-base station association and resource allocation. Based on the current action and state, obtain the current time slot reward and calculate the G for all users. k (t); If the base station is overloaded, a load balancing scheme will be triggered, and step S2 will be executed. At this time, the spatial layer agent will be based on the current state. The optimization actions for the current time slot are given—unloading user-drone association and resource allocation. The agent calculates the current reward; otherwise, the spatial layer agent skips the current iteration; calculates the reward for the current time slot and moves to the next state; until the algorithm converges or the iteration ends.
[0027] S5. Air-to-Ground Imperfect Channel Modeling—Robustness Verification: An air-to-ground imperfect channel model is established based on a Gaussian channel error model to verify the robustness of the S4 learning algorithm under different error conditions.
[0028]
[0029] In the formula, and These represent the actual channel state information of the base station and the drone, respectively. and These are the estimated channel state information known to the base station and the drone, respectively. and The channel errors of the base station-user and UAV-user channels are represented respectively; the common Gaussian channel error is used for modeling, and it follows an independent and identically distributed complex Gaussian distribution. , This represents the variance of the distribution.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] (1) Based on the consideration of differentiated service needs, this invention introduces the performance dissatisfaction rate as an evaluation index, and prioritizes unloading users with high dissatisfaction to the airspace layer. This can avoid the problem of users whose needs are met being forcibly unloaded, resulting in poor performance experience, while improving the performance of each service and ensuring relative fairness for users.
[0032] (2) The optimization algorithm of this invention adopts an independent multi-agent architecture, introduces triggering conditions, reduces invalid decision steps, thereby reducing computational complexity and accelerating the convergence process.
[0033] (3) The present invention can still exhibit good robustness under imperfect channel conditions, providing verification of the practicality of load balancing schemes and optimization methods. Attached Figure Description
[0034] Figure 1 This is a technical roadmap of an embodiment of the present invention;
[0035] Figure 2 This is a schematic diagram showing the locations of the base station, user, and drone in an embodiment of the present invention;
[0036] Figure 3 This is a schematic diagram of the C-IMADDPG algorithm architecture proposed in an embodiment of the present invention;
[0037] Figure 4 This is a schematic diagram of the C-IMADDPG algorithm proposed in an embodiment of the present invention;
[0038] Figure 5 This is a schematic diagram comparing the total throughput results of the load balancing scheme based on performance dissatisfaction rate proposed in the embodiments of the present invention;
[0039] Figure 6 This is a schematic diagram comparing the average latency results of the load balancing scheme based on performance dissatisfaction rate proposed in the embodiments of the present invention;
[0040] Figure 7 This is a schematic diagram comparing the average signal-to-interference-plus-noise ratio results of the load balancing scheme based on performance dissatisfaction rate proposed in the embodiments of the present invention.
[0041] Figure 8 This is a schematic diagram comparing the total throughput results of the C-IMADDPG algorithm proposed in this embodiment of the invention;
[0042] Figure 9 This is a schematic diagram comparing the average latency results of the C-IMADDPG algorithm proposed in this embodiment of the invention;
[0043] Figure 10 This is a schematic diagram comparing the average signal-to-interference-to-noise ratio results of the C-IMADDPG algorithm proposed in this embodiment of the invention;
[0044] Figure 11 This is a schematic diagram comparing the average latency results of fixed / optimized UAV positions according to an embodiment of the present invention;
[0045] Figure 12 This is a schematic diagram comparing the average signal-to-interference-to-noise ratio results for fixed / optimized UAV positions according to an embodiment of the present invention.
[0046] Figure 13 This is a schematic diagram comparing the average signal-to-interference-to-noise ratio results for fixed / optimized UAV positions according to an embodiment of the present invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. A performance gap-driven air-ground integrated load balancing method, such as... Figure 1 As shown:
[0048] S1. Air-to-Ground Network Channel and Performance Indicator Modeling: Under a two-layer air-to-ground network architecture, air-to-ground and ground-to-ground wireless communication channel models are established. Performance indicators for three service types—high throughput, low latency, and wide coverage—are analyzed. This provides a theoretical basis for establishing performance gap quantification indicators and proposing a cross-layer offloading method based on performance dissatisfaction rate. The specific implementation is as follows:
[0049] Based on the research and analysis of service requirements in existing scenarios, the system simultaneously provides randomized high-throughput, low-latency, and wide-coverage services. Considering the downlink air-to-ground integrated network, this network consists of… The ground layer consisting of base stations and the... This is composed of an airspace layer made up of drones. One base station and The set of drone nodes is denoted as follows: and At the same time, remember .
[0050] like Figure 2 As shown in the example, in a three-dimensional space with a ground area of 3000m×3000m, the positions of the base station and the user are randomly generated, and the positions of the drone are fixed at (1000, 2000, 100) and (2000, 1000, 100), where the drone height is fixed at 100m. When optimizing the drone position, only the horizontal movement of the drone is considered.
[0051] Assume that the base station and drone nodes operate in non-overlapping frequency bands to avoid cross-layer interference. The available bandwidth for each layer is... Hz, this bandwidth is divided into There are orthogonal subcarriers, the set of which is: The total power of each base station and drone is set to [specific value]. and Furthermore, assuming Ground users (collected as) Random requests for high-throughput, low-latency, and wide-coverage services, which are respectively provided by... It means, and In addition, time slots Request service The number of users is denoted as This set can be defined as .
[0052] For downlink communication scenarios between users and base stations, the channel fading of the downlink is a concern. Following the Rayleigh channel model, this assumes no line-of-sight path in the communication link and that the amplitude statistics of the multipath signal conform to a Rayleigh distribution, thus constructing a typical flat fading channel environment. Therefore, in the time slot... user Associated subcarriers With base station Channel coefficients It can be calculated as:
[0053]
[0054] in, It is the distance between the base station and the user. It is the path loss index.
[0055] In the air-to-ground wireless communication system considered in this invention, the downlink communication between the user equipment and the UAV uses the Ricean channel model to characterize its channel fading characteristics. Specifically, since UAVs typically fly at line-of-sight, the downlink contains a strong line-of-sight component, superimposed with multiple non-line-of-sight scattering paths caused by the ground surface and surrounding buildings, vegetation, etc. The statistical distribution of the received signal envelope follows a Ricean distribution, and its probability density function is the ratio of the line-of-sight energy to the scattering path energy—that is, the Ricean fading factor. This is determined by the Ricean channel model. Using this model, the amplitude and phase fluctuation characteristics of the UAV downlink signal in complex air-to-ground environments can be accurately simulated. Therefore, in the time slot... user Associated subcarriers With drones Channel coefficients between It can be calculated as:
[0056]
[0057] in, It is the distance between the drone and the user; It is the average channel power gain at a reference distance of 1 meter; For the decline factor; Represents the line-of-sight component, and ; Indicates the non-line-of-sight component;
[0058] definition and Base stations and drones In the time slot Assigned to use subcarrier users The transmit power. Furthermore, node association and subcarrier allocation are jointly defined as an allocation factor, denoted as... If the user In the time slot Assigned to a node subcarrier ,but Otherwise, it is 0; this can be obtained from the user. In the time slot Related nodes and subcarrier The signal-to-interference-to-noise ratio is:
[0059]
[0060] in, Indicates power allocation; Indicates the channel coefficient; This indicates interference with the user; Represents the set of system users; Indicates the power of the interfering user; Indicates the total bandwidth; Indicates the number of subcarriers; Indicates noise power.
[0061] Based on the subcarrier correlation factor, the user In the time slot Signal interference plus noise ratio The calculation is as follows:
[0062]
[0063] Performance metrics for three service categories—high throughput, low latency, and wide coverage—are modeled separately:
[0064] S11. High-throughput service communication model construction: For high-throughput services that need to support large-volume data transmission, the core performance indicator is throughput. Specifically, throughput represents the total data transmission rate of all users requesting the service under given bandwidth, channel conditions, and resource allocation schemes. According to Shannon's formula, the throughput of a user... The throughput is calculated as follows:
[0065]
[0066] Therefore, based on the definition of total throughput, the total throughput of high-throughput services can be obtained as follows:
[0067]
[0068] in, Indicates time slot A set of users requesting high-throughput services.
[0069] S12. Low-Latency Service Communication Model Construction: In wireless communication scenarios, some services (such as online games, voice calls, and industrial automation control) fall into the category of low-latency services. The user experience of these services mainly depends on the magnitude of service latency, including propagation latency, transmission latency, and queuing latency. Therefore, the first step is to establish a queuing model for low-latency users, allowing... The random data arrival process for low-latency services, where Obeying the rate is The Poisson arrival process is assumed, and it is assumed that users are independent of each other; furthermore, the service process can be modeled as an M / D / 1 queue; therefore, users In the time slot Related nodes Service delay It can be represented as:
[0070]
[0071] in, It's the speed of light. User Distance to associated nodes Indicates user Data arrival rate.
[0072] The latency of low-latency services is defined as the average latency of all users requesting low-latency services, calculated as follows:
[0073]
[0074] in, This is an indicator function that outputs 1 if the input is true, and 0 otherwise. Indicates time slot The number of users requesting low-latency service.
[0075] S13. Wide Coverage Service Communication Model Construction: For wide coverage services, this application adopts a coverage determination scheme based on the average signal-to-interference-plus-noise ratio (SNR). Therefore, the average SNR is the core indicator for measuring coverage quality. To further evaluate the overall coverage quality of the wide coverage service, the average SNR index is introduced to characterize the statistical mean of the SNR of all users served by the wide coverage service. Its calculation formula is as follows:
[0076]
[0077] in, Indicates time slot The number of users requesting wide coverage service.
[0078] S2. Implementation of Cross-Layer Load Balancing Method Based on Performance Dissatisfaction Rate: Based on the performance indicator model constructed in S1, the performance gap between users is further quantified and defined as the performance dissatisfaction rate. A cross-layer offloading strategy based on the performance dissatisfaction rate is then designed. This strategy fully considers the service performance experience of offloaded users, prioritizing the offloading of users whose performance needs are not met to the spatial layer, while maintaining the association of users whose performance needs are met. This ensures service continuity while improving the overall performance of the system.
[0079] First, calculate the performance metrics associated with the ground layer for different service users, as follows:
[0080]
[0081]
[0082]
[0083] in, User The downlink rate associated with the ground layer; Indicates user Service latency associated with the ground layer; Indicates user The signal interference-to-noise ratio associated with the ground layer; This represents the joint correlation factor between the subcarrier and the base station; For users In the time slot Associated base stations The signal-to-interference-to-noise ratio;
[0084] definition , and These are the user's expected data rate, latency, and average signal-to-interference-plus-noise ratio (SNR) thresholds, respectively. For users whose actual performance differs significantly from their expected requirements, their load is preferentially offloaded to the spatial domain layer. Therefore, the user performance gap is defined as the difference between the actual performance and the expected threshold, and a performance dissatisfaction rate is introduced. As a quantitative indicator, it is used to evaluate users. The gap between the actual performance and the target performance is defined as:
[0085]
[0086] For each user Normalization is performed to ensure fairness for all users. The higher the value, the greater the gap between the user's actual performance and the expected goals. Therefore, it is necessary to prioritize offloading such users to the drone-assisted airspace layer to alleviate resource competition at the ground layer. Conversely, A lower value indicates a better consistency between the user's actual performance and the target performance, and that the user's needs are more likely to be met at the ground layer or that the requirements have already been met at the ground layer. Therefore, ground layer access should be maintained to ensure service continuity.
[0087] S3. Construction of a utility optimization model under capacity constraints: To jointly optimize the various performance indicators corresponding to different services of the system, a function that can uniformly represent these indicators is constructed based on the scalarization method and defined as system utility. Then, the system utility optimization is established based on resource constraints.
[0088] Since the key performance indicators of various services differ significantly, a scalarization method is used to uniformly represent them and define them as system utility:
[0089]
[0090] in, These represent the weights of total throughput, average latency, and average signal-to-interference-plus-noise ratio, respectively. This represents the initial maximum benefit of latency;
[0091] This leads to the problem of system optimization. , represents the following:
[0092]
[0093] in, Represents binary variables The range of values for; This indicates that in the time slot Each user can be allocated at most one subcarrier; Indicates in time slot Each user can be associated with at most one network component; furthermore, This restricts the nonnegativity of the transmit power, while This indicates that the number of users associated with the same base station cannot exceed the maximum capacity. ; This represents the joint correlation factor between the subcarrier and the base station; This represents the joint correlation factor between subcarriers and UAVs; This indicates the power allocation from the base station to the user; This indicates the power allocation provided by the drone to the user; This is a collection of drone nodes.
[0094] S4. To solve the optimization problem described in S3, and considering the offloading coordination between the ground layer and the airspace layer, a conditionally triggered independent deep reinforcement learning algorithm—the C-IMADDPG algorithm—is proposed. This algorithm consists of an environment, two independent DDPG agents, and a triggering module, as follows: Figure 3The following example illustrates the algorithm proposed in this invention. To clearly explain the algorithm, it is described in three parts: quadruple construction, network update strategy, and algorithm execution flow. The quadruple, as the basic data unit, provides data support for the agent's state, action selection, and state transition learning. The network parameter update part is responsible for optimizing the agent's network parameters based on empirical samples and designing a corresponding update mechanism for the characteristics of multi-agent collaborative decision-making.
[0095] S41, Quadruple Construction: Considering the dynamic characteristics of the proposed transmission scenario, optimization problem This can be described as a Markov Decision Process (MDP). MDPs can be represented using quadruples. It means that, among them and These represent the observation status of the current time slot and the next time slot, respectively. It's an action set. This represents the agent's reward, which can be customized for different systems. This quadruple is stored in the experience replay region, from which a mini-batch of samples is randomly selected for training the neural network.
[0096] Current status and the state at the next moment Intelligent agent The state observation is defined as the channel state, therefore the current state and the next state can be expressed as:
[0097]
[0098] in, Indicates time slot Real-time base station-user channel state information;
[0099] Current status and the state at the next moment Similarly, intelligent agents The current state and the next state can be described as follows:
[0100]
[0101] in, Indicates time slot Channel status information between UAVs and users at that time;
[0102] Intelligent agent It is responsible for ground layer optimization, including subcarrier allocation and power control for all users. Therefore, this set of actions can be represented as:
[0103]
[0104] Intelligent agent It is responsible for spatial layer optimization, including subcarrier offloading and power control for users. Therefore, this action set can be defined as:
[0105]
[0106] To match the continuous action space in the DDPG agent, The binary variables are first converted into continuous variables with a value range of [0,1]. Finally, these variables are converted back into binary variables through a rounding function.
[0107] award and In this system, two agents jointly optimize the system's utility; therefore, the rewards for the two agents are defined as follows:
[0108]
[0109] S42. Network Update Strategy: All agents in the proposed C-IMADDPG algorithm employ the same actor-critic architecture. Each agent consists of two online neural networks and two target neural networks: an online actor network, an online critic network, a target actor network, and a target critic network. The online critic network updates the network through a state-action value function. To evaluate the action, among which These are the weight parameters of the online critic network, and its loss function can be expressed as:
[0110]
[0111] in, The objective value of the value function, An immediate reward for the current observation state and actions; This indicates the current observation state, corresponding to the above. and ; This indicates the action at the current moment, corresponding to the above. and γ represents the discount factor; This indicates the reward for the current state and action, corresponding to the above. and ; Indicates the observed state at the next moment; Indicates the strategic action to be taken in the next moment.
[0112] Online actor network by parameters The parameter is described as being updated using gradient descent, i.e.:
[0113]
[0114] in, express The gradient; Expressing expectations; For a deterministic policy function, i.e., given a state It directly outputs a specific action value; Indicates the action Find the gradient.
[0115] The Targeted Criticism Network and the Targeted Actor Network are respectively composed of and Parameterization. and Using a soft update method and employing a small constant. The update is as follows:
[0116]
[0117] in, It represents the value used to calculate the objective value of the value function. The target of the criticism network parameters, These are the parameters of the target actor network, used to calculate the target action. It is the soft update coefficient, which is usually a very small positive number used to control how quickly the target network moves closer to the current network.
[0118] S43. Algorithm Execution Flow: Based on the above discussion, the proposed C-IMADDPG algorithm can solve the problem of maximizing the overall system utility. The algorithm execution flow is as follows: Figure 4 As shown. First, the algorithm initializes the network parameters, the experience replay pool, and the state space. Ground agent 1 selects an action based on the current policy and state. Agent 1 interacts with the environment and performs actions. It calculates the ground layer load and, if the ground layer base station is overloaded, triggers an offloading scheme based on performance dissatisfaction rate: calculates the performance dissatisfaction rate values of all users and sorts them in ascending order, then sets the sorted values as follows: Each user is unloaded to the airspace layer, and the corresponding associated metric is set to 0. For base stations The current load; the spatial layer agent 2 selects actions based on the current policy and state. Agent 2 interacts with the environment and performs actions. The system obtains rewards for the current state and action; stores four-tuples in Experience Replay 1 and Experience Replay 2 respectively; when the experience replay pool is full or the training conditions are met, random mini-batch samples are extracted from each experience replay for neural network training; the network parameters are updated according to S42, and the system transitions to the next state until the loop ends. If the ground layer base station is not overloaded, the agent in the spatial layer will skip the current iteration.
[0119] S5. Imperfect Air-to-Ground Channel Modeling—Robustness Verification: Given the difficulty in obtaining perfect channel state information under real-world conditions, and the challenge in further evaluating the robustness of the proposed algorithm, imperfect channel state information is considered to evaluate the algorithm's robustness. An imperfect channel model is established by introducing channel error, as follows:
[0120]
[0121] In the formula, and These represent the actual channel state information of the base station and the drone, respectively. and These are the estimated channel state information known to the base station and the drone, respectively. and Let represent the channel errors of the base station-user and UAV-user channels, respectively; Gaussian channel errors are used for modeling, and they follow independent and identically distributed characteristics. .
[0122] All numerical results were implemented using the Tensorflow framework in Python 6. In the proposed algorithm model, each network has three hidden layers containing 400, 300, and 10 neurons respectively, and the learning rates for the actor and critic networks were set to 0.0001 and 0.0002, respectively. The experiential memory for each agent was set to 5000 quadruplets, and 32 batches of samples were used for training. The number of base stations and drones were set to... and The available bandwidth for each network component is... The bandwidth is divided into Subcarriers.
[0123] To verify the effectiveness and superiority of the proposed load balancing method based on performance dissatisfaction rate and the proposed C_IMADDPG algorithm in the air-ground cooperation scenario, simulation experiments and comparative experiments were set up. The benchmark scheme for comparison is as follows: (1) All actions in the comparative algorithm are executed by a single DDPG agent, and the reward function and the agent parameters are completely consistent with the proposed algorithm; (2) All actions in the comparative algorithm are executed by a single TD3 agent, and the reward function and the agent parameters are completely consistent with the proposed algorithm; (3) Priority-based load balancing is optimized using the same algorithm. In the priority-based load balancing, the high-throughput service, low-latency service, and wide-coverage service are offloaded to the airspace layer with low, medium, and high priorities, respectively. That is, users of wide-coverage services are preferentially offloaded to the airspace layer; (4) Optimal UAV position is achieved using the same load balancing method and the same algorithm, but UAV position optimization is added to the actions of the airspace layer agent, that is, the action dimension is expanded to:
[0124]
[0125] in, Indicates drone In the time slot The horizontal coordinates are fixed, and the vertical height of all drones is fixed at 100m.
[0126] Example 1, Figure 5 , Figure 6 and Figure 7 The proposed load balancing method and the priority-based cross-layer load balancing method were compared in terms of total throughput, average latency, and average signal-to-interference-plus-noise ratio (SNR). The priority-based load balancing method achieved higher total throughput than the proposed performance-driven load balancing method when the number of users was small. However, as the number of users increased, the total throughput of the priority-based load balancing method gradually fell below that of the proposed method. The priority-based load balancing method achieved the lowest average SNR. In this scheme, the resource allocation priority for wide-coverage services was the lowest. When ground resources were insufficient, these users were the last to be offloaded to the spatial layer to obtain services, thus achieving the lowest average SNR. At the same time, the average latency of the priority-based load balancing method was also higher than that of the proposed method. This is because the priority-based load balancing method sacrificed the performance of low-priority users (low-latency users have a middle priority) to improve the performance of high-priority users.
[0127] Example 2, Figure 8 , Figure 9 and Figure 10The proposed algorithm was compared with DDPG and TD3 algorithms in terms of total throughput, average latency, and average signal-to-interference-plus-noise ratio. Compared with DDPG and TD3, the proposed C-IMADDPG algorithm achieves better performance in all three metrics. The core of this performance improvement lies in setting a trigger mechanism for the agents in the spatial layer, i.e., when the ground layer is overloaded, effectively reducing redundant updates and meaningless interaction interference. This asymmetric design makes the timing of cooperation more precise, avoiding reward fluctuations caused by the continuous exploration of the second agent, and alleviating the policy drift problem that TD3's dual-latency updates still cannot completely suppress.
[0128] Example 3, Figure 11 , Figure 12 and Figure 13 The performance metrics of the proposed algorithm and the DDPG algorithm were compared in scenarios with fixed UAV positions and optimal UAV positions. In the optimal UAV position scenario, the two-dimensional coordinates of the UAV are incorporated into the joint optimization action space (with a fixed altitude), allowing the algorithm to dynamically adjust the UAV position based on real-time channel conditions and rewards. The comparison results show that, compared with fixed UAV positions, incorporating the UAV position into the action space can effectively exploit deployment freedom and further improve the performance of total throughput, average latency, and average signal-to-interference-plus-noise ratio. It is worth noting that the increased action dimension of the agent usually poses a challenge to the convergence of reinforcement learning algorithms, but the proposed algorithm still maintains a stable advantage over the DDPG algorithm in this scenario.
[0129] Example 4, Table 1 summarizes the performance of the proposed method and the comparison methods under different channel error conditions. As the channel estimation error increases, the performance of the proposed scheme decreases, indicating that imperfect CSI leads to performance degradation. Nevertheless, the proposed scheme maintains good performance even with imperfect CSI, demonstrating the robustness of the proposed load balancing strategy in the face of channel uncertainty. In particular, under imperfect CSI conditions, the performance achieved by the proposed method is comparable to, or even better than, several benchmark schemes with perfect CSI. For example, under imperfect CSI conditions, the total throughput and average latency of the proposed scheme are superior to the performance of the DDPG and TD3 methods under perfect CSI conditions. This further highlights the effectiveness of the proposed load balancing method and the C-IMADDPG algorithm.
[0130] Table 1. Performance comparison under different channel errors: K=27
[0131]
[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A performance gap-driven integrated air-ground load balancing method, characterized in that, Includes the following steps: S1. Based on the characteristics of air-to-ground and ground-to-ground wireless channels, establish a channel model, analyze the performance indicators of high throughput, low latency and wide coverage services, including total throughput, average latency and average signal-to-interference-plus-noise ratio, and establish a performance indicator analysis model. S2. Based on the performance index analysis model, further quantify the user performance gap, define the performance dissatisfaction rate, and design a cross-layer offloading strategy based on the performance dissatisfaction rate. S3. Based on the performance indicators described in S1, establish a scalar-based utility function, and propose a joint optimization problem for the performance indicators based on the utility function. S4. Based on the cross-layer offloading strategy described in S2, a condition-triggered learning algorithm is proposed: the DDPG agent in the ground layer is used to optimize the association between users and base stations and resource allocation. If there is base station overload in the ground layer, the DDPG agent in the airspace layer is used to optimize the offloading of the association between users and UAVs and resource allocation. Otherwise, the agent in the airspace layer will skip the current iteration. S5. Based on the Gaussian channel error model, establish an air-to-ground imperfect channel model to verify the robustness of the learning algorithm described in S4 under different error conditions.
2. The performance gap-driven integrated air-ground load balancing method according to claim 1, characterized in that, The relevant formulas for the channel model mentioned in step S1 are as follows: Channel fading in the downlink between the user and the base station Following the Rayleigh channel model, in time slot user Associated subcarriers With base station Channel coefficient The calculation is as follows: in, It is the distance between the base station and the user. It is the path loss index; Downlink channel fading between user and drone Characterized using the Ricean channel model, in time slots user Associated subcarriers With drones Channel coefficient The calculation is as follows: in, It is the distance between the drone and the user; It is the average channel power gain at a reference distance of 1 meter; For the decline factor; Represents the line-of-sight component, and ; Indicates the non-line-of-sight component; Indicates the modulus; The performance index analysis model is as follows: In the formula, , , These are total throughput, average latency, and average signal-to-interference-plus-noise ratio, respectively. and users respectively In the time slot Data rate and latency; This is an indicator function that outputs 1 if the input is true, and 0 otherwise. Assign joint factors to node associations and subcarriers; For users Associated with nodes and subcarriers The signal-to-interference-to-noise ratio; and These represent the number of users requesting low-latency service and wide-coverage service, respectively. This represents the set of users requesting high-throughput services. This represents the set of users requesting low-latency services. This represents the set of users requesting broad-coverage services; It is a collection of base stations and drone nodes. It is a set of subcarriers.
3. The performance gap-driven integrated air-ground load balancing method according to claim 1, characterized in that, The performance dissatisfaction rate mentioned in step S2 The calculation formula is: in, User The downlink rate associated with the ground layer; This indicates the service latency between the user and the base station; This indicates the signal-to-interference-to-noise ratio (SNR) between the user and the base station. , and These are the user's desired data rate, latency, and signal-to-noise ratio threshold, respectively. The cross-layer offloading strategy is as follows: users are given priority access to the ground layer; when the ground layer is overloaded, all users are offloaded. Sort in ascending order and unload excess users to the airspace layer in descending order.
4. The performance gap-driven integrated air-ground load balancing method according to claim 1, characterized in that, The utility function mentioned in step S3 is: in, These represent the weights of total throughput, average latency, and average signal-to-interference-plus-noise ratio, respectively. This represents the initial maximum benefit of latency; The joint optimization problem is expressed as: in, Represents binary variables The range of values for; This indicates that in the time slot Each user can be allocated at most one subcarrier; Indicates in time slot Each user can be associated with at most one network component; furthermore, This restricts the nonnegativity of the transmit power, while This indicates that the number of users associated with the same base station cannot exceed the maximum capacity. ; Indicates user Associated with base station and subcarriers The joint factor; Indicates user Related to drones and subcarriers The joint factor; Indicates base station For users Power allocation; Indicates drone For users Power allocation; M is the set of base stations; For the set of all users; This is a collection of drone nodes.
5. The performance gap-driven integrated air-ground load balancing method according to claim 1, characterized in that, Step S4 is specifically implemented as follows: The ground-level agent, based on the current state... The optimization actions for the current time slot are given—user-base station association and resource allocation. Based on the current action and state, calculate the total number of users in the current time slot. ; If the base station is overloaded, a load balancing scheme will be triggered, and step S2 will be executed. At this time, the spatial layer agent will be based on the current state. The optimization actions for the current time slot are given—unloading user-drone association and resource allocation. The agent calculates the current reward; otherwise, the spatial layer agent skips the current iteration; calculates the reward for the current time slot and moves to the next state; until the algorithm converges or the iteration ends.
6. The performance gap-driven integrated air-ground load balancing method according to claim 1, characterized in that, The air-to-ground imperfect channel modeling described in step S5 is as follows: In the formula, and These represent the actual channel state information of the base station and the drone, respectively. and These are the estimated channel state information known to the base station and the drone, respectively. and The channel errors of the base station-user and UAV-user channels are represented respectively; the common Gaussian channel error is used for modeling, and it follows an independent and identically distributed complex Gaussian distribution. , Let be the variance of this distribution.