Adaptive load balancing ground user access method for unmanned-aerial-vehicle-assisted network

By converting the ground user access problem into MDP and using the DQN algorithm to optimize UAV deployment, the access and load balancing problems of ground users in UAV-assisted networks are solved, achieving more efficient access and faster convergence.

WO2025213797A1PCT designated stage Publication Date: 2025-10-16NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
PCT/CN2024/136284
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-12
Filing Date
2024-12-03
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

In UAV-assisted networks, the access of ground users faces problems of seamless cooperation and load balancing, especially in the BS-UAV-NTN integrated network. Existing technologies find it difficult to effectively utilize limited channel resources and optimize the energy consumption of data flows.

Method used

A distributed cooperative learning method is adopted to transform the ground user access problem into a Markov decision process (MDP). A UAV deployment algorithm based on a deep Q-learning network (DQN) is used to design the state and action space and reward mechanism to optimize the load balancing access in the BS-UAV-NTN integrated network.

Benefits of technology

It achieves rapid achievement of a stable state in a dynamic and unknown environment, improves the efficiency of ground user access and load balancing capabilities. Simulation results show that compared with the traditional Q-learning algorithm, the number of accessed GUs is significantly increased by 22.95%, and the performance is 22.79% better than the QL algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136284_16102025_PF_FP_ABST
    Figure CN2024136284_16102025_PF_FP_ABST
Patent Text Reader

Abstract

An adaptive load balancing (ALB) ground user (GU) access method for an unmanned-aerial-vehicle-assisted network. By means of an unmanned aerial vehicle deployment algorithm based on a deep Q-learning network (DQN), and ALB for GU access, a GU access problem in a BS-UAV-NTN is converted into a maximization problem, and the maximization problem is converted into a Markov decision process (MDP) problem for unmanned aerial vehicle deployment in an unknown environment. The method comprises an unmanned aerial vehicle deployment algorithm based on a DQN, and an access scheme for performing priority ranking on BSs and unmanned aerial vehicles. A simulation result shows that the access scheme is superior to conventional Q-learning and random schemes in the reward aspect, and the number of access GUs.
Need to check novelty before this filing date? Find Prior Art

Description

Adaptive load balancing ground user access method for uav-assisted networks TECHNICAL FIELD

[0001] The present application relates to the field of unmanned aerial vehicle communication networks, and particularly to an adaptive load balancing ground user access method for uav-assisted networks. BACKGROUND

[0002] Recently, unmanned aerial vehicles (UAVs) have become a key component of modern communication networks, playing a vital role in both ground base stations (BSs) and non-terrestrial networks (NTNs). Their importance stems from their extensive coverage capabilities and ability to facilitate collaborative decision-making. However, the dynamic nature of modern communication networks, with various node types coexisting, changing demands, and fluctuating channel conditions, presents significant challenges, particularly in enabling ground users (GUs) to access. This is particularly challenging when dealing with unknown environments in BS-UAV-NTN integrated networks.

[0003] To address the aforementioned challenges, researchers have proposed a series of methods to address the GUs access problem in uav-assisted networks. Some of these methods prioritize efficient coverage control, introducing distributed learning methods to mitigate interference and enhance data collection capabilities; others focus on facilitating multiple link modes. For example, data relaying in both uav and satellite modes is achieved through Dantzigg-Wolfe decomposition and column generation-based algorithms. Additionally, some research work provides a unique approach, utilizing normal and store-and-forward modes to optimize throughput through thorough probabilistic analysis.

[0004] Therefore, in uav-assisted networks, GUs access faces two key challenges. First, seamless collaboration between ground and non-terrestrial infrastructure to achieve efficient network access with limited channel resources remains an underdeveloped area of current research. Second, the problem of load balancing, which is crucial for minimizing energy consumption and optimizing data flow, is often overlooked in existing work. SUMMARY

[0005] The present application proposes an adaptive load balancing ground user access method for unmanned aerial vehicle assisted network, aiming to solve these challenges by proposing innovative solutions to facilitate better cooperation and improve load balancing strategies. The present application adopts a distributed cooperative learning method to form an adaptive load balancing GU access scheme in the BS-UAV-NTN integrated network. The GU access problem is converted into a minimization task, and it is converted into a Markov decision process (MDP) for unmanned aerial vehicles. The MDP uses a deep Q-learning network (DQN) based unmanned aerial vehicle deployment algorithm to solve it. The state and action space, reward and state transition probability are designed for the MDP. The DQN-based algorithm implemented, and the adaptive and load balancing (ALB) access scheme, emphasize the BS-primary, NTN-secondary and quality of service (QoS)-oriented approach. In the present application, the channel between the unmanned aerial vehicle and the balloon or satellite is named as NTN link, which will not significantly affect the unmanned aerial vehicle deployment or direct GU connection due to power and distance limitations.

[0006] The adaptive load balancing ground user access method for unmanned aerial vehicle assisted network comprises the following steps:

[0007] Step 1, design to use the integrated network of ground base station BS-unmanned aerial vehicle UAV-non-terrestrial network NTN, wherein BSs, UAVs and NTNs jointly serve ground users GUs;

[0008] Step 2, model the system, establish the UAV to GU channel U2G, obtain the path loss model of the inter-channel U2G link, and the signal to interference and noise ratio model and the outage probability model;

[0009] Step 3, establish the BS to GU channel B2G model, give the signal to interference and noise ratio model and the outage probability model;

[0010] Step 4, in the B2G model, each BS provides service to the GU based on the factors of shorter distance, less interference and lower channel noise; assuming that the GUs have the same interference and noise model, a BS tends to establish a link with the nearest GU; the GUs upload or download data with the BS or the UAV, prefer to consider the BS as the primary and the NTN as the secondary, and face the QoS to realize the access management;

[0011] Step 5, based on the optimization of GU access in the unknown environment of BS-UAV-NTN integrated network, set the objective function and the corresponding constraint condition;

[0012] Step 6, design an adaptive and load balancing ALB access scheme for GUs, which gives priority to BSs and NTN, and provides a QoS-oriented solution for the integrated network; the ALB scheme focuses on supporting BS and NTN auxiliary functions to ensure adaptive and load balancing access of GUs in various network scenarios;

[0013] The ALB scheme of GUs access is specifically that, according to the DQN-based UAV deployment algorithm of steps 1-5, the appropriate number of UAV movements is provided to optimize the load balancing of all GUs links, ensure that GUs are connected to the base station first, and then to the UAV, and if the GUs in the coverage range no longer need to be linked, the UAV will leave; at the same time, the UAV flight is divided into cruise-hover time slots, and the decision of the cruise time slot is also made by the DQN-based UAV deployment algorithm.

[0014] The beneficial effects achieved by the present application are:

[0015] (1) The ALB scheme of GUs access supported by the DQN-based UAV deployment algorithm is proposed to solve the dynamic and load imbalance problem of GUs.

[0016] (2) It can quickly reach a stable state and can more effectively store experience in the uncertain and dynamic environment of the BS-UAV-NTN integrated network, and has better adaptability to unknown environments.

[0017] (3) With the increase of the number of UAVs, more access GUs are realized, and the neural network is used for decision-making to realize faster convergence, and the simulation results show that compared with the QL algorithm, the number of accessed GUs is significantly improved by 22.95%.

[0018] (4) The simulation results show that the performance of this scheme is better than that of the QL algorithm by 22.79% under different UAV conditions.

[0019] (5) The performance increases with the number of BSs, and compared with the best scenario, the performance of this method is 10.68% better than that of the QL algorithm and 5 BSs. BRIEF DESCRIPTION OF DRAWINGS

[0020] Fig. 1 is a diagram illustrating the movement of a UAV in an embodiment of the present application

[0021] Fig. 2 is a workflow diagram of the ALB scheme in an embodiment of the present application: the UAV performs a move-hover operation to access more GUs, but the GUs mainly access BSs, and the UAV is second.

[0022] Fig. 3 is a simulation performance diagram of the total reward of each UAV and the reward of different sets in an embodiment of the present application.

[0023] Figure 4 is a simulation performance graph of the number of GUs accessed and different numbers of GUs in an embodiment of the present application.

[0024] Figure 5 is a simulation performance graph of the number of GUs accessed and different numbers of UAVs in an embodiment of the present application.

[0025] Figure 6 is a simulation performance graph of the number of GUs accessed and different numbers of BSs in an embodiment of the present application. DETAILED DESCRIPTION

[0026] The technical solutions of the present application will be further described in detail below in conjunction with the accompanying drawings of the specification.

[0027] The adaptive load balancing ground user access method of the UAV-assisted network is realized by the following technical solutions:

[0028] In the first aspect, the present method introduces a unique GUs access problem tailored for BS-UAV-NTN integrated networks, which initially aims to maximize GUs access. The present method converts this problem into a UAV-centric MDP and proposes an innovative UAV deployment algorithm based on DQN. This algorithm is designed with carefully designed state and action spaces, reward mechanisms, and state transition processes to effectively solve the MDP problem. It includes the following steps:

[0029] Step 1, a traditional access point, such as a 5G / 6G base station, is considered, which can handle normal access of a small group of GUs, but it is difficult to handle large-scale events such as live performances or sports matches. In order to solve the congestion problem in these situations, the present application proposes an alternative transmission path using BS-UAV-NTN integrated networks, as shown in Figure 1.

[0030] In this network, BSs, UAVs, and NTNs jointly serve GUs. When data bursts, GUs can use UAVs to relay data to other BSs or NTNs. However, if GUs have already obtained good coverage from BSs, they will prefer to choose BSs with lower latency.

[0031] Assuming that network administrators lack direct knowledge of GUs' experience of poor QoS, as GUs are not easy to share their location, deploying UAVs to explore and exploit this unknown environment becomes a logical solution. This deployment helps to bridge the gap between BSs and NTNs.

[0032] Step 2, the system is modeled, first U2G is established.

[0033] The U2G or UAV-to-GU channel is modeled as a generalized Nakagami-m fading model, which is found to accurately model the small-scale fading conditions of the UAV-to-GU. Thus, the probability density function (PDF) of the random variable x can be written as:

[0034] where x is the amplitude of the received signal, Γ(m) is the gamma function, and m and Ω are the shape and scale parameters, respectively.

[0035] Thus, the path loss model for the inter-link U2G channel can be expressed as:

[0036] where is the path loss at a reference distance d0m, typically d0= 1, n is the path loss exponent, d is the distance of the U2G link, and X U2G is a lognormal shadowing component with a standard deviation. The path loss exponent n is a function of the Nakagami-m parameter m and the ratio of the carrier frequency to the speed of light f c / c, expressed as:

[0037] where a is a constant that depends on the environment and the antenna heights of the UAV and the GU. Furthermore, the increasing access of GUs makes the carrier frequency f c insufficient.

[0038] The signal-to-interference-and-noise ratio (SINR) expression for the U2G link over a Nakagami-m fading channel is written as

[0039] where P t is the transmit power, N0is the noise power spectral density, Ω UAV and Ω GU are the antenna gains of the UAV and the GU, respectively, and m is the Nakagami-m fading parameter. I U2G is the total interference power, which depends on the number of GUs within its communication range.

[0040] Finally, the outage probability OP of the U2G channel is expressed in terms of the Q function as:

[0041] where is the Gaussian Q function, γ th is the outage threshold SINR, f γ (γ) is the PDF of the SINR, and the parameter γ represents a discount factor from 0 to 1, reflecting whether future rewards are more valuable than immediate rewards when updating the policy. It is obtained as:

[0042] The Gaussian Q function, the PDF of SINR, and the threshold SINR of outage are used to model the reality of U2G links and lay the foundation for the following problems.

[0043] In step 3, the B2G model is established, and the path loss model of the B2G link is given:

[0044] where denotes the path loss at the reference distance d0=1. The parameter n represents the path loss exponent, which takes a value in the range of 2-4. The variable d is the distance between the transmitter and the receiver in the B2G link. B2G denotes the lognormal shadowing with a mean of zero and a standard deviation of σ shadow .

[0045] Therefore, the SINR model is represented as:

[0046] where P t is the transmit power, Ω BS is the antenna gain of the BS, Ω GU is the antenna gain of the GU, which is the same as in U2G. β is the channel gain. N0 is the noise power spectral density, W is the bandwidth, I B2G is the total interference power; PL(d) represents the path loss, which is related to the distance d.

[0047] The OP model is represented as:

[0048] where Q(x) is the Gaussian Q function, γ th is the threshold of SINR, below which the communication is considered to be in outage. Different OP models are designed to consider the unique characteristics of Nakagami-m and 5G BS channels;

[0049] In step 4, considering the B2G model, the BS prioritizes the service to the GUs based on shorter distance, less interference, and lower channel noise. In the B2G model, assuming that the GUs have the same interference and noise model, a BS tends to establish a link with the closest GU.

[0050] In this method, GUs upload or download data with BSs or UAVs. However, BSs only show the best performance under a certain GU capacity, as shown by the light-colored GUs in Figure 1. Beyond this threshold, additional GU access will result in suboptimal performance, causing problems such as packet loss, delay, or retransmission, as shown by the dark-colored GUs in Figure 1. Therefore, dark-colored GUs should access an available drone for data relay. Therefore, the BS-UAV-NTN integrated network, especially as shown in Figure 1, requires a solution that prioritizes BSs as the primary and NTNs as the secondary, and is QoS-oriented to achieve effective access management.

[0051] Then assume that GU accesses P U2G or P B2G is lower than P th , then this link is named as potential access, if there is an available BS or an available UAV, it will turn to valid access, of course, BS is used first. Then write down an and UAV access set.

[0052] OP model is represented as:

[0053] where U i is the ith drone, B j is the jth base station.

[0054] Step 5, the main goal is to optimize the GUs access in the unknown environment in the BS-UAV-NTN integrated network. Although BSs can satisfy a larger number of GUs compared to drones, the flexibility of drones is superior to BSs. Considering these factors, the goal is:

[0055] where C1 and C2 represent that the interruption probability is less than or less than the threshold, C3 and C4 represent that the accessed GUs should not break the maximum set and C5 represents that the location of the drone should be within the network area , C6 represents that the speed V UAV of the drone is limited V msx .

[0056] In the second aspect, the ALB access scheme for GUs is comprehensively described in this method. This scheme prioritizes BSs and NTNs, providing a QoS-oriented solution for the integrated network. The ALB scheme focuses on supporting BS and NTN auxiliary functions, aiming to guarantee adaptive and load-balanced access for GUs in various network scenarios. The following steps are included:

[0057] Step 1, the present invention divides the coverage area into a grid structure to improve the convenience of deployment and the speed of decision-making, and to alleviate problems such as coverage overlap and interference. The central part of Figure 1 gives a schematic diagram of the movement of the UAV, which illustrates the UAV moving in various directions, each corresponding to a grid displacement. Therefore, the basic task solved by the present invention includes ensuring the rapid and accurate movement of the UAV.

[0058] 1) State: First, the UAV needs to move according to its own state in order to make a decision as soon as possible after obtaining the state. Since the UAV has no information about the distribution of GUs, they will grid the network area to quantify the state. Then, the state of the UAV is defined as

[0059] where l U represents the position of the UAV, N grid represents the grid number on one side of the area. In addition, the state space is represented as where t represents the time slot outside the total time T. Since the UAV moves one step at a time, the transition probability is represented as:

[0060] wherein, represents the transition probability. In other words, for the movement of the UAV, the future is independent of the past.

[0061] 2) Action: Based on the state and system assumptions, the UAV action represents the decision or choice made that affects the system. First, define the length space as {1, 2, …, η} × S m , where S m is the movement step length, and η is the maximum multiple of the movement step length, represented by V max in equation (11); then, define the direction space as {N, E, W, S, H}, representing the five directions of north N, east E, west W, south S, and hovering H; finally, define the action space as the Cartesian product of and , represented as

[0062] 3) Reward: The reward function assigns a value to each state-action pair, and the reward represents the immediate desirability or related cost of taking a particular action in a particular state. - It is represented as the difference in IAs that have the greatest impact on the objective in equation (11), and is defined as:

[0063] 4) Purpose: To formulate the problem in the equation; (11) As an MDP problem, use a tuple<S,A,P,R> Rewrite the objective to maximize the expected total discounted reward for each drone as:

[0064] The policy π represents the mapping from state space to action space; the parameter γ represents the discount factor ranging from 0 to 1, reflecting whether future rewards are more valuable than immediate rewards when updating the policy; the optimal policy for drone i is Satisfy the Bellman equation. ;The objective function is expressed as:

[0065] where i is the action a of drone i i In state s′ i The next state after that, V(s′ i ) is the objective function of the next state, p i (s′ i ∣s i ,a i ) refers to conditional probability. In step 2, through the MDP modeling of step 1, each drone has an action A and a state S. If a precise grid and speed multiplier η are set, the resulting complexity is enormous. Furthermore, a drone's movement may interfere with other drones, affecting their iterations. Therefore, a DQN-based algorithm is employed due to its deep processing and adaptive capabilities to handle large-scale convergence problems with inter-agent interaction.

[0066] Therefore, Q-learning is achieved by Q(s,a)=Q(s,a)+α(r+γmax a′ Q(s′,a′)-Q(s,a)) updates the Q value, where the variables can be referenced; however, due to the limited state space, Q-learning must overcome the limitation of the state space and converge quickly by updating the following formula.

[0067] Here because The DQN model naturally generalizes beyond the states and actions it was trained on.

[0068] At each training iteration, the replay memory uniformly sample the experience e t =(s t ,a t ,r t ,s t+1 ), the network loss is determined as follows:

[0069] The parameters θ of the target network- are copied from the parameters of the current policy network and are not updated frequently but at a certain frequency. Among them are the outdated update targets given by the target network . The updates made in this way have been proven to be easy to handle and stable.

[0070] ALB scheme for GUs access: The DQN-based drone deployment algorithm provides appropriate number of drone movement steps to optimize the load balancing of all GUs links. This ensures that GUs can effectively connect to the base station first and then to the drone if the GUs in the coverage no longer need to link. The drone will leave. This scheme divides the drone flight into cruise-hover time slots, and the decision of the cruise time slot is made by the DQN-based drone deployment algorithm. Figure 2 illustrates the workflow of the ALB scheme.

[0071] In this embodiment, the method establishes a fully connected layer, Adam optimization and DQN model of mean square error (MSE) loss function. For this network scenario, the method defines U2G and B2G channels, and the key parameters are shown in Table 1:

[0072] Table 1 Key parameters of U2G and B2G channels

[0073] For comparison, the method uses the Q-learning scheme with parameters α q = 0.8, γ q = 0.9, but in the simulation results, the variables are ε = 0.1 and ε = 0.3, i.e. QL-epsilon = 0.1 and QL-epsilon = 0.3. In addition, the method also proposes a random scheme with random selection in the action space. All other operations are the same, including taking action, moving, hovering, time slot, etc. Finally, the ALB scheme is denoted as DQN-LR = 0.005 and DQN-LR = 0.001, representing different learning rates. Figure 3 configures 1500 GUs, 2 drones and 2 BSs. Figure 4 involves 2 BSs and 2 drones. Figure 5 considers 2 BSs and 1500 GUs. Figure 6 shows 2 drones and 1500 GUs. All GUs visited in Figures 4, 5, 6 represent the sum of more than 10000 sets per set.

[0074] In Fig. 3, two DQN models outperform the Q-learning (QL) models in reaching stability. Among them, DQN-LR=0.005 stabilizes around 2000 episodes, which is 800 episodes faster than DQN-LR=0.001, 500 episodes faster than QL-epsilon=0.1, and 1800 episodes faster than QL-epsilon=0.3. It is worth noting that the random scheme cannot stabilize. In addition, when the two QL models converge, their rewards are significantly lower than those of the two DQN models. This is because DQN can more effectively store experiences in the uncertain and dynamic environment of the BS-UAV-NTN integrated network. When QL-epsilon=0.1 and QL-epsilon=0.3, both of them are superior to the random scheme in adapting to the unknown environment. QL-epsilon=0.1 obtains a higher reward, demonstrating its focus on a wider range of experiences, making it very suitable for this challenging environment.

[0075] In Fig. 4, it is clear that as the number of UAVs increases, all five schemes achieve more visited GUs. This leads to the preliminary conclusion that UAVs make a positive contribution to the BS-UAV-NTN integrated network. It is worth noting that these two DQN schemes outperform the other three schemes, making decisions using neural networks and achieving faster convergence on 10000 episodes. The two Q-learning (QL) schemes also perform well in terms of visited GUs, thanks to their adaptive characteristics in unknown environments. At the same time, the number of visited GUs for the random scheme increases as more GUs result in more visited GUs. Therefore, when the number of GUs is 6000, DQN-LR=0.005 is the best, and compared with QL-epsilon=0.1, the number of visited GUs is significantly improved by 22.95% under the same conditions.

[0076] In FIG. 5, both versions of the DQN algorithm consistently achieve the highest number of GUs visited. In the Q-learning variant, epsilon = 0.1 outperforms epsilon = 0.3, while the Random method exhibits the lowest performance. DQN-LR = 0.005 generally increases with the number of drones, suggesting that DQN-LR is effective in maximizing GUs access as the number of drones increases. In comparison to DQN-LR = 0.005, a lower learning rate (0.001) for DQN-LR can result in a slightly slower learning process. QL-epsilon = 0.3 does increase, but the rise is smaller in comparison to QL-epsilon = 0.1. This can be due to the higher exploration rate (epsilon = 0.3) causing the algorithm to explore suboptimal actions, resulting in slower growth. It is worth noting that the best performing DQN-LR = 0.005 indicates a 22.79% increase in efficiency in visiting GUs in comparison to QL-epsilon = 0.1 with 6 drones.

[0077] In FIG. 6, performance generally rises with the number of BSs in all algorithms. Learning algorithms (DQN and Q-learning) consistently outperform the random selection strategy. DQN-LR = 0.005 exhibits an initial increase, peaks around 5 BSs, and then decreases slightly. This suggests that having more BSs can initially improve GUs access, but beyond a certain point, the benefit decreases. This can be because the BSs have already visited a sufficient number of GUs, limiting further improvements in drone movement. In comparison to the best scenario, DQN-LR = 0.005 and 3 BSs outperform QL-epsilon = 0.1 and 5 BSs by 10.68%.

[0078] The above description is merely that of preferred embodiments of the application, and the protection scope of the application is not limited to the above embodiments. Any equivalent modifications or changes made by those skilled in the art based on the disclosure of the application shall fall within the protection scope of the claims.

Claims

1. An adaptive load balancing ground user access method for a UAV-assisted network, characterized by: The method comprises the following steps: Step 1: Design an alternative transmission path using a ground base station (BS)-unmanned aerial vehicle (UAV)-non-terrestrial network (NTN) integrated network, where BSs, UAVs, and NTNs jointly serve ground users (GUs); Step 2: Model the system, establish the UAV to GU channel U2G, obtain the path loss model of the channel between U2G links, as well as the signal-to-interference-and-noise ratio model and the outage probability model; In step 2, U2G is modeled as a generalized Nakagami-m fading model; The probability density function PDF of the random variable x is written as: Where x is the amplitude of the received signal, Γ(m) is the gamma function, m and Ω are the shape and scale parameters respectively; Therefore, the path loss model of the U2G link channel is expressed as: Where, is the path loss at the reference distance d0 meters, generally d0 = 1, n is the path loss exponent, d is the distance of the U2G link, X U2G is the lognormal shadow component with standard deviation; The path loss exponent n is the product of the Nakagami-m parameter m and the ratio of the carrier frequency to the speed of light f. c The function of / c is expressed as: Where α is a constant that depends on the environment and the antenna heights of the UAV and GU; The SINR expression of the U2G link over the Nakagami-m fading channel is written as: Where P t is the transmission power, N0 is the noise power spectrum density, Ω UAV and Ω GU are the antenna gains of the UAV and GU respectively, m is the Nakagami-m fading parameter; I U2G is the total interference power, which depends on the number of GUs within its communication range; Finally, the outage probability OP of the U2G channel is expressed by the Q function as: Where, is the Gaussian Q function, γ th is the shutdown threshold SINR, f γ (γ) is the PDF of SINR, and the parameter γ represents a discount factor ranging from 0 to 1, reflecting whether future rewards are more valuable than immediate rewards when updating the strategy; we get: The Gaussian Q function, the PDF of SINR, and the threshold SINR for outages are used to simulate the reality of U2G links; Step 3: Establish a B2G model for the BS-to-GU channel, and provide a signal-to-interference-and-noise ratio model and an outage probability model. In step 3, the path loss model of the B2G link is: In the formula represents the path loss at the reference distance d0 = 1; the parameter n represents the path loss exponent, which ranges from 2 to 4; the variable d is the distance between the transmitter and the receiver in the B2G link; X B2G represents a lognormal distribution with mean zero and standard deviation σ shadow ; Therefore, the SINR model is expressed as: PL(d) represents the path loss, which is related to the distance d; Where P t is the transmission power, Ω BS is the BS antenna gain, Ω GU is the antenna gain of GU, which is the same as the model in U2G, β is the channel gain, N0 is the noise power spectral density, W is the bandwidth, I B2G is the total interference power; The OP model is expressed as: Where Q(x) is the Gaussian Q function, γ th is the SINR threshold, below which communication is considered to be in an interrupted state; Step 4: In the B2G model, each BS prioritizes serving the GU based on shorter distance, less interference, and lower channel noise. Assuming GUs have the same interference and noise model, a BS prefers to establish a link with the GU closest to it. When a GU uploads or downloads data to or from a BS or UAV, the BS is prioritized as primary and the NTN as secondary, with QoS-oriented access management. Step 5: Based on optimizing the GUs access in unknown environment in BS-UAV-NTN integrated network, set the objective function and corresponding constraints; In step 5, the objective function is: Where C1 and C2 indicate that the interruption probability is less than or less than the threshold, and C3 and C4 indicate that the accessed GUs should not break the maximum value set and C5 indicates the position of the drone Should be in the network area C6 represents the speed of the drone V UAV Limited V max , P th The threshold value representing the probability of interruption; Step 6: Design an adaptive and load-balanced ALB access scheme for GUs. This scheme prioritizes BSs and NTNs, providing a QoS-oriented solution for the integrated network. The ALB scheme focuses on supporting BS and NTN auxiliary functions to ensure adaptive and load-balanced access to GUs in various network scenarios. Step 6 includes the following steps: Step 6-1: Divide the coverage area into a grid structure, where the UAV traverses in various directions, and each direction corresponds to a grid displacement. Based on the UAV state and action, establish the UAV reward function and obtain the UAV's objective function that maximizes the reward expectation. Step 6-2: Through Markov decision process MDP modeling, each drone has an action A and a state S, and a DQN-based algorithm is used to handle large-scale convergence problems with inter-agent impact; The ALB solution for GUs access is to provide an appropriate number of drone movement steps according to the DQN-based drone deployment algorithm in steps 1-5 to optimize the load balancing of all GUs links, ensuring that GUs are connected to the base station first, followed by drones. If the GUs within the coverage area no longer need to be connected, the drone will leave; at the same time, the drone flight is divided into cruise-hover time slots, and the decision on the cruise time slot is also made by the DQN-based drone deployment algorithm.

2. The method for adaptive load balancing ground user access in a UAV-assisted network according to claim 1, characterized in that: In step 1, in the BS-UAV-NTN integrated network, when data bursts, GUs use drones to relay data to other BSs or NTNs; when GUs have obtained coverage from BSs, BSs with lower latency are preferentially selected.

3. The method for adaptive load balancing ground user access in a UAV-assisted network according to claim 1, characterized in that: In step 4, assume that GU accesses P U2G or P B2G Lower than P th Any BS or UAV, then name this link as a potential access, if there is an available BS or available UAV, it will turn into an effective access, giving priority to the BS; then write a and The access set, The OP model is expressed as: The outage probability OP of the U2G channel is expressed by the Q function as follows: P th The threshold value representing the probability of interruption; Among them U i For the i-th drone, B j is the jth base station.

4. The method for adaptive load balancing ground user access in a UAV-assisted network according to claim 1, characterized in that: In step 6-1, the drones move according to their own states. Since the drones have no information about the distribution of GUs, they grid the network area to quantify the state. Then, the state of the drone is defined as: Among them, l U Indicates the position of the drone, N grid represents the grid number on one side of the region; in addition, the state space is represented as Where t represents the time slot outside the total time T; since the drone moves one step at a time, the transition probability is expressed as: Where, Denote the migration probability; Based on the state and system assumptions, drone actions represent decisions or choices made that affect the system; first define the length space {1,2,…,η}×S m , where S m is the moving step length, η is the maximum multiple of the moving step length, and V in formula (11) is used max Representation; then, define the direction space {N,E,W,S,H}, indicating the five directions of north N, east E, west W, south S and hover H; finally, the action space Expressed as and The Cartesian product of The drone reward function assigns a value to each state-action pair. The reward represents the immediate desirability or associated cost of taking a specific action in a specific state. It is expressed as the difference between the IAs that have the greatest impact on the goal in equation (11) and is defined as: Formulate the problem in equation (11) as an MDP problem with a tuple<S,A,P,R> Representation; rewrite the objective to maximize the total discounted reward expectation of each drone, expressed as: Refers to the total discounted reward expectation of each drone; (11) As an MDP problem, use a tuple<S,A,P,R> Represents state, action, reward, and next state respectively; r k 、s k Represents a tuple<S,A,P,R> elements, k represents the number; The policy π represents the mapping from state space to action space; the parameter γ represents the discount factor ranging from 0 to 1, reflecting whether future rewards are more valuable than immediate rewards when updating the policy; the optimal policy for drone i is Satisfies the Bellman equation; the objective function is expressed as: where s′ i It is the action of drone i a i In state s i The next state after that, V(s′ i ) is the objective function of the next state, p i (s′ i ∣s i ,a i ) refers to conditional probability.

5. The method for adaptive load balancing ground user access in a UAV-assisted network according to claim 1, characterized in that: In step 6-2, Q-learning is done by Q(s,a)=Q(s,a)+α(r+γmax a′ Q(s′,a′)-Q(s,a)) updates the Q value, where the variable can be referenced; Q-learning overcomes the limitations of the state space and converges quickly by updating the following formula: Here because The DQN model generalizes things outside the states and actions it is trained on; α is a constant that depends on the environment and the antenna heights of the UAV and GU; Indicates the gradient; L(θ i ) represents the loss function; At each training iteration i, the replay memory uniformly sample the experience e t =(s t ,a t ,r t ,s t+1 ), the network loss is determined as follows: The parameters θ of the target network - is the parameter θ from the current policy network i Copied and not updated frequently, but updated at a certain frequency; in For the target network The given stale update target.

6. The method for adaptive load balancing ground user access in a UAV-assisted network according to claim 1, characterized in that: In step 6, the steps of the DQN-based drone deployment algorithm are as follows: Input: Drone positions, GUs positions, GUs data traffic, GUs channel parameters, base station positions, running time T; Output: Drone positioning; 1: Initialize Q-learning parameters θ, λ, α; DQN parameters w, b; 2: Start when the system time t < T; 3: Get status 4: Obtain actions from the DQN-based decision; 5: Obtain rewards; 6: Obtain the DQN-based decision; 7: Update state space 8: The drone moves following the action output; 9: Update t; 10: End.

Citation Information

Patent Citations

  • Air-ground non-orthogonal multiple access uplink transmission method based on intelligent reflecting surface

    CN114422056A

  • Unmanned aerial vehicle assisted air-ground communication optimization algorithm based on deep reinforcement learning algorithm

    CN114826380A

  • Self-adaptive load balancing ground user access method of unmanned aerial vehicle auxiliary network

    CN118042528A

Cited By

  • Heterogeneous multi-agent efficient collaboration scheme in air-ground sensing integrated system

    CN121174190A

  • Power equipment defect sensing method and system based on communication signal shielding analysis

    CN122282810A

  • Load-driven unmanned aerial vehicle computing task regulation method and system

    CN122458097A