A user-centered multi-uav fair communication dynamic deployment method

By adopting a user-centric dynamic deployment method for fair multi-UAV communication, and utilizing genetic algorithms and multi-agent reinforcement learning algorithms to optimize UAV trajectories and resource allocation, this approach solves the problem of changing user needs in traditional UAV deployment schemes, realizes an efficient and flexible multi-UAV communication network, and improves fair throughput and adaptability.

CN119485327BActive Publication Date: 2025-11-28NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411371936.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-11-28
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

Traditional UAV-centric drone deployment solutions cannot adapt to changing user needs, especially when users leave the target area, they cannot provide continuous service, and it is difficult to achieve flexible deployment of multiple drones to meet the communication needs of mobile users.

Method used

A user-centric, multi-UAV fair communication dynamic deployment method is adopted. By jointly optimizing UAV trajectories, user connections, communication UAV scheduling and power allocation, and using genetic algorithms and multi-agent reinforcement learning (MADRL) algorithms, dynamic deployment of UAVs and bandwidth resource allocation are achieved to ensure fairness and efficient communication.

Benefits of technology

It improves the system's fair throughput by at least 36%, ensuring the performance advantages of multi-UAV communication networks, adapting to the distribution of mobile users, and reducing inter-cell interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119485327B_ABST
    Figure CN119485327B_ABST
Patent Text Reader

Abstract

The application discloses a user-centered multi-UAV fair communication dynamic deployment method, calculates the cluster fair throughput centered on users, adopts a genetic algorithm to pretreat the relative distance between users in a cell and the relative distance between cells, utilizes a MADRL algorithm to allocate exclusive flight actions for each UAV so that the UAVs reach corresponding positions within a specified time slot, and selects the UAVs serving each UAV according to the corresponding distance between the UAVs and the ground cell to complete one-to-one matching between the UAVs and the ground cell. According to the MADRL algorithm, the bandwidth resource allocation of the corresponding UAVs is realized in each user cell. The above steps are repeated within a specified number of time slots until multiple convergent actor online networks are obtained, and the multi-UAV communication dynamic deployment is realized. The application solves the problem that the traditional UAV-centered deployment strategy not only fails to fully utilize the flexibility of the UAVs, but also cannot guarantee the fairness between multiple mobile users.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of multi-unmanned aerial vehicle dynamic deployment, and particularly relates to a user-centered multi-unmanned aerial vehicle fair communication dynamic deployment method. BACKGROUND

[0002] With the development of low-altitude economy, unmanned aerial vehicles equipped with mobile base stations (BSs) can act as temporary aerial base stations (ABSs) to achieve rapid deployment and provide wireless coverage in designated areas. Compared with ground fixed BSs, ABSs can provide timely communication and high-quality services in the face of emergencies, and have the characteristics of high flexibility and low deployment cost.

[0003] However, with the continuous increase in the number of users, the challenge of effectively deploying multiple UAVs to achieve mobile user access becomes more apparent. In the face of this challenge, traditional UAV-centered strategies include deploying multiple fixed UAVs in a designated area, and each UAV only serves users within its coverage area. Therefore, in this deployment scheme, it is crucial to optimize the coverage area of each UAV, which can be achieved by determining the optimal flight height according to the transmission power of the UAV. Unfortunately, the UAV-centered deployment scheme cannot take advantage of the flexibility of unmanned aerial vehicles, and it is difficult to adapt to changes in user demand, especially when users leave the target area, it cannot provide continuous service.

[0004] Therefore, it is important to build a user-centered multi-unmanned aerial vehicle communication network, dynamically cluster adjacent mobile users, and enable each unmanned aerial vehicle base station to adaptively change its position according to the time-varying user-centered cell. SUMMARY

[0005] The application provides a user-centered multi-unmanned aerial vehicle fair communication dynamic deployment method, which realizes unmanned aerial vehicle trajectory planning for ground user communication coverage by jointly optimizing unmanned aerial vehicle trajectory, user connection, communication unmanned aerial vehicle scheduling and power allocation.

[0006] Technical scheme: The user-centered multi-unmanned aerial vehicle fair communication dynamic deployment method provided by the application comprises the following steps:

[0007] (1) According to the average path loss model of communication between unmanned aerial vehicles and users, and combining the Jain fairness index, the user-centered cell fair throughput is obtained;

[0008] (2) Model the dynamic deployment optimization problem of user-centered multi-UAV considering fairness;

[0009] (3) using genetic algorithm to preprocess the relative distance between users in the cell and the relative distance between cells to clarify the user-centered cell distribution in each time slot;

[0010] (4) using multi-agent reinforcement learning algorithm MADRL to assign a dedicated flight action for each UAV according to the distribution of ground user cells so that it reaches the corresponding position within the specified time slot;

[0011] (5) the ground user cell selects the UAV that serves it according to the corresponding distance between each UAV and it, completing the one-to-one matching of UAVs and ground cells;

[0012] (6) according to the MADRL algorithm, the bandwidth resource allocation of the corresponding UAV in each user cell is realized in combination with the user-centered cell fair throughput;

[0013] (7) repeating steps (3) to (6) within the specified number of time slots until multiple converged actor online networks are obtained, wherein each actor online network corresponds to a UAV;

[0014] (8) each UAV selects its flight action and bandwidth allocation action according to its converged actor online network and the distribution of ground mobile user cells, thereby realizing user-centered multi-UAV communication dynamic deployment considering fairness.

[0015] Further, the step (1) is implemented as follows:

[0016] In the nth time slot, the three-dimensional coordinates of the jth UAV and the two-dimensional coordinates of the ith user are represented as and In each time slot, the mobile users are divided into K user-centered clusters, wherein the center coordinates of the kth user-centered cluster are n The number of users in each user-centered cluster where I represents the total number of users, and the users in the kth user-centered cluster in the nth time slot are represented as ; define the time slot set N = {1,...,n,...,N}, the UAV set The user-centered cell set K n = {1,...,k n ,...,K}, the set of all users I = {1,...,i,...,I}, and the user set in each time slot in different user-centered clusters

[0017] The channel model in the considered network is based on the average path loss model, wherein the probability of LOS path is as follows:​

[0018]

[0019] Where α and β are constant values ​​that depend on the environment. The horizontal distance between the drone and the user is the NLOS probability, denoted as: P LOS =1-P NLOS ;

[0020] Therefore, the average path loss model is expressed as follows:

[0021]

[0022] Among them, f c and v c These represent the carrier frequency and the speed of light, respectively; furthermore, the distance between the drone and the user m is determined by l. i,j Give, η LOS and η NLOS The value of k represents the average additional loss value of LOS and NLOS propagation of the connection; n In the first user-centric community Channel gain between user j and drone j Represented as:

[0023] Finally, the kth n In the first user-centric cluster The throughput per user is:

[0024]

[0025] in, They represent the nth time slot and the kth time slot, respectively. n The j-th drone in the user-centric cluster matching is assigned to the first... Bandwidth per user; This indicates whether the i-th user has been assigned to the k-th user-centric cell in the n-th time slot; Let σ represent the transmit power of the j-th UAV. 2 Indicates noise power; M ICI This indicates inter-cell interference. j'∈J / J represents the j'th UAV in the nth time slot, excluding the one associated with the kth time slot. n The j-th UAV matched in the cell;

[0026] Therefore, in the nth time slot, the kth time slot n In the first user-centric cluster The throughput rate for a single user is defined as: wherein, denotes the total throughput of the kth user in the nth time slot in the kth user-centric cell, n denotes the total throughput of all users in the kth user-centric cluster in the nth time slot, denotes the total throughput of all users in the kth user-centric cluster in the nth time slot; n

[0027] The Jain fairness index is defined to measure the throughput difference between users in the kth cell in the nth time slot: n

[0028]

[0029] Therefore, based on the Jain-FI, the throughput of the kth user-centric cell in the nth time slot is defined as the fair throughput of the user-centric cell: n

[0030]

[0031] Further, the step (2) is implemented as follows:

[0032] For quadrotor UAVs, the thrust of each rotor is a function of the UAV's velocity and acceleration , where bold indicates their vectors:

[0033]

[0034] wherein, and denote the velocity direction vector and velocity magnitude of the UAV, respectively, g denotes the gravity acceleration vector, n r and m denote the number of rotors and the mass of the UAV, respectively, p and S FP denote the air density and the equivalent flat area of the body, respectively; the propulsion power is expressed as:

[0035]

[0036] wherein, d, c T , p, c sol , c fac , d0 represent the local blade section drag coefficient, the thrust coefficient based on the valve area, the air density, the rotor stability, the incremental correction coefficient of the induced power with respect to the linear induced velocity, the disc area of each rotor, and the body resistance ratio of each rotor, respectively, and t c denotes the climb angle of the aircraft; the residual energy of the jth UAV in the nth time slot is expressed as: E max ​​​​Δmaxis the maximum storage energy of the UAVs n is the time slot length;

[0037] Therefore, in order to achieve a user-centric multi-UAV dynamic deployment scheme considering fairness, the optimization problem P1 is modeled as follows:

[0038]

[0039] wherein C1 represents whether the jth UAV serves the kth user-centric cell in the nth time slot, C2 represents the number of users in each user-centric cell; under the guarantee of C3, the total bandwidth allocated to users in each user-centric cell by each UAV is equal to the total bandwidth B of UAVBS; C3 is the guarantee of energy of UAVs not depleted during the task by C4; the movement of the UAV is decomposed into x, y, z three directions, C5 and C6 reflect the maximum values of acceleration and speed in three directions; finally, C7 is to prevent collision between multiple UAVs; the three-dimensional matrix respectively represent the acceleration of the UAV in three directions, the bandwidth allocation of the UAV when serving the matched cluster, the matching of the UAV and the cell, and the matching of the user and the cell.

[0040] Further, the step (3) is implemented as follows:

[0041] The preprocessed users are used to form user-centric cells, and the preprocessed problem P2 is modeled as:

[0042]

[0043] wherein represents the sum of the relative distances between users in the cell, and represents the sum of the relative distances between cells;

[0044] For the preprocessed problem P2, in the genetic algorithm, a chromosome is an effective solution to the preprocessed problem P2, so a complete chromosome is shown as follows: The element ch in the chromosome belongs to the set {1,..., i,..., I,..., I+K-1}, when the value of ch is in [1, I] represents the number of ground mobile users, and when the value of ch is in [I+1, I+K-1] represents the separation point between cells in the chromosome; according to the constraint C2 in P1, there will be an element ch with a value in the range of [I+1, I+K-1] after every I k position in the chromosome, and the value of the element in other positions is in the range of [1, I];

[0045] The chromosomes meeting the requirements will form an initial population pop, the size of which is represented by Ω; for the population pop, the fitness function value fit(ω) of each chromosome is obtained according to the optimization objective function P2; based on the distribution of the fitness function values fit(ω) of all chromosomes in the population, the stochastic universal sampling (SUS) method is introduced, so as to select the chromosomes with higher fitness function values, the number of which is Ω1, where Ω1<Ω; next, for the selected chromosomes, the partial match crossover strategy completes the crossover operation of these chromosomes; if the random number is less than the mutation probability, then the elements of any two points on the chromosome are randomly exchanged; finally, the chromosomes subjected to the above operation are combined with the Ω-Ω1 chromosomes with higher fitness function values in the initial population to form a new generation population; the above steps are repeated until the optimal solution is converged, and finally the chromosome with the highest fitness function value, i.e. the user clustering strategy under the current time slot, is obtained.

[0046] Further, the step (4) is implemented as follows:

[0047] The preprocessing clarifies the user-centered cell distribution of each time slot, and then optimizes the multiple UAV positions and network resource allocation, so as to convert P1 into P3 problem:

[0048]

[0049] For the problem P3, the MADRL algorithm is adopted to solve the challenge of user-centered multi-UAV dynamic deployment, while considering fairness, wherein the multi-agent deep deterministic policy MADDPG is adopted in the MADRL; the MADDPG involves four neural networks, including actor online network, critic online network, actor target network and critic target network, and the update of these networks needs to know the state, action and reward of multiple UAVs; the observation space of each agent is composed of the positions of all UAVs, the velocities in three directions, the positions of ground users and the cell clustering of ground users, and is represented as:

[0050]

[0051] Therefore, the state space of all agents is:

[0052] The action of each UAV contains two parts: the acceleration of the UAV in three directions and the bandwidth allocation of the UAV, and the work is represented as: And the action set of all UAVs is: A={A1,...,A j};

[0053] Since each UAV is uniformly accelerated in each direction, after the actor outputs the action of three direction accelerations of each UAV, the coordinate and velocity of each UAV are updated by the following equations:

[0054]

[0055] Based on the above coordinate and velocity update equations, the position and velocity of each UAV in the specified time slot are obtained. For such actions, the reward function first prevents the UAV velocity or position from exceeding the limit, which is expressed as:

[0056]

[0057] To ensure that all UAVs do not collide, the criterion is that any two UAVs at the same time, as long as the three direction coordinates of the UAV are not equal, meet the constraint C6 in P3, and the reward function is expressed as:

[0058]

[0059] For the flight of the UAV, the following reward function reflects whether it can meet the energy constraint: λ j_3 = E j (n); in the above reward function, is used to adjust the size of the reward value.

[0060] Further, the step (5) is implemented as follows:

[0061] According to the coordinates of the user cell center and the coordinates of each UAV, the distance between different user cells and UAVs is calculated, and the calculation formula is described as follows:

[0062]

[0063] According to the distance formula between the cell and the UAV, starting from the first user cell, each cell is traversed to find the nearest UAV to the center until the matching is completed.

[0064] The reward function is designed to measure the matching effect: wherein, is used to adjust the size of the reward value, is used to limit the distance between each UAV and its matched cell.

[0065] Further, the step (6) is implemented as follows:

[0066] Firstly, the users in the user cell matched with each UAV are sorted by their sum of throughput before time slot n from low to high, and the bandwidth allocation value output by the actor online network is sorted from high to low; Next, the sorted users in the cell and the sorted bandwidth allocation values are matched one by one to achieve the bandwidth resource allocation of each UAV; Finally, in order to improve the user-centric cell fair throughput, a reward function is designed to measure the influence of bandwidth resource allocation on user-centric cell fair throughput, which is expressed as: λ j_3 =CFT.

[0067] Further, the step (7) is implemented as follows:

[0068] Firstly, according to all the reward functions, the reward of each UAV is expressed as:

[0069] RE j =λ j_1 +λ j_2 +λ j_3 +λ j_4 +λ j_5

[0070] All UAV rewards are expressed as:

[0071] RE={RE1,RE2,RE3,RE4,RE5}

[0072] Next, {S, A, RE, D, S'} will be stored as experience in the replay buffer, where D = {D1,...,D j} indicates whether the training of each UAV ends at each time slot, and S' indicates the state set of all UAVs in the next time slot;

[0073] The parameters θ Q of the critic online network and the parameters θ μ of the actor online network are updated using the gradient descent method:

[0074]

[0075] Wherein, the update target is ζ φ =REφ+γ(Q'(S' φ ,A'|θ μ' ),θ Q' ), and Φ and γ represent the size of the experience sampled from the experience pool and the discount factor, respectively; S φ , S' φ , A φ and RE φ represent the data sampled from the experience pool; A' represents the actor target network in the next state S'φ The action output of the lower output, whose parameter is denoted as θ μ' ; A represents the action output of the actor online network in the current state S φ ; Q(·) is the value function output by the critic online network, and Q'(·) is the value function output by the critic target network, whose parameter is denoted as θ Q '; for the parameter update of the target network, a soft update method is adopted, which is specifically given by the following formula:

[0076] θ Q' = μτθ Q + (1-τ)θ Q'

[0077] θ μ' = τθ μ + (1-τ)θ μ'

[0078] Wherein, τ represents the update frequency.

[0079] Further, the step (8) is implemented as follows:

[0080] After training, each UAV will have four converged networks; in the implementation phase, the preprocessed user-centered cell distribution is input into the actor online network of each UAV, and then the actor online network outputs the operation to be performed by the UAV, finally realizing the user-centered multi-UAV communication dynamic deployment considering fairness.

[0081] Advantages: compared with the prior art, the advantages of the present application: first, the user-centered multi-UAV deployment scheme proposed in the present application not only can guide multiple UAVs to adapt to the distribution of mobile users, but also can effectively improve the system fairness throughput while reducing the interference between cells; in addition, compared with the traditional UAV-centered deployment scheme, the method proposed in the present application improves the fairness throughput by at least 36%, thereby ensuring the performance advantage of the user-centered multi-UAV communication network. BRIEF DESCRIPTION OF DRAWINGS

[0082] Figure 1 is the flowchart of the present application;

[0083] Figure 2 is the schematic diagram of the UAV communication coverage system scene proposed in the present application;

[0084] Figure 3 is the result diagram of dynamically deploying multiple UAVs in two consecutive time slots in a user-centered manner;

[0085] Figure 4 is the fairness throughput comparison diagram of the deployment scheme proposed in the present application and the traditional deployment scheme;

[0086] Figure 5 A comparison graph of the cumulative distribution function of fair throughput for a user-centric cell under different access methods;

[0087] Figure 6 A graph comparing the fair throughput of networks under different clustering parameters. Detailed Implementation

[0088] The present invention will now be described in further detail with reference to the accompanying drawings.

[0089] like Figure 1 As shown, this invention proposes a user-centric dynamic deployment method for fair communication among multiple unmanned aerial vehicles (UAVs), specifically including the following steps:

[0090] Step 1: Based on the average path loss model of communication between the UAV and the user, and combined with the Jain Fairness Index (Jain-FI), a user-centric cell fair throughput model is obtained.

[0091] In the nth time slot, the three-dimensional coordinates of the jth UAV and the two-dimensional coordinates of the ith user can be represented as follows: and Within each time slot, mobile users are divided into K user-centric clusters, where the k-th cluster is... n The center coordinates of a user-centric cluster are represented as follows: The number of users in each user-centric cluster is I. k It can be used Where I represents the total number of users, and the users in the k-th user-centric cluster in the n-th time slot are... This represents the total number of users, while the number of users in the k-th user-centric cluster in the n-th time slot is represented by [data / data / etc.]. Finally, for ease of description of the present invention, the time slot set N = {1,...,n,...,N} is defined, and the set of unmanned aerial vehicles (UAVs) is defined. User-centric community set K n ={1,...,k n Let I = {1,...,i,...,I} be the set of all users, and let I be the set of users in different user center clusters for each time slot. The channel model in the considered network is based on the average path loss model, where the probabilities of LOS paths are as follows:

[0092]

[0093] Where α and β are constant values ​​that depend on the environment. This is the horizontal distance between the drone and the user. The NLOS probability can be expressed as: P LOS =1-P NLOS .

[0094] Therefore, the average path loss equation model is expressed as follows:

[0095]

[0096] Among them, f c and v c Let l represent the carrier frequency and the speed of light, respectively. Furthermore, the distance between the drone and the user m is determined by l. i,j Give, η LOS and η NLOS represents the average additional loss value of LOS and NLOS propagation of the connection. Furthermore, the k-th... n In the first user-centric community Channel gain between user j and drone j Represented as:

[0097] Finally, the kth n In the first user-centric cluster The throughput of a single user can be given by the following formula:

[0098]

[0099] in, They represent the nth time slot and the kth time slot, respectively. n The j-th drone in the user-centric cluster matching is assigned to the first... Bandwidth per user; This indicates whether the i-th user has been assigned to the k-th user-centric cell in the n-th time slot; Let σ represent the transmit power of the j-th UAV. 2 Indicates noise power; M ICI Inter-cell interference can be represented as: j'∈J / J represents the j'th UAV in the nth time slot, excluding the one associated with the kth time slot. n The j-th UAV matched to the cell.

[0100] Therefore, in the nth time slot, the kth time slot n In the first user-centric cluster The throughput rate for a single user is defined as: in Indicates the kth time slot before time slot n. n In the first user-centric community Total throughput per user denotes the sum of the throughputs of all users in the kth user-centric cluster before the nth time slot. n denotes the sum of the throughputs of all users in the kth user-centric cluster before the nth time slot.

[0101] The Jain Fairness index (Jain-FI) is then defined to measure the throughput difference among users in the kth cell in the nth time slot: n The Jain Fairness index (Jain-FI) is then defined to measure the throughput difference among users in the kth cell in the nth time slot:

[0102]

[0103] Therefore, based on the Jain-FI, the throughput of the kth user-centric cell in the nth time slot is defined as the fair throughput of the user-centric cell, as follows: n Therefore, based on the Jain-FI, the throughput of the kth user-centric cell in the nth time slot is defined as the fair throughput of the user-centric cell, as follows:

[0104] Step 2: Model the optimization problem of dynamic deployment of multi-UAVs considering fairness.

[0105] The energy consumption of a UAV is mainly composed of two parts: communication-related energy and UAV propulsion-related energy, and the communication-related energy is much smaller than the propulsion-related energy, so it is often ignored. For a quadrotor UAV, the thrust of each rotor is a function of the UAV speed and acceleration , where bold indicates that they are vectors:

[0106]

[0107] where and denote the direction vector and the magnitude of the speed of the UAV, respectively, and g denotes the gravity acceleration vector. Here, n r and m denote the number of rotors and the mass of the UAV, respectively, and p and S FP denote the air density and the equivalent flat area of the body, respectively. Therefore, the propulsion power can be expressed as:

[0108]

[0109] where d, c T , p, c sol , c fac , d0 are some related parameters in the formula, representing the local blade section drag coefficient, the thrust coefficient based on the valve area, the air density, the rotor stability, the incremental correction coefficient of the induced power with respect to the linear induction speed, the disc area of each rotor, and the resistance ratio of the body rotors, respectively, and t cThe climb angle of the aircraft is denoted. Thus the residual energy of the jth UAV at the nth time slot can be denoted as: where E max is the maximum storage energy of the UAV, Δ n is the time slot length.

[0110] Therefore, in order to achieve a user-centric multi-UAV dynamic deployment scheme considering fairness, the optimization problem P1 is modeled as follows:

[0111]

[0112] where C1 denotes whether the jth UAV serves the kth user-centric cell at the nth time slot, and C2 denotes the number of users in each user-centric cell. Under the guarantee of C3, the total bandwidth allocated to users in each user-centric cell by each UAV is equal to the total bandwidth B of UAVBS. In addition, it is guaranteed by C4 that the energy of the UAV is not depleted during the task. In this paper, the motion of the UAV is decomposed into x, y, z three directions, so C5 and C6 reflect the maximum values of acceleration and speed in the three directions. Finally, the collision between multiple UAVs is prevented in C7. Three-dimensional matrix respectively represent the acceleration of the UAV in three directions, the bandwidth allocation of the UAV when serving the matched cluster, the matching of the UAV and the cell, and the matching of the user and the cell.

[0113] Step 3: Preprocess the ground mobile users using a genetic algorithm to obtain suitable user cells, where the relative distance between users in a cell and the relative distance between cells are used to measure the effect of preprocessing.

[0114] Problem P1 contains discrete variables and and continuous variables and is a non-convex optimization problem that is difficult to solve directly. First, the users are preprocessed to form user-centric cells, so the preprocessed problem P2 can be modeled as:

[0115]

[0116] where, denotes the sum of the relative distances between users in a cell, and denotes the sum of the relative distances between cells.

[0117] For the pre-processing problem P2, an innovative genetic algorithm for solving the user-centric grouping problem is proposed. In fact, in the genetic algorithm, a chromosome is an effective solution to the pre-processing problem P2, so a complete chromosome is shown as follows: where the element ch belongs to the set {1,..., i,..., I,..., I+K-1}, and when the value of ch is in [1, I], it represents the number of the ground mobile user, while when the value of ch is in [I+1, I+K-1], it represents the separation point between two cells in the chromosome. In fact, according to the constraint C2 in P1, there will be an element ch with the value in the range of [I+1, I+K-1] after every I k position in the chromosome, while the value of the other elements is in the range of [1, I].

[0118] Therefore, a series of qualified chromosomes will form the initial population pop, and the size of the population can be represented by Ω. For the population pop, the fitness function value fit(ω) of each chromosome can be obtained according to the optimization objective function of P2. Based on the distribution of the fitness function values fit(ω) of all chromosomes in the population, the stochastic universal sampling (SUS) method is introduced to select the chromosomes with higher fitness function values, and the number of selected chromosomes is Ω1, where Ω1<Ω. Next, for the selected chromosomes, the partial match crossover (PMX) strategy can complete the crossover operation of these chromosomes. If the random number is less than the mutation probability, then randomly exchange the elements of any two points in the chromosome. Finally, the chromosomes after the above operation are combined with the Ω-Ω1 chromosomes with higher fitness function values in the initial population to form a new generation of population. Repeat the above steps until the optimal solution is converged, and finally obtain the chromosome with the highest fitness function value, i.e. the user clustering strategy under the current time slot.

[0119] Step 4: Using the multi-agent deep reinforcement learning (MADRL) algorithm, assign a dedicated flight action to each UAV according to the distribution of the ground user cells to reach the corresponding position within the specified time slot.

[0120] In step 3, the pre-processing explicitly determines the user-centric cell distribution in each time slot, so that in this step, the positions of multiple UAVs and the allocation of network resources can be optimized, thereby converting P1 into P3 problem:

[0121]

[0122] For problem P3, the GA-MADRL algorithm is used to solve the challenge of user-centric multi-UAV dynamic deployment, while considering fairness, in which MADRL adopts multi-agent deep deterministic policy gradient (MADDPG). Since MADDPG involves four neural networks, including actor online network, critic online network, actor target network, and critic target network, and updating these networks requires knowing the state, action, and reward of multiple UAVs, the state in the input neural network needs to be understood first. The observation space of each agent is composed of the positions of all UAVs, the velocities in three directions, the positions of ground users, and the cell clustering of ground users, and can be expressed as:

[0123]

[0124] Therefore, the state space of all agents is:

[0125] The action of each UAV includes two parts: the acceleration of the UAV in three directions and the bandwidth allocation of the UAV, which is expressed as: And the action set of all UAVs can be expressed as: A={A1,...,A j}. In each actor online network, the two actions are output simultaneously, but there is a sequence in the execution process, and the action of the acceleration of the UAV in three directions is discussed first in this step.

[0126] Since each UAV is uniformly accelerated in each direction, after the actor online network outputs the action of the acceleration of the UAV in three directions, the coordinates and velocities of each UAV in three directions can be obtained by the following formula:

[0127]

[0128] Based on the above coordinate and velocity update formula, the position and velocity of the UAV reached within a specified time slot can be obtained. For such actions, the reward function first prevents the velocity or position of the UAV from exceeding the limit, which can be expressed as:

[0129]

[0130] In addition, in order to ensure that all UAVs do not collide, the criterion is that any two UAVs at the same time, as long as their three-direction coordinates are not equal, can meet the constraint C6 in P3, and the reward function can be expressed as:

[0131]

[0132] For the flight of UAVs whether to meet its energy constraints, there is a reward function as follows: λ j_3 = E j (n). In the reward function, to adjust the size of the reward value.

[0133] Step 5: The ground user cell selects the UAV to serve according to the distance between each UAV and the corresponding ground cell, thereby completing the one-to-one matching of the UAV and the ground cell.

[0134] Each UAV reaches the corresponding position according to its exclusive action, and at this time the ground mobile user has also completed the construction of the ground user cell according to step four, so each user cell will find the nearest UAV to the center of the cell to complete the matching in this step.

[0135] First, according to the coordinates of the user cell center and the coordinates of each UAV, the distance between different user cells and UAVs can be calculated, and the calculation formula is described as follows:

[0136]

[0137] Then, according to the distance formula between the cell and the UAV, starting from the first user cell, each cell traverses the idle UAVs one by one until the nearest UAV to the center of the cell is found to complete the matching. In order to measure the effect of matching, a reward function is designed: wherein, to adjust the size of the reward value, and to limit the distance between each UAV and the cell it matches.

[0138] Step 6: According to the MADRL algorithm, the bandwidth resource allocation of the corresponding UAV in each user cell is realized in combination with the user-centered cell fair throughput, including:

[0139] According to step 4, the action of each UAV contains two parts: the acceleration of the three directions of the UAV and the bandwidth allocation of the UAV, and the action of the acceleration of the three directions of the UAV is executed in step 4, so the action of the bandwidth allocation of the UAV will be executed in this step.

[0140] Since the UAVs and the user-centric cells are one-to-one matched, and the matching problem between the UAVs and the user-centric cells has been solved in step five, for each UAV, it first sorts the users in the user-centric cell matched with the UAV from low to high according to the sum of the throughputs of the users before time slot n, and sorts the bandwidth allocation values output by the actor online network from high to low. Next, the sorted users in the user-centric cell and the sorted bandwidth allocation values are matched one-to-one, so as to realize the allocation of the bandwidth resources of each UAV. Finally, in order to improve the fair throughput of the user-centric cell, a reward function is designed to measure the influence of the allocation of the bandwidth resources on the fair throughput of the user-centric cell, and the reward function can be expressed as: λ j_3 = CFT.

[0141] Step 7: repeating steps 3 to 6 for a training round in a specified number of time slots, training multiple rounds until multiple converged actor online networks are obtained, wherein each actor online network corresponds to a UAV, including:

[0142] First, according to all the reward functions described above, the reward of each UAV can be expressed as: RE j = λ j_1 + λ j_2 + λ j_3 + λ j_4 + λ j_5 , and the reward of all the UAVs can be expressed as: RE = {RE1, RE2, RE3, RE4, RE5}. Next, {S, A, RE, D, S'} will be stored as experience in the replay buffer, where D = {D1,..., D j} indicates whether the training of each UAV ends at the end of each time slot, and S' indicates the state set of all the UAVs in the next time slot.

[0143] Therefore, the parameters θ Q of the critic online network and the parameters θ μ of the actor online network can be updated using a gradient descent method, and the related update formula is expressed as:

[0144]

[0145] wherein the update target is ζ φ = RE φ + γ (Q'(S' φ , A'| θ μ' ), θ Q' ), and Φ and γ respectively represent the size of the experience sampled from the experience pool and the discount factor. In addition, S φ , S' φ , Aφ and RE φ denotes the data sampled from the experience pool, while A' denotes the action outputted by the actor target network at the next state S' φ , whose parameter is denoted as θ μ' . Similarly, A denotes the action outputted by the actor online network at the current state S φ . In addition, Q(·) is the value function outputted by the critic online network, and Q'(·) is the value function outputted by the critic target network, whose parameter is denoted as θ Q '. Therefore, for the parameter update of the target network, a soft update method is adopted, which is specifically given by the following formula:

[0146] θ Q' = τθ Q + (1 - τ)θ Q'

[0147] θ μ' = τθ μ + (1 - τ)θ μ'

[0148] wherein τ denotes the update frequency.

[0149] Finally, after obtaining the above network-related update formula, it is stipulated that repeating steps 3 to 6 within a set number of time slots is one training round, and an "episode" is defined as the number of training rounds. In each training round, the ground users are preprocessed to form user-centered cells. Then, each drone starts to explore according to the action value outputted by the actor online network, at this time, in order to increase the exploratory nature, the action of each drone can be updated as: A j = A j | θ μ + Λ, wherein Λ denotes an action value obtained through Gaussian noise. Subsequently, the drones update their positions and velocities. Subsequently, each user-centered cell is matched with the nearest drone that is not assigned to other clusters. Then each drone completes the allocation of bandwidth resources within the user cell to which it is matched. After the action is executed, {S, A, RE, D, S'} is stored in the experience pool, wherein the size of the data stored in the experience pool is denoted as Γ, and the minimum training data size of the experience pool is set to Network training occurs every fixed number of times, and the frequency is set to When the training condition is met, that is, and i_episode% i_episode and the current time slot is the initial time slot of the current training episode, where i_episode represents the current training number, and the four networks of each UAV will randomly extract data from the experience pool and train according to the network update formula. If the UAV violates C3 or C5 in P3 or moves out of the boundary, the action execution of the current time slot will end, resulting in a transition to the next training episode. Finally, when the training phase ends, each UAV will obtain four converged networks, including the online network of the actor.

[0150] Step 8: Each UAV selects its flight action and bandwidth allocation action according to its converged actor online network and the distribution of the ground mobile user cell, thereby realizing the user-centric multi-UAV communication dynamic deployment considering fairness.

[0151] After training, each UAV will have four converged networks. In the implementation phase, the preprocessed user-centric cell distribution is input into the actor online network of each UAV, and then the actor online network outputs the operation to be performed by the UAV, finally realizing the user-centric multi-UAV communication dynamic deployment considering fairness.

[0152] Figure 2 is the scenario provided by the present application to realize the user-centric multi-UAV communication dynamic deployment considering fairness under the condition of meeting the constraints of UAV flight trajectory, user scheduling, energy consumption, etc. Figure 3 The distribution of mobile users and the position of UAVs in two consecutive time slots randomly selected within a task period are shown, where the length of each time slot is set to 5 seconds. From Figure 3 It can be seen that as the users move, the user-centric cell will change, and the position of each UAV will also be adjusted accordingly. Therefore, the proposed scheme can guide multiple UAVs to adapt to the distribution of mobile users. Figure 4 To compare the fairness throughput of the proposed deployment scheme with the traditional deployment scheme. Compared with the traditional one-UAV-centric deployment scheme, the proposed scheme is a user-centric, fair and dynamic deployment scheme, which can realize fair on-demand deployment as the ground users move, so in the case of different user numbers, the fairness throughput obtained by the proposed scheme is better than that obtained by the traditional fixed base station deployment scheme. Figure 5A comparison chart of the cumulative distribution function of the fair throughput of the user-centric cell under different access methods, including the GA-MADRL with fairness, the FDMA method, and the GA-MADRL without fairness. The GA-MADRL with fairness provides fair services for users in the user-centric cell based on the Jain-FI and the adaptive bandwidth allocation strategy. In contrast, the FDMA method only allocates bandwidth evenly, and the GA-MADRL without fairness does not consider the Jain-FI when serving users. Therefore, the access methods used in this paper can achieve higher fair throughput of the user-centric cell. Figure 6 A comparison chart of the fair throughput of the envisioned network under different clustering parameters, wherein The value of the relative distance between users in the user-centric cell indicates whether the relative distance between users in the user-centric cell is considered, and The value of the relative distance between user-centric cells indicates whether the relative distance between user-centric cells is considered. As the number of mobile users increases, the interference between different user-centric cells also increases, which reduces the fair throughput of the system. Therefore, compared with considering only a single factor, the appropriate and can effectively reduce the interference between clusters, thereby improving the fair throughput of the system.

[0153] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present application is defined by the appended claims rather than the above description, and all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be regarded as limiting the claims involved.

Claims

1. A user-centric, multi-UAV fair communication dynamic deployment method, characterized in that, Includes the following steps: (1) Based on the average path loss model of communication between UAV and user, combined with the Jain fairness index, the cell fair throughput centered on the user is obtained. (2) Model the dynamic deployment optimization problem of multiple UAVs centered on fairness; (3) Use a genetic algorithm to preprocess the relative distance between users within a cell and the relative distance between cells to clarify the cell distribution centered on the user in each time slot; (4) Using the multi-agent reinforcement learning algorithm MADRL, each UAV is assigned a dedicated flight maneuver based on the distribution of ground user cells so that it can reach the corresponding position within the specified time slot. (5) The ground user cell selects the drone that serves it based on the corresponding distance of each drone, thus completing the one-to-one matching between the drone and the ground cell; (6) Based on the MADRL algorithm and combined with the user-centric cell fair throughput, bandwidth resource allocation for the corresponding UAV is implemented in each user cell; (7) Repeat steps (3) to (6) within the specified number of time slots until multiple converged actor online networks are obtained, where each actor online network corresponds to a drone; (8) Each UAV selects its flight actions and bandwidth allocation actions based on the distribution of its converged actor online network and ground mobile user cells, thereby achieving a user-centric dynamic deployment of multi-UAV communication that takes fairness into account. The implementation process of step (2) is as follows: For a quadcopter drone, the thrust of each rotor is the UAV speed. and acceleration The functions, where bold represents their vectors: in, and These represent the velocity direction vector and velocity magnitude of the UAV, respectively; g represents the gravitational acceleration vector; n r ρ and m represent the number of rotors and the mass of the UAV, respectively. FP These represent air density and equivalent flat surface area of ​​the fuselage, respectively; propulsion power is expressed as: Among them, δ, c T , ρ, c sol c fac , d0 represents the local blade section drag coefficient, thrust coefficient based on valve disc area, air density, rotor stability, incremental correction coefficient of induction power relative to linear induction speed, disk area of ​​each rotor, and drag ratio of each rotor in the fuselage, respectively, while τ c The climb angle of the aircraft is represented by: The remaining energy of the j-th UAV in the nth time slot is represented as: E max Δ is the maximum stored energy of a UAV. n The time slot length; The optimization problem P1 is modeled as follows: Where C1 indicates whether the j-th UAV serves the k-th user-centric cell in the n-th time slot, and C2 indicates the number of users in each user-centric cell; under the guarantee of C3, the total bandwidth allocated to users by each UAV in each user-centric cell is equal to the total bandwidth B of the UAVBS; C3 is guaranteed by C4 that the UAV's energy is not exhausted during the mission; the motion of the UAV is decomposed into three directions: x, y, and z, and C5 and C6 reflect the maximum values ​​of acceleration and velocity in the three directions; finally, C7 is used to prevent collisions between multiple UAVs; a three-dimensional matrix. These represent the acceleration of the UAV in three directions, the bandwidth allocation of the UAV when serving a matched cluster, the matching status of the UAV with the cell, and the matching status of the user with the cell, respectively.

2. The user-centric multi-UAV fair communication dynamic deployment method according to claim 1, characterized in that, The implementation process of step (1) is as follows: In the nth time slot, the three-dimensional coordinates of the jth UAV and the two-dimensional coordinates of the i-th user are respectively represented as: and Within each time slot, mobile users are divided into K user-centric clusters, where the k-th cluster is... n The center coordinates of the user-centric cluster are Number of users in each user-centric cluster Where I represents the total number of users, and the number of users in the k-th user-centric cluster in the n-th time slot is... Representation; Define the time slot set N = {1,...,n,...,N}, and the set of unmanned aerial vehicles (UAVs). User-centric community set K n ={1,...,k n Let I = {1,...,i,...,I} be the set of all users, and let I be the set of users in different user center clusters for each time slot. The channel model in the considered network is based on the average path loss model, where the probabilities of LOS paths are as follows: Where α and β are constant values ​​dependent on the environment, and represents the horizontal distance NLOS probability between the drone and the user, expressed as: P LOS =1-P NLOS ; Therefore, the average path loss model is expressed as follows: Among them, f c and v c These represent the carrier frequency and the speed of light, respectively; furthermore, the distance between the drone and the user m is determined by l. i,j Give, η LOS and η NLOS The value of k represents the average additional loss value of LOS and NLOS propagation of the connection; n In the first user-centric community Channel gain between user j and drone j Represented as: Finally, the kth n In the first user-centric cluster The throughput per user is: in, They represent the nth time slot and the kth time slot, respectively. n The j-th drone in the user-centric cluster matching is assigned to the first... Bandwidth per user; This indicates whether the i-th user has been assigned to the k-th user-centric cell in the n-th time slot; Let σ represent the transmit power of the j-th UAV. 2 Indicates noise power; M ICI This indicates inter-cell interference. j'∈J / J represents the j'th UAV in the nth time slot, excluding the one associated with the kth time slot. n The j-th UAV matched in the cell; Therefore, in the nth time slot, the kth time slot n In the first user-centric cluster The throughput rate for a single user is defined as: in, Indicates the kth time slot before time slot n. n In the first user-centric community Total throughput per user Indicates the kth time slot before time slot n. n The sum of throughput for all users in a user-centric cluster; Define the Jain fairness index to measure the k-th time slot within the n-th time slot. n Throughput differences among users in each cell: Therefore, based on Jain-FI, the k-th time slot of the nth time slot is... n The throughput of a user-centric cell is defined as the fair throughput of the user-centric cell:

3. The user-centric multi-UAV fair communication dynamic deployment method according to claim 1, characterized in that, The implementation process of step (3) is as follows: Preprocessing users to form user-centric cells, the preprocessing problem P2 is modeled as: Wherein, represents the sum of the relative distances between users within the community, and represents the sum of the relative distances between communities; For the preprocessing problem P2, in the genetic algorithm, a chromosome represents a valid solution to the preprocessing problem P2. Therefore, a complete chromosome is shown below: The element ch belongs to the set {1,...,i,...,I,...,I+K-1}. When the value of ch is in [1,I], it represents the number of the ground mobile user, and when the value of ch is in [I+1,I+K-1], it represents the dividing point between cells in the chromosome. According to constraint C2 in P1, every I in the chromosome k After a certain position, the value of element ch will be in the range [I+1, I+K-1], while the value of elements in other positions will be in the range [1, I]. Chromosomes meeting the requirements will form the initial population pop, whose size is represented by Ω. For the population pop, the fitness function value fit(ω) of each chromosome is obtained according to the optimization objective function of P2. Based on the distribution of the fitness function values ​​fit(ω) of all chromosomes in the population, the random traversal sampling (SUS) method is introduced to sort all chromosomes from high to low according to their fitness function values ​​and select chromosomes of number Ω1, where Ω1 < Ω. Next, for the selected chromosomes, the partial matching crossover strategy is used to perform the crossover operation on these chromosomes. If the random number is less than the mutation probability, then the elements of any two points on the chromosome are randomly swapped. Finally, all chromosomes in the initial population are sorted from high to low according to their fitness function values ​​and chromosomes of number Ω-Ω1 are selected and combined with the chromosomes that have undergone the above operation to form a new generation of population. The above steps are repeated until the optimal solution is converged, and finally the chromosome with the highest fitness function value is obtained, which is the user clustering strategy in the current time slot.

4. The user-centric multi-UAV fair communication dynamic deployment method according to claim 1, characterized in that, The implementation process of step (4) is as follows: Preprocessing clarifies the user-centric cell distribution for each time slot, and then optimizes the allocation of multiple UAV locations and network resources, thereby transforming P1 into a P3 problem: For problem P3, the MADRL algorithm is used to address the challenge of user-centric dynamic deployment of multiple UAVs, while also considering fairness. MADRL employs a multi-agent deep deterministic strategy, MADDPG. MADDPG involves four neural networks: an online actor network, an online critic network, an actor-target network, and a critic-target network. Updating these networks requires knowledge of the states, actions, and rewards of multiple UAVs. Each agent's observation space consists of the positions of all UAVs, their velocities in three directions, the positions of ground users, and the clustering of ground user segments, represented as follows: Therefore, the state space of all agents is: The motion of each drone consists of two parts: the drone's acceleration in three directions and the drone's bandwidth allocation. The set of motions for each drone is represented as: The set of actions for all drones is: A = {A1,...,A} j }; Since each drone undergoes uniform acceleration in every direction, after the actor online network outputs the corresponding drone's acceleration in the three directions, the coordinates and velocity updates for each drone in the three directions are obtained by the following formulas: Based on the above coordinate and velocity update formulas, the position and velocity of the drone within the specified time slot are obtained. For this type of action, the reward function first prevents the drone's speed or position from exceeding the limit, expressed as: To ensure that no drones collide, the criterion is that any two drones, at the same time, satisfy constraint C6 in P3 as long as their coordinates in the three directions are not simultaneously equal. The reward function is expressed as: Whether a drone's flight can meet its energy constraints is reflected by the following reward function: λ j_3 =E j (n); In the above reward function Used to adjust the size of the reward value.

5. The user-centric multi-UAV fair communication dynamic deployment method according to claim 1, characterized in that, The implementation process of step (5) is as follows: Based on the coordinates of the user cell center and the coordinates of each drone, the distance between different user cells and drones is calculated. The calculation formula is described as follows: Based on the distance formula between cells and drones, starting from the first user cell, each cell iterates through the idle drones one by one until it finds the drone closest to its center to complete the match. Design a reward function to measure matching effectiveness: in, Used to adjust the size of the reward value. Used to limit the distance between each UAV and its matched cell.

6. The user-centric multi-UAV fair communication dynamic deployment method according to claim 1, characterized in that, The implementation process of step (6) is as follows: First, users within the user cell matched with each drone are sorted from low to high according to the sum of their throughput before time slot n, and the bandwidth allocation values ​​of their actor online network output are sorted from high to low. Next, the bandwidth allocation values ​​for the sorted cell users are matched one-to-one to achieve bandwidth resource allocation for each drone. Finally, to improve user-centric cell fair throughput, a reward function is designed to measure the impact of bandwidth resource allocation on user-centric cell fair throughput. This reward function is expressed as: λ j_3 =CFT.

7. The user-centric multi-UAV fair communication dynamic deployment method according to claim 1, characterized in that, The implementation process of step (7) is as follows: First, based on all the reward functions, the reward for each drone is expressed as follows: RE j =λ j_1 +λ j_2 +λ j_3 +λ j_4 +λ j_5 All drone rewards are represented as follows: RE = {RE1, RE2, RE3, RE4, RE5} Next, {S,A,RE,D,S'} will be stored as experience in the replay buffer, where D = {D1,...,D}. j } indicates whether the training of each drone ends in each time slot, and S' represents the state set of all drones in the next time slot; The parameter θ of the critic online network Q The parameter θ of the actor online network μ Update using a gradient descent-based method: The update target is ζ φ =RE φ +γ(Q'(S' φ ,A'|θ μ' ),θ Q' ), where Φ and γ represent the empirical size and discount factor sampled from the empirical pool, respectively; S φ 、S'φ、A φ and RE φ A' represents the data sampled from the experience pool; A' represents the next state S' of the actor target network. φ The output action is denoted by θ. μ' A represents the current state S of the actor online network. φ The action output is defined as follows: Q(·) is the value function output by the online critic network, and Q'(·) is the value function output by the target critic network, with parameters denoted as θ. Q' For parameter updates of the target network, a soft update method is used, specifically given by the following formula: i Q' =tθ Q +(1-τ)θ Q' i μ' =tθ μ +(1-τ)θ μ' Where τ represents the update frequency.

8. A user-centric, multi-UAV fair communication dynamic deployment method according to claim 1, characterized in that, The implementation process of step (8) is as follows: After training, each drone will have four converged networks; in the implementation phase, the preprocessed user-centric cell distribution is input into the actor online network of each drone, and then the actor online network outputs the operation to be performed by the drone, ultimately realizing a user-centric dynamic deployment of multi-drone communication that takes fairness into account.

Citation Information

Patent Citations

  • Cooperative transmission method for unmanned aerial vehicle network dynamic cache deployment

    CN111970046A

  • Multi-unmanned aerial vehicle dynamic deployment method

    CN114567888A