Unmanned aerial vehicle group intelligent resource scheduling and interference coordination optimization method and device
By building a system model and channel model of the UAV cluster service user system, using a near-end strategy optimization algorithm to solve the problem of time-frequency resource allocation optimization, the optimization of resource allocation and interference coordination in the UAV network is achieved, the problems of resource allocation complexity and performance degradation in the existing technology are solved, and the anti-interference ability and dynamic adaptability of the system are improved.
Patent Information
- Application Number
- CN202510102760.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art is difficult to effectively manage resource allocation and interference coordination in drone networks, especially in the case of dynamic changes in user locations and resource constraints, resulting in system performance degradation.
By building a system model of the UAV cluster service user system, the channel model and rate model are determined, the time-frequency resource allocation optimization problem is constructed, and the near-end strategy optimization algorithm is used to solve it, the UAV cluster resource allocation scheme is obtained to maximize the total user transmission rate.
Reduces the complexity of resource allocation, improves the stability of convergence, enhances the anti-interference ability of the communication system, can adapt to the dynamic changes in the deployment environment of the drone base station, and maintains high performance in complex network environments.
Smart Images

Figure CN120018293A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technology, and specifically relates to a method and device for optimizing intelligent resource scheduling and interference coordination of a swarm of unmanned aerial vehicles. Background Art
[0002] Drones are widely used in many fields such as post-disaster recovery, surveillance, and providing communication services in rural and underdeveloped areas. With the development of mobile communication technology and the emergence of various new services, the explosive growth of mobile data traffic has become increasingly obvious. However, due to factors such as complex terrain, the coverage capacity of existing ground base stations is still insufficient to meet the growing communication needs.
[0003] Therefore, the development of integrated space-air-ground networks has become a key trend in future communication systems, and drone networks are an important component of space networks. Drones equipped with wireless communication modules can be quickly deployed as airborne base stations to provide universal access to ground user equipment (UE) during temporary events (such as hot spots or large public gatherings), becoming a promising solution. The deployment of drones is seen as an effective complement to existing cellular systems, which can improve wireless capacity and extend coverage in scenarios with ultra-high density traffic demand. By leveraging their mobility and adaptability, drone-based networks provide a transformative opportunity to address the limitations of traditional ground communication infrastructure. However, integrating drones into wireless communication systems faces significant challenges, especially in terms of spectrum management. As airborne base stations or relay nodes, drones must efficiently share limited wireless resources with ground users. These challenges are further exacerbated by the high mobility of drones and the dynamic changes in user locations, resulting in time-varying channel conditions and interference patterns.
[0004] Resource allocation in UAV communications faces several key challenges. First, the dynamic changes in user locations introduce strong spatiotemporal correlations in the communication environment. Second, UAVs have limited resources, such as battery power and spectrum, which require efficient utilization strategies to balance communication quality and resource efficiency. Finally, resource competition and interference among multiple users require effective scheduling mechanisms to ensure fairness and performance. The complexity of these issues makes it difficult for traditional static resource allocation methods to cope with the dynamic characteristics of UAV networks. Therefore, developing innovative spectrum allocation strategies for UAV networks is crucial to fully realize their potential and address the unique challenges they face.
[0005] With the widespread application of aerial drone base stations (UAV-BS) in emergency communications, post-disaster recovery, and temporary large-scale activities, interference coordination and resource allocation technology has become one of the core issues in serving ground users. However, existing interference coordination and resource allocation schemes still have significant deficiencies in practical applications. First, traditional rule-based resource allocation methods rely on pre-set fixed strategies, which are difficult to cope with the dynamic changes of user needs, channel conditions, and interference environments. This static resource allocation method often cannot provide sufficient flexibility and efficiency in complex network scenarios, resulting in a decrease in system performance. Secondly, although methods based on optimization theory (such as convex optimization and non-convex optimization) can obtain theoretical optimal solutions to a certain extent, they rely on accurate modeling of the problem and complex calculation processes. Especially when UAV base stations serve high-density users or high-speed mobile scenarios, the optimization problem may become extremely complex, resulting in a long solution time or convergence to a suboptimal solution. In addition, traditional optimization methods are usually difficult to handle nonlinear, high-dimensional, and highly dynamic environments, and lack real-time and scalability.
[0006] Therefore, how to provide a drone swarm intelligent resource scheduling and interference coordination optimization method that can adapt to the dynamic changes of drone base station deployment environment, maintain high performance in a complex and highly uncertain network environment, and meet the real-time requirements of resource management has become an urgent problem to be solved. Summary of the invention
[0007] In order to solve the above problems existing in the prior art, the present invention provides a method and device for intelligent resource scheduling and interference coordination optimization of a drone swarm.
[0008] The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0009] In a first aspect, the present invention provides a method for intelligent resource scheduling and interference coordination optimization of a drone swarm, comprising:
[0010] Construct a system model of a drone swarm service user system; the system model includes N drones and U users; each drone serves U users. N Users;
[0011] Determine a channel model and a rate model according to the system model; the channel model is used to represent the free path loss between the drone and the user it serves; the rate model is used to represent the transmission rate obtained by the user after the channel is allocated;
[0012] A time-frequency resource allocation optimization problem is constructed based on the channel model and the rate model, and the proximal strategy optimization algorithm is used to solve the time-frequency resource allocation optimization problem to obtain a resource allocation plan for the drone swarm. In the time-frequency resource allocation optimization problem, the maximum transmission power and limited bandwidth of each drone are used as constraints, and maximizing the total user transmission rate is used as the optimization goal. The maximum transmission power indicates that the total power transmitted by each drone to the users it serves cannot exceed its own total transmission power. The limited bandwidth indicates that the total bandwidth transmitted by each drone to the users it serves cannot exceed its own total bandwidth.
[0013] Optionally, a proximal strategy optimization algorithm is used to solve the time-frequency resource allocation optimization problem and obtain a drone swarm resource allocation solution, including:
[0014] Modeling the time-frequency resource allocation optimization problem as an MDP;
[0015] By using the proximal strategy optimization algorithm, the maximum transmission power and limited bandwidth of each UAV are constrained, and the maximum total user transmission rate is taken as the optimization goal to update the MDP, and the resource allocation scheme of the UAV swarm is obtained;
[0016] Among them, the state space, action space and reward in MDP are defined as follows:
[0017] Each action of the drone in the action space is mapped to a column in a U×M binary matrix; the element in the mth column of the uth row in the binary matrix is used to describe the action of the uth user being assigned to the mth channel;
[0018] The state space is defined as a one-dimensional array for reflecting the current channel allocation state; each element in the one-dimensional array is used to describe the state of a user being allocated to a channel;
[0019] The reward is defined as the sum of the total transmission rates of all users under the current channel allocation state.
[0020] Optionally, the channel model is:
[0021]
[0022] in, Represents the nth drone and the uth drone it serves n The free path loss between users; f c Indicates the carrier frequency; Represents the nth drone and the uth drone it serves n The distance between users; c = 3 × 10 8 m / s.
[0023] Optionally, the rate model is:
[0024]
[0025] Among them, R total represents the total transmission rate of the user; represents time slot t, the uth n The transmission rate obtained after each user is assigned a channel; B represents the total bandwidth of the UAV; Indicates time slot t, the mth channel is assigned to the uth n Allocation plan for each user; Indicates the uth node served by the nth drone n The signal-to-noise ratio of a user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n ; m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of drones; U n represents the total number of users served by the nth UAV; M represents the total number of channels.
[0026] Optionally, the time-frequency resource allocation optimization problem is:
[0027]
[0028] in, Indicates the nth drone’s service to the uth drone n The transmission power of each user; Represents the nth drone and the uth drone it serves n Channel parameters between users; represents the internal interference of the nth UAV at time slot t; represents the time slot t, the u-th node served by the n-th drone n The interference of other drones to each user; 2 (t) represents Gaussian white noise; SINR threshold represents the signal-to-noise ratio threshold; Indicates the uth n The demand rate of each user; P n represents the total transmission power of the nth UAV; B represents the total bandwidth of the UAV; Indicates time slot t, the mth channel is assigned to the uth n Allocation plan for each user; Indicates the uth node served by the nth drone n The signal-to-noise ratio of a user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,Un ; m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of drones; U n represents the total number of users served by the nth UAV; M represents the total number of channels; represents a collection of drones; represents the set of users served by the nth drone; Represents a collection of channels.
[0029] In a second aspect, the present invention provides a drone swarm intelligent resource scheduling and interference coordination optimization device, comprising:
[0030] A construction module is used to construct a system model of a drone swarm service user system; the system model includes N drones and U users; each drone serves U users. N Users;
[0031] A determination module, used to determine a channel model and a rate model according to the system model; the channel model is used to represent the free path loss between the drone and the user it serves; the rate model is used to represent the transmission rate obtained by the user after the channel is allocated;
[0032] A solution module is used to construct a time-frequency resource allocation optimization problem based on the channel model and the rate model, and use a proximal strategy optimization algorithm to solve the time-frequency resource allocation optimization problem to obtain a drone swarm resource allocation solution; in the time-frequency resource allocation optimization problem, the maximum transmission power and limited bandwidth of each drone are used as constraints, and maximizing the total user transmission rate is used as the optimization goal; the maximum transmission power means that the total power transmitted by each drone to the users it serves cannot exceed its own total transmission power; the limited bandwidth means that the total bandwidth transmitted by each drone to the users it serves cannot exceed its own total bandwidth.
[0033] The present invention provides an intelligent resource scheduling and interference coordination optimization method for a swarm of unmanned aerial vehicles. Taking into account the problem that resources are limited under interference conditions and that multiple ground users share limited wireless resources, a time-frequency resource allocation optimization problem is constructed to maximize the total user transmission rate under the constraints of each unmanned aerial vehicle's maximum transmission power and limited bandwidth. A proximal strategy optimization algorithm is used to solve the time-frequency resource allocation optimization problem. Compared with a dynamic resource allocation algorithm, the complexity of resource allocation is reduced. While ensuring the efficiency of strategy improvement, the stability of convergence is improved, and the dimensional disaster problem that may be faced by traditional optimization algorithms in a high-dimensional environment is effectively avoided. Compared with a static allocation algorithm, the anti-interference ability of the communication system is enhanced, and it can adapt to the dynamic changes in the deployment environment of unmanned aerial vehicle base stations, and can maintain high performance in a complex network environment with high uncertainty. It is closer to the actual situation and has strong applicability.
[0034] The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a flow chart of a method for intelligent resource scheduling and interference coordination optimization of a drone swarm provided by an embodiment of the present invention;
[0036] Figure 2 This is a schematic diagram of the scenario of the drone swarm service user system;
[0037] Figure 3 is a schematic diagram of the interference scenario;
[0038] Figure 4 It is a process diagram of the PPO algorithm;
[0039] Figure 5 This is a schematic diagram of the original reward value during PPO training;
[0040] Figure 6 It is a schematic diagram of the deviation of the PPO reward value relative to the mean;
[0041] Figure 7 This is a schematic diagram of the comparison between the original PPO reward value and the 50-round moving average;
[0042] Figure 8 It is a schematic diagram of Coefficient of Variation;
[0043] Fig. 9 It is a schematic diagram of the PPO training dynamics and the comparison results with the greedy algorithm and the polling strategy;
[0044] Fig.10 It is a structural schematic diagram of a drone swarm intelligent resource scheduling and interference coordination optimization device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0045] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0046] In order to solve the problems of low flexibility, poor efficiency, complex solution, lack of real-time and scalability in the existing interference coordination and resource allocation methods, the embodiment of the present invention provides a method for intelligent resource scheduling and interference coordination optimization of drone swarms, see Figure 1 , Figure 1 : is a flowchart of a method for intelligent resource scheduling and interference coordination optimization of a drone swarm provided by an embodiment of the present invention, which specifically includes the following steps:
[0047] Step S101, construct a system model of a drone swarm service user system; the system model includes N drones and U users; each drone serves U users. N users.
[0048] See also Figure 2 , Figure 2 This is a schematic diagram of the scenario of a drone swarm serving a user system. The drone stays above the user and transmits high-definition video to each user on the ground. Based on the user's channel conditions and current resource allocation status, it selects time and frequency resources for multiple users to maximize the total throughput of the multi-drone network.
[0049] In this scenario, the time resource is divided into T time slots, denoted as the set There are N drones, recorded as a set Each drone is equipped with an omnidirectional antenna array. There are U users, denoted as the set Each user is equipped with an omnidirectional antenna array, and each drone serves a user called U N , recorded as a set And the coverage radius of each drone is r n ;
[0050] The total frequency band of all drones is B = f up -f down Divided into M sub-channels, denoted as The total power of each drone is P n , drone n to u n The transmission power of each user Assuming that the user positions in the scene change slowly over time, the position of the u-th user is recorded as L u =(x u (t),y u (t),h u (t)), and the user’s height remains unchanged, i.e. h u (t) = h u (0), h u (t) represents the height of the u-th user at time t, h u (0) represents the initial height of the u-th user, x u (t) represents the x-axis coordinate of the u-th user in time slot t, y u (t) represents the y-axis coordinate of the u-th user in time slot t; the position of the n-th UAV in the t-th time slot is recorded as L n (t) = (x n (t),y n (t),h n (t)), x n (t) represents the x-axis coordinate of the nth UAV at time slot t; y n(t) represents the y-axis coordinate of the nth UAV at time slot t; h n (t) represents the height of the nth UAV at time slot t.
[0051] In the embodiment of the present invention, users can be further divided into ground users and interfered users. When a user is in the coverage of two drones at the same time, he will be interfered. Figure 2 In the figure, there is an interfered user where the coverage of drone A and drone C overlaps, and there is also an interfered user where the coverage of drone A and drone B overlaps.
[0052] Step S102, determining a channel model and a rate model according to the system model; the channel model is used to represent the free path loss between the drone and the user it serves; the rate model is used to represent the transmission rate obtained by the user after being assigned a channel.
[0053] In the embodiment of the present invention, assuming that the channel condition is good, in the tth time slot, the nth UAV and the uth UAV n The channel model between users adopts the free path loss model to construct the free path loss of a single transmission.
[0054]
[0055] Among them, f c is the carrier frequency, c = 3 × 10 8 m / s, The calculation formula is as follows:
[0056]
[0057] in, For the nth drone and the uth drone n The distance between users.
[0058] In the embodiment of the present invention, the channel allocation scheme in time slot t is used express, Indicates the channel indicator. Indicates time slot t, the mth channel is assigned to the uth n In time slot t, if the mth subchannel is allocated to the uth user n users, then otherwise,
[0059] In this embodiment of the present invention, the rate model is as follows:
[0060]
[0061] Among them, Rtotal Indicates the total user transmission rate; represents time slot t, the uth n The transmission rate obtained after each user is assigned a channel; B represents the total bandwidth of the UAV; Indicates time slot t, the mth channel is assigned to the uth n Allocation plan for each user; Indicates the uth node served by the nth drone n The signal-to-noise ratio of a user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n ; m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of drones; U n represents the total number of users served by the nth UAV; M represents the total number of channels. The uth drone serving the nth drone n The signal-to-noise ratio of a user in time slot t is expressed as follows:
[0062]
[0063] Among them, δ 2 (t) is additive white Gaussian noise, represents the internal interference of the nth UAV at time slot t, that is, the interference of other users served by the nth UAV to the uth UAV. n Interference of users, represents the time slot t, the u-th node served by the n-th drone n The following will discuss the interference of other drones on each user. and The calculation process is explained below.
[0064] See also Figure 3 , Figure 3 This is a schematic diagram of the interference scenario. User 1 is served by drone A, but the distance between user 1 and drone B is d. UAV_B ,Right now Smaller than the coverage radius r of drone B UAV_B ,Right now That is, drone B will interfere with user 1, so a variable To represent the interference of other drones j to the current user:
[0065]
[0066] in, Indicates the uth n The x-axis coordinate of the user; j(t) represents the x-axis coordinate of the jth UAV; Indicates the uth n The x-axis coordinate of the user; j (t) represents the y-axis coordinate of the jth UAV; Indicates the uth n The height of a user in time slot t; h j (t) represents the altitude of the jth UAV at time slot t.
[0067] From this, the internal interference of the drone and the interference of the drone to the user can be calculated:
[0068]
[0069] Among them, k m,i (t) indicates whether the i-th user occupies the m-th channel, 1 indicates occupied, and 0 indicates unoccupied; P n,i represents the transmission power of the nth UAV to the i-th user; Indicates the relationship between the nth drone and the uth drone n Channel gain between users; Indicates the uth j Whether the user occupies the mth channel; Indicates the uth n Whether the user receives interference from the jth UAV; P j,i represents the transmission power of the j-th UAV to the i-th user.
[0070] Step S103, construct a time-frequency resource allocation optimization problem based on the channel model and the rate model, and use the proximal strategy optimization algorithm to solve the time-frequency resource allocation optimization problem to obtain the drone swarm resource allocation plan; in the time-frequency resource allocation optimization problem, the maximum transmission power and limited bandwidth of each drone are used as constraints, and maximizing the total user transmission rate is the optimization goal; the maximum transmission power means that the total power transmitted by each drone to the users it serves cannot exceed its own total transmission power; the limited bandwidth means that the total bandwidth transmitted by each drone to the users it serves cannot exceed its own total bandwidth.
[0071] In an embodiment of the present invention, a time-frequency resource allocation optimization problem based on a channel model and a rate model is solved, and the goal of the optimization problem is to maximize the total transmission rate under the constraints of the maximum transmission power and limited bandwidth of each UAV.
[0072] Therefore, the resource allocation problem of each time slot is formulated as a time-frequency resource allocation optimization problem:
[0073]
[0074] in, Indicates the nth drone’s service to the uth drone n The transmission power of each user; Represents the nth drone and the uth drone it serves n Channel parameters between users; represents the internal interference of the nth UAV at time slot t; represents the time slot t, the u-th node served by the n-th drone n The interference of other drones to each user; 2 (t) represents Gaussian white noise; SINR threshold represents the signal-to-noise ratio threshold; Indicates the uth n The demand rate of each user; P n represents the total transmission power of the nth UAV; B represents the total bandwidth of the UAV; Indicates time slot t, the mth channel is assigned to the uth n Allocation plan for each user; Indicates the uth node served by the nth drone n The signal-to-noise ratio of a user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n ; m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of drones; U n represents the total number of users served by the nth UAV; M represents the total number of channels; represents a collection of drones; represents the set of users served by the nth drone; Represents a collection of channels.
[0075] In the above constraints, C1 is used to ensure that the signal-to-noise ratio received by the user is greater than the specified threshold; C2 means that the rate allocated to the user in each time slot must be greater than the user's required rate; C3 and C4 mean that the number of subchannels and bandwidth allocated to a single user are lower than the total number of channels and total bandwidth; C5 means that a channel of a certain drone can only be allocated to one user; C6 means that the power allocated to the user is greater than 0 and less than the maximum power; C7 means that the sum of the powers allocated by the drone to the user is less than the maximum power of the drone.
[0076] In the embodiment of the present invention, the proximal policy optimization (PPO) algorithm is an emerging policy gradient (PG) algorithm, which uses proximal policy optimization to make the optimization process more stable and efficient. The PPO algorithm avoids drastic performance fluctuations during the optimization process by limiting the step size of the policy update.
[0077] In an embodiment of the present invention, taking into account the problem of limited resources and multiple users on the ground sharing limited wireless resources under interference conditions, a time-frequency resource allocation optimization problem is constructed to maximize the total user transmission rate under the constraints of the maximum transmission power and limited bandwidth of each UAV, and a proximal strategy optimization algorithm is adopted to solve the time-frequency resource allocation optimization problem. Compared with the dynamic resource allocation algorithm, the complexity of resource allocation is reduced, and while ensuring the efficiency of strategy improvement, the stability of convergence is improved, and the dimensionality disaster problem that may be faced by traditional optimization algorithms in high-dimensional environments is effectively avoided. Compared with the static allocation algorithm, the anti-interference ability of the communication system is enhanced, which can adapt to the dynamic changes of the deployment environment of the UAV base station, and can maintain high performance in a complex and highly uncertain network environment, which is closer to the actual situation and has strong applicability.
[0078] In one implementation, a proximal strategy optimization algorithm is used to solve the time-frequency resource allocation optimization problem and obtain a resource allocation solution for a drone swarm, including:
[0079] Model the time-frequency resource allocation optimization problem as an MDP (Markov Decision Process);
[0080] Using the proximal strategy optimization algorithm, the maximum transmission power and limited bandwidth of each UAV are constrained, and the maximum total user transmission rate is used as the optimization goal to update the MDP and obtain the UAV swarm resource allocation plan;
[0081] Among them, the state space, action space and reward in MDP are defined as follows:
[0082] Each action of the drone in the action space is mapped to a column in a U×M binary matrix; the element in the mth column of the uth row in the binary matrix is used to describe the action of the uth user being assigned to the mth channel;
[0083] The state space is defined as a one-dimensional array used to reflect the current channel allocation state; each element in the one-dimensional array is used to describe the state of a user being allocated to a channel;
[0084] The reward function is defined as the sum of the total transmission rates of all users under the current channel allocation state.
[0085] In the embodiment of the present invention, the drone is used as an intelligent agent. MDP describes the interaction process between the intelligent agent and the environment, which consists of a 4-tuple definition. represents the transition probability. Secondly, the state space under the optimization problem is given Action Space And the reward function Detailed definition of .
[0086] (1) Action space design
[0087] In this environment, the design of the action space follows the form of discrete action space, taking into account the allocation of users and channels. The specific description is as follows:
[0088] 1) Action definition:
[0089] The action is represented as an integer, and its corresponding binary code represents the user's channel allocation matrix. For example, the size of the action space is 2 U×M , where U is the total number of users and M is the total number of channels.
[0090] 2) Action space range:
[0091] The action space is defined by the software built-in function, and its size is 2 U×M This means that each action can be uniquely mapped to a U×M binary matrix, which represents the user's channel allocation status. In this matrix, if the element in the mth column of the uth row is 1, it means that the uth user is assigned to the mth channel.
[0092] 3) Action Mapping:
[0093] In the environment, the integer action is mapped to a binary matrix through a function to represent the specific channel allocation status. For example, the integer action is padded to a string of length after binary encoding and finally converted into a matrix.
[0094] (2) State space design
[0095] The design of the state space reflects the current channel allocation status of the system. The specific design is as follows:
[0096] 1) Status definition:
[0097] The state represents the current configuration of channel allocation. Specifically, the state is a one-dimensional array of length U×M, each element of which is 0 or 1, indicating whether the user is assigned to a channel. The state space is defined by a built-in function and its range is [0,1] U×M The upper and lower limits here represent the binary nature of the allocation state (0 or 1). For example, if the state is [1, 0, 0, 1, ...], it means that the first user and the fourth user are allocated certain channels respectively.
[0098] 2) State initialization:
[0099] In the reset function, the state is initialized by randomly assigning channels. The initial state is a U×M matrix, which is then flattened into a one-dimensional array. The reset function is a function used to reset the state of an object to the default value.
[0100] (3) Reward function design
[0101] The design goal of the reward function is to maximize the total data transmission rate of the system, that is, to encourage efficient channel allocation schemes. The specific design is as follows:
[0102] 1) Reward definition:
[0103] The reward is expressed as the sum of the total transmission rates of all users under the current channel allocation state. The transmission rate is calculated by Shannon's theorem:
[0104]
[0105] 2) Signal-to-noise ratio calculation:
[0106] The signal-to-noise ratio (SINR) of a channel is defined as:
[0107]
[0108] 3) Reward function definition:
[0109] The total reward is the sum of the rates of all users on all channels in all time slots:
[0110]
[0111] In the embodiment of the present invention, the proximal strategy optimization algorithm is used to update the MDP with the maximum transmission power and limited bandwidth of each drone as constraints and the maximum total user transmission rate as the optimization goal. The specific process of obtaining the drone swarm resource allocation scheme is as follows:
[0112] The PPO algorithm is an emerging policy gradient (PG) algorithm. A novel objective function is designed to achieve small batch updates, which solves the problem that the PG algorithm is sensitive to step size and difficult to determine a reasonable step size. PPO is derived from the trust region policy optimization (TRPO) algorithm and introduces a clipping proxy objective. The basic idea is to use importance sampling to measure the difference before and after the policy update, thereby controlling the amplitude of the policy update and ensuring that the new policy remains within the "trust region" of the old policy.
[0113] See also Figure 4 , Figure 4 This is a schematic diagram of the PPO algorithm process. Figure 4 The operation process of the PPO algorithm is explained. The policy gradient algorithm optimizes the policy function π(a|s) to π(a|s; θ) through a parameter θ. Where a represents the action and s represents the state. In the policy gradient-based method, the original objective function is:
[0114]
[0115] where π θ is a random strategy, τ=(s 1 ,a 1 ,...,s T ,a T ) is an iterative trajectory with a step size of T. represents the cumulative reward of the iteration, and γ is the discount factor.
[0116] Estimate the policy gradient based on the differential J(θ) and get the loss function L PG :
[0117]
[0118] Among them, A t is the estimator of the advantage function at step t; Indicates the desired operation.
[0119] According to the policy gradient method, the parameter update equation can be written as:
[0120]
[0121] Among them, θ new represents the new parameters of the policy network; θ old Represents the original strategy network parameters; Represents the gradient operation; α is the update step size. It can be seen that α directly determines the quality of the new strategy. If α is inappropriate, the updated strategy is a worse strategy. When learning on a bad strategy, the updated strategy will become worse, which will cause the whole learning process to deteriorate.
[0122] Therefore, in order to ensure that the new strategy can make its reward function monotonically non-decreasing and the learning process will not deteriorate, the TRPO (Trust Region Policy Optimization) algorithm is proposed. In the TRPO algorithm, a trust region constraint is added to the policy gradient method using KL divergence (a method for measuring the difference between two probability distributions), indicating that the current policy π θ (a t |s t ) and the old policy The probability ratio The new objective function L TRPO (θ) is defined as:
[0123]
[0124] Here, ψ(θ) represents the probability ratio.
[0125] The TRPO algorithm is a constrained optimization problem, which requires a lot of computation to calculate the conjugate gradient in order to perform constrained optimization. Therefore, the proximal policy optimization algorithm is proposed. The PPO algorithm introduces a penalty to penalize large policy updates, thus avoiding the calculation of the conjugate gradient. Finally, the objective function used in the embodiment of the present invention can be expressed as:
[0126]
[0127] Here, ε is a hyperparameter and clip constrains ψ(θ) to be within (1-ε,1+ε) to avoid excessive policy updates.
[0128] In the drone swarm service user system, the drone is regarded as an intelligent agent, which contains a policy network π θ and a value assessment network At each time step t, the agent takes the current state s t Select an action a from the environment t , the selection is made through the policy network π θ Generated.
[0129] When the agent performs action a t After that, the environment returns a reward r t , and transfer to the new state s t+1 , the conversion t ,a t ,r t ,s t+1 > are stored in the experience replay buffer. As more experience data accumulates, the agent will randomly draw a batch of samples from the buffer to update the actor network and the critic network to optimize the policy and value function.
[0130] The actor network minimizes the loss function L PPO To optimize its policy parameter θ, the truncated generalized advantage estimate is used to calculate A t .
[0131]
[0132] In the formula, is the state value function of time slot t; is the state value function at time slot t+1.
[0133] Critics Network critic The empirical average A of each batch of samples t To update. The loss function of the critic network Defined as:
[0134]
[0135] The simulation experiment of the UAV swarm intelligent resource scheduling and interference coordination optimization method provided by the embodiment of the present invention is as follows:
[0136] In this simulation, three areas of 30km×30km are considered, in which a UAV base station (UAV-BS) is deployed in each area to provide wireless communication services to users. Here, the path loss model in the environment is a free path loss model.
[0137] In the policy network of the PPO algorithm, a two-layer feedforward fully connected neural network structure is used to generate the action probability distribution under the current state. The input layer accepts a state vector with a dimension of the number of states, which represents the state information of the current environment; the hidden layer contains 16 neurons, and the nonlinear mapping of features is realized through linear transformation and ReLU activation function; the dimension of the output layer is the number of actions, and the probability distribution of each action in the action space is generated through the Softmax function to meet the requirements of probability normalization. The main function of the policy network is to realize the mapping from state to action probability and guide the decision-making process of the intelligent agent in the environment. In the PPO algorithm, this network is used to represent the policy function, and the intelligent agent samples actions or selects the optimal action to interact with the environment according to the output action probability distribution.
[0138] The value network uses a two-layer feedforward fully connected neural network structure similar to the policy network, but its function is to estimate the value of the input state. The input layer accepts a state vector of the same size as the number of states. The hidden layer consists of 16 neurons and uses the ReLU activation function for nonlinear feature extraction. The output layer is a single scalar node that directly outputs the scalar value of the current state without the need for an additional activation function. The main function of the value network is to map from state to state value, that is, to estimate the long-term expected return of the agent in the current state. In the PPO algorithm, the value network provides a benchmark for policy optimization by calculating the state value, which is used to estimate the advantage function, thereby guiding the direction of policy updates and improving the stability and efficiency of training.
[0139] In addition, the Adam optimizer is used to train the neural network parameters. The simulation is completed in Python 3.7 and TensorFlow 2.6. Other parameters in the simulation are given in Table 1:
[0140] Table 1
[0141]
[0142] See also Figure 5 , Figure 5 This is a schematic diagram of the original reward value during PPO training, where the X-axis represents the training phase and the Y-axis represents the reward value. In the early stages of training (0–520 rounds), the PPO reward value fluctuates violently, which is caused by the randomness of the strategy update during the exploration phase. However, as the training progresses, after 520 rounds, the PPO reward gradually increases and eventually stabilizes at around 882.885Mbps. This result verifies the performance advantage of the PPO algorithm in obtaining higher returns and achieving convergence.
[0143] See also Figure 6 , Figure 6 This is a schematic diagram of the deviation of the PPO reward value relative to the mean, where the X-axis represents the training stage and the Y-axis represents the deviation rate. In the initial stage, due to the exploration process, the deviation value shows a large negative fluctuation, indicating that the reward result is unstable. As the number of training rounds increases, the deviation gradually converges and stabilizes within the ±3% threshold (red dotted line) after about 760 rounds. This phenomenon shows that the PPO algorithm gradually converges as the training progresses, and the reward value fluctuates around the mean and tends to be stable.
[0144] See also Figure 7 , Figure 7 This is a schematic diagram of the comparison between the original reward value of PPO and the 50-round moving average, where the X-axis represents the training stage and the Y-axis represents the reward value. The original reward data of PPO fluctuates greatly in the early stage, which is due to the instability caused by the exploration and learning process. In order to reveal the overall trend of the data, a 50-round moving average (green solid line) is applied to the original data.
[0145] See also Figure 8 , Figure 8 The figure is a schematic diagram of Coefficient of Variation. Coefficient of Variation (CV) is a statistical indicator used to measure the ratio of the standard deviation of a data set to the mean. It is used to evaluate the relative discreteness of data and is particularly suitable for comparing the stability of data of different scales or units. The X-axis represents the training stage and the Y-axis represents the CV value. In the initial stage of training, the CV value is high, indicating that the reward value fluctuates violently. This is mainly due to the randomness in the exploration process and the instability of the strategy update, which makes the standard deviation larger than the mean. As the number of training rounds increases, the CV gradually decreases. This shows that the PPO algorithm gradually converges, the reward values are stably distributed around a higher mean, and the standard deviation is significantly reduced. At this time, CV converges to a lower level, reflecting the high stability and low volatility of the PPO training results.
[0146] See also Fig. 9 , Fig. 9 It is a schematic diagram of the PPO training dynamics and the comparison results with the greedy algorithm and the polling strategy. Fig. 9 (a) in the figure shows the comparison between PPO training results and other algorithms. (a) shows the reward value obtained during PPO training and is compared with Algorithm 1 (the red dashed line represents the greedy algorithm) and Algorithm 2 (the green dashed line represents the polling strategy). In the initial stage, the PPO reward fluctuates greatly, which is caused by the exploration process. As the training progresses, the PPO reward gradually increases and stabilizes around 870Mbps, with an average of 882.885Mbps, which is significantly higher than Algorithm 1 (reward is 680.387Mbps) and Algorithm 2 (reward is 568.249Mbps). Fig. 9 (b) in Figure 1 shows the moving average reward value comparison. By taking the moving average of 50 training rounds (green solid line), the PPO raw data (light blue) is smoothed. The results show that, especially after 900 rounds, the PPO reward is basically stable, far exceeding Algorithm 1 and Algorithm 2. Fig. 9 (c) in the figure shows the deviation ratio of the PPO reward value relative to the mean. Starting from round 750, the deviation stabilizes within the ±3% threshold range (red dashed line), indicating that the PPO reward value gradually converges and the volatility decreases. Fig. 9 (d) in the figure shows the final performance comparison of the algorithms, and compares the final performance of PPO, Algorithm 1, and Algorithm 2 through a bar chart. The final average reward of PPO is 882.885Mbps, which is 29.8% higher than Algorithm 1 (680.387Mbps) and 55.4% higher than Algorithm 2 (568.249Mbps).
[0147] In summary, the PPO algorithm is superior to the two benchmark algorithms in terms of reward value, and has good stability and fast convergence speed, demonstrating its good performance advantages.
[0148] In the embodiment of the present invention, with respect to the resource allocation problem of UAV-assisted communication, the present invention considers the problem that resources are limited under interference conditions and multiple users on the ground share limited wireless resources, and proposes a joint solution to construct a resource allocation problem that maximizes the total rate under the constraints of the maximum transmission power and limited bandwidth of each UAV, and adopts a solution based on a proximal strategy optimization algorithm to solve the joint optimization problem. In addition, compared with the dynamic resource allocation algorithm, the present invention can reduce the complexity of resource allocation, improve the stability of convergence while ensuring the efficiency of strategy improvement, and effectively avoid the dimensional disaster problem that traditional optimization algorithms may face in high-dimensional environments. Compared with the static allocation algorithm, the enhanced anti-interference ability of the communication system can adapt to the dynamic changes in the deployment environment of the UAV base station, and can maintain high performance in a complex and highly uncertain network environment, which is closer to the actual situation and has strong applicability.
[0149] Based on the same inventive concept, the embodiment of the present invention also provides a drone swarm intelligent resource scheduling and interference coordination optimization device, see Fig.10 , Fig.10 : is a schematic diagram of a structure of a drone swarm intelligent resource scheduling and interference coordination optimization device provided by an embodiment of the present invention, comprising:
[0150] Construction module 1001 is used to construct a system model of a drone swarm service user system; the system model includes N drones and U users; each drone serves U users. N Users;
[0151] The determination module 1002 is used to determine a channel model and a rate model according to the system model; the channel model is used to represent the free path loss between the drone and the user it serves; the rate model is used to represent the transmission rate obtained by the user after the channel is allocated;
[0152] The solution module 1003 is used to construct a time-frequency resource allocation optimization problem based on the channel model and the rate model, and use the proximal strategy optimization algorithm to solve the time-frequency resource allocation optimization problem to obtain a drone swarm resource allocation plan; the time-frequency resource allocation optimization problem uses the maximum transmission power and limited bandwidth of each drone as constraints, and takes maximizing the total user transmission rate as the optimization goal; the maximum transmission power means that the total power transmitted by each drone to the users it serves cannot exceed its own total transmission power; the limited bandwidth means that the total bandwidth transmitted by each drone to the users it serves cannot exceed its own total bandwidth.
[0153] In an embodiment of the present invention, considering the problem of limited resources and multiple ground users sharing limited wireless resources under interference conditions, a time-frequency resource allocation optimization problem is constructed to maximize the total user transmission rate under the constraints of the maximum transmission power and limited bandwidth of each UAV, and a proximal strategy optimization algorithm is adopted to solve the time-frequency resource allocation optimization problem. Compared with the dynamic resource allocation algorithm, the complexity of resource allocation is reduced, and while ensuring the efficiency of strategy improvement, the stability of convergence is improved, and the dimensionality disaster problem that may be faced by traditional optimization algorithms in high-dimensional environments is effectively avoided. Compared with the static allocation algorithm, the anti-interference ability of the communication system is enhanced, which can adapt to the dynamic changes of the deployment environment of the UAV base station, and can maintain high performance in a complex and highly uncertain network environment, which is closer to the actual situation and has strong applicability.
[0154] Optionally, the solution module uses a proximal strategy optimization algorithm to solve the time-frequency resource allocation optimization problem and obtain a drone swarm resource allocation solution, including:
[0155] Modeling the time-frequency resource allocation optimization problem as an MDP;
[0156] By using the proximal strategy optimization algorithm, the maximum transmission power and limited bandwidth of each UAV are constrained, and the maximum total user transmission rate is taken as the optimization goal to update the MDP, and the resource allocation scheme of the UAV swarm is obtained;
[0157] Among them, the state space, action space and reward in MDP are defined as follows:
[0158] Each action of the drone in the action space is mapped to a column in a U×M binary matrix; the element in the mth column of the uth row in the binary matrix is used to describe the action of the uth user being assigned to the mth channel;
[0159] The state space is defined as a one-dimensional array for reflecting the current channel allocation state; each element in the one-dimensional array is used to describe the state of a user being allocated to a channel;
[0160] The reward is defined as the sum of the total transmission rates of all users under the current channel allocation state.
[0161] Optionally, the channel model is:
[0162]
[0163] in, Represents the nth drone and the uth drone it serves n The free path loss between users; f c Indicates the carrier frequency; Represents the nth drone and the uth drone it serves n The distance between users; c = 3 × 10 8 m / s.
[0164] Optionally, the rate model is:
[0165]
[0166] Among them, R total represents the total transmission rate of the user; represents time slot t, the uth n The transmission rate obtained after each user is assigned a channel; B represents the total bandwidth of the UAV; Indicates time slot t, the mth channel is assigned to the uth n Allocation plan for each user; Indicates the uth node served by the nth drone n The signal-to-noise ratio of a user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n; m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of drones; U n represents the total number of users served by the nth UAV; M represents the total number of channels.
[0167] Optionally, the time-frequency resource allocation optimization problem is:
[0168]
[0169] in, Indicates the nth drone’s service to the uth drone n The transmission power of each user; Represents the nth drone and the uth drone it serves n Channel parameters between users; represents the internal interference of the nth UAV at time slot t; represents the time slot t, the u-th node served by the n-th drone n The interference of other drones to each user; 2 (t) represents Gaussian white noise; SINR threshold represents the signal-to-noise ratio threshold; Indicates the uth n The demand rate of each user; P n represents the total transmission power of the nth UAV; B represents the total bandwidth of the UAV; Indicates time slot t, the mth channel is assigned to the uth n Allocation plan for each user; Indicates the uth node served by the nth drone n The signal-to-noise ratio of a user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n ; m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of drones; U n represents the total number of users served by the nth UAV; M represents the total number of channels; represents a collection of drones; represents the set of users served by the nth drone; Represents a collection of channels.
[0170] It should be noted that the terms "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention.
[0171] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.
[0172] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other changes to the disclosed embodiments by viewing the drawings and the disclosed content. In the description of the present invention, the term "comprising" does not exclude other components or steps, "one" or "an" does not exclude multiple situations, and "multiple" means two or more, unless otherwise clearly and specifically limited. In addition, certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0173] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0174] It should be noted that the device of an embodiment of the present invention is a device that applies the above-mentioned method for intelligent resource scheduling and interference coordination optimization of a swarm of unmanned aerial vehicles. All embodiments of the above-mentioned method for intelligent resource scheduling and interference coordination optimization of a swarm of unmanned aerial vehicles are applicable to the device and can achieve the same or similar beneficial effects.
[0175] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A method for intelligent resource scheduling and interference coordination optimization of drone swarms, characterized in that: include: Construct a system model of a drone swarm service user system; the system model includes N drones and U users; each drone serves U users. N Users; Determine a channel model and a rate model according to the system model; the channel model is used to represent the free path loss between the drone and the user it serves; the rate model is used to represent the transmission rate obtained by the user after the channel is allocated; A time-frequency resource allocation optimization problem is constructed based on the channel model and the rate model, and the proximal strategy optimization algorithm is used to solve the time-frequency resource allocation optimization problem to obtain a resource allocation plan for the drone swarm. In the time-frequency resource allocation optimization problem, the maximum transmission power and limited bandwidth of each drone are used as constraints, and maximizing the total user transmission rate is used as the optimization goal. The maximum transmission power indicates that the total power transmitted by each drone to the users it serves cannot exceed its own total transmission power. The limited bandwidth indicates that the total bandwidth transmitted by each drone to the users it serves cannot exceed its own total bandwidth.
2. The method for intelligent resource scheduling and interference coordination optimization of drone swarms according to claim 1 is characterized in that: The proximal strategy optimization algorithm is used to solve the time-frequency resource allocation optimization problem and obtain the resource allocation scheme for the drone swarm, including: Modeling the time-frequency resource allocation optimization problem as an MDP; By using the proximal strategy optimization algorithm, the maximum transmission power and limited bandwidth of each UAV are constrained, and the maximum total user transmission rate is taken as the optimization goal to update the MDP, and the resource allocation scheme of the UAV swarm is obtained; Among them, the state space, action space and reward in MDP are defined as follows: Each action of the drone in the action space is mapped to a column in a U×M binary matrix; the element in the mth column of the uth row in the binary matrix is used to describe the action of the uth user being assigned to the mth channel; The state space is defined as a one-dimensional array for reflecting the current channel allocation state; each element in the one-dimensional array is used to describe the state of a user being allocated to a channel; The reward is defined as the sum of the total transmission rates of all users under the current channel allocation state.
3. The method for intelligent resource scheduling and interference coordination optimization of drone swarm according to claim 1 is characterized in that: The channel model is: in, Represents the nth drone and the uth drone it serves n The free path loss between users; f c Indicates the carrier frequency; Represents the nth drone and the uth drone it serves n The distance between users; c = 3 × 10 8 m / s.
4. The method for intelligent resource scheduling and interference coordination optimization of drone swarms according to claim 1 is characterized in that: The rate model is: Among them, R total represents the total transmission rate of the user; represents time slot t, the uth n The transmission rate obtained after each user is assigned a channel; B represents the total bandwidth of the UAV; Indicates time slot t, the mth channel is assigned to the uth n Allocation plan for each user; Indicates the uth node served by the nth drone n The signal-to-noise ratio of a user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n ; m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of drones; U n represents the total number of users served by the nth UAV; M represents the total number of channels.
5. The method according to claim 1, characterized in that: The time-frequency resource allocation optimization problem is: in, Indicates the nth drone’s service to the uth drone n The transmission power of each user; Represents the nth drone and the uth drone it serves n Channel parameters between users; represents the internal interference of the nth UAV at time slot t; represents the time slot t, the u-th node served by the n-th drone n The interference of other drones to each user; 2 (t) represents Gaussian white noise; SINR threshold represents the signal-to-noise ratio threshold; Indicates the uth n The demand rate of each user; P n represents the total transmission power of the nth UAV; B represents the total bandwidth of the UAV; Indicates time slot t, the mth channel is assigned to the uth n Allocation plan for each user; Indicates the uth node served by the nth drone n The signal-to-noise ratio of a user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n ; m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of drones; U n represents the total number of users served by the nth drone; M represents the total number of channels; represents a collection of drones; represents the set of users served by the nth drone; Represents a collection of channels.
6. A drone swarm intelligent resource scheduling and interference coordination optimization device, characterized in that: include: A construction module is used to construct a system model of a drone swarm service user system; the system model includes N drones and U users; each drone serves U users. N Users; A determination module, used to determine a channel model and a rate model according to the system model; the channel model is used to represent the free path loss between the drone and the user it serves; the rate model is used to represent the transmission rate obtained by the user after the channel is allocated; A solution module is used to construct a time-frequency resource allocation optimization problem based on the channel model and the rate model, and use a proximal strategy optimization algorithm to solve the time-frequency resource allocation optimization problem to obtain a drone swarm resource allocation solution; in the time-frequency resource allocation optimization problem, the maximum transmission power and limited bandwidth of each drone are used as constraints, and maximizing the total user transmission rate is used as the optimization goal; the maximum transmission power means that the total power transmitted by each drone to the users it serves cannot exceed its own total transmission power; the limited bandwidth means that the total bandwidth transmitted by each drone to the users it serves cannot exceed its own total bandwidth.
7. The drone swarm intelligent resource scheduling and interference coordination optimization device according to claim 6 is characterized in that: The solution module uses a proximal strategy optimization algorithm to solve the time-frequency resource allocation optimization problem and obtain a resource allocation solution for the drone swarm, including: Modeling the time-frequency resource allocation optimization problem as an MDP; By using the proximal strategy optimization algorithm, the maximum transmission power and limited bandwidth of each UAV are constrained, and the maximum total user transmission rate is taken as the optimization goal to update the MDP, and the resource allocation scheme of the UAV swarm is obtained; Among them, the state space, action space and reward in MDP are defined as follows: Each action of the drone in the action space is mapped to a column in a U×M binary matrix; the element in the mth column of the uth row in the binary matrix is used to describe the action of the uth user being assigned to the mth channel; The state space is defined as a one-dimensional array for reflecting the current channel allocation state; each element in the one-dimensional array is used to describe the state of a user being allocated to a channel; The reward is defined as the sum of the total transmission rates of all users under the current channel allocation state.
8. The drone swarm intelligent resource scheduling and interference coordination optimization device according to claim 6 is characterized in that: The channel model is: in, Represents the nth drone and the uth drone it serves n The free path loss between users; f c Indicates the carrier frequency; Represents the nth drone and the uth drone it serves n The distance between users; c = 3 × 10 8 m / s.
9. The drone swarm intelligent resource scheduling and interference coordination optimization device according to claim 6, characterized in that: The rate model is: Among them, R total represents the total transmission rate of the user; represents time slot t, the uth n The transmission rate obtained after each user is assigned a channel; B represents the total bandwidth of the UAV; Indicates time slot t, the mth channel is assigned to the uth n Allocation plan for each user; Indicates the uth node served by the nth drone n The signal-to-noise ratio of a user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n ; m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of drones; U n represents the total number of users served by the nth UAV; M represents the total number of channels.
10. The drone swarm intelligent resource scheduling and interference coordination optimization device according to claim 6, characterized in that: The time-frequency resource allocation optimization problem is: in, Indicates the nth drone’s service to the uth drone n The transmission power of each user; Represents the nth drone and the uth drone it serves n Channel parameters between users; represents the internal interference of the nth UAV at time slot t; represents the time slot t, the u-th node served by the n-th drone n The interference of other drones to each user; 2 (t) represents Gaussian white noise; SINR threshold represents the signal-to-noise ratio threshold; Indicates the uth n The demand rate of each user; P n represents the total transmission power of the nth UAV; B represents the total bandwidth of the UAV; Indicates time slot t, the mth channel is assigned to the uth n Allocation plan for each user; Indicates the uth node served by the nth drone n The signal-to-noise ratio of a user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n ; m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of drones; U n represents the total number of users served by the nth drone; M represents the total number of channels; represents a collection of drones; represents the set of users served by the nth drone; Represents a collection of channels.
Citation Information
Patent Citations
Unmanned aerial vehicle auxiliary communication anti-interference method based on multi-agent reinforcement learning
CN118921099A
Resource allocation and cooperative unloading method based on federal deep reinforcement learning under multi-unmanned aerial vehicle assisted internet of vehicles
CN119316879A
Managing satellite bearer resources
US20220158719A1