An unmanned aerial vehicle swarm intelligent resource scheduling and interference coordination optimization method and device
By constructing a UAV swarm service user system model and using a near-end policy optimization algorithm to solve the time-frequency resource allocation problem, the problem of insufficient flexibility and real-time performance in resource allocation in UAV communication systems is solved, achieving efficient spectrum management and interference coordination, and improving the system's dynamic adaptability and performance.
Patent Information
- Application Number
- CN202510102760.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Existing UAV communication systems suffer from low flexibility, poor efficiency, insufficient real-time performance, and inadequate scalability in spectrum management and resource allocation. They struggle to cope with dynamic changes in user needs, channel conditions, and interference environments, especially experiencing performance degradation in high-density user and high-speed mobile scenarios.
A system model for a drone swarm service user system is constructed. The near-end policy optimization algorithm (PPO) is used to solve the time-frequency resource allocation optimization problem. By constructing an MDP model with the goal of maximizing the total transmission rate of users, and considering the maximum transmit power and limited bandwidth of each drone as constraints, dynamic resource allocation is achieved.
It improves the anti-interference capability of the communication system, reduces the complexity of resource allocation, enhances the system's performance in complex and highly uncertain network environments, adapts to the dynamic changes in the deployment environment of UAV base stations, and improves the real-time performance and scalability of resource management.
Smart Images

Figure CN120018293B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of communication technology, specifically relating to a method and apparatus for intelligent resource scheduling and interference coordination optimization of unmanned aerial vehicle (UAV) swarms. Background Technology
[0002] Drones are widely used in various fields, including disaster recovery, surveillance, and providing communication services in rural and underdeveloped areas. With the development of mobile communication technology and the emergence of various new services, the explosive growth of mobile data traffic has become increasingly apparent. However, due to factors such as complex terrain, the coverage capacity of existing ground base stations is still insufficient to meet the ever-increasing communication demands.
[0003] Therefore, the development of integrated space-air-terrestrial networks has become a key trend in future communication systems, and unmanned aerial vehicle (UAV) networks are an important component of space networks. UAVs equipped with wireless communication modules can be rapidly deployed as airborne base stations, providing ubiquitous access to ground user equipment (UEs) during temporary events such as hotspots or large public gatherings, representing a promising solution. UAV deployment is seen as an effective complement to existing cellular systems, capable of enhancing wireless capacity and extending coverage in scenarios with ultra-high-density traffic demands. By leveraging their mobility and adaptability, UAV-based networks offer a transformative opportunity to address the limitations of traditional terrestrial communication infrastructure. However, integrating UAVs into wireless communication systems presents significant challenges, particularly in spectrum management. As airborne base stations or relay nodes, UAVs must efficiently share limited wireless resources with ground users. These challenges are further exacerbated by the high mobility of UAVs and the dynamic changes in user locations, leading to time-varying channel conditions and interference patterns.
[0004] Resource allocation in UAV communications faces several key challenges. First, the dynamic changes in user locations introduce strong spatiotemporal correlations into the communication environment. Second, UAVs have limited resources, such as battery power and spectrum, requiring efficient utilization strategies to balance communication quality and resource efficiency. Finally, resource contention and interference among multiple users necessitate effective scheduling mechanisms to ensure fairness and performance. The complexity of these issues makes traditional static resource allocation methods ill-suited to the dynamic characteristics of UAV networks. Therefore, developing innovative spectrum allocation strategies for UAV networks is crucial to fully realizing their potential and addressing their unique challenges.
[0005] With the widespread application of UAV-BS (Unmanned Aerial Vehicle Base Stations) in emergency communications, disaster recovery, and temporary large-scale events, interference coordination and resource allocation technologies have become one of the core issues in serving ground users. However, existing interference coordination and resource allocation schemes still have significant shortcomings in practical applications. First, traditional rule-based resource allocation methods rely on pre-set fixed strategies, making it difficult to cope with dynamic changes in user needs, channel conditions, and interference environments. This static resource allocation approach often fails to provide sufficient flexibility and efficiency in complex network scenarios, leading to a decline in system performance. Second, while optimization theory-based methods (such as convex and non-convex optimization) can obtain theoretically optimal solutions to some extent, they rely on precise problem modeling and complex computational processes. Especially when UAV-BS serves high-density users or high-speed mobile scenarios, the optimization problem may become extremely complex, resulting in excessively long solution times or convergence to suboptimal solutions. Furthermore, traditional optimization methods are often difficult to handle nonlinear, high-dimensional, and highly dynamic environments, lacking real-time performance and scalability.
[0006] Therefore, how to provide a method for intelligent resource scheduling and interference coordination optimization of UAV swarms that can adapt to the dynamic changes in the deployment environment of UAV base stations, maintain high performance in complex and highly uncertain network environments, and meet the real-time requirements of resource management has become an urgent problem to be solved. Summary of the Invention
[0007] To address the aforementioned problems in the existing technology, this invention provides a method and apparatus for intelligent resource scheduling and interference coordination optimization of unmanned aerial vehicle (UAV) swarms.
[0008] The technical problem to be solved by this invention is achieved through the following technical solution:
[0009] In a first aspect, the present invention provides a method for intelligent resource scheduling and interference coordination optimization of unmanned aerial vehicle (UAV) swarms, comprising:
[0010] A system model for a drone swarm service user system is constructed; the system model includes N drones and U users; wherein each drone serves U users. N One user;
[0011] Based on the system model, a channel model and a rate model are determined; the channel model represents the free path loss between the UAV and the users it serves; the rate model represents the transmission rate obtained by the users after being assigned a channel.
[0012] Based on the channel model and the rate model, a time-frequency resource allocation optimization problem is constructed, and the near-end policy optimization algorithm is used to solve the time-frequency resource allocation optimization problem to obtain a UAV swarm resource allocation scheme. The time-frequency resource allocation optimization problem is constrained by the maximum transmit power and limited bandwidth of each UAV, and the optimization objective is to maximize the total transmission rate of users. The maximum transmit power means that the total power transmitted by each UAV to the users it serves cannot exceed its own total transmit power. The limited bandwidth means that the total bandwidth transmitted by each UAV to the users it serves cannot exceed its own total bandwidth.
[0013] Optionally, a near-end policy optimization algorithm is used to solve the time-frequency resource allocation optimization problem to obtain a UAV swarm resource allocation scheme, including:
[0014] The time-frequency resource allocation optimization problem is modeled as an MDP.
[0015] Using a near-end strategy optimization algorithm, the MDP is updated with the maximum transmit power and limited bandwidth of each UAV as constraints and the goal of maximizing the total transmission rate of users, to obtain a UAV swarm resource allocation scheme.
[0016] The definitions of state space, action space, and reward in MDP are as follows:
[0017] Each action of the UAV in the action space is mapped to a column of a U×M binary matrix; the element in the u-th row and m-th column of the binary matrix is used to describe the action of the u-th user being assigned to the m-th channel;
[0018] The state space is defined as a one-dimensional array that reflects the current channel allocation state; each element in the one-dimensional array describes the state in which a user is allocated to a channel.
[0019] The reward is defined as the sum of the total transmission rates of all users under the current channel allocation state.
[0020] Optionally, the channel model is:
[0021]
[0022] in, This represents the relationship between the nth drone and the uth drone it serves. n Free path loss between users; f c Indicates the carrier frequency; This represents the nth drone and the uth drone it serves. n The distance between users; c = 3 × 10 8 m / s.
[0023] Optionally, the rate model is:
[0024]
[0025] Among them, R total This represents the total transmission rate of the user; Represents time slot t, the u-th n The transmission rate obtained by each user after being assigned a channel; B represents the total bandwidth of the UAV. In time slot t, the m-th channel is assigned to the u-th channel. n Allocation scheme for individual users; This indicates that the u-th drone serves the nth drone. n Signal-to-noise ratio of each user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of UAVs; U n M represents the total number of users served by the nth drone; M represents the total number of channels.
[0026] Optionally, the time-frequency resource allocation optimization problem is:
[0027]
[0028] in, This indicates that the nth drone serves the uth drone. n Transmit power of each user; This represents the relationship between the nth drone and the uth drone it serves. n Channel parameters between users; Indicates the internal interference of the nth UAV in time slot t; Indicates time slot t, where the u-th time slot is served by the nth drone. n Interference from other drones experienced by individual users; δ 2 (t) represents Gaussian white noise; SINR threshold Indicates the signal-to-noise ratio threshold; Indicates the uth n The demand rate of each user; P n B represents the total transmit power of the nth UAV; B represents the total bandwidth of the UAV. In time slot t, the m-th channel is assigned to the u-th channel. n Allocation scheme for individual users; This indicates that the u-th drone serves the nth drone. n Signal-to-noise ratio of each user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,Un m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of UAVs; U n M represents the total number of users served by the nth drone; M represents the total number of channels. A collection of drones; This represents the set of users served by the nth drone; A set representing channels.
[0029] Secondly, the present invention provides an intelligent resource scheduling and interference coordination optimization device for unmanned aerial vehicle (UAV) swarms, comprising:
[0030] The construction module is used to build a system model for a drone swarm service user system; the system model includes N drones and U users; wherein each drone serves U users. N One user;
[0031] The determination module is used to determine the channel model and the rate model based on the system model; the channel model is used to represent the free path loss between the UAV and the users it serves; the rate model is used to represent the transmission rate obtained by the users after being assigned a channel.
[0032] The solution module is used to construct a time-frequency resource allocation optimization problem based on the channel model and the rate model, and to solve the time-frequency resource allocation optimization problem using a near-end policy optimization algorithm to obtain a UAV swarm resource allocation scheme. The time-frequency resource allocation optimization problem is constrained by the maximum transmit power and limited bandwidth of each UAV, with the optimization objective being to maximize the total transmission rate of users. The maximum transmit power means that the total power transmitted by each UAV to the users it serves cannot exceed its own total transmit power; the limited bandwidth means that the total bandwidth transmitted by each UAV to the users it serves cannot exceed its own total bandwidth.
[0033] This invention provides an intelligent resource scheduling and interference coordination optimization method for UAV swarms. Considering the resource constraints under interference conditions and the problem of multiple ground users sharing limited wireless resources, it constructs a time-frequency resource allocation optimization problem to maximize the total transmission rate of users under the constraints of the maximum transmit power and limited bandwidth of each UAV. Furthermore, it adopts a near-end policy optimization algorithm to solve the time-frequency resource allocation optimization problem. Compared with the dynamic resource allocation algorithm, it reduces the resource allocation complexity, improves the convergence stability while ensuring the efficiency of policy improvement, and effectively avoids the dimensionality curse problem that traditional optimization algorithms may face in high-dimensional environments. Compared with the static allocation algorithm, it enhances the anti-interference capability of the communication system, can adapt to the dynamic changes of the UAV base station deployment environment, maintains high performance in complex and highly uncertain network environments, is closer to real-world conditions, and has strong applicability.
[0034] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description
[0035] Figure 1 This is a flowchart illustrating an intelligent resource scheduling and interference coordination optimization method for unmanned aerial vehicle (UAV) swarms provided in an embodiment of the present invention.
[0036] Figure 2 This is a schematic diagram of a drone swarm service user system scenario;
[0037] Figure 3 This is a schematic diagram of the interference scenario;
[0038] Figure 4 This is a schematic diagram of the PPO algorithm process;
[0039] Figure 5 This is a diagram illustrating the original reward values during the PPO training process;
[0040] Figure 6 This is a diagram illustrating the deviation of the PPO reward value from the mean;
[0041] Figure 7 This is a diagram showing the comparison between the original PPO reward value and the 50-round moving average.
[0042] Figure 8 This is a diagram illustrating the Coefficient of Variation;
[0043] Figure 9 This is a diagram illustrating the dynamics of PPO training and its comparison with greedy algorithms and polling strategies.
[0044] Figure 10 This is a schematic diagram of the structure of an intelligent resource scheduling and interference coordination optimization device for unmanned aerial vehicle (UAV) swarms provided in an embodiment of the present invention. Detailed Implementation
[0045] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0046] To address the shortcomings of existing interference coordination and resource allocation methods, such as low flexibility, poor efficiency, complex solutions, and lack of real-time performance and scalability, this invention provides an intelligent resource scheduling and interference coordination optimization method for unmanned aerial vehicle (UAV) swarms. (See [link to relevant documentation]). Figure 1 , Figure 1 This is a flowchart illustrating an intelligent resource scheduling and interference coordination optimization method for unmanned aerial vehicle (UAV) swarms provided by an embodiment of the present invention, specifically including the following steps:
[0047] Step S101: Construct a system model for the UAV swarm service user system; the system model includes N UAVs and U users; wherein each UAV serves U users. N One user.
[0048] See Figure 2 , Figure 2 This is a schematic diagram of a drone swarm service user system scenario. The drones hover over the users and transmit high-definition video to each user on the ground. Based on the users' channel conditions and the current resource allocation, time and frequency resources are selected for multiple users to maximize the total throughput of the multi-drone network.
[0049] In this scenario, time resources are divided into T time slots, denoted as set. There are N drones, denoted as set. Each drone is equipped with an omnidirectional antenna array, and there are U users, denoted as set. Each user is equipped with an omnidirectional antenna array, and each drone serves a U number of users. N Let there be a set of . And the coverage radius of each drone is r n ;
[0050] The total frequency band of all drones B = f up -f down Divided into M sub-channels, denoted as The total power of each drone is P. n The drone n pairs with the uth n Transmit power of individual users Assume that the user's position changes slowly over time in the scenario, and the position of the u-th user is denoted as L. u =(x u (t),y u (t),h u (t)), and the user's height remains unchanged, i.e., h u (t)=h u (0), h u (t) represents the altitude of the u-th user at time t, h u (0) represents the initial height of the u-th user, x u (t) represents the x-axis coordinate of the u-th user in time slot t, and y-axis coordinate of the u-th user. u (t) represents the y-coordinate of the u-th user in time slot t; the position of the n-th drone in time slot t is denoted as L. n (t)=(x n (t),y n (t),h n (t)), x n (t) represents the x-axis coordinate of the nth UAV in time slot t; y n(t) represents the y-coordinate of the nth UAV in time slot t; h n (t) represents the altitude of the nth UAV in time slot t.
[0051] In this embodiment of the invention, users can be further divided into ground users and interfered users. When a user is simultaneously within the coverage area of two drones, they will be interfered with. For example, in... Figure 2 In the case of the overlapping coverage areas of drones A and C, there is a user being interfered with; similarly, there is a user being interfered with where the coverage areas of drones A and B overlap.
[0052] Step S102: Determine the channel model and rate model based on the system model; the channel model is used to represent the free path loss between the UAV and the users it serves; the rate model is used to represent the transmission rate obtained by the user after being assigned a channel.
[0053] In this embodiment of the invention, assuming good channel conditions, in the t-th time slot, the n-th UAV and the u-th UAV... n The channel model between individual users adopts the free path loss model, constructing the free path loss of a single transmission.
[0054]
[0055] Among them, f c Where the carrier frequency is c = 3 × 10 8 m / s, The calculation formula is as follows:
[0056]
[0057] in, For the nth drone and the uth n The distance between users.
[0058] In this embodiment of the invention, the channel allocation scheme in time slot t uses express, Indicates the channel indicator. In time slot t, the m-th channel is assigned to the u-th channel. n The allocation scheme for each user. In time slot t, if the m-th subchannel is allocated to the u-th subchannel... n For each user, otherwise,
[0059] In this embodiment of the invention, the rate model is as follows:
[0060]
[0061] Among them, Rtotal Indicates the total data transfer rate for the user; Represents time slot t, the u-th n The transmission rate obtained by each user after being assigned a channel; B represents the total bandwidth of the UAV. In time slot t, the m-th channel is assigned to the u-th channel. n Allocation scheme for individual users; This indicates that the u-th drone serves the nth drone. n Signal-to-noise ratio of each user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of UAVs; U n M represents the total number of users served by the nth drone; M represents the total number of channels. The uth serving the nth drone n The signal-to-noise ratio of a user in time slot t is expressed by the following formula:
[0062]
[0063] Where, δ 2 (t) represents additive white Gaussian noise. In time slot t, the internal interference of the nth UAV refers to the interference between the nth UAV and other users served by the uth UAV. n Interference from individual users Indicates time slot t, where the u-th time slot is served by the nth drone. n Interference from other drones experienced by individual users. The following will discuss... and The calculation process will be explained.
[0064] See Figure 3 , Figure 3 This is a diagram illustrating an interference scenario. User 1 is served by drone A, but the distance d between User 1 and drone B is... UAV_B ,Right now The coverage radius r of drone B is smaller than that of drone B. UAV_B ,Right now Drone B will interfere with User 1, therefore a variable is used. To represent the interference of other drones j on the current user:
[0065]
[0066] in, Indicates the uth n x-axis coordinates of each user; x j(t) represents the x-axis coordinate of the j-th UAV; Indicates the uth n x-axis coordinates of each user; y j (t) represents the y-axis coordinate of the j-th UAV; Indicates the uth n The height of a user in time slot t; h j (t) represents the altitude of the j-th UAV in time slot t.
[0067] This allows us to calculate the internal interference of the drone and the interference the drone causes to the user:
[0068]
[0069] Where, k m,i (t) indicates whether the i-th user occupies the m-th channel, 1 indicates occupancy, 0 indicates non-occupancy; P n,i This represents the transmission power of the nth drone to the i-th user; This indicates that the nth drone and the uth drone... n Channel gain between users; Indicates the uth j Does each user occupy the m-th channel? Indicates the uth n Does user P receive interference from the j-th drone? j,i This represents the transmission power of the j-th drone to the i-th user.
[0070] Step S103: Construct a time-frequency resource allocation optimization problem based on the channel model and rate model, and solve the time-frequency resource allocation optimization problem using the near-end policy optimization algorithm to obtain the UAV swarm resource allocation scheme. In the time-frequency resource allocation optimization problem, the maximum transmit power and limited bandwidth of each UAV are used as constraints, and the optimization objective is to maximize the total transmission rate of users. The maximum transmit power means that the total power transmitted by each UAV to the users it serves cannot exceed its own total transmit power. The limited bandwidth means that the total bandwidth transmitted by each UAV to the users it serves cannot exceed its own total bandwidth.
[0071] In this embodiment of the invention, a time-frequency resource allocation optimization problem based on a channel model and a rate model is solved. The goal of this optimization problem is to maximize the total transmission rate under the constraints of the maximum transmit power and limited bandwidth of each UAV.
[0072] Therefore, the resource allocation problem for each time slot can be formulated as a time-frequency resource allocation optimization problem:
[0073]
[0074] in, This indicates that the nth drone serves the uth drone. n Transmit power of each user; This represents the relationship between the nth drone and the uth drone it serves. n Channel parameters between users; Indicates the internal interference of the nth UAV in time slot t; Indicates time slot t, where the u-th time slot is served by the nth drone. n Interference from other drones experienced by individual users; δ 2 (t) represents Gaussian white noise; SINR threshold Indicates the signal-to-noise ratio threshold; Indicates the uth n The demand rate of each user; P n B represents the total transmit power of the nth UAV; B represents the total bandwidth of the UAV. In time slot t, the m-th channel is assigned to the u-th channel. n Allocation scheme for individual users; This indicates that the u-th drone serves the nth drone. n Signal-to-noise ratio of each user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of UAVs; U n M represents the total number of users served by the nth drone; M represents the total number of channels. A collection of drones; This represents the set of users served by the nth drone; A set representing channels.
[0075] In the above constraints, C1 is used to ensure that the signal-to-noise ratio received by the user is greater than the specified threshold value; C2 indicates that the rate allocated to the user in each time slot is greater than the user's required rate; C3 and C4 indicate that the number of sub-channels and bandwidth allocated to a single user are less than the total number of channels and total bandwidth; C5 indicates that a channel of a certain UAV can only be allocated to one user; C6 indicates that the power allocated to the user is greater than 0 and less than the maximum power; C7 indicates that the sum of the power allocated to users by the UAV is less than the maximum power of the UAV.
[0076] In this embodiment of the invention, the Proximal Policy Optimization (PPO) algorithm is an emerging policy gradient (PG) algorithm that utilizes proximal policy optimization to make the optimization process more stable and efficient. The PPO algorithm avoids drastic performance fluctuations during the optimization process by limiting the step size of policy updates.
[0077] In this embodiment of the invention, considering the problem of limited resources under interference conditions and the sharing of limited wireless resources by multiple ground users, a time-frequency resource allocation optimization problem is constructed to maximize the total transmission rate of users under the constraints of the maximum transmit power and limited bandwidth of each UAV. A near-end policy optimization algorithm is adopted to solve the time-frequency resource allocation optimization problem. Compared with the dynamic resource allocation algorithm, the resource allocation complexity is reduced. While ensuring the efficiency of policy improvement, the convergence stability is improved. It effectively avoids the dimensionality curse problem that traditional optimization algorithms may face in high-dimensional environments. Compared with the static allocation algorithm, it enhances the anti-interference capability of the communication system and can adapt to the dynamic changes of the UAV base station deployment environment. It can maintain high performance in complex and highly uncertain network environments, is closer to the real situation, and has strong applicability.
[0078] In one implementation, a near-end policy optimization algorithm is used to solve the time-frequency resource allocation optimization problem, resulting in a UAV swarm resource allocation scheme, including:
[0079] The time-frequency resource allocation optimization problem is modeled as an MDP (Markov Decision Process).
[0080] By using a near-end policy optimization algorithm, with the maximum transmit power and limited bandwidth of each UAV as constraints, and the optimization objective of maximizing the total user transmission rate, the MDP is updated to obtain a UAV swarm resource allocation scheme.
[0081] The definitions of state space, action space, and reward in MDP are as follows:
[0082] Each action of the UAV in the action space is mapped to a column of a U×M binary matrix; the element in the u-th row and m-th column of the binary matrix is used to describe the action of the u-th user being assigned to the m-th channel;
[0083] The state space is defined as a one-dimensional array that reflects the current channel allocation state; each element in the one-dimensional array describes the state of a user being allocated to a channel.
[0084] The reward function is defined as the sum of the total transmission rates of all users under the current channel allocation state.
[0085] In this embodiment of the invention, a drone is used as an intelligent agent. The MDP describes the interaction process between the agent and its environment, which is represented by a 4-tuple. definition. The transition probability is represented by [the variable name]. Secondly, the state space under the optimization problem is given. Action space and reward function Detailed definition.
[0086] (1) Action space design
[0087] In this environment, the action space is designed following the form of a discrete action space, taking into account the allocation method between users and channels. A detailed description follows:
[0088] 1) Action definition:
[0089] An action is represented as an integer, and its corresponding binary code represents the user's channel allocation matrix. For example, the size of the action space is 2^32. U×M Where U is the total number of users and M is the total number of channels.
[0090] 2) Range of motion:
[0091] The action space is defined using built-in software functions, and its size is 2. U×M This means that each action can be uniquely mapped to a U×M binary matrix, representing the user's channel allocation status. Specifically, if the element in the u-th row and m-th column of the matrix is 1, it indicates that the u-th user is assigned to the m-th channel.
[0092] 3) Action mapping:
[0093] In this environment, integer actions are mapped to a binary matrix via a function to represent specific channel allocation states. For example, integer actions are encoded in binary, padded with a string of length, and ultimately transformed into a matrix.
[0094] (2) State-space design
[0095] The state space design reflects the current channel allocation state of the system, and its specific design is as follows:
[0096] 1) State definition:
[0097] The state represents the current channel allocation configuration. Specifically, the state is a one-dimensional array of length U×M, where each element is either 0 or 1, indicating whether a user has been allocated a channel. The state space is defined by built-in functions and ranges from [0,1]. U×M The upper and lower limits here represent the binary nature (0 or 1) of the allocation state. For example, if the state is [1,0,0,1,…], it means that the first user and the fourth user have been allocated certain channels respectively.
[0098] 2) State initialization:
[0099] In the `reset` function, the state is initialized by randomly assigning channels. The initial state is a U×M matrix, which is then flattened into a one-dimensional array. The `reset` function is used to reset the state of an object to its default value.
[0100] (3) Reward function design
[0101] The reward function is designed to maximize the system's total data transmission rate, i.e., to encourage efficient channel allocation schemes. The specific design is as follows:
[0102] 1) Reward definition:
[0103] The reward is represented as the sum of the total transmission rates for all users under the current channel allocation state. The transmission rate is calculated using Shannon's theorem:
[0104]
[0105] 2) Signal-to-noise ratio calculation:
[0106] The signal-to-noise ratio (SINR) of a channel is defined as:
[0107]
[0108] 3) Definition of reward function:
[0109] The total reward is the sum of the rates of all users across all channels in all time slots:
[0110]
[0111] In this embodiment of the invention, the specific process of updating the MDP (Multi-Demand Plane) to obtain the drone swarm resource allocation scheme is as follows: The algorithm utilizes a near-end policy optimization method, with constraints on the maximum transmit power and limited bandwidth of each drone, and aims to maximize the total user transmission rate.
[0112] The Policy Gradient (PPO) algorithm is an emerging policy gradient (PG) algorithm. It designs a novel objective function to achieve mini-batch updates, addressing the PG algorithm's sensitivity to step size and difficulty in determining a reasonable step size. PPO originates from the Trust Region Policy Optimization (TRPO) algorithm, introducing a shearing agent objective. Its basic idea is to use importance sampling to measure the difference before and after the policy update, thereby controlling the magnitude of the policy update and ensuring that the new policy remains within the "trust region" of the old policy.
[0113] See Figure 4 , Figure 4 This is a schematic diagram of the PPO algorithm process, which will be explained below. Figure 4 The computational process of the PPO algorithm is explained. The policy gradient algorithm optimizes the policy function π(a|s) into π(a|s;θ) using a parameter θ. Here, a represents the action, and s represents the state. In policy gradient-based methods, the original objective function is:
[0114]
[0115] Where π θ For a random policy, τ=(s1,a1,...,s T ,a T ) is the iterative trajectory with a step size of T. γ represents the cumulative reward of the iteration, and γ is the discount factor.
[0116] Based on the differential J(θ) estimation of the policy gradient, the loss function L is obtained. PG :
[0117]
[0118] Among them, A t It is an estimator of the dominance function at step t; This indicates the operation of seeking the expected value.
[0119] According to the policy gradient method, the parameter update equation can be written as:
[0120]
[0121] Where, θ new Represents new parameters of the policy network; θ old This represents the original strategy network parameters; This indicates the gradient calculation operation; α is the update step size. It can be seen that α directly determines the quality of the new policy. If α is not appropriate, the updated policy will be a worse policy. When learning on a bad policy, the updated policy will become even worse, thus making the entire learning process worse.
[0122] Therefore, to ensure that the new policy's reward function remains monotonically constant and the learning process does not deteriorate, the TRPO (Trust Region Policy Optimization) algorithm is proposed. In the TRPO algorithm, KL divergence (a method for measuring the difference between two probability distributions) is used to add a trust region constraint to the policy gradient method, representing the current policy π. θ (a t |s t ) and old policies The probability ratio New objective function L TRPO (θ) is defined as:
[0123]
[0124] Where ψ(θ) represents the probability ratio.
[0125] The TRPO algorithm is a constrained optimization problem, requiring significant computation to calculate conjugate gradients before constraint optimization can be performed. Therefore, the Proximal Policy Optimization (PPO) algorithm was proposed. The PPO algorithm avoids the computation of conjugate gradients by introducing a penalty for larger policy updates. Finally, the objective function used in this embodiment can be expressed as:
[0126]
[0127] Here, ε is a hyperparameter, and the clip constraint ψ(θ) is within (1-ε, 1+ε) to avoid excessive policy updates.
[0128] In this drone swarm service user system, the drone is treated as an intelligent agent, which includes a policy network π. θ and a value assessment network At each time step t, the agent bases its actions on the current state s. t Select an action a from the environment. t This choice is made through the policy network π. θ Generated.
[0129] When the agent performs action a t Afterwards, the environment returns a reward r. t and transition to the new state. t+1 The conversion t ,a t ,r t ,s t+1 The data is stored in an experience replay buffer. As more experience data accumulates, the agent will randomly draw a batch of samples from the buffer to update the actor network and the critic network, thereby optimizing the policy and value function.
[0130] Actor networks minimize the loss function L PPO To optimize its policy parameters θ, A is calculated using truncated generalized advantage estimation. t .
[0131]
[0132] In the formula, It is the state value function of time slot t; It is the state value function of time slot t+1.
[0133] Critics Network critic A is derived from the empirical average of each batch of samples. t Update. Loss function of the commentator network. Defined as:
[0134]
[0135] The simulation experiment of the intelligent resource scheduling and interference coordination optimization method for UAV swarms provided in the embodiments of the present invention is as follows:
[0136] In this simulation, three regions with an area of 30km × 30km were considered, each equipped with a UAV-BS to provide wireless communication services to users. The path loss model in this environment is the free path loss model.
[0137] The Policy Network (PPO) in the PPO algorithm employs a two-layer feedforward fully connected neural network structure to generate the action probability distribution in the current state. The input layer accepts a state vector with dimensions equal to the number of states, representing the current state information of the environment. The hidden layer contains 16 neurons, achieving non-linear feature mapping through linear transformation and the ReLU activation function. The output layer, with dimensions equal to the number of actions, generates the probability distribution of each action in the action space using the Softmax function, satisfying the probability normalization requirement. The main function of the Policy Network is to realize the mapping from state to action probabilities, guiding the agent's decision-making process in the environment. In the PPO algorithm, this network represents the policy function, and the agent samples actions or selects the optimal action based on the output action probability distribution to interact with the environment.
[0138] The Value Network employs a two-layer feedforward fully connected neural network structure similar to the Policy Network, but its function is to estimate the value of the input state. The input layer accepts a state vector of the same size as the number of states. The hidden layer consists of 16 neurons and uses the ReLU activation function for non-linear feature extraction. The output layer is a single scalar node that directly outputs the scalar value of the current state without requiring an additional activation function. The main function of the Value Network is the mapping from state to state value, i.e., estimating the agent's long-term expected return in the current state. In the PPO algorithm, the Value Network provides a benchmark for policy optimization by calculating state values, used to estimate the Advantage Function, thereby guiding the direction of policy updates and improving the stability and efficiency of training.
[0139] Furthermore, the Adam optimizer was used to train the neural network parameters. The simulation was performed using Python 3.7 and TensorFlow 2.6. Other parameters used in the simulation are given in Table 1.
[0140] Table 1
[0141]
[0142] See Figure 5 , Figure 5This diagram illustrates the initial reward values during the PPO training process, where the X-axis represents the training phase and the Y-axis represents the reward value. In the early stages of training (0–520 rounds), the PPO reward value fluctuates significantly due to the randomness of policy updates during the exploration phase. However, as training progresses, after 520 rounds, the PPO reward gradually increases, eventually stabilizing at around 882.885 Mbps. This result validates the performance advantages of the PPO algorithm in achieving higher rewards and convergence.
[0143] See Figure 6 , Figure 6 This diagram illustrates the deviation of the PPO reward value from the mean, where the X-axis represents the training phase and the Y-axis represents the deviation rate. Initially, due to the exploration process, the deviation value exhibits significant negative fluctuations, indicating unstable reward results. As the number of training rounds increases, the deviation gradually converges and stabilizes within ±3% of the threshold (red dashed line) after approximately 760 rounds. This phenomenon demonstrates that the PPO algorithm gradually converges as training progresses, and the reward value tends to stabilize around the mean.
[0144] See Figure 7 , Figure 7 This is a diagram comparing the original PPO reward values with a 50-round moving average. The X-axis represents the training phase, and the Y-axis represents the reward value. The original PPO reward data fluctuated significantly in the early stages due to the instability introduced during the exploration and learning process. To reveal the overall trend of the data, a 50-round moving average (solid green line) was applied to the original data.
[0145] See Figure 8 , Figure 8 This is a diagram illustrating the Coefficient of Variation (CV). The Coefficient of Variation (CV) is a statistical metric used to measure the ratio of the standard deviation to the mean of a dataset. It assesses the relative dispersion of data and is particularly suitable for comparing the stability of data at different scales or units. The X-axis represents the training phase, and the Y-axis represents the CV value. In the initial stages of training, the CV value is high, indicating that the reward values fluctuate wildly. This is mainly due to the randomness of the exploration process and the instability of policy updates, making the standard deviation relatively large compared to the mean. As the number of training epochs increases, the CV gradually decreases. This indicates that the PPO algorithm gradually converges, and the reward values are stably distributed around a higher mean, with a significant decrease in the standard deviation. At this point, the CV converges to a lower level, reflecting the high stability and low volatility of the PPO training results.
[0146] See Figure 9 , Figure 9 This is a diagram illustrating the dynamics of PPO training and its comparison with greedy algorithms and polling strategies. Figure 9Figure (a) shows the PPO training results compared to other algorithms, and (b) shows the reward values obtained during PPO training, compared with Algorithm 1 (red dashed line represents the greedy algorithm) and Algorithm 2 (green dashed line represents the polling strategy). In the initial stage, the PPO reward fluctuated significantly due to the exploration process. As training progressed, the PPO reward gradually increased and stabilized around 870 Mbps, with an average of 882.885 Mbps, significantly exceeding Algorithm 1 (reward of 680.387 Mbps) and Algorithm 2 (reward of 568.249 Mbps). Figure 9 (b) shows the comparison of moving average reward values, smoothed by taking the moving average over 50 training rounds (green solid line) to smooth the original PPO data (light blue). The results show that, especially after 900 rounds, the PPO reward is basically stable, far exceeding that of Algorithm 1 and Algorithm 2. Figure 9 (c) shows the deviation ratio of the PPO reward value from the mean. Starting from round 750, the deviation stabilizes within the ±3% threshold range (red dashed line), indicating that the PPO reward value gradually converges and the volatility decreases. Figure 9 (d) shows the final performance comparison of the algorithms, comparing the final performance of PPO, Algorithm 1, and Algorithm 2 using a bar chart. The final average reward of PPO is 882.885 Mbps, which is 29.8% higher than Algorithm 1 (680.387 Mbps) and 55.4% higher than Algorithm 2 (568.249 Mbps).
[0147] In summary, the PPO algorithm outperforms the two benchmark algorithms in terms of reward value, and also demonstrates good stability and fast convergence speed, showcasing its superior performance advantages.
[0148] In this embodiment of the invention, addressing the resource allocation problem for UAV-assisted communication, the invention considers the issue of limited resources under interference conditions and the sharing of limited wireless resources by multiple ground users. A joint solution scheme is proposed, constructing a resource allocation problem that maximizes the total rate under the constraints of maximum transmit power and limited bandwidth for each UAV. Furthermore, a solution based on a near-end policy optimization algorithm is employed to solve the joint optimization problem. Moreover, compared to dynamic resource allocation algorithms, this invention reduces resource allocation complexity, improves convergence stability while ensuring policy improvement efficiency, and effectively avoids the curse of dimensionality that traditional optimization algorithms may face in high-dimensional environments. Compared to static allocation algorithms, it enhances the anti-interference capability of the communication system, adapts to the dynamic changes in the UAV base station deployment environment, maintains high performance in complex and highly uncertain network environments, is closer to real-world conditions, and has strong applicability.
[0149] Based on the same inventive concept, embodiments of the present invention also provide an intelligent resource scheduling and interference coordination optimization device for unmanned aerial vehicle (UAV) swarms, see [link to relevant documentation]. Figure 10 , Figure 10 This is a schematic diagram of the structure of an intelligent resource scheduling and interference coordination optimization device for unmanned aerial vehicle (UAV) swarms provided in an embodiment of the present invention, comprising:
[0150] Module 1001 is used to construct a system model for a drone swarm service user system; the system model includes N drones and U users; wherein each drone serves U users. N One user;
[0151] The determining module 1002 is used to determine the channel model and the rate model based on the system model; the channel model is used to represent the free path loss between the UAV and the users it serves; the rate model is used to represent the transmission rate obtained by the user after being assigned a channel.
[0152] The solution module 1003 is used to construct a time-frequency resource allocation optimization problem based on the channel model and the rate model, and to solve the time-frequency resource allocation optimization problem using a near-end policy optimization algorithm to obtain a UAV swarm resource allocation scheme. The time-frequency resource allocation optimization problem is constrained by the maximum transmit power and limited bandwidth of each UAV, and the optimization objective is to maximize the total transmission rate of users. The maximum transmit power means that the total power transmitted by each UAV to the users it serves cannot exceed its own total transmit power. The limited bandwidth means that the total bandwidth transmitted by each UAV to the users it serves cannot exceed its own total bandwidth.
[0153] In this embodiment of the invention, considering the problem of limited resources under interference conditions and the sharing of limited wireless resources by multiple ground users, a time-frequency resource allocation optimization problem is constructed to maximize the total transmission rate of users under the constraints of the maximum transmit power and limited bandwidth of each UAV. A near-end policy optimization algorithm is adopted to solve the time-frequency resource allocation optimization problem. Compared with the dynamic resource allocation algorithm, the resource allocation complexity is reduced. While ensuring the efficiency of policy improvement, the convergence stability is improved. It effectively avoids the dimensionality curse problem that traditional optimization algorithms may face in high-dimensional environments. Compared with the static allocation algorithm, it enhances the anti-interference capability of the communication system and can adapt to the dynamic changes of the UAV base station deployment environment. It can maintain high performance in complex and highly uncertain network environments, is closer to the real situation, and has strong applicability.
[0154] Optionally, the solution module uses a near-end policy optimization algorithm to solve the time-frequency resource allocation optimization problem to obtain a UAV swarm resource allocation scheme, including:
[0155] The time-frequency resource allocation optimization problem is modeled as an MDP.
[0156] Using a near-end strategy optimization algorithm, the MDP is updated with the maximum transmit power and limited bandwidth of each UAV as constraints and the goal of maximizing the total transmission rate of users, to obtain a UAV swarm resource allocation scheme.
[0157] The definitions of state space, action space, and reward in MDP are as follows:
[0158] Each action of the UAV in the action space is mapped to a column of a U×M binary matrix; the element in the u-th row and m-th column of the binary matrix is used to describe the action of the u-th user being assigned to the m-th channel;
[0159] The state space is defined as a one-dimensional array that reflects the current channel allocation state; each element in the one-dimensional array describes the state in which a user is allocated to a channel.
[0160] The reward is defined as the sum of the total transmission rates of all users under the current channel allocation state.
[0161] Optionally, the channel model is:
[0162]
[0163] in, This represents the relationship between the nth drone and the uth drone it serves. n Free path loss between users; f c Indicates the carrier frequency; This represents the nth drone and the uth drone it serves. n The distance between users; c = 3 × 10 8 m / s.
[0164] Optionally, the rate model is:
[0165]
[0166] Among them, R total This represents the total transmission rate of the user; Represents time slot t, the u-th n The transmission rate obtained by each user after being assigned a channel; B represents the total bandwidth of the UAV. In time slot t, the m-th channel is assigned to the u-th channel. n Allocation scheme for individual users; This indicates that the u-th drone serves the nth drone. n Signal-to-noise ratio of each user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U nm = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of UAVs; U n M represents the total number of users served by the nth drone; M represents the total number of channels.
[0167] Optionally, the time-frequency resource allocation optimization problem is:
[0168]
[0169] in, This indicates that the nth drone serves the uth drone. n Transmit power of each user; This represents the relationship between the nth drone and the uth drone it serves. n Channel parameters between users; Indicates the internal interference of the nth UAV in time slot t; Indicates time slot t, where the u-th time slot is served by the nth drone. n Interference from other drones experienced by individual users; δ 2 (t) represents Gaussian white noise; SINR threshold Indicates the signal-to-noise ratio threshold; Indicates the uth n The demand rate of each user; P n B represents the total transmit power of the nth UAV; B represents the total bandwidth of the UAV. In time slot t, the m-th channel is assigned to the u-th channel. n Allocation scheme for individual users; This indicates that the u-th drone serves the nth drone. n Signal-to-noise ratio of each user in time slot t; t = 1, 2, ..., T; n = 1, 2, ..., N; u n =1,2,...,U n m = 1, 2, ..., M; T represents the total number of time slots; N represents the total number of UAVs; U n M represents the total number of users served by the nth drone; M represents the total number of channels. A collection of drones; This represents the set of users served by the nth drone; A set representing channels.
[0170] It should be noted that the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention.
[0171] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0172] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings and the disclosure in carrying out the claimed invention. In the description of the invention, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "a plurality" means two or more, unless otherwise explicitly specified. Furthermore, while different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce good results.
[0173] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0174] It should be noted that the device in this embodiment of the invention is an apparatus that applies the above-described intelligent resource scheduling and interference coordination optimization method for UAV swarms. Therefore, all embodiments of the above-described intelligent resource scheduling and interference coordination optimization method for UAV swarms are applicable to this device and can achieve the same or similar beneficial effects.
[0175] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A method for intelligent resource scheduling and interference coordination optimization of unmanned aerial vehicle (UAV) swarms, characterized in that, include: Construct a system model for a drone swarm service user system; the system model includes drones and Each drone serves one user; One user; Based on the system model, a channel model and a rate model are determined; the channel model represents the free path loss between the UAV and the users it serves; the rate model represents the transmission rate obtained by the users after being assigned a channel. Based on the channel model and the rate model, a time-frequency resource allocation optimization problem is constructed. A near-end policy optimization algorithm is then used to solve this problem, yielding a UAV swarm resource allocation scheme. The time-frequency resource allocation optimization problem is constrained by the maximum transmit power and limited bandwidth of each UAV, with the optimization objective being to maximize the total user transmission rate. The maximum transmit power means that the total power transmitted by each UAV to the users it serves cannot exceed its own total transmit power; the limited bandwidth means that the total bandwidth transmitted by each UAV to the users it serves cannot exceed its own total bandwidth. The time-frequency resource allocation optimization problem is as follows: ; in, Indicates the first The first drone serves the first Transmit power of each user; Indicates the first The drone and the first one it serves Channel parameters between users; Indicates time slot , No. Internal interference with the drone; Indicates time slot , No. The first drone service Interference from other drones experienced by individual users; Indicates Gaussian white noise; Indicates the signal-to-noise ratio threshold; Indicates the first The demand rate of each user; Indicates the first Total launch power of the drones; This indicates the total bandwidth of the drone; Indicates time slot , No. The first channel is assigned to the second... Allocation scheme for individual users; Indicates the first The first drone service Individual users in time slots Signal-to-noise ratio within; ; ; ; ; Indicates the total number of time slots; Indicates the total number of drones; Indicates the first The total number of users served by each drone; Indicates the total number of channels; A collection of drones; Indicates the first The collection of users served by a drone; A set representing channels.
2. The intelligent resource scheduling and interference coordination optimization method for unmanned aerial vehicle (UAV) swarms according to claim 1, characterized in that, The time-frequency resource allocation optimization problem is solved using a near-end policy optimization algorithm to obtain a UAV swarm resource allocation scheme, including: The time-frequency resource allocation optimization problem is modeled as an MDP. Using a near-end strategy optimization algorithm, the MDP is updated with the maximum transmit power and limited bandwidth of each UAV as constraints and the goal of maximizing the total transmission rate of users, to obtain a UAV swarm resource allocation scheme. The definitions of state space, action space, and reward in MDP are as follows: Each action of the drone in the action space is mapped as follows: A column in a binary matrix; the first column in the binary matrix The first line The elements of the column are used to describe the first The user was assigned to the first Actions of each channel; The state space is defined as a one-dimensional array that reflects the current channel allocation state; each element in the one-dimensional array describes the state in which a user is allocated to a channel. The reward is defined as the sum of the total transmission rates of all users under the current channel allocation state.
3. The intelligent resource scheduling and interference coordination optimization method for unmanned aerial vehicle (UAV) swarms according to claim 1, characterized in that, The channel model is as follows: ; in, Indicates the first The drone and the first one it serves Free path loss between users; Indicates the carrier frequency; Indicates the first The drone and the first one it serves The distance between users; .
4. The intelligent resource scheduling and interference coordination optimization method for unmanned aerial vehicle (UAV) swarms according to claim 1, characterized in that, The rate model is as follows: ; in, This represents the total transmission rate of the user; Indicates time slot , No. The transmission rate obtained by each user after being assigned a channel; This indicates the total bandwidth of the drone; Indicates time slot , No. The first channel is assigned to the second... Allocation scheme for individual users; Indicates the first The first drone service Individual users in time slots Signal-to-noise ratio within; ; ; ; ; Indicates the total number of time slots; Indicates the total number of drones; Indicates the first The total number of users served by each drone; This indicates the total number of channels.
5. A device for intelligent resource scheduling and interference coordination optimization of unmanned aerial vehicle (UAV) swarms, characterized in that, include: The building module is used to construct a system model for a drone swarm service user system; the system model includes... drones and Each drone serves one user; One user; The determination module is used to determine the channel model and the rate model based on the system model; the channel model is used to represent the free path loss between the UAV and the users it serves; the rate model is used to represent the transmission rate obtained by the users after being assigned a channel. The solution module is used to construct a time-frequency resource allocation optimization problem based on the channel model and the rate model, and to solve the time-frequency resource allocation optimization problem using a near-end policy optimization algorithm to obtain a UAV swarm resource allocation scheme. The time-frequency resource allocation optimization problem is constrained by the maximum transmit power and finite bandwidth of each UAV, with the optimization objective being to maximize the total user transmission rate. The maximum transmit power means that the total power transmitted by each UAV to the users it serves cannot exceed its own total transmit power; the finite bandwidth means that the total bandwidth transmitted by each UAV to the users it serves cannot exceed its own total bandwidth. The time-frequency resource allocation optimization problem is as follows: ; in, Indicates the first The first drone serves the first Transmit power of each user; Indicates the first The drone and the first one it serves Channel parameters between users; Indicates time slot , No. Internal interference with the drone; Indicates time slot , No. The first drone service Interference from other drones experienced by individual users; Indicates Gaussian white noise; Indicates the signal-to-noise ratio threshold; Indicates the first The demand rate of each user; Indicates the first Total launch power of the drones; This indicates the total bandwidth of the drone; Indicates time slot , No. The first channel is assigned to the second... Allocation scheme for individual users; Indicates the first The first drone service Individual users in time slots Signal-to-noise ratio within; ; ; ; ; Indicates the total number of time slots; Indicates the total number of drones; Indicates the first The total number of users served by each drone; Indicates the total number of channels; A collection of drones; Indicates the first The collection of users served by a drone; A set representing channels.
6. The intelligent resource scheduling and interference coordination optimization device for unmanned aerial vehicle (UAV) swarms according to claim 5, characterized in that, The solution module uses a near-end policy optimization algorithm to solve the time-frequency resource allocation optimization problem, obtaining a UAV swarm resource allocation scheme, including: The time-frequency resource allocation optimization problem is modeled as an MDP. Using a near-end strategy optimization algorithm, the MDP is updated with the maximum transmit power and limited bandwidth of each UAV as constraints and the goal of maximizing the total transmission rate of users, to obtain a UAV swarm resource allocation scheme. The definitions of state space, action space, and reward in MDP are as follows: Each action of the drone in the action space is mapped as follows: A column in a binary matrix; the first column in the binary matrix The first line The elements of the column are used to describe the first The user was assigned to the first Actions of each channel; The state space is defined as a one-dimensional array that reflects the current channel allocation state; each element in the one-dimensional array describes the state in which a user is allocated to a channel. The reward is defined as the sum of the total transmission rates of all users under the current channel allocation state.
7. The intelligent resource scheduling and interference coordination optimization device for unmanned aerial vehicle (UAV) swarms according to claim 5, characterized in that, The channel model is as follows: ; in, Indicates the first The drone and the first one it serves Free path loss between users; Indicates the carrier frequency; Indicates the first The drone and the first one it serves The distance between users; .
8. The intelligent resource scheduling and interference coordination optimization device for unmanned aerial vehicle (UAV) swarms according to claim 5, characterized in that, The rate model is as follows: ; in, This represents the total transmission rate of the user; Indicates time slot , No. The transmission rate obtained by each user after being assigned a channel; This indicates the total bandwidth of the drone; Indicates time slot , No. The first channel is assigned to the second... Allocation scheme for individual users; Indicates the first The first drone service Individual users in time slots Signal-to-noise ratio within; ; ; ; ; Indicates the total number of time slots; Indicates the total number of drones; Indicates the first The total number of users served by each drone; This indicates the total number of channels.
Citation Information
Patent Citations
Unmanned aerial vehicle auxiliary communication anti-interference method based on multi-agent reinforcement learning
CN118921099A
Resource allocation and cooperative unloading method based on federal deep reinforcement learning under multi-unmanned aerial vehicle assisted internet of vehicles
CN119316879A