Distributed unmanned aerial vehicle base station low-altitude wireless signal coverage method

By modeling the drone base station collaborative coverage problem as a decentralized partially observable Markov decision process, and adopting a distributed multi-agent deep reinforcement learning algorithm and a spatiotemporal graph attention collaborative network, the strategy learning problem in drone-assisted wireless communications is solved, and efficient collaborative coverage and system optimization of drone swarms in complex environments are achieved.

CN120751395APending Publication Date: 2025-10-03CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510901657.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing drone-assisted wireless communications have difficulty in effectively learning strategies in some observable scenarios, resulting in slow or difficult convergence. In addition, traditional centralized training and distributed execution frameworks have significant communication overhead and delay problems in complex dynamic environments, affecting algorithm performance.

Method used

The UAV base station collaborative coverage optimization problem is modeled as a decentralized partially observable Markov decision process. A distributed multi-agent deep reinforcement learning algorithm is adopted, and a spatiotemporal graph attention collaborative network (STGAS-Net) is designed for strategy optimization. The network is trained using the training framework of the distributed multi-agent deep reinforcement learning algorithm.

Benefits of technology

Efficient collaborative coverage of drone base stations was achieved in complex dynamic environments, optimizing coverage score, fairness, energy consumption, communication quality, and collision times, thereby improving system robustness and autonomy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751395A_ABST
    Figure CN120751395A_ABST
Patent Text Reader

Abstract

The invention relates to a distributed unmanned aerial vehicle base station low-altitude wireless signal coverage method, and belongs to the technical field of wireless communication. The method comprises the following steps: S1, modeling an unmanned aerial vehicle base station coverage optimization scene and proposing an optimization problem; s2, modeling the optimization problem into a Markov decision process with an observable decentralized part; s3, designing a network structure of a distributed multi-agent deep reinforcement learning algorithm; and S4, designing a training framework of a distributed multi-agent deep reinforcement learning algorithm. The method aims to solve the problem of unmanned aerial vehicle base station collaborative coverage optimization in a complex dynamic environment, and aims to overcome the defects of dependence on a global state, adaptability to a part of observable scenes and the like in a traditional method. According to the invention, an unmanned aerial vehicle base station collaborative coverage optimization problem in a partially observable environment is modeled into a decentralized partially observable Markov decision process, and a network structure and a training framework of a distributed multi-agent deep reinforcement learning algorithm are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a low-altitude wireless signal coverage method for a distributed unmanned aerial vehicle (UAV) base station, and belongs to the technical field of wireless communications. Background Art

[0002] In recent years, unmanned aerial vehicles (UAVs) have become increasingly prominent in wireless communications due to their flexibility, maneuverability, and high potential for establishing Line of Sight (LoS) communication links. Currently, research on UAV-assisted wireless communications focuses on three key areas: using them as relay base stations to assist in building aerial networks, using them as terminals to assist in data collection, and using them as aerial base stations to assist in ubiquitous coverage.

[0003] Existing research on aerial base stations as assisted ubiquitous coverage generally assumes full observability, assuming that drones can obtain global environmental states in real time. Due to the uncertainty and dynamic nature of real-world environments, agents rarely obtain complete state information, leading to research on partially observable scenarios. Partial observability raises issues such as information loss and sparse rewards. Sparse rewards require that specific actions or tasks be completed within a certain time step to receive a reward. This results in sparse reward signals received by agents during exploration, making it difficult to learn optimal strategies and ultimately leading to slow or even impossible convergence.

[0004] Currently, frameworks based on centralized training with decentralized execution (CTDE) are widely used to address partial observability issues in multi-agent systems. These frameworks rely on global state information during training, while each agent makes independent decisions based solely on its own observations during execution, thus reducing reliance on global information. However, in practical applications, drone communications are affected by factors such as distance and obstacles, making it difficult to fully acquire the global state. Furthermore, in dynamic flight environments, the demand for real-time state updates is high, and communication overhead and latency can significantly impact algorithm performance. In contrast, in a fully distributed architecture, each agent independently updates its strategy and selects actions based solely on its own observations, eliminating the need for central coordination. This makes it suitable for scenarios with limited communication or high autonomy. This architecture reduces overhead by eliminating the need for frequent communication, and the failure of a single agent does not affect the overall system operation, significantly improving system robustness. Summary of the Invention

[0005] Aiming at the problem of optimizing the coordinated coverage of UAV base stations in complex dynamic environments, and addressing the shortcomings of traditional methods in terms of dependence on global state and adaptability to partially observable scenarios, this paper models the problem of optimizing the coordinated coverage of UAV base stations in partially observable environments as a decentralized partially observable Markov decision process, and proposes a network structure and training framework for a distributed multi-agent deep reinforcement learning algorithm.

[0006] To achieve the above objectives, the present invention provides a distributed UAV base station low-altitude wireless signal coverage method, which is characterized by mainly comprising the following steps:

[0007] Step 1: Model the drone base station coverage optimization scenario and propose the optimization problem;

[0008] Step 2: Model the optimization problem as a decentralized partially observable Markov decision process;

[0009] Step 3: Design the network structure of the distributed multi-agent deep reinforcement learning algorithm;

[0010] Step 4: Design a training framework for a distributed multi-agent deep reinforcement learning algorithm.

[0011] Furthermore, in step 1, the optimization scenario of drone base station coverage is modeled and the optimization problem is proposed. The specific steps are as follows: the continuous space L×L is divided by the scaling factor s using the equal-spaced gridding method, and the length of each grid is The task cycle is discretized into T time slots, each time slot lasts for l, that is, The scenario contains I UAVs with fixed beam widths and K randomly roaming ground users. The sets of UAVs and users are At time slot t, the position of UAV i is expressed as The position of the kth ground user is expressed as The downtilt angle of the drone base station antenna is 90°-θ. At the same time, a drone can serve multiple users, while a user can only be served by one drone. The positions of ground users follow a non-uniform distribution, and their motion model is described by a random walk process with a maximum velocity constraint. The drone perceives the circular observation range r through its onboard sensors. obs The user information in the data is limited by the quality of service (QoS) requirement for the communication rate. The effective communication coverage of UAV i is constrained to be In order to overcome the problem of incomplete regional information caused by the limited observation range of a single node, the UAVs are connected based on the signal-to-noise ratio threshold γ thBuild a communication topology. Taking into account multiple conditions in the three-dimensional trajectory planning of drones, including boundary restrictions, collision avoidance constraints, communication quality constraints, energy consumption constraints, and kinematic constraints, we systematically optimize various evaluation indicators to achieve coordinated coverage optimization of drone swarms in complex three-dimensional spaces.

[0012] The evaluation indicators of the drone base station coverage mission specifically include: coverage score, fairness score, energy consumption score, system throughput and number of collisions, covering five aspects: coverage, fairness, energy consumption, communication quality and safety, aiming to comprehensively evaluate the coverage mission.

[0013] Furthermore, in step 2, the optimization problem is modeled as a decentralized partially observable Markov decision process (Decentralized POMDP, Dec-POMDP), which is specifically defined as a seven-tuple: in: For the drone set i∈{1,...,I}, is the global state, is the action space, is the observation space, is the state transition probability function, is the local reward function, and γ∈[0, 1] is the discount factor.

[0014] The action space Features: Action Space Includes 13 discrete options: where e * represents a unit direction vector, and 0 represents a zero vector.

[0015] The observation space Features: UAV i observation It includes the following parts: (1) Polar coordinate partition statistical information. In particular, in order to keep the observation dimension consistent at different heights, the polar coordinate partition method is used to count user information. With the drone itself as the center, [0,2π) is evenly divided into K θ sectors, will Evenly divided into K r annular area, we get [K θ ×K r ] polar coordinate grids, count the number of users in each partition, and input the normalized number of users in each partition into the neural network. (2) UAV status: the normalized position of the current UAV Normalized speed Energy consumption The lengths are 3, 3, and 1 respectively. iIt is represented by [x]-bit binary ID. UAV coverage density and observation density: The length is 2 dimensions.

[0016] The local reward function The features include: proposing a heuristic reward function, taking the evaluation indicators into consideration, balancing individual and group rewards, guiding the agents to optimize their own behavior and promote collaboration, and maximizing system performance. It consists of the following components:

[0017]

[0018] in For coverage items, is the fairness item, is the communication throughput term, is the collision penalty term, is the energy consumption item. The Dec-POMDP model uses a distributed strategy Generate actions, and optimize the goal to maximize the expected cumulative reward,

[0019]

[0020] Furthermore, in step 3, the network structure of the distributed multi-agent deep reinforcement learning algorithm is designed, specifically including: proposing a spatio-temporal graph attention synergy network (STGAS-Net). The STGAS-Net network structure includes four modules: spatial feature encoding, graph attention aggregation, temporal state modeling, and action decision decoding. The network structure extracts node features through the spatial encoder, and the graph attention network captures the dependencies between nodes. t Constraining effective connections, GRU processes time series information and finally outputs action values ​​through the decoder.

[0021] The adjacency matrix A t Specifically, in the dynamic UAV network topology modeling, the Flying Self-Organizing Network (FANET) architecture is used to describe the interaction between UAVs. FANET achieves autonomous node collaboration through a decentralized communication mechanism, and its topology can be formally represented as an unweighted undirected graph G t = {V, E t}, where V represents the set of drone nodes, is a set of communication links, where Represents the communication signal-to-noise ratio between UAV i and UAV j. Adjacency matrix A based on channel perception t ∈{0, 1}I×I By integrating the geometric relationship of three-dimensional space with the characteristics of wireless channels, it can meet the following requirements:

[0022]

[0023] Furthermore, in step 4, a training framework for a distributed multi-agent deep reinforcement learning algorithm is designed, specifically including: during the interaction process, each drone generates a state observation sequence through environmental interaction. After the network encoder generates an abstract state representation, the action is selected through the ε-greedy strategy The transfer tuple (o t , h t , A t , a t , r t , o t+1 , h t+1 , A t+1 ) will be stored in the experience pool, where h represents the GRU hidden state. During training, the system periodically randomly samples batches of data from the buffer pool. The current network Q processes the spatiotemporal state features using the STGAS-Net architecture to calculate the current Q value distribution. The Adam optimizer is used to train the neural network, minimizing the TD error using the mean squared error loss function. The target network Q′ is maintained stable through a delayed parameter update mechanism, and its parameters θ′ are periodically synchronized through hard updates.

[0024] The beneficial effects of the present invention are:

[0025] Aiming at the problem of optimizing the coordinated coverage of UAV base stations in complex dynamic environments, and addressing the shortcomings of traditional methods in terms of dependence on global state and adaptability to partially observable scenarios, this paper models the problem of optimizing the coordinated coverage of UAV base stations in partially observable environments as a decentralized partially observable Markov decision process, and proposes a network structure and training framework for a distributed multi-agent deep reinforcement learning algorithm.

[0026] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings.

[0028] A detailed description of the selection, including:

[0029] Figure 1 This is a flow chart of the low-altitude wireless signal coverage method of the distributed UAV base station of the present invention;

[0030] Figure 2 This is a schematic diagram of the network structure of the distributed multi-agent deep reinforcement learning algorithm of the present invention;

[0031] Figure 3 Schematic diagram of the training framework of the distributed multi-agent deep reinforcement learning algorithm described in the present invention;

[0032] Figure 4 This is a heat map of coverage strength of the algorithm described in the embodiment of the present invention in a simulation environment;

[0033] FIG5 is a diagram showing the experimental results provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following describes the embodiments of the present invention through specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0035] See also Figures 1 to 3 , which is the distributed UAV base station low-altitude wireless signal coverage method provided by the present invention.

[0036] Figure 1 The flowchart of the method for three-dimensional coverage of multiple UAV base stations in a partially observable environment of the present invention mainly includes the following steps:

[0037] Step 1: Model the drone base station coverage optimization scenario and propose the optimization problem;

[0038] Step 2: Model the optimization problem as a decentralized partially observable Markov decision process;

[0039] Step 3: Design the network structure of the distributed multi-agent deep reinforcement learning algorithm;

[0040] Step 4: Design a training framework for a distributed multi-agent deep reinforcement learning algorithm.

[0041] Furthermore, in step 1, the drone base station coverage optimization scenario is modeled and the optimization problem is proposed, which specifically includes: considering the drone coverage task under partially observable environment, establishing Figure 1The scene model shown uses an equal-spaced gridding method to divide the continuous space L×L according to the scaling factor s. The length of each grid is The task cycle is discretized into T time slots, each time slot lasts for l, that is, The scenario contains I UAVs with fixed beam widths and K randomly roaming ground users. The sets of UAVs and users are At time slot t, the position of UAV i is expressed as The position of the kth ground user is expressed as The downtilt angle of the drone base station antenna is 90°-θ. In order to comprehensively evaluate the system performance, five global indicators are used, namely, coverage score C t , Fairness score F t , energy consumption fraction E t , communication index R t and the collision index χ t Under the constraints of boundary, collision, communication quality and motion, the three-dimensional flight trajectory of the UAV is optimized. represents the position of UAV i at time t. The optimization problem can be expressed as,

[0042]

[0043] Where λ n ∈[0,1], n=1,2,...,5, is the weight coefficient of each indicator, satisfying ∑λ n = 1. Constraint C1 is about the restriction of the range of UAV activities, and constraint C2 is about the setting of UAV collision avoidance. The positions of any two UAVs are greater than the collision threshold D col ,Constraint C3 requires that the communication rate between the UAV and the service user is greater than R QoS , constraint C4 is the association variable between the drone and the user Constraints,Constraint C5 ensures that any user is associated with only one drone at any,time.,Constraint C6 is about the limits on acceleration and,speed.

[0044] Furthermore, in step 2, the optimization problem is modeled as a decentralized partially observable Markov decision process, specifically including: since each UAV can only obtain local observations, the observation range and communication range are limited, the coverage task is modeled as a Dec-POMDP, which is specifically defined as a seven-tuple: in: For the drone set i∈{1,...,I}, is the global state, is the action space, is the observation space, is the state transition probability function, is the local reward function, and γ∈[0, 1] is the discount factor.

[0045] For i∈{1,...,I},k∈{1,...,K},t∈{1,...,T}, the global state Include:

[0046] (1) Drone status: location speed Energy consumption

[0047] (2) User status: location speed

[0048] (3) Adjacency matrix A t .

[0049] In the dynamic UAV network topology modeling, the Flying Ad-hoc Network (FANET) architecture is used to describe the interaction between UAVs. FANET achieves autonomous node collaboration through a decentralized communication mechanism, and its topology can be formally represented as an unweighted undirected graph G t ={V,E t}, where V represents the set of drone nodes, is the set of communication links. Adjacency matrix A based on channel perception t ∈{0,1} I×I By integrating the geometric relationship of three-dimensional space with the characteristics of wireless channels, it can meet the following requirements:

[0050]

[0051] Action Space Includes 13 discrete options: where e * represents a unit direction vector, and 0 represents a zero vector.

[0052] Local observation by drone i include:

[0053] (1) Polar coordinate partitioning statistics of user information. In particular, in order to keep the observation dimensions of different heights consistent, the polar coordinate partitioning method is used to count user information. With drone i as the center, [0, 2π) is evenly divided into K θ sectors, will Evenly divided into K r annular area, we get [K θ ×K r ] polar coordinate grids, counting the number of users in each partition μ m, the partition normalized user number set of drone i is expressed as

[0054] (2) The drone's own status. The current normalized position of the drone Normalized speed Energy consumption The lengths are 3, 3, and 1 respectively. i It is represented by [x]-bit binary ID. UAV coverage density and observation density: The length is 2 dimensions.

[0055] Aggregate the above information into observations of drone i Expressed as:

[0056]

[0057] Among them, Concat(·) represents the feature concatenation operation, The total dimension D is [K θ ×K r +x+9].

[0058] Due to partial observability, rewards are obtained based on local states. To quantify the contribution of a single drone, this section proposes a multi-dimensional heuristic reward mechanism to balance individual contribution, group collaboration, communication efficiency, and fairness. Specifically, the instantaneous reward for each drone i is It consists of the following components:

[0059] (1) Coverage status

[0060] The zero coverage penalty mechanism is used to avoid the lazy behavior of the agent, while the effective coverage number is linearly rewarded.

[0061]

[0062] The connected subgraph to which drone i belongs The collaboration reward is expressed as:

[0063]

[0064] The above formula is |G i |>1, where β group Balance weights for group collaboration.

[0065] The coverage reward is expressed as:

[0066]

[0067] (2) Fairness

[0068] The coverage of low-coverage users is tilted, and its reward function is expressed as:

[0069]

[0070] in is the coverage rate of user k at time t, given by calculate.

[0071] The reward function for collaborative compensation for users with low coverage covered by other drones in the group is expressed as:

[0072]

[0073] The fairness reward is expressed as:

[0074]

[0075] (3) Communication throughput item

[0076] Normalized communication throughput reward, expressed as:

[0077]

[0078] in is the total throughput of UAV i,

[0079] (4) Collision penalty:

[0080]

[0081] Among them, β collision is the collision penalty coefficient.

[0082] (5) Energy consumption:

[0083]

[0084] in, is the cumulative energy consumption of UAV i,

[0085] Combining the above components, the total reward of drone i at time t is:

[0086]

[0087] The above Dec-POMDP model uses a distributed strategy Generate actions, and the optimization goal is to maximize the expected cumulative reward, expressed as

[0088]

[0089] Furthermore, in step 3, the network structure of the distributed multi-agent deep reinforcement learning algorithm is designed, specifically including: the proposed Spatio-Temporal Graph Attentive Synergy Network (STGAS-Net) is used for the coordinated coverage control of the drone swarm in a partially observable environment. The network structure includes four modules: spatial feature encoding, graph attention aggregation, temporal state modeling, and action decision decoding. Let H be the hidden layer dimension, A be the action dimension, B be the batch, D be the observation dimension, and M be the number of attention heads. Figure 2 Demonstrates the hierarchical design of the network structure.

[0090] Input observation vector Extract high-dimensional features through spatial encoder and introduce position embedding Enhanced spatial perception, expressed as:

[0091] h spatial =Encoder(o t )+P. (18)

[0092] Where Encoder(·) is a two-layer linear transformation, It is a globally shared location embedding that is extended to all nodes through a broadcast mechanism.

[0093] A multi-head graph attention mechanism is used to aggregate neighborhood information. For each attention head m∈{1, 2, ..., M}, let For a single head dimension, the latent features are projected into the M head space, expressed as:

[0094]

[0095] in is the linear projection parameter. Using the decomposition attention parameterization method, the attention coefficient of drone node i, j in the attention head m is defined as:

[0096]

[0097] in is the attention parameter vector of head m, decomposed into left and right components Adjacency matrix A t Constrain effective connectivity, expressed as:

[0098]

[0099] Perform feature aggregation after normalization:

[0100]

[0101] Furthermore, the multi-head outputs are concatenated and fused through a linear layer:

[0102]

[0103] in The output projection matrix is ​​used to integrate multi-head information. The residual connection is introduced to enhance gradient propagation:

[0104] h attn ←h attn +W skip h spatial (twenty four)

[0105] in Projection matrix for skip connections.

[0106] GRU is used to model temporal dependencies in dynamic environments and convert spatial feature sequences into temporal processing formats:

[0107]

[0108] Based on the current observation t Update hidden state:

[0109]

[0110] The hidden state Encode historical observation information to achieve implicit reasoning from partially observable to global state. The decoder finally outputs the Q-value function estimate Q(o t , h t-1 ).

[0111] The STGAS-Net network structure extracts node features through the spatial encoder, the graph attention network captures the dependencies between nodes, the GRU processes the time series information, and finally outputs the action value through the decoder.

[0112] Furthermore, in step 4, a training framework for a distributed multi-agent deep reinforcement learning algorithm is designed, specifically including: the training framework for a distributed multi-agent deep reinforcement learning algorithm designed by the present invention is as follows: Figure 3 During the interaction process, each UAV generates a state observation sequence through environmental interaction After the network encoder generates an abstract state representation, the action is selected through the ε-greedy strategy Obtained by the following formula:

[0113]

[0114] Where ∈ is the exploration rate, which is set to decay exponentially from ∈0 to ∈ min The formula is:

[0115]

[0116] where n episode Indicates the number of training epochs. In the early stages of training, random actions are encouraged to explore the environment, and as training progresses, the Q-value maximizing action is gradually favored.

[0117] After the drone performs an action based on the current observation and hidden state, the generated transfer tuple (o t , h t , A t , a t , r t , o t +1 , h t+1 , A t+1 ) will be stored in the experience pool.

[0118] During training, the system periodically randomly samples batches of data from the buffer pool. The current network Q is processed through the STGAS-Net architecture to process the spatiotemporal state features and calculate the current Q value distribution. The Adam optimizer is used to train the neural network, and the mean squared error loss function is used to minimize the TD error:

[0119]

[0120] The target network Q′ maintains stability through a delayed parameter update mechanism, and its parameters θ′ are periodically synchronized via θ′←θ. Table 1 shows the network and parameter settings during training.

[0121] Table 1 Network training parameter settings

[0122]

[0123] like Figure 4 The figure shows a heat map of coverage intensity of the algorithm described in the embodiment in a simulation environment. The coverage task was performed for 100 steps in an environment with 10 drones and 100 users. The analysis results reveal the following key features: (1) In terms of spatial correlation, the UE cluster area is effectively incorporated into the coverage network, and the coverage duration of its surrounding coordinate points generally exceeds 80 time steps, indicating that the drone successfully locks on the user activity hotspot area through dynamic path planning. (2) In terms of temporal continuity, the drone forms a high coverage duration band (light-colored area) where UEs are concentrated in the area. The gradient from dark to light in the figure corresponds to the user distribution density from sparse to dense, which intuitively confirms the drone's optimization ability in spatial resource allocation.

[0124] Figure 5 shows the experimental results of the embodiment of the present invention. The simulation was performed using Pytorch 1.4.0 and Python 3.8. Consider an area of ​​500m×500m. At the initial moment, users are distributed in the service area according to certain rules. The initial position of the drone is at any position on a plane with a height of 50m. The maximum flight altitude is 150m, the initial speed is 0m / s, the collision range of the drone is 10m, and the drone observation angle is partitioned into K. θ =8, UAV observation radius partition K r =5, communication rate threshold R QoS =200kbps, signal-to-noise ratio threshold γ th =10dB, the maximum acceleration of the drone is a max =5m / s 2 , the maximum speed of the drone v max =20m / s. Figure 5(a) is the reward curve, Figure 5(b) is the coverage score curve, Figure 5(c) is the fairness score curve, Figure 5(d) is the energy consumption score curve, Figure 5(e) is the system throughput curve, and Figure 5(f) is the collision count curve. The compared algorithms include:

[0125] (1) Deep Q-Network (DQN): A distributed decision-making framework based on the Independent Q-Learning (IQL) architecture.

[0126] (2) Deep Recurrent Q-Network (DRQN): For some observable scenarios, it introduces RNN to process action-observation history sequences to enhance the agent's temporal reasoning ability.

[0127] (3) Deep Graph Network (DGN): Models the interaction relationship between drones based on the graph attention mechanism and uses the feature aggregation of neighborhood nodes to achieve collaborative strategy learning.

[0128] Figure 5(a) shows the dynamic evolution of the global reward of different deep reinforcement learning algorithms over the number of training epochs. As training progresses, the STGAS-Net algorithm achieves rapid reward growth in the early stages and maintains a stable, peak level in the later stages, with relatively small fluctuations, demonstrating excellent stability. The DQN, DRQN, and DGN algorithms experience more limited reward growth and significant curve fluctuations, reflecting the instability of their strategies during training.

[0129] Figures 5(b), 5(c), and 5(d) respectively show the dynamic evolution of coverage, fairness, and energy consumption scores with the number of training steps. In Figure 5(b), the STGAS-Net algorithm rapidly rises and stabilizes at a high level, significantly outperforming the DQN and DRQN ​​algorithms. In Figure 5(c), the STGAS-Net algorithm curve remains high, demonstrating its ability to effectively optimize resource allocation balance and improve fairness during training. In Figure 5(d), the STGAS-Net algorithm maintains a reasonable range, ensuring coverage and fairness while avoiding excessive energy consumption, highlighting the advantages of multi-dimensional joint optimization. In contrast, other comparison algorithms exhibit limited performance gains or significant fluctuations, making it difficult to achieve such coordinated performance. This fully demonstrates the superior effectiveness of the proposed algorithm for multi-objective optimization in complex scenarios.

[0130] Figure 5(e) shows that as training progresses, the system throughput of each algorithm increases and eventually stabilizes. STGAS-Net's system throughput curve remains consistently high, demonstrating a more robust growth process. Although numerically second only to the DQN algorithm, it should be noted that the DQN solution has a lower coverage score, which allows users to obtain more bandwidth and power resources allocated by the drone base station.

[0131] As shown in Figure 5(f), the number of collisions for each algorithm decreases as training progresses. STGAS-Net's collision count decreases rapidly in the early stages of training and subsequently remains stable at an extremely low level, demonstrating the algorithm's obstacle avoidance capabilities and collaborative mechanism in dynamic environments. This result further validates that the proposed algorithm, through its multi-dimensional joint optimization framework, effectively balances coverage, fairness, energy consumption, system throughput, and collision risk, highlighting the synergistic advantages of three-dimensional trajectory optimization and adaptive decision-making under partial observability.

Claims

1. A distributed UAV base station low-altitude wireless signal coverage method, characterized in that: The following steps are involved: Step 1: Model the drone base station coverage optimization scenario and propose the optimization problem; Step 2: Model the optimization problem as a decentralized partially observable Markov decision process; Step 3: Design the network structure of the distributed multi-agent deep reinforcement learning algorithm; Step 4: Design a training framework for a distributed multi-agent deep reinforcement learning algorithm.

2. The distributed UAV base station low-altitude wireless signal coverage method according to claim 1 is characterized in that: In step 1, the optimization scenario of drone base station coverage is modeled and the optimization problem is proposed. Specifically, the continuous space L×L is divided by the scaling factor s using the equidistant gridding method, and the length of each grid is The task cycle is discretized into T time slots, each time slot lasts for l, that is, The scenario contains I UAVs with fixed beam widths and K randomly roaming ground users. The sets of UAVs and users are At time slot t, the position of UAV i is expressed as The position of the kth ground user is expressed as The downtilt angle of the drone base station antenna is 90°-θ. At the same time, a drone can serve multiple users, while a user can only be served by one drone. The positions of ground users follow a non-uniform distribution, and their motion model is described by a random walk process with a maximum velocity constraint. The drone perceives the circular observation range r through its onboard sensors. obs User information in, limited by the QoS requirement for communication rate, the effective communication coverage of UAV i is constrained to be In order to overcome the problem of incomplete regional information caused by the limited observation range of a single node, the UAVs are connected based on the signal-to-noise ratio threshold γ th Build a communication topology. Taking into account multiple conditions in the three-dimensional trajectory planning of drones, including boundary restrictions, collision avoidance constraints, communication quality constraints, energy consumption constraints, and kinematic constraints, we systematically optimize various evaluation indicators to achieve coordinated coverage optimization of drone swarms in complex three-dimensional spaces.

3. The evaluation indicator for the UAV base station coverage task according to claim 2 is characterized by: The evaluation indicators include coverage score, fairness score, energy consumption score, system throughput and number of collisions, covering five aspects: coverage, fairness, energy consumption, communication quality and security, aiming to comprehensively evaluate the coverage task.

4. The distributed UAV base station low-altitude wireless signal coverage method according to claim 1 is characterized in that: In step 2, the optimization problem is modeled as a decentralized partially observable Markov decision process (DecentralizedPOMDP, Dec-POMDP), which is specifically defined as a seven-tuple: in: For the drone set i∈{1,...,I}, is the global state, is the action space, is the observation space, is the state transition probability function, is the local reward function, and γ∈[0, 1] is the discount factor.

5. The action space according to claim 4 Its characteristics are: Action Space Includes 13 discrete options: where e * represents a unit direction vector, and 0 represents a zero vector.

6. The observation space according to claim 4 Its characteristics are: Observation by drone i It includes the following parts: (1) Polar coordinate partition statistical information. In particular, in order to keep the observation dimension consistent at different heights, the polar coordinate partition method is used to count user information. With the drone itself as the center, [0, 2π) is evenly divided into K θ sectors, will Evenly divided into K r annular area, we get [K θ ×K r ] polar coordinate grids, count the number of users in each partition, and input the normalized number of users in each partition into the neural network. (2) UAV status: the current UAV’s normalized position, normalized speed, energy consumption, UAV identity code, UAV coverage density, and observation density.

7. The local reward function according to claim 4 It is characterized in that Specifically include: A heuristic reward function is proposed to take the above evaluation indicators into consideration and balance individual and group rewards, guiding the agents to optimize their own behavior and promote collaboration to maximize system performance. for, in, For coverage items, is the fairness item, is the communication throughput term, is the collision penalty term, The Dec-POMDP model adopts a distributed strategy Generate actions, and the optimization goal is to maximize the expected cumulative reward.

8. The distributed UAV base station low-altitude wireless signal coverage method according to claim 1 is characterized in that: In step 3, the network structure of the distributed multi-agent deep reinforcement learning algorithm is designed, specifically including: proposing a spatio-temporal graph attention synergy network (STGAS-Net) for the coordinated coverage control of drone swarms in partially observable environments. The network structure includes four modules: spatial feature encoding, graph attention aggregation, temporal state modeling, and action decision decoding. The STGAS-Net network structure extracts node features through the spatial encoder, and the graph attention network captures the dependencies between nodes. t Constraining effective connections, GRU processes time series information and finally outputs action values ​​through the decoder.

9. The adjacency matrix A according to claim 8 t , characterized in that, Specifically include: In the dynamic UAV network topology modeling, the Flying Ad-Hoc Network (FANET) architecture is used to describe the interaction between UAVs. FANET achieves autonomous node collaboration through a decentralized communication mechanism, and its topology can be formally represented as an unweighted undirected graph G t ={V,E t }, where V represents the set of drone nodes, is a set of communication links, where Represents the communication signal-to-noise ratio between UAV i and UAV j. Adjacency matrix A based on channel perception t ∈{0,1} I×I It is constructed by integrating three-dimensional spatial geometric relationships with wireless channel characteristics.

10. The distributed UAV base station low-altitude wireless signal coverage method according to claim 1 is characterized in that: In step 4, a training framework for a distributed multi-agent deep reinforcement learning algorithm is designed, specifically including: during the interaction process, each drone generates a state observation sequence through environmental interaction After the network encoder generates an abstract state representation, the action is selected through the ε-greedy strategy The transition tuple generated by executing the action will be stored in the experience pool, where h represents the GRU hidden state. During training, the system periodically randomly samples batches of data from the buffer pool. The current network Q processes the spatiotemporal state features using the STGAS-Net architecture to calculate the current Q value distribution. The Adam optimizer is used to train the neural network, minimizing the TD error using the mean squared error loss function. The target network Q′ is maintained stable through a delayed parameter update mechanism, and its parameters θ′ are periodically synchronized through hard updates.

Citation Information

Cited By

  • Distributed unmanned aerial vehicle measurement and control service system and method

    CN121722133A