Access control and trajectory planning method and device for emergency communication network
By constructing a limited sensing, channel, and communication model for an emergency communication network, and employing a QMIX trajectory planning algorithm enhanced with priority access control and weight generation network, the problems of complex communication requirements and poor convergence in multi-UAV systems are solved, achieving efficient emergency communication resource management.
Patent Information
- Application Number
- CN202510044078.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-01-10
AI Technical Summary
In emergency communication scenarios, multi-UAV systems face challenges such as high demand for differentiated communication, computational complexity, and poor convergence. In particular, in large-scale networks, existing multi-agent deep reinforcement learning algorithms struggle to effectively optimize UAV trajectories and communication resources.
A finite sensing model, channel model, and communication model for an emergency communication network are constructed. The model is described as a deneutralized partially observable Markov process through a joint optimization model. The model is solved using a priority-based access control algorithm and a QMIX trajectory planning algorithm based on weighted generation network enhancement, thereby optimizing the differentiated communication requirements of uplink and downlink.
It achieves efficient communication and resource management in emergency communication scenarios, reduces the state action space, improves convergence and training efficiency, and can quickly converge to better communication and resource management results.
Smart Images

Figure CN119854769B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of emergency communication technology for unmanned aerial vehicles (UAVs), and in particular to an access control and trajectory planning method and apparatus for emergency communication networks. Background Technology
[0002] Unmanned aerial vehicles (UAVs) are an indispensable component of future mobile networks, designed to establish seamless 3D communication. UAVs can serve as aerial base stations, extending network coverage and capacity. Compared to traditional ground base stations, UAV base stations offer high mobility, flexible deployment, and low cost, making them ideal for emergency communications in disaster relief and other emergencies. When ground infrastructure is paralyzed, UAVs can act as communication relays between remote control centers and ground nodes, providing reliable and low-cost communication services to ground nodes.
[0003] In disaster relief scenarios, ensuring Quality of Service (QoS) in communications is crucial, as it directly impacts the progress of rescue operations. Currently, much work focuses on optimizing achievable data rates or latency. However, in some specific scenarios, evaluating the performance of the entire network from a single dimension is insufficient. For example, in disaster and rescue scenarios, ground robots can act as ground nodes to perform specific tasks. Equipped with sensors such as cameras and radar, ground robots can replace humans in tasks such as disaster area monitoring and life detection. Due to the destruction of ground infrastructure, drone swarms are often used as communication relays to transmit operational information between remote control centers and ground robots. In this scenario, ground robots transmit collected environmental information back to the control center via UAVs. Simultaneously, UAVs also need to forward control commands issued by the control center to the ground robots. Therefore, ensuring QoS between the remote control center and the ground robots is critical for disaster relief. Specifically, the link from the ground robot to the UAV needs to provide sufficient channel capacity to transmit environmental information, while the link from the UAV to the ground robot needs to ensure the timeliness and reliability of control information. Therefore, simultaneously ensuring the differentiated communication needs of the uplink and downlink is essential.
[0004] However, the communication resources and coverage of a single UAV are limited. To overcome these limitations, multiple UAVs are employed to collaboratively provide services to ground nodes, thereby expanding network coverage and improving communication capabilities. Unlike a single UAV, using the same spectrum resources with multiple UAV base stations introduces additional interference. UAV location, communication resource allocation, and node scheduling can be adjusted to improve the quality of service to ground nodes. Therefore, joint optimization of UAV trajectories, communication resources, and node scheduling is necessary. However, this problem exhibits high complexity. Traditional optimization-based methods rely on global information to calculate UAV trajectories and resource allocation schemes offline. However, this approach has significant limitations. On the one hand, obtaining global information in real time is impractical in disaster scenarios where existing ground infrastructure is destroyed. On the other hand, when node locations change, complex solution processes need to be repeatedly calculated, resulting in significant time overhead.
[0005] Fortunately, the emergence of reinforcement learning has brought more possibilities to multi-UAV assisted communication. Multi-agent deep reinforcement learning (MADRL) allows each UAV to act as an independent decision-maker to adapt to the distributed needs of wireless networks. Currently, many researchers have applied MADRL to the joint optimization problem of UAV trajectory and communication. Although most MADRL algorithms show satisfactory performance in small-scale networks, the state-action space grows exponentially with the number of UAVs and nodes as the number of network nodes increases. Furthermore, due to the limited perception capabilities of UAV agents, each agent typically observes locally. The nodes they observe change at different times, leading to poorer convergence of MADRL. Existing MADRL algorithms require a large number of empirical samples to achieve convergence in large-scale scenarios when solving the joint optimization problem of UAV trajectory and communication. Summary of the Invention
[0006] Based on this, it is necessary to provide an access control and trajectory planning method and device for emergency communication networks to address the technical problems of high differentiated communication requirements, computational complexity and poor convergence in the joint optimization problem of uplink and downlink under the condition of limited frequency resources in the above-mentioned emergency communication scenarios, so as to achieve efficient communication and resource management in emergency communication scenarios.
[0007] An access control and trajectory planning method for emergency communication networks, the method comprising:
[0008] Construct an emergency communication network assisted by multiple UAVs and obtain the uplink and downlink information within the network; the emergency communication network consists of multiple UAVs, multiple ground nodes, and a remote control center;
[0009] Construct a limited perception model, channel model, and communication model between the UAV and ground nodes based on the emergency communication network;
[0010] By jointly considering the uplink's throughput requirements and the downlink's timeliness and reliability requirements, and based on the finite sensing model, channel model, and communication model, a joint optimization model for the access control and UAV trajectory planning problem is constructed.
[0011] The joint optimization model is described as a deneutralized partially observable Markov process. A priority-based access control algorithm and a QMIX trajectory planning algorithm based on weighted generative networks are used to solve the joint optimization model to obtain the optimal access decision and the optimal trajectory of the UAV.
[0012] In one embodiment, in the emergency communication network, each ground node is used to collect and acquire environmental information and transmit it to a remote control center. The remote control center is used to make decisions based on the received environmental information and send command and control information, including movement control and mission instructions, to the ground nodes. Each UAV acts as an airborne base station and is used to forward the transmitted information between the ground nodes and the remote control center. The uplink in the network is defined as the link for transmitting environmental information from the ground nodes to the UAVs, and the downlink is defined as the link for transmitting command and control information from the remote control center to the ground nodes.
[0013] In one embodiment, the finite perception model is described using perception variables between the UAV and the ground node, where the perception variable between the UAV m and the ground node u at time slot t is defined as c. m,u (t), and c m,u (t)∈{0,1}; where, if ground node u is within the perception range of UAV m at time slot t, c m,u (t) = 1; otherwise c m,u (t) = 0.
[0014] In one embodiment, the channel model is described using the channel gain between the ground node and the UAV. The calculation process for the channel gain is as follows:
[0015] The probabilities of line-of-sight (LoS) and non-line-of-sight (NLoS) links between ground node u and UAV m are calculated based on the pitch angle from ground node u to UAV m at time slot t, and are expressed as follows:
[0016]
[0017] in, This represents the probability of a Loss of Sight (LoS) link between ground node u and drone m at time slot t. This represents the probability of an NLoS link between ground node u and UAV m at time slot t. 'a' and 'b' are parameters determined by the environment type. The pitch angle from ground node u to UAV m at time slot t represents the fixed altitude of the UAV flight, and q represents the pitch angle from ground node u to UAV m. u Let q be the position of ground node u. m (t) represents the position of UAV m in time slot t;
[0018] The path loss between ground node u and drone m is calculated based on the positions of ground node u and drone m at time slot t, and is expressed as follows:
[0019]
[0020] Where, η ξ f represents the average additional path loss of a LosS link or an NLoS link. c ξ is the carrier frequency, c is the speed of light, and ξ represents a LosS link or an NLoS link.
[0021] according to The path loss for each link is calculated to obtain the channel gain between ground node u and UAV m, expressed as:
[0022]
[0023] in, This represents the path loss (LoS) of the link between ground node u and UAV m at time slot t. This represents the path loss of the NLoS link between ground node u and UAV m at time slot t.
[0024] In one embodiment, the communication model includes:
[0025] For the uplink, the communication model is described using the uplink signal-to-interference-plus-noise ratio (SINR), uplink data rate, and throughput; where the uplink SINR from ground node u to UAV m at time slot t is... Represented as
[0026]
[0027] Among them, g u,m (t) represents the channel gain between ground node u and UAV m at time slot t, g u′,m (t) represents the channel gain between another ground node u′ and the UAV m at time slot t, α u,m (t) represents the access control variable between ground node u and UAV m at time slot t, α u′,m′ (t) represents the access control variable between another ground node u′ and another UAV m′ at time slot t, n0 represents the power spectral density of additive white Gaussian noise, and p uLet B be the transmit power of ground node u on each subchannel, B be the subchannel bandwidth, M be the number of UAVs, and U be the number of ground nodes.
[0028] Uplink data rate from ground node u to UAV m at time slot t Represented as
[0029]
[0030] The total throughput of ground node u and the entire emergency communication network over the past t time slots are respectively expressed as: and Where t' = {1, 2, ..., t} is any one of the past t time slots, and δ is the time slot length; meanwhile, the average throughput of ground node u over the past t time slots is expressed as:
[0031] For the downlink, the communication model is described by the downlink signal-to-interference-plus-noise ratio (SINR), downlink data rate, and decoding error probability; where the downlink SINR from UAV m to ground node u at time slot t is... Represented as
[0032]
[0033] Where, α m,u (t) represents the access control variable between UAV m and ground node u at time slot t, α m',u' (t) represents the access control variable between another UAV m′ and another ground node u′ at time slot t, g m,u (t) represents the channel gain between the UAV m and the ground node u at time slot t, g m',u (t) represents the channel gain between another UAV m′ and ground node u at time slot t, p m Let m be the transmit power of the UAV on each sub-channel. For drones, For the set of ground nodes;
[0034] Downlink data rate from UAV m to ground node u in time slot t Represented as
[0035]
[0036] Where τ is the transmission delay of the command and control packet (the packet containing command and control information), and Q... -1 (·) is the inverse function of the Q function, ε max The maximum decoding error probability is given by V, which represents the channel dispersion, specifically expressed as:
[0037] When UAV m transmits an instruction / command packet to ground node u in time slot t, the decoding error probability of ground node u is expressed as:
[0038]
[0039] Where, τ max Let ε represent the maximum tolerable delay, S represent the given packet length, and ε represent the maximum tolerable delay. u (t)≤ε max .
[0040] In one embodiment, the joint optimization model for the access control and UAV trajectory planning problem is expressed as follows:
[0041]
[0042] The joint optimization model is a mixed-integer nonlinear programming problem consisting of continuous variable q and discrete variable α; where λ is a weighting coefficient used to balance the relative importance of uplink and downlink. This indicates the service period of the ground node for the drone. Total fair throughput, which takes into account the uplink's communication requirements for throughput; service cycle. It is divided into T equal-length time slots, each time slot having a length of δ; This represents the fairness coefficient at time slot t. Let U be the average throughput of ground node u over the past t time slots, where U represents the number of ground nodes; This represents the uplink data rate of ground node u at time slot t; This represents the number of downlinks that meet QoS (QoS-DL), which takes into account the timeliness and reliability requirements of downlink communication; 1(·) represents the indicator function, specifically expressed as ε max ε represents the maximum decoding error probability. u (t) represents the decoding error probability of ground node u at time slot t; M represents the number of UAVs;
[0043] The joint optimization model satisfies the following constraints:
[0044]
[0045] Wherein, constraint (1) indicates that only ground nodes within the UAV's perception range can connect to the UAV, α m,u (t) represents the access control variable between UAV m and ground node u at time slot t. For drones, For the set of ground nodes, The constraint (2) indicates that each ground node can only access one UAV at most; the constraint (3) indicates that the number of ground nodes accessed by each UAV cannot exceed the number of sub-channels, where K is the total number of sub-channels; the constraint (4) indicates the range of values for the access control variable and the sensing variable, where c m,u (t) represents the sensing variables between UAV m and ground node u at time slot t; constraint (5) indicates that the position of the UAV cannot exceed the boundary of the target area, where x m (t) and y m (t) represents the x and y coordinates of UAV m at time slot t, respectively, and X and Y represent the x and y boundaries of the target area, respectively; constraint (6) indicates that the distance between any two UAVs is greater than the minimum spacing D. min To avoid collisions, where q i (t) and q j (t) represents the positions of UAV i and UAV j in time slot t, respectively.
[0046] In one embodiment, the joint optimization model is described as a deneutralized partially observable Markov process, including:
[0047] The feature vectors of the UAV m and the ground node u at time slot t are obtained respectively, and are expressed as follows:
[0048]
[0049] Wherein, the feature vector of UAV m at time slot t Including the current position q of drone m m (t) and the total number of ground nodes accessed by UAV m in the previous time slot. The feature vector of ground node u at time slot t Including the position q of ground node n u Was the drone connected in the previous time slot? Average uplink throughput up to time slot t-1 And the decoding error probability ε of time slot t-1 u (t-1);
[0050] Based on the feature vectors of both UAVs and ground nodes, the joint optimization model is described as a deneutralized partially observable Markov process. This deneutralized partially observable Markov process is described using environmental state, UAV local observation information, UAV actions, and a reward function. Specifically, the environmental state is described using the set of feature vectors of all entities; the UAV local observation information is described using the UAV's own feature vector, the set of feature vectors of neighboring UAVs within its perception range, and the set of feature vectors of ground nodes; the UAV actions are described using the UAV's movement direction decision and the ID of the ground node accessing the UAV; and the reward function is described using a global reward and an individual UAV penalty value, with the global reward being a weighted sum of total fair throughput and QoS-DL.
[0051] In one embodiment, the priority-based access control algorithm includes:
[0052] For ground nodes, in each time slot, each ground node prioritizes accessing the nearest drone;
[0053] For UAVs, a throughput-first principle is applied to the uplink, meaning that ground nodes with lower throughput are prioritized to connect to the UAV within the first t time slots. A distance-first principle is applied to the downlink, meaning the nearest ground node is prioritized to connect to the UAV. The priority ranking of the UAV's access to the set of ground nodes within its perception range is obtained by weighted summing of the normalized throughput and distance, denoted as:
[0054]
[0055] Among them, P m Let m be the set of ground nodes within the perception range of the UAV. The priority sorting, sort(·), means that each UAV sorts the ground nodes according to the weighted sum of normalized throughput and distance. Ground nodes with higher priority are connected to the UAV first, and for each ground node connected to the UAV, the corresponding UAV will assign it the sub-channel with the least uplink and downlink interference at the current time; Γ u (t) represents the total throughput obtained by ground node u in the past t time slots. λ represents the maximum total throughput, λ is the weighting coefficient, and q is the maximum value of the total throughput. u Let q be the position of ground node u. m (t) represents the position of UAV m in time slot t, and u′ represents the set Any ground node in the network, A collection of drones.
[0056] In one embodiment, the QMIX trajectory planning algorithm based on weighted generation network enhancement consists of several agent networks, a hybrid network, and a super network;
[0057] The agent network is used to acquire local observation information of the corresponding UAV and fit the output of the individual value of the UAV agent. The agent network consists of three parts: input layer, intermediate layer, and output layer. The input layer consists of two weight generation networks and a multilayer perceptron. The two weight generation networks are used to extract features from the feature vector sets of neighboring UAVs and ground nodes within the local observation information of the UAV. The multilayer perceptron is used to extract features from the UAV's own feature vectors within the local observation information of the UAV. The input layer finally merges the outputs of the two weight generation networks and the multilayer perceptron to output an embedded feature. The intermediate layer uses a gated recursive unit, taking the hidden state of the previous time slot and the embedded feature output by the input layer as input, and outputs the hidden state of the current time slot. The output layer uses a multilayer perceptron, which calculates the individual value of the UAV by acquiring the hidden state of the current time slot. Based on its own individual value, the UAV makes action decisions using an ∈-greedy strategy.
[0058] Hybrid networks are used to take the individual value of each drone as input and obtain the joint action value through forward propagation;
[0059] Hypernetworks are used to calculate the weights and biases of hybrid networks based on environmental conditions.
[0060] An access control and trajectory planning device for emergency communication networks, the device comprising:
[0061] The network construction module is used to build an emergency communication network assisted by multiple UAVs and to acquire the uplink and downlink within the network; the emergency communication network consists of multiple UAVs, multiple ground nodes, and a remote control center;
[0062] The model building module is used to construct a limited perception model, channel model, and communication model between the UAV and ground nodes based on the emergency communication network.
[0063] The problem modeling module is used to jointly consider the communication requirements of uplink throughput and downlink timeliness and reliability, and to construct a joint optimization model for the access control and UAV trajectory planning problem based on the finite sensing model, channel model and communication model.
[0064] The joint solution module is used to describe the joint optimization model as a deneutralized partially observable Markov process, and uses a priority-based access control algorithm and a QMIX trajectory planning algorithm based on weighted generative networks to solve the joint optimization model, thereby obtaining the optimal access decision and the optimal trajectory of the UAV.
[0065] The aforementioned access control and trajectory planning method and apparatus for emergency communication networks comprehensively consider the communication requirements of uplink throughput and downlink timeliness and reliability in emergency communication networks. A joint optimization model for access control and UAV trajectory planning is modeled. Solving this joint optimization model enables joint optimization of differentiated communication requirements of uplink and downlink. Furthermore, a priority-based access control algorithm and a QMIX trajectory planning algorithm enhanced by a weight generation network are proposed to solve this joint optimization model. The priority-based access control algorithm uses a heuristic method to achieve node access, effectively reducing the state-action space of the QMIX algorithm and solving the problem of poor convergence. The QMIX trajectory planning algorithm enhanced by a weight generation network improves the QMIX algorithm, solving the problem of low training efficiency caused by dynamic observation and enabling faster convergence to better results. This method can achieve efficient communication and resource management in emergency communication scenarios. Attached Figure Description
[0066] Figure 1 This is a flowchart illustrating an access control and trajectory planning method for an emergency communication network in one embodiment.
[0067] Figure 2 This is a schematic diagram of a multi-UAV-assisted emergency communication network in one embodiment.
[0068] Figure 3 This is a schematic diagram of the weight generation network in one embodiment;
[0069] Figure 4 This is a schematic diagram of the architecture of the QMIX trajectory planning algorithm based on a weight generation network enhancement in one embodiment;
[0070] Figure 5 This is a flowchart illustrating the PW-QMIX algorithm in one embodiment;
[0071] Figure 6 This is a schematic diagram of the long-term reward convergence curves under different learning rate conditions in one embodiment.
[0072] Figure 7 This is a schematic diagram comparing the decision performance of the PW-QMIX algorithm and the QMIX algorithm in one embodiment;
[0073] Figure 8This is a schematic diagram of the convergence curves for schemes with / without dynamic observation processing under different numbers of ground nodes in one embodiment.
[0074] Figure 9 This is a schematic diagram of the convergence curves for different observation processing schemes in one embodiment;
[0075] Figure 10 This is a schematic diagram illustrating the relationship between λ and fair throughput under different charge packet lengths in one embodiment.
[0076] Figure 11 This is a schematic diagram illustrating the relationship between λ and QoS-DL under different charge packet lengths in one embodiment;
[0077] Figure 12 This is a schematic diagram illustrating the relationship between the number of drones and fair throughput for different algorithms in one embodiment.
[0078] Figure 13 This is a schematic diagram illustrating the relationship between the number of drones and QoS-DL for different algorithms in one embodiment. Detailed Implementation
[0079] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0080] In one embodiment, such as Figure 1 As shown, an access control and trajectory planning method for emergency communication networks is provided, including the following steps:
[0081] Step S1: Construct an emergency communication network assisted by multiple UAVs and obtain the uplink and downlink within the network; wherein, the emergency communication network consists of multiple UAVs, multiple ground nodes and a remote control center.
[0082] Specifically, such as Figure 2As shown, this application considers a multi-UAV-assisted emergency communication network in disaster relief scenarios. When ground communication infrastructure (such as 5G base stations) is destroyed, UAVs are deployed as airborne base stations to quickly provide emergency communication services. Rescue robots, acting as ground nodes (GNs), can collect frontline environmental information and transmit it in image or video format to a remote control center. The remote control center makes decisions based on the received environmental information and sends command and control information, including movement control and mission instructions, to the ground nodes. The rescue robots then execute search and rescue tasks based on the received command and control information. Each UAV, acting as an airborne base station, forwards the transmitted information between the ground nodes and the remote control center. The uplink within the network is defined as the link for transmitting environmental information from the ground nodes to the UAVs, and the downlink is defined as the link for transmitting command and control information from the remote control center to the ground nodes.
[0083] Assume there are M drones and U ground nodes in the emergency communication network. Let... and Let X and Y represent the set of UAVs and the set of ground nodes, respectively. This application considers that UAVs provide time-period [time period] to the ground nodes within a square target area of size X×Y. Service. Service period. The system is divided into T equal-length time slots, each with a length of δ. The UAV flies at a fixed altitude H and a constant speed v. A 3D Cartesian coordinate system is used to describe the positions of the UAV and the ground nodes. In time slot t, the position of UAV m is represented as... The position of ground node u is represented as In a given time slot, a ground node can only communicate with one drone. To avoid collisions, drones need to maintain a safe distance D. min , represented as
[0084] Step S2: Construct a limited perception model, channel model, and communication model between the UAV and ground nodes based on the emergency communication network.
[0085] Limited Perception Model: When temporarily deployed drones provide communication services to a large, unknown area, they typically cannot know the status information of all nodes in advance. In actual emergency communication networks, ground nodes can periodically send location and communication request information through a common channel. Furthermore, due to signal attenuation, drones can only receive messages from ground nodes within a certain distance. Similarly, drones can only perceive neighboring drones within a certain range. Assume that the perception distance of each drone to ground nodes and neighboring drones is D. GN and D UAV In time slot t, only when Only when the drone m can it sense the ground node u. Similarly, when Only when time slot t is reached can UAV i and UAV j perceive each other. This shows that each UAV can only acquire partial state information of the ground node in real time. Due to changes in the UAV's position, its observation area also constantly changes. To characterize the perception relationship between the UAV and the ground node, the finite perception model constructed in this application uses perception variables between the UAV and the ground node for description. Specifically, the perception variable between UAV m and ground node u at time slot t is defined as c. m,u (t), and c m,u (t)∈{0,1}; where, if ground node u is within the perception range of UAV m at time slot t, c m,u (t) = 1; otherwise c m,u (t) = 0.
[0086] Channel Model: The channel model constructed in this application is described using the channel gain between the ground node and the UAV. The calculation process of the channel gain is as follows:
[0087] This application considers a general air-to-ground (A2G) model that depends on the probability of line-of-sight (LoS) and non-line-of-sight (NLoS) links occurring between the UAV and ground nodes.
[0088] The probabilities of line-of-sight (LoS) and non-line-of-sight (NLoS) links between ground node u and UAV m are calculated based on the pitch angle from ground node u to UAV m at time slot t, and are expressed as follows:
[0089]
[0090] in, This represents the probability of a Loss of Sight (LoS) link between ground node u and drone m at time slot t. This represents the probability of an NLoS link between ground node u and UAV m at time slot t. 'a' and 'b' are parameters determined by the environment type. The pitch angle from ground node u to UAV m at time slot t represents the fixed altitude of the UAV flight, and q represents the pitch angle from ground node u to UAV m. u Let q be the position of ground node u. m (t) represents the position of UAV m in time slot t;
[0091] The path loss between ground node u and drone m is calculated based on the positions of ground node u and drone m at time slot t, and is expressed as follows:
[0092]
[0093] Where, η ξf represents the average additional path loss of a LosS link or an NLoS link. c ξ is the carrier frequency, c is the speed of light, and ξ represents a LosS link or an NLoS link.
[0094] according to The path loss for each link is calculated to obtain the channel gain between ground node u and UAV m, which is the mathematical expectation for both LosS and NLoS cases, specifically expressed as follows:
[0095]
[0096] in, This represents the path loss (LoS) of the link between ground node u and UAV m at time slot t. This represents the path loss of the NLoS link between ground node u and UAV m at time slot t.
[0097] Communication Model: To minimize interference, each UAV provides uplink and downlink access to the ground nodes it connects to via Orthogonal Frequency Division Multiple Access (OFDMA). Simultaneously, the uplink and downlink operate in Frequency Division Duplexing (FDD) mode, meaning each ground node's uplink and downlink operate on different frequency bands. Assume each UAV has K... UL =K DL =K uplink sub-channels and downlink sub-channels, each with a bandwidth of B, denoted as follows: The transmit power of each drone and ground node on each subchannel is p m and p u Define the access control variable α. u,m (t) describes the connection relationship between UAV m and ground node u. When ground node u connects to UAV m in time slot t, α u,m (t) = 1, otherwise α u,m (t) = 0.
[0098] For the uplink, the communication model is described using the uplink signal-to-interference-plus-noise ratio (SINR), uplink data rate, and throughput; where SINR is the uplink signal-to-interference-plus-noise ratio (SINR) from ground node u to UAV m at time slot t. Represented as
[0099]
[0100] Among them, gu,m (t) represents the channel gain between ground node u and UAV m at time slot t, g u′,m (t) represents the channel gain between another ground node u′ and the UAV m at time slot t, α u,m (t) represents the access control variable between ground node u and UAV m at time slot t, α u′,m′ (t) represents the access control variable between another ground node u′ and another UAV m′ at time slot t, n0 represents the power spectrum density (PSD) of additive white Gaussian noise (AWGN), p u Let B be the transmit power of ground node u on each subchannel, and let B be the subchannel bandwidth.
[0101] Considering the environmental information of uplink transmission, its achievable data rate can be calculated using Shannon's formula. The uplink data rate from ground node u to UAV m at time slot t is... Represented as
[0102]
[0103] According to the above formula, the total throughput obtained by ground node u and the entire emergency communication network in the past t time slots are respectively expressed as: and Where t' = {1, 2, ..., t} is any one of the past t time slots, and δ is the time slot length; meanwhile, the average throughput of ground node u over the past t time slots is expressed as:
[0104] For the downlink, the communication model is described using the downlink signal-to-interference-plus-noise ratio (SINR), downlink data rate, and decoding error probability. For the downlink, command and control information sent to ground nodes typically has strict requirements for timeliness and reliability. The QoS requirements of command and control information can be expressed as the maximum tolerable delay τ. max and maximum decoding error probability ε max To quantify this, the command and control packet length is typically short. In short packet communication, the results of information theory derived from Shannon capacity are inapplicable because the law of large numbers does not apply. Therefore, for any signal-to-noise ratio, the decoding error probability ε is non-zero. To accurately characterize the transmission performance of command and control packets, this application uses short packets to achieve an approximation of the downlink data rate.
[0105] Specifically, the downlink signal-to-interference-plus-noise ratio (SINOR) from UAV m to ground node u at time slot t. Represented as
[0106]
[0107] Where, α m,u (t) represents the access control variable between UAV m and ground node u at time slot t, α m',u' (t) represents the access control variable between another UAV m′ and another ground node u′ at time slot t, g m,u (t) represents the channel gain between the UAV m and the ground node u at time slot t, g m',u (t) represents the channel gain between another UAV m′ and ground node u at time slot t, p m Let m be the transmit power of the UAV on each sub-channel.
[0108] Downlink data rate from UAV m to ground node u in time slot t Represented as
[0109]
[0110] Where τ is the transmission delay of the charge packet, and Q -1 (·) is the inverse function of the Q function, ε max The maximum decoding error probability is given by V, which represents the channel dispersion, specifically expressed as:
[0111] The end-to-end delay of the command ignores queuing delay and only considers transmission delay. For a given packet length S, the maximum tolerable delay τ... max This application can obtain the following by transmitting a packet internally: When UAV m transmits an instruction / command packet to ground node u in time slot t, the decoding error probability of ground node u is expressed as:
[0112]
[0113] Where, τ max Let S represent the maximum tolerable delay, and τ represent the given packet length. The meaning of the above formula is that, given a packet length S and a maximum tolerable delay τ, ... max Under the given conditions, the decoding error probability ε can be calculated using the formula described above. u (t). Therefore, in order to simultaneously meet the timeliness and reliability requirements of the accusation group, the decoding error probability must not exceed the threshold ε. max , can be represented as ε u (t)≤ε max .
[0114] Step S3: By jointly considering the uplink's throughput requirements and the downlink's timeliness and reliability requirements, and based on the finite sensing model, channel model, and communication model, a joint optimization model for the access control and UAV trajectory planning problem is constructed.
[0115] Specifically, for the uplink, solely pursuing maximum emergency communication network throughput may result in only a few ground nodes continuously receiving communication services for most of the time, while other nodes remain in communication dead zones. To avoid this situation, the Jain fairness coefficient is introduced to evaluate the fairness of the communication services received by ground nodes. This fairness coefficient is calculated based on the average throughput of each ground node in the time slots preceding time slot t. In time slot t, the formula for calculating the fairness coefficient is as follows:
[0116]
[0117] According to the Cauchy-Schwarz inequality, η(t) always satisfies 1 / U ≤ η(t) ≤ 1. A higher fairness coefficient indicates a smaller difference in average throughput between different ground nodes, and better service fairness within the emergency communication network. To maximize the throughput of the emergency communication network while ensuring fairness in ground node services, fair throughput is defined as the product of the throughput in each time slot and the fairness coefficient. Therefore, ground nodes in the service period... The total fair throughput can be expressed as
[0118]
[0119] When the downlink satisfies ε u (t)≤ε max This allows for the simultaneous fulfillment of the latency and reliability requirements of command and control packets. Within a service cycle, this application aims to establish more links that meet communication needs. Therefore, the number of downlinks satisfying QoS (QoS-DL) is chosen as the evaluation metric, expressed as:
[0120]
[0121] Where 1(·) represents the indicator function, specifically expressed as
[0122] Because it needs to meet the differentiated communication requirements of uplink and downlink, the utility of the emergency communication network is defined as a weighted sum of fair throughput and QoS-DL. Both fair throughput and QoS-DL are affected by UAV location and access control policies. To maximize both metrics simultaneously, the UAV trajectory planning and access control problem is modeled as follows:
[0123]
[0124] Wherein, λ is a weighting coefficient used to balance the relative importance of uplink and downlink; total fair throughput is used to consider the communication requirements of uplink for throughput; QoS-DL is used to consider the communication requirements of downlink for timeliness and reliability.
[0125] The joint optimization model satisfies the following constraints:
[0126]
[0127] Wherein, constraint (1) indicates that only ground nodes within the UAV's perception range can connect to the UAV, α m,u (t) represents the access control variable between UAV m and ground node u at time slot t. For drones, For the set of ground nodes, The constraint (2) indicates that each ground node can only access one UAV at most; the constraint (3) indicates that the number of ground nodes accessed by each UAV cannot exceed the number of sub-channels, where K is the total number of sub-channels; the constraint (4) indicates the range of values for the access control variable and the sensing variable, where c m,u (t) represents the sensing variables between UAV m and ground node u at time slot t; constraint (5) indicates that the position of the UAV cannot exceed the boundary of the target area, where x m (t) and y m (t) represents the x and y coordinates of UAV m at time slot t, respectively, and X and Y represent the x and y boundaries of the target area, respectively; constraint (6) indicates that the distance between any two UAVs is greater than the minimum spacing D. min To avoid collisions, where q i (t) and q j (t) represents the positions of UAV i and UAV j in time slot t, respectively.
[0128] The joint optimization model is a mixed-integer nonlinear programming problem consisting of continuous variable q and discrete variable α. Due to the destruction of basic communication infrastructure, remote control centers often struggle to obtain the location information of all ground nodes. Therefore, this application requires the design of a distributed solution scheme.
[0129] Step S4: The joint optimization model is described as a deneutralized partially observable Markov process, and the joint optimization model is solved using a priority-based access control algorithm and a QMIX trajectory planning algorithm enhanced by a weighted generative network to obtain the optimal access decision and the optimal trajectory of the UAV.
[0130] The decentralized partially observable markov decision process (Dec-POMDP) consists of a quintuple. Composition, in which Indicates the environmental state. This represents the joint action space of M agents. Represents local observation information. This is the state transition function. The reward function is given. In the Dec-POMDP model, environmental information at time slot t is represented as... Each drone acts as an independent intelligent agent. Obtain local observation information from the current environmental information s(t). Then, each agent, based on local observation information, m (t) Select an action The actions of all agents form a combined action. Then, the environment is determined according to the state transition probability function. Transition to the next state In a fully cooperative setting, all agents share the same reward function. That is, r(t) = r1(t) = ... = r M (t). Ultimately, the goal of MADRL is to learn the optimal policy to maximize long-term returns. Where γ∈[0,1] represents the discount factor.
[0131] Specifically, the joint optimization model is described as a deneutralized partially observable Markov process, including:
[0132] In scenarios where multiple drones provide communication services to ground nodes, there are two types of entities: drones and ground nodes. The feature vector of each entity contains environmental information related to its own category. The feature vectors of drone m and ground node u at time slot t are obtained respectively, and are represented as follows:
[0133]
[0134] Wherein, the feature vector of UAV m at time slot t Including the current position q of drone m m (t) and the total number of ground nodes accessed by UAV m in the previous time slot. The feature vector of ground node u at time slot t Including the position q of ground node u u Was the drone connected in the previous time slot? Average uplink throughput up to time slot t-1 And the decoding error probability ε of time slot t-1 u (t-1). and Each feature has a crucial impact on the optimization objective of the joint optimization model. Below, we will define the key components of Dec-POMDP based on the feature vectors of the aforementioned entities.
[0135] Specifically, based on the feature vectors of the two entities, UAV and ground node, the joint optimization model is described as a deneutralized partially observable Markov process. The deneutralized partially observable Markov process is described by environmental state, local observation information of UAV, UAV actions, and reward function.
[0136] The environmental state is described using the set of feature vectors of all entities. Specifically, in time slot t, the environmental state s(t) contains all entity state information relevant to the problem solution, and is composed of the entity feature vectors. Let... and Let represent the sets of feature vectors for all UAVs and all ground nodes, respectively. Therefore, the environmental state can be represented as the set of feature vectors for all entities:
[0137] s(t)=[s UAV (t),s GN (t)].
[0138] The local observation information of the UAV is described using the UAV's own feature vector, the set of feature vectors of neighboring UAVs within its perception range, and the set of feature vectors of ground nodes. Specifically, each UAV can only perceive distances of D. GN and D UAV The detection range includes nodes and neighboring drones. Therefore, each drone agent can only observe a subset of ground nodes and exchanges observation information with drones within its perception range. Thus, in time slot t, drone m observes its neighboring drone. and ground nodes They are respectively represented as follows
[0139]
[0140] make This represents the eigenvector of the UAV θ itself. This represents the set of feature vectors of neighboring drones within the θ-sensing range of the drone. Let represent the set of feature vectors of ground nodes within the perception range of the UAV m. Therefore, the local observation information of the UAV θ can be all the features of these three elements, and can be represented as...
[0141]
[0142] The UAV's actions are described using action decisions based on the UAV's movement direction and the ID of the ground node accessing the UAV. Specifically, in each time slot, the actions the UAV can take include movement and ground node access control. Since the UAV flies at a constant speed v and a fixed altitude H, the UAV's movement can be simplified to movement in a certain direction in a two-dimensional plane. This application discretizes the horizontal movement direction, which is represented as Ψ = {0, π / 2, π, 3π / 2}. Furthermore, the scheduled ground node must be a ground node within the perception range. The UAV θ's action in time slot t can be represented as...
[0143]
[0144] in, This indicates the action decision for the UAV m moving in that direction. This represents the ID of the ground node that connected to UAV m. The displacement of UAV m in this time slot is...
[0145] △q m (t)=(△x m (t),△y m (t),0)=(δ×vcosψ, δ×vsinψ,0).
[0146] The reward function is described by a global reward and individual UAV penalties. The global reward is a weighted sum of total fair throughput and QoS-DL. Specifically, all UAVs share a common goal: to improve the overall utility of the emergency communication network. Therefore, this application selects the sum of the utilities generated by all UAVs as the global reward. Similar to the optimization objective of the joint optimization model, considering the differentiated communication requirements of uplink and downlink, their weighted sum is chosen as the reward. In time slot t, the global reward can be expressed as a weighted sum of total fair throughput and QoS-DL, with the following expression:
[0147]
[0148] Furthermore, to avoid collisions between drones and considering that drones need to always fly within the service area, a binary vector ξ(t)∈{0,1} is defined to describe situations where a drone collides or flies out of the service area. Each element ξ... m (t) indicates whether the drone m exceeds the collision limit or the area limit in time slot t. If ξ m If (t) = 1, a penalty value μ is added to the reward value; otherwise, the penalty value is 0. In summary, combining the global reward and the individual drone penalty value yields the complete reward function, which can be expressed as:
[0149] r(t) = rg (t)-ξ m (t)·μ.
[0150] Priority-based access control algorithms address situations where, during UAV service, a ground node *u* may be within the perception range of multiple UAVs, and the number of nodes within the perception range of UAV *m* may exceed the number of available channels. Therefore, an effective access control strategy is needed to meet optimization objectives and ensure optimal communication service under limited resources. The core idea of the access control strategy proposed in this application is that, in each time slot, each ground node prioritizes accessing the nearest UAV, while each UAV prioritizes selecting a higher-priority node to provide service.
[0151] For ground nodes, in each time slot, each ground node prioritizes connecting to the nearest drone. Specifically, when ground node u is within the coverage area of multiple drones, it will prioritize connecting to the nearest drone m. ★ Send access request information, i.e.
[0152]
[0153] in, This represents the set of drones within the sensing range of ground node u.
[0154] For UAVs, prioritizing ground nodes requires a comprehensive consideration of both uplink throughput and downlink QoS requirements. For the uplink, considering throughput requirements and service fairness, the node with the lowest throughput at the current moment is prioritized for service. For the downlink, considering the reliability and timeliness requirements of command and control information transmission, ground nodes closer to the UAV have greater channel gain and are more likely to meet these requirements. Specifically, a throughput priority principle is adopted for the uplink, meaning that in the first t time slots, the ground node with the lowest throughput is prioritized for access to the UAV; a distance priority principle is adopted for the downlink, meaning the nearest ground node is prioritized for access to the UAV. To optimize consistency, a weighted sum is used to define the priority. Furthermore, since "throughput" and "distance" have different dimensions, they are normalized. The throughput of ground nodes within the UAV's perception range (m) is divided by the maximum throughput. Similarly, the distance is divided by the maximum distance from the ground node to the UAV. By weighted summing the normalized throughput and distance, the priority ranking of the set of ground nodes within the UAV's perception range is obtained, denoted as:
[0155]
[0156] Among them, P m Let m be the set of ground nodes within the perception range of the UAV. The priority sorting, sort(·) means that each UAV sorts the ground nodes according to the weighted sum of normalized throughput and distance. Ground nodes with higher priority are given priority to access UAVs. For each ground node that is connected to the UAV, the corresponding UAV will allocate the sub-channel with the least uplink and downlink interference at the current time.
[0157] The architecture of the QMIX trajectory planning algorithm enhanced by a weight generation network is as follows: Figure 4 As shown.
[0158] Among these considerations, while Multilayer Perceptrons (MLPs) possess powerful data fitting capabilities, they struggle to adapt to changes in input dimensions. Therefore, this application designs an automatic weight generation mechanism to dynamically adjust the weight matrix to accommodate variations in the number of input elements. Since each feature vector x... i Each corresponds to a sub-weight matrix w i To ensure that the dimension of the weight matrix changes synchronously with the number of input elements, a matrix such as... Figure 3 The diagram shows a Weight Generation Network (WGN) consisting of a single hidden layer. The WGN processes each entity feature x... i As input, the corresponding sub-weight matrix w is generated. i Therefore, the sub-weight matrix can be represented as:
[0159]
[0160] Furthermore, the output of the entire network can be represented as
[0161]
[0162] Where L represents the number of entities.
[0163] like Figure 3 As shown, the features of all entities are fed into a shared neural network, and after forward propagation, the output weight matrix w is obtained. i Then, each weight matrix is compared with its corresponding input feature x. i Perform matrix multiplication (matmul). Then, use a summation operation to aggregate all outputs to obtain the feature representation y. It's worth noting the weight generation network. The input and output dimensions of a weight generation network depend only on the feature dimension of each entity, not on the number of input entities. Therefore, weight generation networks can flexibly handle dynamic changes in input dimensions caused by variations in the number of input entities, thus maintaining the network's robustness and generalization ability. Furthermore, weight generation networks can generate weight matrices based on the feature vectors of each entity to measure the impact of each observed entity on the current agent.
[0164] Figure 4 As shown, the QMIX trajectory planning algorithm based on weighted generation network enhancement consists of several agent networks, a hybrid network, and a super network.
[0165] Among them, the agent network is used to acquire local observation information of the corresponding UAV. m Fit the output of the individual value Q of the drone intelligent agent. m Based on the observation vectors, it can be seen that each UAV's observations include feature vectors of neighboring UAVs and ground nodes within its perception range. In different time slots, o m The dimension of the object changes with the number of entities within the perception range. The agent network consists of three parts: an input layer, an intermediate layer, and an output layer. In the classic QMIX algorithm, the agent network uses an MLP to process the observation information. However, as mentioned earlier, MLP struggles to handle the problem of local observation dimension. To address this issue, this application chooses WGN to generate weights for features with a variable number of entities. The input layer consists of two weight generation networks and a multilayer perceptron. The two weight generation networks are used to extract features from the feature vector sets of neighboring drones within the drone's perception range and the feature vector sets of ground nodes, respectively, in the drone's local observation information. Furthermore, since the dimension of the agent's own observation vector remains constant, this application still uses an MLP to extract features from the drone's own feature vectors in the drone's local observation information. The input layer ultimately outputs an embedded feature by merging the outputs of the two weight generation networks and the multilayer perceptron. This process can be represented as...
[0166]
[0167] in, These are functions implemented by neural networks for the corresponding entity types.
[0168] To maintain a memory of historical observations and actions, the intermediate layer employs a Gate Recurrent Unit (GRU) to store the hidden state h from the previous time slot. m (t-1) and the embedded features e of the input layer output m (t) is taken as input, and the hidden state h of the current time slot is output. m(t). The GRU unit introduces historical information by establishing the dependency between each time slot and the previous time slot. Based on this historical information, each UAV agent can better infer changes in the global environmental state, thereby making better decisions. This process can be represented as...
[0169]
[0170] in, This represents a function implemented by GRU.
[0171] The output layer uses a multilayer perceptron, which obtains the hidden state h of the current time slot. m (t) is used to calculate and output the individual value of the drone, expressed as:
[0172]
[0173] in, θ represents the output function implemented by the MLP. a This represents the network parameters of the agent. Based on the derived individual value Q... m The agent makes action decisions using an ∈-greedy strategy, which can be expressed as the following formula.
[0174]
[0175] After each agent selects an action, the utility of each agent to the environment is calculated. Then, the reward r(t) is calculated according to the reward function described above, and the environment transitions to the next state s(t+1).
[0176] During the training phase, the hybrid network is used to assign the individual value Q of each drone. m As input, the joint action value Q is obtained through forward propagation. tot Hypernetworks, a special type of neural network, are used to calculate the weights and biases of a hybrid network based on environmental states. Specifically, the hybrid network is a two-layer feedforward neural network that calculates the local action value Q of each agent. m (o m (t),a m (t); θ a As input, the joint action value Q is obtained through forward propagation. tot (s(t), a(t); θ mix ,θ h ), where θ mix and θ h These represent the parameters of the hybrid network and the hypernetwork, respectively. Joint action value Q tot This represents the expected long-term cumulative return achieved under a joint strategy. In contrast, the individual value Q... mThe contribution of each agent to this cumulative reward is quantified. When Q tot When hybrid networks more accurately approximate the true value of joint actions, they can adjust individuals more effectively through backpropagation.
[0177] To ensure consistency between the globally optimal policy and the locally optimal policy of a single agent, the IGM (Individual-Global-Max) criterion must be satisfied:
[0178]
[0179] The above formula shows that the value of the combined action Q tot and individual value Q m There is a monotonic relationship between them. Therefore, Q tot and each Q m The following monotonicity constraints also need to be satisfied:
[0180]
[0181] Specifically, such as Figure 4 As shown, QMIX ensures the value of joint actions Q through the design of hybrid networks and hypernetworks. tot The individual value Q of each intelligent agent m A nonnegative monotonic function. Hypernetwork Using the environmental state s(t) as input, parameters θ are generated for the hybrid network. mix = {w1, w2, b1, b2}, where w1, w2, b1, b2 represent the weights and biases of the first and second layers of the hybrid network, respectively. Therefore, this application can obtain The output values of the supernetwork are processed through an absolute value activation function to obtain weights w1 and w2. This means that w1 and w2 are non-negative, thus ensuring the monotonicity constraint. Bias b1 is generated by a single-layer MLP, while bias b2 is generated by a two-layer MLP with ReLU activation functions. Finally, the joint action-value function can be expressed as...
[0182] Q tot (s(t), a(t); θ mix ,θ h )=w2·φ ELU (w1·Q+b1)+b2;
[0183] Where Q = [Q1(o1(t), a1(t); θ] a ),…,Q M (o M (t),a M (t); θ a )],φ ELU (·) represents the ELU activation function.
[0184] Furthermore, this application employs the DQN (Deep Q-Network) method to train all networks, that is, simultaneously training two neural networks with the same structure. One neural network is an online network responsible for real-time evaluation of Q-values, and its weights are updated every time slot; the other is a target network responsible for calculating the target Q-value, and its weights are periodically synchronized with the online network. The parameters of the agent network, hybrid network, and supernetwork of the target network are respectively expressed as follows: All networks can be trained by minimizing the following loss function:
[0185]
[0186] in, This indicates a sample randomly sampled from the playback buffer. This represents the global objective Q value.
[0187] The aforementioned priority-based access control algorithm and WGN-enhanced QMIX trajectory planning algorithm are collectively referred to as the PW-QMIX algorithm (priority-based access control and WGN-enhanced QMIX-based trajectory planning algorithm). This PW-QMIX algorithm achieves efficient communication services to ground nodes by dynamically selecting the access node for each time slot and the UAV's movement direction for each time slot. The PW-QMIX algorithm process is as follows: Figure 5 As shown, the specific steps are as follows:
[0188] Step 1, Parameter Initialization: Randomly initialize the online network parameter θ a ,θ mix ,θ h and target network parameters Initialize playback buffer
[0189] Step 2: Initialize the environment state: Before entering each round, initialize the environment state s(t).
[0190] Step 3: Obtain Local Observations: During the training round, each agent obtains its local observation information in each time slot t. m (t).
[0191] Step 4: Make action decisions: Each agent makes a movement direction decision based on the ∈-greedy policy.
[0192] Step 5: Calculate the joint action value: Combine the environmental state s(t) with the individual value Q of each UAV agent. m The input is fed into a hybrid network to calculate the joint action value Q. tot .
[0193] Step 6: Make access control decisions: Each ground node calculates the distance to the UAV within its sensing range; then it selects the closest UAV, i.e., m. ★ The priority P of each UAV within its computational perception range for ground nodes. m If the UAV has an available channel, then a sub-channel is allocated to the ground node; after all UAVs have made their control decisions, the access control variable α is updated.
[0194] Step 7: Execute actions and transition the environment state: Each agent executes a movement action and an access action. The environment state transitions from s(t) to the next state s(t+1), and each agent updates its local observation information o. m (t+1).
[0195] Step 8: Calculate the global reward: Calculate the reward function r(t).
[0196] Step 9: Store experience samples: End a training round and store the experience samples of one round in the replay buffer.
[0197] Step 10: Update online network parameters: If the playback buffer size... Larger than the mini-batch size The sampling size from the playback buffer is The samples are then used to update the online network parameters based on the loss function.
[0198] Step 11: Update target network parameters: If the target network's parameter update frequency is reached, then assign the online network's parameters to the target network.
[0199] Step 12: Output agent network parameters: After training, output the parameters of the agent network.
[0200] Furthermore, to verify the beneficial effects of this application, the PW-QMIX algorithm proposed in this application was compared with other algorithms, and a series of simulation experiments were conducted to verify the significance of the modeled joint optimization problem and the superiority of the proposed algorithm.
[0201] The network simulation scenario is set as follows: UAV nodes are randomly deployed in a rectangular area of 1500m × 1500m. This area is divided into four non-overlapping hotspot areas, each containing an equal number of ground nodes. At the beginning of each service period, the positions of U ground nodes and M UAVs are reset and randomly distributed within the area. The positions of the ground nodes remain unchanged within each service period. The service period is divided into 150 time slots. At the end of each time slot, the UAVs make decisions regarding their flight direction and ground node access control until the end of the service period. The path loss parameters are set according to the urban scenario model. Environmental parameters a and b are set to 9.61 and 0.16 respectively, and the additional path loss η... LoS η NLoS The settings are 1dB, 20dB, path loss exponent ζ = 2, UAV speed v = 10m / s, flight altitude H = 100m, and carrier frequency f. c The frequency is 2 GHz, the sub-channel bandwidth B is 200 kHz, the time slot length δ is 10 s, the number of sub-channels K is 5, the power spectral density n0 of additive white Gaussian noise is -174 dBm / Hz, and the maximum tolerable delay τ is... max The maximum decoding error probability is ε, which is 2ms. max 10 -5 The size of the charge packet S is 64 bytes, and the transmit power p of the drone and ground node are... m p u The sensing ranges D for the drone and ground node are 0.5W and 0.2W respectively. UAV D GN The safe distance D between the drones is 200m and 200m respectively. min It is 5m.
[0202] First, convergence verification is performed: To illustrate the convergence of the PW-QMIX algorithm, this application simulates the relationship between long-term reward and training epochs under different learning rate settings. The long-term reward convergence curves under different learning rate conditions are shown below. Figure 6 As shown, the reward gradually increases with the number of training epochs. When lr = 0.001 and 0.0001, the two cases converged after 15,000 and 30,000 training epochs, respectively. However, the training efficiency was lower when lr = 0.00001. Considering that lr = 0.001 converged to a higher reward value faster, this application sets the learning rate lr = 0.001 to implement the subsequent simulations.
[0203] Comparative verification was conducted for different access methods: This application compared the proposed PW-QMIX algorithm with the QMIX algorithm alone for trajectory and node access decisions. The performance comparison results of the PW-QMIX algorithm and the QMIX algorithm are as follows: Figure 7As shown, from Figure 7 As can be seen, the priority-based access control algorithm converges after approximately 15,000 episodes of training, while QMIX fails to converge. This is because the QMIX algorithm relies entirely on reinforcement learning for node access and UAV trajectory decisions. The enormous state-action space prevents the algorithm from learning an effective policy within a finite number of rounds, thus preventing convergence. In contrast, the priority-based access control algorithm proposed in this application uses a heuristic approach for node access. Compared to the QMIX algorithm, which only requires trajectory optimization decisions, the proposed algorithm significantly reduces the state-action space. Therefore, the algorithm proposed in this application is clearly superior to using only the QMIX algorithm.
[0204] Verification of dynamic observation processing capabilities: To verify the proposed algorithm's ability to handle local observation problems, ablation experiments were first conducted under different node numbers. Furthermore, the proposed PW-QMIX algorithm was compared with other state-of-the-art (SOTA) algorithms. The number of UAVs was set to 4, and the number of ground nodes was set to 20, 40, and 60 respectively. MLP was chosen as the input layer of QMIX to process observation features, serving as the benchmark algorithm. Only the priority-based access control algorithm proposed in this application, denoted as P-QMIX, was used. Then, PW-QMIX and P-QMIX were compared to verify the performance with and without local observation processing schemes. For ground node access, both algorithms used the priority-based access control algorithm proposed in this application.
[0205] Figure 8 Convergence curves are described for schemes with and without dynamic observation processing under different numbers of ground nodes. From Figure 8 As can be seen, all algorithms converged with increasing training rounds. Regardless of the number of nodes (20, 40, or 60), PW-QMIX consistently outperformed P-QMIX without local observation processing. This is because WGN dynamically generates weight matrices for each entity's features, enhancing the model's representational power. Specifically, the performance advantage of PW-QMIX becomes more pronounced with increasing node count. Specifically, when the number of nodes is 20, 40, and 60, PW-QMIX outperforms P-QMIX by approximately 6%, 9%, and 13%, respectively. This is because, within a finite area, as the number of nodes increases, the density of nodes within the perception range leads to significant changes in the local observations of the UAV in each time slot, as the local observation scheme becomes more densely packed. The inefficient MLP structure struggles to effectively handle the highly dynamic changes in input features. Conversely, PW-QMIX can make better, more reasonable decisions in each time slot, improving the UAV's service quality to ground nodes.
[0206] To further evaluate the performance of WGN, this application selected three commonly used dynamic observation processing schemes for comparison: Attention, GNN, and Sets. Details are as follows:
[0207] (1) PA-QMIX: This application uses an Attention mechanism to enhance the input layer of the QMIX algorithm. The attention mechanism processes input features through flexible weighted summation. The main advantage of the attention mechanism is that it can operate independently of a fixed input dimension. This allows it to dynamically assign weights to the features of any number of entities, effectively solving the problem of dynamic observation.
[0208] (2) PG-QMIX: This algorithm is the result of integrating GNN into QMIX in this application. GNN uses a graph structure to model UAVs and ground nodes and their relationships. It is inherently robust to changes in the number of entities because its performance is not affected by changes in the input dimension.
[0209] (3) PS-QMIX: This algorithm is the result of integrating the Sets mechanism into QMIX in this application. The Sets mechanism uses a shared neural network to extract feature representations of entities. These feature representations are then aggregated through a summation operation for subsequent processing.
[0210] Furthermore, the aforementioned baseline employs the priority-based access control algorithm proposed in this application. The number of ground nodes is set to 80.
[0211] Figure 9 The convergence curves of several different observation processing schemes, including PW-QMIX, PA-QMIX, PG-QMIX, PS-QMIX, and P-QMIX, are shown. Figure 9 As shown, algorithms with dynamic observation processing schemes outperform P-QMIX. Specifically, PW-QMIX achieves the best convergence performance, with a final return value of approximately 1110, which is about 5.1% higher than PA-QMIX and PS-QMIX, and about 20% higher than the less efficient P-QMIX and PG-QMIX.
[0212] Table 1 Results of different observation processing schemes
[0213] algorithm Fair throughput (Gbits) QoS-DL P-QMIX 9.58 2391 PA-QMIX 11.4 2670 PS-QMIX 11.0 2722 PG-QMIX 8.91 2415 PW-QMIX 12.2 2773
[0214] As shown in Table 1, this performance improvement stems from enhancements in uplink fair throughput and QoS-DL. First, PW-QMIX achieved the highest values across all metrics, validating its overall advantage in improving system utility. Second, PW-QMIX exhibited faster convergence. It converged after approximately 15,000 rounds, while PA-QMIX and PS-QMIX required approximately 25,000 rounds to converge. Furthermore, PW-QMIX surpassed the convergence values of all other algorithms after only 10,000 training iterations. These results demonstrate that WGN possesses superior convergence performance in dynamic observation processing.
[0215] To verify the importance of the optimization objective, this application modifies the weighting factor λ in the optimization objective. λ is set to 0, 0.5, and 1, representing the cases considering only the downlink, both uplink and downlink, and only the uplink, respectively. Considering that downlink QoS requirements are a key factor affecting the optimization objective, simulations are performed under different command and control packet lengths. Figure 10 This diagram illustrates the relationship between λ and fair throughput for different charge group lengths. Figure 11 This diagram illustrates the relationship between λ and QoS-DL for different charge packet lengths.
[0216] from Figure 10 and Figure 11It can be seen that as λ increases, the fair throughput for all packet lengths improves, while QoS-DL gradually decreases. These results indicate that when optimizing only one type of link, the performance of unconsidered links degrades significantly. Specifically, when the metrics of the considered links reach their optimum, the metrics of the unconsidered links reach their worst. However, for the disaster relief scenario considered in this application, focusing on only one type of link is unacceptable. Joint optimization of uplink and downlink is a more reasonable choice. Joint optimization may result in a performance loss for one type of metric compared to optimizing only one type of link requirement. However, considering the gains achieved, this loss is acceptable. For example, when S = 64 bytes, the fair throughput at λ = 0.5 is about 9% lower than at λ = 1, but about 37.6% higher than at λ = 0. Similarly, the QoS-DL at λ = 0.5 is about 0.6% lower than at λ = 0, but 183% higher than at λ = 1. Furthermore, the charge packet length has a critical impact on system utility. For example, when λ = 0.5, both uplink and downlink metrics decrease as the command and control packet length increases. This is because larger command and control packets require higher quality links to meet quality of service (QoS) requirements. Since the UAV makes decisions based on a comprehensive consideration of both link requirements, changes in downlink QoS also affect uplink QoS. Through this simulation, this application analyzes the relationship between system utility and λ under different command and control packet lengths, verifying the importance of joint uplink and downlink optimization.
[0217] The scalability of the algorithm is compared and verified: the scalability of the proposed PW-QMIX algorithm under different numbers of agents is further analyzed. The VDN (Value Decomposition Network) algorithm is a classic value-based algorithm widely used to solve multi-agent cooperative tasks. However, different algorithms endow agents with different cooperative capabilities. Therefore, this simulation selects the VDN algorithm as a comparison algorithm to verify the scalability of the proposed algorithm in multi-UAV cooperation. To ensure a fair comparison, the VDN algorithm also adopts the same dynamic observation processing scheme as PW-QMIX. Therefore, this algorithm is named PW-VDN. The number of ground nodes is set to 80.
[0218] To demonstrate the scalability of the proposed algorithm, this application compares the two algorithms under different numbers of drones. Figure 12 and Figure 13The results for the PW-QMIX and PW-VDN algorithms, considering the number of drones and their fair throughput and QoS-DL, are presented separately. When M = 4, 6, 8, and 10, the fair throughput of PW-QMIX is approximately 4.3%, 3.2%, 8.6%, and 10.1% higher than that of PW-VDN, respectively. The number of downlinks meeting QoS requirements is also higher for PW-QMIX in all cases. This result directly verifies the scalability advantage of the PW-QMIX algorithm.
[0219] Furthermore, for individual algorithms, the fair throughput of both algorithms gradually increases with the increase in the number of drone agents. However, the growth rate gradually decreases. For example, when the number of drones increases from 4 to 6, the fair throughput of PW-QMIX and PW-VDN increases by approximately 15.5% and 16.8%, respectively. When increasing from 6 to 8, the fair throughput of the two algorithms increases by approximately 14% and 8.9%, respectively. When increasing from 8 to 10, the increases are approximately 7.4% and 6.1%, respectively. The growth rate of fair throughput decreases to varying degrees. On the one hand, due to the increase in the number of drones, the number of accessing nodes also increases, leading to increased interference on the same channel. On the other hand, the increase in the number of drones leads to an exponential increase in the state space, and the cooperation between drones becomes more complex. Therefore, the performance of decision-making deteriorates. Similarly, QoS-DL shows a similar trend to fair throughput.
[0220] In summary, this application comprehensively considers the differentiated communication requirements of an emergency communication system composed of UAVs and ground nodes. Addressing the communication requirements of uplink throughput and downlink timeliness and reliability, this application models a multi-UAV trajectory planning and access control problem. By optimizing the UAV trajectories and accessing nodes, the system's fair throughput and QoS-DL are maximized. Then, this application proposes a priority-based access control algorithm and a QMIX trajectory planning algorithm enhanced by a weighted generation network. The priority-based access control algorithm achieves node access through a heuristic method, effectively reducing the state-action space. Simultaneously, the weighted generation network improves QMIX, solving the problem of low training efficiency caused by dynamic observation. Simulation results verify the significance of the optimization objectives of this application. Furthermore, compared with other algorithms, the algorithm proposed in this application can converge to better results faster and has stronger scalability.
[0221] It should be understood that, although Figure 1 and Figure 5 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 and Figure 5 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0222] In one embodiment, an access control and trajectory planning device for emergency communication networks is provided, comprising:
[0223] The network construction module is used to build an emergency communication network assisted by multiple UAVs and to acquire the uplink and downlink within the network; the emergency communication network consists of multiple UAVs, multiple ground nodes, and a remote control center;
[0224] The model building module is used to construct a limited perception model, channel model, and communication model between the UAV and ground nodes based on the emergency communication network.
[0225] The problem modeling module is used to jointly consider the communication requirements of uplink throughput and downlink timeliness and reliability, and to construct a joint optimization model for the access control and UAV trajectory planning problem based on the finite sensing model, channel model and communication model.
[0226] The joint solution module is used to describe the joint optimization model as a deneutralized partially observable Markov process, and uses a priority-based access control algorithm and a QMIX trajectory planning algorithm based on weighted generative networks to solve the joint optimization model, thereby obtaining the optimal access decision and the optimal trajectory of the UAV.
[0227] Specific limitations regarding the access control and trajectory planning device for emergency communication networks can be found in the limitations of the access control and trajectory planning method for emergency communication networks described above, and will not be repeated here. Each module in the aforementioned access control and trajectory planning device for emergency communication networks can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0228] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0229] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for access control and trajectory planning for emergency communication networks, characterized in that, The method includes: Construct an emergency communication network assisted by multiple UAVs and obtain the uplink and downlink information within the network; wherein the emergency communication network consists of multiple UAVs, multiple ground nodes, and a remote control center; Based on the aforementioned emergency communication network, a limited perception model, a channel model, and a communication model between the UAV and ground nodes are constructed. By jointly considering the uplink's throughput requirements and the downlink's timeliness and reliability requirements, and based on the finite sensing model, channel model, and communication model, a joint optimization model for the access control and UAV trajectory planning problem is constructed. The joint optimization model is described as a deneutralized partially observable Markov process, and the joint optimization model is solved using a priority-based access control algorithm and a QMIX trajectory planning algorithm based on weighted generative network enhancement to obtain the optimal access decision and the optimal trajectory of the UAV. The joint optimization model for the access control and UAV trajectory planning problem is expressed as follows: ; The joint optimization model is a function consisting of continuous variables. and discrete variables The mixed-integer nonlinear programming problem is composed of; where, This is a weighting coefficient used to balance the relative importance of uplink and downlink; This indicates the service period of the ground node for the drone. Total fair throughput, which takes into account the uplink's communication requirements for throughput; service period Classified as There are 3 time slots of equal length, each time slot having a length of 1. ; Indicates time slot Fairness coefficient at that time ground nodes in the past Average throughput within each time slot Indicates the number of ground nodes; Indicates time slot Time ground node Uplink data rate; This indicates the number of downlink QoS-DLs that meet QoS requirements, where the QoS-DLs are used to consider the downlink's communication requirements for timeliness and reliability. This indicates an indicator function, specifically represented as ; To maximize the decoding error probability, Indicates time slot Time ground node The probability of decoding errors; Indicates the number of drones; The joint optimization model satisfies the following constraints: ; Among them, constraint (1) means that only ground nodes within the UAV's perception range can connect to the UAV. Indicates time slot drones With ground nodes Access control variables between them For drones, For the set of ground nodes, The constraint (2) indicates that each ground node can only access one UAV; the constraint (3) indicates that the number of ground nodes accessed by each UAV cannot exceed the number of sub-channels. The total number of sub-channels; constraint (4) represents the range of values for the access control variable and the sensing variable, where, Indicates time slot drones With ground nodes The perceived variables between; constraint (5) indicates that the position of the UAV cannot exceed the boundary of the target area, where, and They represent time slots respectively. drones The x and y coordinates, and The horizontal and vertical boundaries of the target area are represented respectively; constraint (6) indicates that the distance between any two UAVs is greater than the minimum spacing. To avoid collisions, among which, and They represent drones With drones In the time slot The position at that time.
2. The method according to claim 1, characterized in that, In the emergency communication network, each ground node is used to collect and acquire environmental information and transmit it to the remote control center. The remote control center is used to make decisions based on the received environmental information and send command and control information, including movement control and mission instructions, to the ground nodes. Each UAV acts as an airborne base station and is used to forward the transmitted information between the ground nodes and the remote control center. The uplink in the network is defined as the link for transmitting environmental information from the ground nodes to the UAVs, and the downlink is defined as the link for transmitting command and control information from the remote control center to the ground nodes.
3. The method according to claim 1, characterized in that, The finite perception model is described using perception variables between the UAV and ground nodes, defining time slots. drones With ground nodes The perceptual variables between them are ,and Among them, if time slot Time ground node In the drone Within the range of perception, ;otherwise .
4. The method according to claim 1, characterized in that, The channel model is described using the channel gain between the ground node and the UAV. The calculation process for the channel gain is as follows: According to time slot Time ground node To drones The pitch angle is calculated to obtain the ground node. With drones The probabilities of line-of-sight (LoS) links and non-line-of-sight (NLoS) links between them are expressed as follows: ; ; in, Indicates time slot Time ground node With drones The probability of a Loss link between them. Indicates time slot Time ground node With drones The probability of NLoS links between them. and These are parameters determined by the environment type. Indicates time slot Time ground node To drones pitch angle, Indicates the fixed altitude at which the drone flies. ground nodes Location, For drones In the time slot The position at that time; According to time slot Time ground node With drones Calculate the location to obtain the ground node. With drones The path loss between them is expressed as: ; in, This represents the average additional path loss for a LosS link or an NLoS link. It is the carrier frequency. It's the speed of light. Indicates a LoS link or an NLoS link; according to , The path loss for each link is calculated to obtain the ground node. With drones The channel gain between them is expressed as: ; in, Indicates time slot Time ground node With drones Path loss of the Loss link between them Indicates time slot Time ground node With drones The path loss of the NLoS link between them.
5. The method according to claim 1, characterized in that, The communication model includes: For the uplink, the communication model is described using uplink signal-to-interference-plus-noise ratio (SINR), uplink data rate, and throughput; where time slots... Time ground node To drones uplink signal-to-interference-plus-noise ratio , represented as: ; in, Indicates time slot Time ground node With drones Channel gain between Indicates time slot Another ground node With drones Channel gain between Indicates time slot Time ground node With drones Access control variables between them Indicates time slot Another ground node With another drone Access control variables between them This represents the power spectral density of additive white Gaussian noise. ground nodes Transmit power on each sub-channel Sub-channel bandwidth, For drones, For the set of ground nodes; Time slot Time ground node To drones uplink data rate , represented as: ; ground nodes and the entire emergency communication network in the past The total throughput obtained within each time slot is expressed as follows: and ;in, For the past Any one of the time slots This refers to the time slot length; simultaneously, the ground node... in the past The average throughput within each time slot is expressed as: , This refers to the number of ground nodes. For the downlink, the communication model is described using downlink signal-to-interference-plus-noise ratio (SINR), downlink data rate, and decoding error probability; where, time slot drones to ground node downlink signal-to-interference-plus-noise ratio , represented as: ; in, Indicates time slot drones With ground nodes Access control variables between them Indicates time slot Another drone With another ground node Access control variables between them Indicates time slot drones With ground nodes Channel gain between Indicates time slot Another drone With ground nodes Channel gain between For drones Transmit power on each sub-channel; Time slot drones to ground node downlink data rate , represented as: ; in, For the transmission delay of the charge packet, It is the inverse function of the Q function. To maximize the decoding error probability, Channel dispersion is represented as follows: ; drones In the time slot Time-to-ground node When transmitting an accusation packet, the ground node The decoding error probability is expressed as: ; in, Indicates the maximum tolerable delay. Indicates the given group length; and .
6. The method according to claim 1, characterized in that, The joint optimization model is described as a deneutralized partially observable Markov process, including: Acquire drones separately and ground nodes In the time slot The eigenvectors at time t are respectively expressed as: ; ; Among them, drones In the time slot eigenvectors at time Including drones Current location and drones Total number of ground nodes connected in the previous time slot Ground nodes In the time slot eigenvectors at time Including ground nodes Location Was the drone connected in the previous time slot? Up to Average uplink throughput up to the time slot as well as Decoding error probability of time slot ; Based on the feature vectors of both UAVs and ground nodes, the joint optimization model is described as a deneutralized partially observable Markov process. This deneutralized partially observable Markov process is described using environmental state, UAV local observation information, UAV actions, and a reward function. Specifically, the environmental state is described using the set of feature vectors of all entities; the UAV local observation information is described using the UAV's own feature vector, the set of feature vectors of neighboring UAVs within its perception range, and the set of feature vectors of ground nodes; the UAV actions are described using the UAV's movement direction decision and the ID of the ground node accessing the UAV; and the reward function is described using a global reward and an individual UAV penalty value, where the global reward is a weighted sum of total fair throughput and QoS-DL.
7. The method according to claim 1, characterized in that, Priority-based access control algorithms include: For ground nodes, in each time slot, each ground node prioritizes accessing the nearest drone; For drones, a throughput-first principle is adopted for uplink, that is, prioritizing throughput in the first step. In each time slot, ground nodes with lower throughput are prioritized for connection to the UAV; for downlink, a distance-first principle is applied, meaning the nearest ground node is prioritized for connection to the UAV; by weighted summing the normalized throughput and distance, the priority ranking of the set of ground nodes within the UAV's perception range is obtained, expressed as: ; in, For drones For the set of ground nodes within the sensing range Priority sorting, This means that each UAV sorts the ground nodes according to the weighted sum of normalized throughput and distance. Ground nodes with higher priority are given priority to access the UAV. For each ground node that is connected to the UAV, the corresponding UAV will allocate the sub-channel with the least uplink and downlink interference at the current time. ground nodes in the past Total throughput obtained within each time slot This represents the maximum value of the total throughput. These are the weighting coefficients. ground nodes Location, For drones In the time slot The position at that time Represents a set Any ground node in the network, A collection of drones.
8. The method according to claim 1, characterized in that, The QMIX trajectory planning algorithm based on weighted generation network enhancement consists of several agent networks, a hybrid network, and a super network. The intelligent agent network is used to acquire local observation information of the corresponding UAV and fit and output the individual value of the UAV agent. The intelligent agent network consists of three parts: an input layer, an intermediate layer, and an output layer. The input layer consists of two weight generation networks and a multilayer perceptron. The two weight generation networks are used to extract features from the feature vector sets of neighboring UAVs and the feature vector sets of ground nodes within the local observation information of the UAV. The multilayer perceptron is used to extract features from the UAV's own feature vectors within the local observation information of the UAV. The input layer finally merges the outputs of the two weight generation networks and the multilayer perceptron to output an embedded feature. The intermediate layer uses a gated recursive unit, taking the hidden state of the previous time slot and the embedded feature output by the input layer as input, and outputs the hidden state of the current time slot. The output layer uses a multilayer perceptron, which calculates the individual value of the UAV by acquiring the hidden state of the current time slot. Based on its own individual value, the UAV makes action decisions using an ε-greedy strategy. The hybrid network is used to take the individual value of each drone as input and obtain the joint action value through forward propagation; The hypernetwork is used to calculate and obtain the weights and biases of the hybrid network based on the environmental conditions.
9. An access control and trajectory planning device for emergency communication networks, characterized in that, The device includes: The network construction module is used to build an emergency communication network assisted by multiple UAVs and to acquire the uplink and downlink within the network; wherein, the emergency communication network consists of multiple UAVs, multiple ground nodes and a remote control center; The model building module is used to build a limited perception model, a channel model, and a communication model between the UAV and ground nodes based on the emergency communication network. The problem modeling module is used to construct a joint optimization model for the access control and UAV trajectory planning problem by jointly considering the uplink's throughput requirements and the downlink's timeliness and reliability requirements, and based on the finite sensing model, channel model, and communication model. The joint solution module is used to describe the joint optimization model as a deneutralized partially observable Markov process, and to solve the joint optimization model using a priority-based access control algorithm and a QMIX trajectory planning algorithm based on weighted generation network enhancement, so as to obtain the optimal access decision and the optimal trajectory of the UAV. The joint optimization model for the access control and UAV trajectory planning problem is expressed as follows: ; The joint optimization model is a function consisting of continuous variables. and discrete variables The mixed-integer nonlinear programming problem is composed of; where, This is a weighting coefficient used to balance the relative importance of uplink and downlink; This indicates the service period of the ground node for the drone. Total fair throughput, which takes into account the uplink's communication requirements for throughput; service period Classified as There are 3 time slots of equal length, each time slot having a length of 1. ; Indicates time slot Fairness coefficient at that time ground nodes in the past Average throughput within each time slot Indicates the number of ground nodes; Indicates time slot Time ground node Uplink data rate; This indicates the number of downlink QoS-DLs that meet QoS requirements, where the QoS-DLs are used to consider the downlink's communication requirements for timeliness and reliability. This indicates an indicator function, specifically represented as ; To maximize the decoding error probability, Indicates time slot Time ground node The probability of decoding errors; Indicates the number of drones; The joint optimization model satisfies the following constraints: ; Among them, constraint (1) means that only ground nodes within the UAV's perception range can connect to the UAV. Indicates time slot drones With ground nodes Access control variables between them For drones, For the set of ground nodes, The constraint (2) indicates that each ground node can only access one UAV; the constraint (3) indicates that the number of ground nodes accessed by each UAV cannot exceed the number of sub-channels. The total number of sub-channels; constraint (4) represents the range of values for the access control variable and the sensing variable, where, Indicates time slot drones With ground nodes The perceived variables between; constraint (5) indicates that the position of the UAV cannot exceed the boundary of the target area, where, and They represent time slots respectively. drones The x and y coordinates, and The horizontal and vertical boundaries of the target area are represented respectively; constraint (6) indicates that the distance between any two UAVs is greater than the minimum spacing. To avoid collisions, among which, and They represent drones With drones In the time slot The position at that time.
Citation Information
Patent Citations
Track optimization method for auxiliary communication of unmanned aerial vehicle
CN116614827A
Flight path planning and channel selection method oriented to multi-unmanned aerial vehicle assisted Internet of Things data collection
CN118968819A