Dynamic resource management method of heterogeneous air-ground collaborative network under energy constraint

By employing genetic algorithms and MADDPG algorithms combined with SWIPT and NOMA techniques in the air-ground cooperative network, clusters are dynamically divided and resource allocation is optimized, solving the lifecycle bottleneck problem caused by energy differences in UAV and unmanned vehicle networks, and realizing long-term autonomous operation and efficient communication of the network.

CN121842071APending Publication Date: 2026-04-10MIANYANG NETOP TELECOM EQUIP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MIANYANG NETOP TELECOM EQUIP
Filing Date
2025-12-25
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In air-ground collaborative networks composed of drones and unmanned vehicles, existing technologies have failed to effectively address the bottleneck effect and dynamic complexity of network lifecycle caused by energy differences. In particular, under comprehensive energy constraints, it is difficult to achieve long-term autonomous operation of the network and dynamic balance and continuous replenishment of resources.

Method used

An improved genetic algorithm and a multi-agent deep deterministic policy gradient algorithm (MADDPG) are adopted, combined with wireless power-carrying technology (SWIPT) and non-orthogonal multiple access technology (NOMA), to dynamically divide clusters and optimize the resource allocation of UAVs and unmanned vehicles in real time under energy balance constraints, including transmit power, spectrum bandwidth, latency and UAV trajectory.

Benefits of technology

It significantly extends the network lifetime, achieves dynamic global optimization of throughput and energy efficiency, enhances the system's adaptability and robustness, improves the efficiency of spectrum and energy resource utilization, and avoids network fragmentation and communication interruption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121842071A_ABST
    Figure CN121842071A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic resource management method for an unmanned air-ground collaborative network under energy constraint, which comprises the following steps of: under comprehensive energy constraint, adopting a framework of'hierarchical collaboration and dynamic optimization ', firstly, dividing clusters for balancing air-ground network energy, and laying a foundation for prolonging the life cycle of the network from the topology level; further, in each cluster, a communication mode integrating SWIPT and NOMA is constructed, and resources such as unmanned vehicle transmitting power, unmanned aerial vehicle transmitting power, tracks and power division factors are jointly optimized in real time based on a multi-agent learning framework of MADDPG. According to the method, the long-term interaction throughput of the network can be maximized while the energy sustainability is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication networks. More specifically, this invention relates to a dynamic resource management method for an air-ground cooperative heterogeneous network composed of unmanned aerial vehicles (UAVs) and unmanned vehicles (UAVs) under comprehensive energy constraints. Background Technology

[0002] With the rapid development of unmanned systems technology, air-ground collaborative heterogeneous networks composed of drones and unmanned vehicles (UAVs) are demonstrating significant application potential in scenarios such as emergency rescue, wide-area environmental monitoring, and large-scale smart agriculture. In this network, drones, with their high mobility and flexible deployment capabilities, can quickly establish air-to-ground line-of-sight communication links, effectively acting as aerial mobile relays or temporary access points, thereby significantly expanding network coverage. Meanwhile, UAVs, typically equipped with larger batteries and payloads, provide stable and persistent communication support and data processing capabilities on the ground. Together, they can construct a three-dimensional, flexible, and highly resilient operational network.

[0003] However, the efficient operation of such heterogeneous networks faces a fundamental constraint: the overall energy constraint of the system. Given a total energy budget, due to inherent differences in platform physical conditions, the onboard energy reserves of UAVs are typically far lower than those of unmanned vehicles. If a crude or static resource management strategy is adopted, energy-limited UAV nodes will prematurely exhaust their energy and fail, leading to network topology disruption, gaps in coverage areas, and potentially causing the interruption of the entire collaborative mission. This "weakest link" effect in network lifecycle caused by energy differences among heterogeneous nodes is a key bottleneck restricting their long-term autonomous operation. Existing resource optimization schemes mostly focus on static or short-term scales, failing to systematically address the dynamic balance and continuous replenishment of energy between air-to-ground nodes from a global, long-term, and dynamic perspective.

[0004] Furthermore, air-to-ground cooperative networks inherently possess high dynamic complexity. The mobility of UAV nodes leads to rapid time-varying changes in network topology and air-to-ground channel conditions. Inter-node communication, especially when employing non-orthogonal multiple access technologies to improve spectral efficiency, introduces severe co-channel interference. The system's adjustable variables, such as transmit power, spectral bandwidth, UAV flight trajectory, and communication scheduling delay, are tightly coupled across multiple dimensions of resources. These factors pose significant challenges to theoretical methods based on traditional convex optimization or static programming when solving such problems. These challenges include the difficulty in accurately characterizing the dynamic environment using mathematical models, the high complexity of solutions due to the non-convex nature of the problem, and the inability to meet the demands of real-time online decision-making.

[0005] To address the aforementioned challenges, several cutting-edge technologies offer partial solutions. Wireless power-carrying technology provides a new approach to overcoming the energy bottleneck of drones, allowing them to harvest energy from radio frequency electromagnetic waves while receiving communication signals, thus enabling online replenishment of onboard batteries. Non-orthogonal multiple access technology, through multi-user multiplexing in the power domain, can serve multiple nodes on the same spectrum resource, significantly improving the system's spectral efficiency. Deep reinforcement learning, particularly its multi-agent version of the deep deterministic policy gradient algorithm, provides a powerful framework for solving multi-agent distributed cooperative decision-making problems in high-dimensional continuous action spaces. It learns through trial and error interaction with the environment, does not rely on precise system models, and exhibits good adaptability to dynamic environmental changes.

[0006] Nevertheless, current research and practice still lack a unified, adaptive, and dynamically adjustable overall framework. This framework needs to organically integrate macro-level network-level energy balance management with micro-level link-level multi-dimensional resource optimization. How to collaboratively utilize wireless power-carrying technology to continuously power UAVs, leverage non-orthogonal multiple access technology to improve spectrum utilization, and employ multi-agent deep deterministic policy gradient algorithms to achieve real-time intelligent decision-making, thereby simultaneously maximizing the overall network lifetime and long-term average throughput under strict comprehensive energy constraints, remains an open problem to be solved. Summary of the Invention

[0007] One object of the present invention is to solve at least the above-mentioned problems and / or defects, and to provide at least the advantages described below.

[0008] To achieve these objectives and other advantages of the present invention, a dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints is provided, comprising:

[0009] S1. Construct a heterogeneous air-ground collaborative network using drones and unmanned vehicles as nodes, set the total energy budget E_total for network operation, and divide the entire mission time into continuous time slots τ.

[0010] S2. An improved genetic algorithm is used, with energy balance as a constraint, to partition the heterogeneous air-ground cooperative network to obtain multiple energy-balanced clusters C. K Each cluster C K It includes at least one drone node and at least one unmanned vehicle node. Each node is equipped with a single antenna and simultaneously supports the communication modes of Wireless Powered Communication (SWIPT) and Non-Orthogonal Multiple Access (NOMA).

[0011] S3, transfer each cluster C KThe drones and unmanned vehicles in the model are all regarded as intelligent agents. A partially observable Markov decision process model is constructed. For each time slot τ, the multi-agent learning framework in the multi-agent deep deterministic policy gradient (MADDPG) algorithm is applied. At the same time, SWIPT and NOMA are combined to perform real-time and distributed joint optimization of resources in the cluster.

[0012] S4. Check if the triggering conditions for network re-division have been met. If they have been met, pause real-time optimization, collect the current network energy and location information, and return to S2 until the task time ends or the total energy is exhausted.

[0013] The resources in S3 include: vehicle transmit power, drone transmit power, power splitting factor, network energy, spectrum bandwidth, latency, and drone trajectory.

[0014] Preferably, in S1, the total energy budget E total It is characterized by the following formula:

[0015]

[0016] In the above formula, Let be the initial onboard energy of the i-th drone. Let be the initial onboard energy of the j-th autonomous vehicle, and M represents the number of drones, and N represents the number of unmanned vehicles.

[0017] Preferably, in S2, the chromosome encoding of the genetic algorithm uses integer encoding, and each gene bit represents a cluster label of a node. The fitness function of the genetic algorithm is characterized by the following formula:

[0018]

[0019] In the above formula, f is the objective function, and d ij D represents the Euclidean distance between node i and node j. max P represents the maximum reliable communication distance between node i and node j. s For a heterogeneous air-ground collaborative network structure, K is the number of clusters, and ρ1 and ρ2 represent the weight coefficients of the penalty term.

[0020] Preferably, in S2, under the SWIPT communication mode, the UAV can harvest energy from the downlink radio frequency signal sent by the unmanned vehicle, and the harvested power P harv,i (t) is:

[0021] P harv,i (t)=η(1-α i (t))β i (t)P tx,j (t)i,j∈CK

[0022] Where η is the energy conversion efficiency, α i (t) represents the received signal power division ratio used for information decoding when UAV i receives signals, and 0 ≤ α. i (t)≤1i∈C k ∩{M},β i (t) is the proportion of power allocated by node j (the transmitter) to node i (the receiver), and S i Let {M} be the set of receiving nodes served by the transmitting node i on the current resource block, and {M} be the set of drones. tx,j (t) represents the unmanned vehicle signal transmission power serving the unmanned vehicle i.

[0023] Preferably, in S2, under NOMA communication mode, the achievable rate R of the target receiving node i is... i (t) is characterized by the following formula:

[0024] R i (t)=B log2(1+SINR i (t))

[0025] In the above formula, B is the transmission bandwidth, and SINR is... i (t) represents the signal-to-interference-plus-noise ratio (SINR) of the target receiving node i in time slot τ, and SINR i (t) is characterized by the following formula:

[0026]

[0027] In the above formula, h ji (t) is the channel gain from node j to node i, N0 is the noise power spectral density, and P tx,j (t) represents the transmit power of the unmanned vehicle, S i β is the set of receiving nodes served by transmitting node i on the current resource block. k (t) represents the proportion of power allocated by node k (the transmitter) to node i (the receiver), where β is the power ratio. i (t) is the proportion of power that node j, acting as the transmitter, allocates to node i, acting as the receiver.

[0028] Preferably, in S3, the method for real-time, distributed joint optimization of resources is as follows:

[0029] S30. In the multi-agent deep deterministic policy gradient (MADDPG) algorithm, the network parameters of MADDPG are initialized.

[0030] S31. For each time slot τ, each agent collects its own local observation state s. i(t), the s i (t) is the sum of its remaining energy E. i (t), its own position coordinates q i (t) Distance d to all other nodes j in the cluster ij (t) Channel State Information (CSI) of the previous time slot ij (t-1) estimated value, current length of the data queue to be transmitted Q i (t), Energy harvesting power P in the previous time slot harv,i The vector of (t-1);

[0031] S32, s i (t) Input to the action network, output a joint action that can be performed by the UAV and the unmanned vehicle. i (t), joint action a i (t) is the action a of the drone agent. uav (t), Autonomous vehicle intelligent agent action a uav The vector of (t), and a uav (t)=[v uav ,P tx ,{β1,β2,...β M}],a ugv (t)=[P tx ,α,{β1,β2,...β N}],v uav P represents the horizontal speed command used to determine the flight path of the drone. tx β represents the transmit power of a drone or unmanned vehicle. M β represents the power allocation ratio coefficient for the drone. N This represents the power distribution ratio coefficient for autonomous vehicles.

[0032] S33. The UAV communicates between nodes in SWIPT communication mode, updates the energy, position, and queue of the corresponding node by calculating the throughput of each link, and the central controller calculates the total cluster reward R according to the formula. total ;

[0033] S34, s i (t), a i (t), R toal s i (t+1) is stored as experience in the experience pool. When the predetermined number of learning steps is reached, the network is updated using the MADDPG algorithm by sampling small batches of data, where s i (t+1) represents the observation value at the next time step.

[0034] Preferably, in S30, in the multi-agent learning framework, each agent i maintains an Actor network and a Critic network. After training, each agent makes independent decisions based on its local Actor network, thereby achieving distributed real-time optimization.

[0035] During the training phase, each agent's Critic network can acquire global state information s and the joint action a of all agents, and output a Q-value to evaluate the quality of the global state-action profile. During the network application phase, each agent's Actor network relies solely on its own locally observed state s. i Output deterministic action a i ;

[0036] In S34, by sampling mini-batch data, the Critic network updates the network by minimizing the temporal difference error, and the Actor network updates the network by policy gradient ascent.

[0037] Preferably, in S33, the total cluster reward R total The global reward function obtained in time slot τ is characterized by the following equation:

[0038]

[0039] In the above formula, R sum (t) characterizes the overall spectral efficiency of all working communication links in time slot τ, and Let D be the signal-to-interference-plus-noise ratio (SIR) from node i to node j in time slot τ. E (t) is the set of nodes within each cluster that violate physical and behavioral constraints, H sum (τ) represents the energy actively harvested by the drone using SWIPT technology, and Let m be the actual energy power collected by the UAV in time slot τ. A normalization factor to ensure training stability.

[0040] Preferably, in S4, the trigger condition for network repartitioning is: when it is detected that the energy of any cluster is about to be exhausted and is severely unbalanced with other clusters, the network repartitioning process is triggered.

[0041] This invention offers at least the following beneficial effects: It comprehensively utilizes Simultaneous Wireless Information and Power Transfer (SWIPT), Non-Orthogonal Multiple Access (NOMA) technology, and the Multiagent Deep Deterministic Policy Gradient (MADDPG) algorithm to jointly optimize network energy, transmit power, spectrum bandwidth, latency, and UAV trajectories, thereby maximizing both network lifetime and overall communication throughput. Specifically, it achieves the following effects:

[0042] First, this invention significantly extends the overall network lifespan. By using "energy-balanced" network partitioning, it proactively avoids the risk of network fragmentation caused by the rapid depletion of energy by a few drones at the topology level. Combined with intra-cluster SWIPT technology to continuously "recharge" drones, it directly replenishes drone energy at the physical level. This two-pronged approach synchronizes the effective working time of air and ground nodes, maximizing the utilization of the fixed total energy budget.

[0043] Secondly, this invention enables dynamic global optimization of throughput and energy efficiency. Specifically, it employs the MADDPG framework, which responds in real-time to changes in network conditions such as channel, location, energy, and queue size, jointly optimizing continuous and discrete actions to maximize long-term cumulative throughput. The reward function cleverly integrates energy balance and constraints into the optimization objective, allowing the algorithm to automatically learn the optimal strategy while satisfying various physical limitations.

[0044] Third, this invention can enhance the system's adaptability and robustness. That is, the deep reinforcement learning-based method does not rely on accurate and hard-to-obtain system models, but learns by interacting with the environment, which has a stronger ability to adapt to network dynamics and uncertainty. At the same time, its distributed execution mechanism ensures low-latency real-time decision-making.

[0045] Fourth, this invention can improve the efficiency of spectrum and energy resource utilization. Specifically, by integrating NOMA technology into intra-cluster communication, the spectrum reuse rate is improved. By jointly optimizing the trajectory and SWIPT, the UAV can receive energy and information at the optimal location, thereby improving the efficiency of wireless energy transmission. The synergistic optimization of multi-dimensional resources produces an overall gain of "1+1>2".

[0046] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description

[0047] Figure 1 This is a flowchart illustrating the dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints according to the present invention. Detailed Implementation

[0048] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description.

[0049] This invention provides a dynamic resource management method for unmanned air-ground cooperative networks under energy constraints. The core idea of ​​this method is: under comprehensive energy constraints, a "hierarchical cooperation and dynamic optimization" framework is adopted. First, clusters are divided with the aim of balancing the energy of the air-ground network, laying the foundation for extending the network lifetime at the topological level. Then, within each cluster, a communication mode integrating SWIPT and NOMA is constructed, and based on the MADDPG multi-agent learning framework, resources such as unmanned vehicle transmission power, UAV transmission power, trajectory, and power partitioning factor are jointly optimized in real time, maximizing the long-term interactive throughput of the network while maintaining energy sustainability.

[0050] A dynamic resource management method for an energy-constrained unmanned air-ground cooperative network is proposed. In a heterogeneous network consisting of M unmanned aerial vehicles (UAVs) and N unmanned vehicles (UAVs), under the strict constraint of a pre-set total energy budget E_total, a two-stage hierarchical optimization framework combining "dynamic partitioning of macroscopic network topology" and "real-time optimization of microscopic intra-cluster resources" is employed to synergistically maximize the overall network lifecycle and the long-term interactive throughput of the system. The method includes the following steps:

[0051] Step 1: Problem Modeling

[0052] Consider a rectangular mission area containing M UAV nodes and N unmanned vehicle nodes. All nodes are equipped with a single antenna and simultaneously support SWIPT and NOMA communication modes. The UAVs fly at a fixed altitude of H, and their maximum horizontal speed is... The unmanned vehicle moves along a predetermined path on the ground to carry out its work, with an average speed of [missing information]. Since air-to-ground network communication is free from excessive obstruction, the channel model employs a combination of LoS probability air-to-ground path loss and small-scale Rayleigh fading. SWIPT technology utilizes power partitioning; after receiving the signal, the UAV divides it into different proportions for information decoding and energy harvesting. All UAVs and unmanned vehicles form an air-to-ground cooperative network, where the total energy budget for network operation is defined as E. total The system runtime is discretized into a time-slot sequence of length τ, and the total duration of this network collaborative operation is T. The initial onboard energy of each UAV i is... The initial onboard energy of each autonomous vehicle j is and much smaller The total energy constraint is satisfied between the two pairs:

[0053]

[0054] Due to energy differences, drones and unmanned vehicles need to be divided into different clusters (C). k They also transmit information to each other, thereby effectively coordinating to improve internal data transmission efficiency and synchronously balance system energy. Therefore, the objective function is established as follows:

[0055]

[0056] Among them, R ij (t) represents the achievable transmission rate from node i to node j in a single time slot, where {M} is the set of drones and {N} is the set of unmanned vehicles.

[0057] After determining the objective function, constraints need to be placed on the physical characteristics and operational capabilities of the mobile nodes to conform to practical realities. For each node, its net energy consumption in any time slot must not exceed its available energy. For drones, this also applies. in P is the remaining energy of UAV i at the beginning of time slot τ. c,i (t) is the total power consumption of UAV i, P harv,i (t) is the power it obtains through wireless power acquisition.

[0058] For driverless cars, their existence P c,j (t) is the total power consumption of the autonomous vehicle j, E j (t) represents.

[0059] At the same time, the transmit power of each node must be limited to a safe and hardware-permissible range, which has a power limit of 0 ≤ P tx,i / j (t) represents the operating power of the drone or the drone itself. This refers to the maximum launch power of the drone or drone.

[0060] In addition, the drone's flight speed must not exceed its maximum speed limit to ensure flight controllability. A minimum safe distance d must be maintained between any two drones. min That is, ||q i (t)-q i′ (t)||≥d min i∈C k ∩{M},q i (t) represents the position coordinates of drone i, q i′(t) represents the position coordinates of UAV i′, thus ensuring the safety of nodes in the network collaborative operation. Regarding energy harvesting, the power partitioning factor must be within a valid range, i.e., 0 ≤ α. i (t)≤1i∈C k ∩{M},α i (t) represents the received signal power segmentation ratio (also known as the segmentation factor) used for information decoding when UAV i receives signals. Power allocation when using power domain multiplexing on the same resource block must also satisfy normalization and positivity, i.e. β i (t) is the power allocation ratio from node j to node i, S i This is the set of communication services for node j. In an orthogonal frequency division multiplexing (OFDM) system, subcarriers must be allocated orthogonally to avoid intra-cluster interference. ρ i,j,c (τ)∈{0,1} is a binary indicator variable, indicating whether subcarrier c is assigned to link (i,j) in time slot τ.

[0061] In addition, the energy consumption model, synchronous transmission model, and communication model of the collaborative network nodes are established as follows:

[0062] 1) Power consumption model:

[0063] The power consumption of a drone mainly includes propulsion power consumption and communication power consumption. For a drone M, the total power consumption P in time slot τ is... tx,i (t) is:

[0064]

[0065] Among them, P prop (.) represents the propulsion power of the UAV, which is its speed. With acceleration a i The two-dimensional function of (t) is expressed as:

[0066]

[0067] Where c1 and c2 are constants related to the aerodynamic parameters of the UAV, and g is the acceleration due to gravity. It is the sum of its communication circuit and signal transmission power.

[0068] The main energy consumption of autonomous vehicles is communication energy consumption. It consists of the unmanned vehicle's transmission power and circuit power consumption.

[0069] 2) SWIPT transmission model:

[0070] The drone can harvest energy from the downlink radio frequency signal transmitted by the unmanned vehicle. A power-divided receiver structure is adopted. In time slot τ, the signal power of the unmanned vehicle j serving the drone i is P after channel attenuation.rx,j (t)H ij (t), where H ij (t) represents the total transmission channel loss, P rx,j (t) represents the signal received by UAV i according to the power division factor α. i (t) is divided, where α i (t) The power ratio is used for information decoding, 1-α i The power proportional to (t) is used for energy harvesting. Therefore, the harvested power P harv,i (t) is:

[0071] P harv,i (t)=η(1-α i (t))β i (t)P rx,j (t)i,j∈C K (5)

[0072] Where η∈(0,1) is the energy conversion efficiency.

[0073] 3)NOMA communication model:

[0074] Within each resource block, downlink power domain NOMA technology is employed. Assume that within a cluster, L nodes simultaneously serve on a single resource block; that is, one autonomous vehicle simultaneously transmits signals to multiple drones, or one drone simultaneously transmits signals to multiple autonomous vehicles. The transmitter assigns different power levels to signals from different nodes, and the receiver uses serial interference cancellation technology for decoding. For the target receiving node i, its signal-to-interference-plus-noise ratio (SINR) within time slot τ is... i (t) is:

[0075]

[0076] Among them, h ji (t) is the channel gain from node j to node i, and N0 is the noise power spectral density. Therefore, the achievable rate of node i is R. i (t)=B log2(1+SINR i (t)), where B is the transmission bandwidth.

[0077] The aforementioned problems involve multidimensional discrete and continuous variables, and complex linear and nonlinear constraints. To effectively address the solution of this complex optimization problem, ensure solution accuracy, and maintain solution efficiency, this invention proposes a two-step processing method. First, the collaborative network is divided into several clusters using the energy of the balanced air-ground network as the main partitioning index, thereby achieving range locking of the SWIPT technique. Then, the dynamic resources in the collaborative network are jointly optimized, thereby improving the network collaboration level and work efficiency.

[0078] Step 2: Dynamic Network Partitioning Based on Energy Balance

[0079] To pre-balance energy load at the network topology level and overcome the energy limitations of UAVs, the global network is dynamically divided into K energy-balancing clusters, each with a Cenergy value. k (k = 1, 2, ..., K) must include at least one UAV node and at least one unmanned vehicle node to ensure air-to-ground collaborative capabilities within the cluster. This minimizes the difference in total available energy between clusters while satisfying the constraints of node uniqueness and maximum communication distance within the cluster, achieving pre-balancing of network energy. For cluster C... k Its total available energy E cluster,k Defined as:

[0080]

[0081] The network partitioning problem can be formalized as a combinatorial optimization problem, and its mathematical expression is as follows:

[0082]

[0083] Where Var() represents variance, which measures the degree of balance in total available energy among clusters. This represents the maximum energy difference among nodes in cluster k. Var() is specifically calculated as follows:

[0084]

[0085] In the above formula, K represents the number of clusters. This represents the average energy of the cluster;

[0086] To ensure that each node must belong to one and only one cluster, constraints need to be established:

[0087]

[0088] Where, x ik ∈{0,1} is a binary decision variable. When node i belongs to cluster k, x ik =1, otherwise 0. Simultaneously, each cluster must contain at least one drone node and one autonomous vehicle node to ensure air-to-ground coordination capabilities within the cluster. Therefore:

[0089]

[0090] in, N total For cluster C k The number of nodes in the system.

[0091] To ensure reliable communication between nodes within a cluster, the distance between any two nodes within the cluster must not exceed the maximum reliable communication distance. Therefore:

[0092]

[0093] Where, d ij Let x represent the Euclidean distance between node i and node j. x is true if and only if nodes i and j belong to the same cluster k. ik =x jk =1, the constraint is activated, requiring the distance between the two to not exceed D. max Furthermore, during network operation, the cooperative network structure P also needs to be adjusted. s With constraints imposed, its modeling is as follows:

[0094]

[0095] Among them, |C K ∩{M}| represents cluster C K The number of drone nodes in China, |C K ∩{N}| represents cluster C K The number of autonomous vehicle nodes. This penalty term stipulates that if a certain cluster C... K If there are no drone nodes, the first item is 1; if there are no autonomous vehicle nodes, the second item is 1. The two items respectively penalize violations of the constraints "each cluster must contain at least one drone node" and "each cluster must contain at least one autonomous vehicle node".

[0096] To solve this mixed-integer nonlinear optimization problem, this invention employs an improved genetic algorithm. The algorithm uses integer encoding for chromosomes, with each gene bit representing a cluster label for a node. The fitness function is designed as follows:

[0097]

[0098] The algorithm negates the sum of the objective function f and the penalty term for constraint violation, thus transforming the minimization problem into a fitness maximization problem. Iteratively searches for the optimal partitioning scheme through operations such as selection, crossover, and mutation. This network partitioning is not fixed but rather determined based on the cumulative communication timescale when an energy imbalance in a cluster exceeds a threshold. At that time, recalculation and adjustment are dynamically triggered.

[0099] Step 3: Joint resource optimization within the cluster based on MADDPG and SWIPT-NOMA

[0100] Within each defined cluster, each drone and unmanned vehicle is treated as an agent, and a partially observable Markov decision process model is constructed. A multi-agent deep deterministic policy gradient algorithm is then applied for real-time, distributed joint resource optimization. The action network employs a three-layer fully connected neural network, with the local state s as input. i (t), the output is action a. iThe mean vector of (t) is scaled to [-1, 1] for continuous actions using the tanh activation function and then mapped to the actual range. For discrete actions, the Softmax technique is used for approximate sampling to facilitate gradient backpropagation. The input to the policy network is the concatenation of all agent states and all agent actions. After passing through several fully connected layers, it outputs a scalar Q-value. The specific training parameter design is as follows:

[0101] 1) State space:

[0102] The local observation state s of agent i in time slot τ i (t) is a vector containing: energy state, position and motion state, network and channel state, and energy harvesting state. That is, its remaining energy E. i (t), its own position coordinates q i (t) Distance d to all other nodes j in the cluster ij (t) Channel State Information (CSI) of the previous time slot ij (t-1) estimated value, current length of the data queue to be transmitted Q i (t), Energy harvesting power P in the previous time slot harv_i (t-1).

[0103] 2) Motion space:

[0104] The joint action a performed by agent i i (t) is a high-dimensional vector, including: the drone agent's action a uav (t)=[v uav ,P tx ,{β1,β2,...β M}], where v uav The horizontal speed command, P, determines the drone's flight path. tx P represents tx β represents the transmit power of the drone. M This represents the power allocation ratio coefficient for the drone. (Action a of the unmanned vehicle intelligent agent) ugv (t)=[P tx ,α,{β1,β2,...β N}], α represents joint action, β N This represents the power distribution ratio coefficient for autonomous vehicles.

[0105] 3) Reward function

[0106] The global reward function R obtained by the cluster (also called the cluster) in time slot τ total (τ) is designed as follows, and this function aims to simultaneously optimize spectral efficiency, promote energy balance, and ensure that system constraints are met:

[0107]

[0108] in, SINR is the overall spectral efficiency of the cluster in time slot τ, i.e., the sum rate of all working communication links. ij (τ) represents the signal-to-interference-plus-noise ratio (SIR) from node i to node j in time slot τ, and its calculation takes into account the interference introduced by NOMA and D. E (t) represents the set of physical constraints between nodes, used to penalize nodes within the cluster for violating physical and behavioral constraints. This is intended to encourage drones to actively harvest energy using SWIPT technology to alleviate their energy bottlenecks. harv,m (t) represents the actual energy power collected by the UAV m in time slot τ. This is a normalization factor used to adjust the magnitude of this reward item to be comparable to other sub-items, ensuring training stability. Reward function R total (τ) is the direct learning objective of the agent in the MADDPG algorithm. By maximizing the long-term cumulative reward, the agent will learn to collaboratively optimize the UAV trajectory, resource allocation, and SWIPT parameters while satisfying all physical and security constraints, thereby achieving the best trade-off between maximizing spectral efficiency and network energy balance.

[0109] 4) Training Process

[0110] The simulation environment employs offline training. First, all Actor and Critic networks, along with their target network, are initialized. In each round, the environment is reset, and the agent interacts a certain number of steps according to the current policy, storing experience tuples in the experience pool. Small batches of data are periodically sampled from the experience pool to calculate the TD error and update the Critic network. Then, the gradients calculated by the Critic network are used to update each Actor network. Finally, the target network is softly updated. Training continues until the cumulative reward converges. After convergence, the trained model can be deployed. During runtime, each agent independently perceives its own state, inputs it into its local Actor network, obtains specific action commands, and controls the drone and unmanned vehicle to perform tasks. The central controller only needs to periodically collect a small amount of data for model fine-tuning or triggering network re-partitioning, such as... Figure 1 As shown, specifically:

[0111] 1. System initialization: Deploy nodes and set the total energy E. total Initialize the MADDPG network parameters and perform the first energy-balanced network partitioning.

[0112] 2. Enter the main loop, for each time slot:

[0113] 2.1 State Awareness: Each agent collects its own locally observed state s i(t).

[0114] 2.2 Decision Generation: Each agent will generate s i (t) Input its action network and output action a i (t).

[0115] 2.3 Action Execution: The UAV and unmanned vehicle perform actions, such as adjusting the transmission power, segmentation factor, resource block, speed, etc.

[0116] 2.4 Environmental Interaction and Reward Calculation: Inter-node communication, UAVs perform SWIPT, calculate the throughput of each link, update node energy, position, and queue, and the central controller calculates the total cluster reward R according to the formula. total .

[0117] 2.5 Experience Storage and Learning: Storing and learning experiences (s) i (t),a i (t),R total ,s i (t+1) is stored in the experience pool. If the designed number of learning steps is reached, the sample is used to update the MADDPG network.

[0118] 2.6 Check re-partitioning conditions: Check whether the triggering conditions for network re-partitioning have been met.

[0119] 2.7 If repartitioning is required: Pause real-time optimization, collect current network energy and location information, execute the improved genetic algorithm of the steps, generate a new network partitioning scheme, update the agent composition of each cluster, and then continue.

[0120] 3. Repeat step 2 until the task time expires or network energy is exhausted.

[0121] Through the above implementation methods, the present invention achieves integrated and adaptive dynamic optimization management of the lifecycle and communication performance of unmanned air-ground cooperative networks under strict energy constraints.

[0122] 5) MADDPG Algorithm Framework

[0123] Each agent i maintains an Actor network and a Critic network. During the training phase, the Critic networks of all agents can acquire global state information s = (s1,...,s...). N The joint action a = (a1,...,a2) and all agents. N The global state-action pair is evaluated using a Q-value. Each agent's Actor network is based solely on its own locally observed states. i Output deterministic action a iThe algorithm adopts a "centralized training, distributed execution" paradigm. During the training phase, the central controller collects experience tuples (s) generated by the interactions of all agents. t ,a t ,R total (t),s t+1 The data is stored in a shared experience replay pool. By sampling small batches of data, the Critic network updates itself by minimizing the temporal difference error. The gradient formula is:

[0124]

[0125] in, y represents the target Q value, r represents the immediate reward, represents , γ represents the discount factor, and Q represents . i′ This represents the Q-value estimate of the target Critic network at the next time step, where s′ represents the state at the next time step. (Table...) Indicates the parameters of the Critic network. Network parameters representing replication, Indicates to Find the gradient. E represents the loss function of the Critic network. (s,a,r,s′) Let Q represent the sample expectation. i This represents the current network prediction value;

[0126] Subsequently, the Actor network is updated via policy gradient ascent. The gradient is provided by the Critic network, and its expression is:

[0127]

[0128] Among them, Q' i and μ' i These are the target Critic and target Actor networks, used for stable training. Indicates to gradient, Represents the Actor network parameters. This represents the loss function of the Actor network. Indicates that for a i Find the gradient, μ i Indicates network output, a N This represents the network action value. After training, each agent independently makes decisions based on its local Actor network, achieving distributed real-time optimization.

[0129] Step 4: Dynamic Updates and Collaborative Operations

[0130] The network forms a cluster structure based on the partitioning results of step two, and each cluster executes the MADDPG optimizer in step three in parallel. In each communication time slot, the agent makes resource allocation and movement decisions based on the current state. The network state is updated accordingly. When the preset repartitioning trigger condition is met, the process returns to step two to repartition and adjust the network structure, thus forming a closed-loop management framework of "dynamic topology partitioning at the macro level - real-time optimization of micro-level resources."

[0131] Example

[0132] The core of this embodiment lies in achieving intelligent management of unmanned air-ground cooperative network resources through a dynamic cyclical process. The entire process can be divided into the following key stages and steps:

[0133] Phase 1: System Initialization and Initial Network Partitioning

[0134] 1. Network deployment and parameter settings:

[0135] Deploy a predetermined number of drones and unmanned vehicles within the target area, and set the total energy budget for the entire network operation, which equals the sum of the initial energy of all drones and unmanned vehicles. Divide the entire mission time into consecutive time slots, and the system will operate on a time-slot basis. Initialize the core parameters of all neural networks in the MADDPG algorithm.

[0136] 2. Perform the initial network partitioning:

[0137] An improved genetic algorithm is used to solve this grouping problem, dividing the entire network into several smaller subclusters. The goal is to make the total available energy of each subgroup, including current energy and potential future replenishment energy, as similar as possible, thereby preventing network paralysis due to premature power loss in some nodes. The algorithm generates a large number of possible grouping schemes, each of which needs to be evaluated. The primary criterion is whether the energy of each group is balanced. Secondly, each group must contain at least one drone and one unmanned vehicle, and the distance between any two nodes within a group must not be too far to ensure smooth communication. The algorithm simulates an "evolutionary" process, including selection, crossover, and mutation operations, continuously eliminating inferior schemes and retaining and optimizing good ones, ultimately finding a near-optimal network partitioning result.

[0138] Phase Two: Real-time Collaborative Optimization of Intra-cluster Resources

[0139] 1. Local state awareness:

[0140] After the network is partitioned, each cluster runs MADDPG's intelligent decision-making algorithm independently and in parallel. Each drone and unmanned vehicle is regarded as an intelligent agent, and at the beginning of each time slot, it collects its own local information, including: energy state, position and motion state, network and communication state.

[0141] 2. Intelligent decision generation:

[0142] Each node inputs its own state information into its local action decision network, which is pre-trained and can quickly calculate the optimal action instruction based on the current state.

[0143] 3. Action execution and interaction with the environment:

[0144] All nodes simultaneously execute the action commands they receive. The drone adjusts its flight path and communication parameters, and the unmanned vehicle transmits signals based on SWIPT technology. These signals simultaneously provide power and transmit data to the drone. The nodes communicate with each other using NOMA technology, which allows multiple nodes to be served on the same channel, thus improving efficiency.

[0145] 4. System status update and performance evaluation:

[0146] Based on the actions and interactions of the nodes, the state of the entire network is updated, including the energy consumed and collected by each node, its new location, changes in the data queue, etc. Then, the central controller evaluates the performance of each node within this time slot using a comprehensive reward function. This reward function considers dimensions such as communication efficiency, energy balance, and safe operation, and simultaneously stores experience data in an experience pool.

[0147] 5. Artificial intelligence model updates:

[0148] When the experience pool contains a sufficient number of samples, a batch of samples is periodically randomly selected to train the two core networks: the evaluation network and the action decision network. Through this periodic learning process, the entire system becomes increasingly intelligent and better able to adapt to environmental changes.

[0149] Phase 3: Dynamic monitoring and reconstruction of the network

[0150] 1. Continuous monitoring and reclassification:

[0151] The system continuously monitors the status of the entire network, especially whether the energy levels of each cluster remain balanced. When it detects that a cluster is about to run out of energy and is severely imbalanced with other clusters, the system will trigger a network re-partitioning process.

[0152] 2. Perform repartitioning and updates:

[0153] Pause the real-time optimization in the second stage, collect the latest energy and position information of all nodes, jump back to step 2 in the first stage, use the improved genetic algorithm and based on the latest network state to recalculate and generate a new network partitioning scheme, update the membership of each cluster according to the new grouping, and drive each cluster to run the real-time optimization process of the second stage independently and in parallel again.

[0154] The above solution is merely an illustration of a preferred example and is not limited thereto. When implementing this invention, appropriate substitutions and / or modifications can be made according to the user's needs.

[0155] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. It can be applied to various fields suitable for the present invention. Other modifications can be readily made by those skilled in the art. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and examples shown and described herein.

Claims

1. A dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints, characterized in that, include: S1. Construct a heterogeneous air-ground collaborative network using drones and unmanned vehicles as nodes, set the total energy budget E_total for network operation, and divide the entire mission time into continuous time slots τ. S2. An improved genetic algorithm is used, with energy balance as a constraint, to partition the heterogeneous air-ground cooperative network to obtain multiple energy-balanced clusters C. K Each cluster C K It includes at least one drone node and at least one unmanned vehicle node. Each node is equipped with a single antenna and simultaneously supports the communication modes of Wireless Powered Communication (SWIPT) and Non-Orthogonal Multiple Access (NOMA). S3, transfer each cluster C K The drones and unmanned vehicles in the model are all regarded as intelligent agents. An observable Markov decision process model is constructed. For each time slot τ, the multi-agent deep deterministic policy gradient (MADDPG) algorithm is applied. At the same time, SWIPT and NOMA are combined to perform real-time and distributed joint optimization of resources in the cluster. S4. Check if the triggering conditions for network re-division have been met. If they have been met, pause real-time optimization and collect the current network energy and location information. Return to S2 until the task time ends or the total energy is exhausted. The resources in S3 include: vehicle transmission power, drone transmission power, power division factor, network energy, and drone trajectory.

2. The dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints as described in claim 1, characterized in that, In S1, the total energy budget E total It is characterized by the following formula: In the above formula, Let be the initial onboard energy of the i-th drone. Let be the initial onboard energy of the j-th autonomous vehicle, and M represents the number of drones, and N represents the number of unmanned vehicles.

3. The dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints as described in claim 1, characterized in that, In S2, the chromosome encoding of the genetic algorithm uses integer encoding, and each gene bit represents a cluster label of a node. The fitness function of the genetic algorithm is represented by the following formula: In the above formula, f is the objective function, and d ij D represents the Euclidean distance between node i and node j. max P represents the maximum reliable communication distance between node i and node j. s For a heterogeneous air-ground collaborative network structure, K is the number of clusters, and ρ1 and ρ2 represent the weight coefficients of the penalty term.

4. The dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints as described in claim 1, characterized in that, In S2, under the SWIPT communication mode, the UAV can harvest energy from the downlink radio frequency signal sent by the unmanned vehicle. The harvested power P... harv,i (t) is: P harv,i (t)=η(1-α i (t))b i (t)P tx,j (t)i,j∈C K Where η is the energy conversion efficiency, α i (t) represents the received signal power division ratio used for information decoding when UAV i receives signals, and 0 ≤ α. i (t)≤1i∈C k ∩{M},β i (t) is the proportion of power allocated by node j (the transmitter) to node i (the receiver), and S i Let {M} be the set of receiving nodes served by the transmitting node i on the current resource block, and {M} be the set of drones. tx,j (t) represents the unmanned vehicle signal transmission power serving the unmanned vehicle i.

5. The dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints as described in claim 1, characterized in that, In S2, under NOMA communication mode, the achievable rate R of the target receiving node i is... i (t) is characterized by the following formula: R i (t)=B log2(1+SINR i (t)) In the above formula, B is the transmission bandwidth, and SINR is... i (t) represents the signal-to-interference-plus-noise ratio (SINR) of the target receiving node i in time slot τ, and SINR i (t) is characterized by the following formula: In the above formula, h ji (t) is the channel gain from node j to node i, N0 is the noise power spectral density, and P tx,j (t) represents the transmit power of the unmanned vehicle, S i β is the set of receiving nodes served by transmitting node i on the current resource block. k (t) represents the proportion of power allocated by node k (the transmitter) to node i (the receiver), where β is the power ratio. i (t) is the proportion of power that node j, acting as the transmitter, allocates to node i, acting as the receiver.

6. The dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints as described in claim 1, characterized in that, In S3, the method for real-time, distributed joint optimization of resources is as follows: S30. In the multi-agent deep deterministic policy gradient (MADDPG) algorithm, the network parameters of MADDPG are initialized. S31. For each time slot τ, each agent collects its own local observation state s. i (t), the s i (t) is the sum of its remaining energy E. i (t), its own position coordinates q i (t) Distance d to all other nodes j in the cluster ij (t) Channel State Information (CSI) of the previous time slot ij (t-1) estimated value, current length of the data queue to be transmitted Q i (t), Energy harvesting power P in the previous time slot harv_i The vector of (t-1); S32, s i (t) Input to the action network, output a joint action that can be performed by the UAV and the unmanned vehicle. i (t), joint action a i (t) is the action a of the drone agent. ugv (t), Autonomous vehicle intelligent agent action a ugv The vector of (t), and a ugv (t)=[v uav ,P tx ,{β1,β2,...β M }],a ugv (t)=[P tx ,α,{β1,β2,...β N }],v uav P represents the horizontal speed command used to determine the flight path of the drone. tx β represents the transmit power of a drone or unmanned vehicle. M β represents the power allocation ratio coefficient for the drone. N This represents the power distribution ratio coefficient for autonomous vehicles; S33. The UAV communicates between nodes in SWIPT communication mode. By calculating the throughput of each link, it updates the energy, position, and queue of the corresponding node. The optimization center calculates the total cluster reward R according to the formula. total ; S34, s i (t), a i (t), R total s i (t+1) is stored as experience in the experience pool. When the predetermined number of learning steps is reached, the network is updated using the MADDPG algorithm by sampling small batches of data, where s i (t+1) represents the observation value at the next time step.

7. The dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints as described in claim 6, characterized in that, In S30, in the multi-agent learning framework, each agent i maintains an Actor network and a Critic network. After training, each agent makes independent decisions based on its local Actor network, achieving distributed real-time optimization. During the training phase, each agent's Critic network can acquire global state information s and the joint action a of all agents, and output a Q-value to evaluate the quality of the global state-action profile. During the network application phase, each agent's Actor network relies solely on its own locally observed state s. i Output deterministic action a i ; In S34, by sampling mini-batch data, the Critic network updates the network by minimizing the temporal difference error, and the Actor network updates the network by policy gradient ascent.

8. The dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints as described in claim 1, characterized in that, In S33, the total cluster reward R total The global reward function obtained in time slot τ is characterized by the following equation: In the above formula, R sum (t) characterizes the overall spectral efficiency of all working communication links in time slot τ, and Let D be the signal-to-interference-plus-noise ratio (SIR) from node i to node j in time slot τ. E (t) is the set of nodes within each cluster that violate physical and behavioral constraints, H sum (τ) represents the energy actively harvested by the drone using SWIPT technology, and P harv,m (t) represents the actual energy power collected by the UAV m in time slot τ. A normalization factor to ensure training stability.

9. The dynamic resource management method for heterogeneous air-ground cooperative networks under energy constraints as described in claim 1, characterized in that, In S4, the network repartitioning is triggered when it is detected that the energy of any cluster is about to be exhausted and is severely unbalanced with other clusters.