Spectrum resource allocation and power control method and system for multi-unmanned aerial vehicle networking
By building a spectrum resource allocation and power control method for multi-UAV networks, using Markov decision-making algorithm and reinforcement learning to optimize channel and power distribution, the problems of link coordination and resource competition in multi-UAV systems are solved, and efficient information transmission and flexible operation capabilities are achieved.
Patent Information
- Application Number
- CN202510549216.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-01
AI Technical Summary
In multi-UAV systems, how to ensure efficient and stable A2G links and U2U links in urban environments with crowded spectrum, coordinate the competition for communication resources of internal links of multi-UAV systems, and achieve efficient and collaborative information transmission capabilities. Traditional methods are difficult to adapt to changing tasks and environments in real time when facing high-dimensional and complex state spaces.
The spectrum resource allocation and power control method of multi-unmanned aerial network is adopted, and the dynamic observations of the U2U communication link are constructed as an agent by constructing a communication resource allocation model, and the dynamic observations of the U2U communication link are used as an agent. The Markov decision algorithm and reinforcement learning generate channel and power joint allocation strategy are used to realize centralized learning and iterative optimization and output the optimal strategy.
It improves the operational flexibility and efficiency of multiple drones under limited communication resources, enhances collaborative modeling capabilities, adapts to dynamically changing environments, and overcomes the limitations of traditional methods.
Smart Images

Figure CN120417063A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of communication resource allocation for multi - UAVs, and particularly to a method and system for spectrum resource allocation and power control in multi - UAV networking. Background Art
[0002] With the rapid development of UAV technology, UAVs are increasingly widely used in multiple fields such as military, agriculture, logistics, and monitoring. Especially in multi - UAV systems, the ability of multiple UAVs to work together to complete complex tasks makes them an important part of future intelligence and automation. However, multi - UAVs face many challenges when performing tasks, and among them, channel allocation and power management are key factors for achieving efficient communication and ensuring the success of tasks.
[0003] In traditional wireless communication systems, the contradiction between the limited spectrum resources and the growing user demands has emerged, especially in multi - user environments, and how to effectively allocate channels and power has become a research hotspot. Due to its high dynamicity and complex network topology, the multi - UAV system further highlights the complexity of this problem. An effective joint channel and power allocation strategy can not only improve the overall performance of the system but also extend the operation time of UAVs and enhance the reliability of task completion.
[0004] In an intelligent transportation supervision system, the high - altitude perspective and flexible mobility of UAVs can achieve real - time monitoring, data collection, and dynamic scheduling of urban traffic. Through the collaborative operation of multiple UAVs, it is possible to monitor traffic conditions in real time in various urban areas, promptly detect traffic accidents or emergencies, assist the traffic management system in monitoring road traffic conditions, analyzing traffic flow, and predicting traffic jams, etc., to improve the management efficiency and safety of urban traffic. Among them, the communication between UAVs (A2G) and between UAVs and the ground control center (U2U) is the core of the efficient operation of the system. The A2G link is used for data aggregation, issuance of control instructions, and transmission of real - time monitoring information to ensure the timely upload of traffic monitoring data and the accurate execution of UAV tasks. The U2U link enables multiple UAVs to share collaborative information in a D2D manner through providing collaborative communication between UAVs, so as to optimize route and task allocation. The A2G and U2U links often play roles simultaneously, which can effectively improve monitoring efficiency and information transparency.
[0005] Although the application prospect of UAVs in intelligent urban traffic monitoring is broad, in the actual deployment and operation process, how to ensure the high - efficiency and stability of A2G and U2U links in a spectrum - crowded urban environment, how to coordinate the communication resource competition within the multi - UAV system, and how to maximize the information transfer ability in an efficient collaborative manner are still technical problems to be urgently solved.
[0006] At present, many works have studied the spectrum dynamic allocation problem: a spectrum allocation algorithm based on improved color-sensitive graph theory coloring has been proposed; a quantum genetic algorithm based on graph theory has been proposed for dynamic spectrum allocation; a spectrum dynamic allocation algorithm based on simulated annealing algorithm has been proposed; a channel selection and power control method combining relevant vector regression (RVR) has been proposed for the problem that the reliability of the UAV data link is seriously threatened in a complex electromagnetic and geographical natural environment. These traditional mathematical methods usually rely on explicit models and assumptions, and need to preset parameters and structures in advance, which seems powerless in the face of high-dimensional and complex state spaces, and usually assume that the dynamics of the system are known, which often makes it difficult to adapt to changing tasks and environments in real time. In a multi-UAV working environment, it is difficult to obtain channel state information (CSI) in a timely and accurate manner, and the rapidly changing channel conditions will significantly increase the uncertainty of resource allocation. On the other hand, in order to support new U2X applications, service demands are also becoming more and more diverse. In the application scenarios mentioned above, it is necessary to maximize the throughput and transmission reliability of U2X at the same time, and such demands are difficult to model in a mathematically precise way. Although a heuristic subchannel allocation algorithm and a power allocation algorithm based on Taylor series and successive convex approximation have been proposed for imperfect CSI conditions to solve the joint subchannel allocation and power allocation problems of a single UAV. However, for a multi-UAV environment, the scalability and collaborative modeling ability of traditional methods are relatively weak, and expert knowledge is often required for parameter tuning and model assumptions, which increases the complexity and uncertainty of modeling and leads to a decrease in the flexibility and efficiency of multi-UAV joint operations. Summary of the Invention
[0007] Based on this, in view of the above technical problems, it is necessary to provide a spectrum resource allocation and power control method and system for multi-UAV networking that can improve the operation flexibility and efficiency of multi-UAVs under limited communication resources.
[0008] A spectrum resource allocation and power control method for multi-UAV networking, the method includes:
[0009] Construct a communication resource allocation model according to the U2U communication link between UAVs and the A2G communication link between UAVs and ground stations in the multi-UAV networking.
[0010] Take the dynamic observation value of the U2U communication link as an agent, and each agent conducts centralized learning through the local observation function generated by the Markov decision algorithm using the communication resource allocation model. According to the observation value of the agent in the current iteration round, obtain the channel state and data packet transmission action of the U2U link corresponding to the agent, and obtain the global reward value based on reinforcement learning and the optimized reward function. Each agent generates a corresponding channel and power joint allocation strategy according to the global reward value.
[0011] The channel states of all agents and the data packet transmission actions are combined to form a global resource allocation action set. The global resource allocation action set and the local observations of the agents in the next iteration round are used as the state in the next iteration round and input into the communication resource allocation model for iterative learning of the channel and power joint allocation strategy until the optimal channel and power joint allocation strategy is output.
[0012] A spectrum resource allocation and power control system for multi - UAV networking, the system comprising:
[0013] A model construction module, configured to construct a communication resource allocation model according to the U2U communication links between UAVs and the A2G communication links between UAVs and ground stations in multi - UAV networking.
[0014] An allocation strategy generation module, configured to use the dynamic observations of the U2U communication links as agents. Each agent performs centralized learning through the local observation function generated by the Markov decision algorithm using the communication resource allocation model, obtains the channel state and data packet transmission actions of the U2U link corresponding to the agent according to the observations of the agent in the current iteration round, and obtains the global reward value based on reinforcement learning and the optimized reward function. Each agent generates the corresponding channel and power joint allocation strategy according to the global reward value.
[0015] A resource allocation module, configured to combine the channel states of all agents and the data packet transmission actions to form a global resource allocation action set. The global resource allocation action set and the local observations of the agents in the next iteration round are used as the state in the next iteration round and input into the communication resource allocation model for iterative learning of the channel and power joint allocation strategy until the optimal channel and power joint allocation strategy is output.
[0016] The above-mentioned spectrum resource allocation and power control method and system for multi-UAV networking first constructs a communication resource allocation model based on the U2U communication links between UAVs and the A2G communication links between UAVs and ground stations. This model abstracts the multi-UAV communication environment and provides a basic framework for subsequent optimization, avoiding the cooperation difficulties caused by the lack of a unified model in traditional methods. Then, the dynamic observation values of the U2U communication links are used as agents, and the agents perform centralized learning using the local observation functions generated by the Markov decision algorithm. This process does not require the intervention of complex expert knowledge. The agents can autonomously learn the variation rules of the communication links from the dynamic observation values, reducing the dependence on manual parameter tuning and model assumptions compared with traditional methods and enhancing the scalability of the scheme. At the same time, the global reward value is obtained based on reinforcement learning and the optimized reward function. The optimized design of the reward function can guide the agents to act in the direction of improving the overall communication efficiency. Each agent generates a joint channel and power allocation strategy according to the global reward value, realizing the cooperative optimization among UAVs and enhancing the cooperative modeling ability. Finally, by forming a global resource allocation action set from the channel states and data packet transmission actions of the agents and using it as the state of the next iteration round to input into the communication resource allocation model for iteration, the joint channel and power allocation strategy is continuously optimized until the optimal strategy is output. This iterative optimization mechanism enables the scheme to adapt to the dynamic changes of the multi-UAV environment, effectively improving the operation flexibility and efficiency of multi-UAVs under limited communication resources and overcoming the limitations of traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 FIG. is an application scenario diagram of a spectrum resource allocation and power control method for multi-UAV networking in an embodiment;
[0018] Figure 2 FIG. is a schematic flowchart of a spectrum resource allocation and power control method for multi-UAV networking in an embodiment;
[0019] Figure 3 FIG. is a schematic diagram of a reinforcement learning framework with multi-U2U links as agents in an embodiment;
[0020] Figure 4 FIG. is a schematic diagram of a D3QN-VAV network structure in an embodiment;
[0021] Figure 5 FIG. is a block diagram of the structure of a spectrum resource allocation and power control system for multi-UAV networking in an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0022] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0023] The spectrum resource allocation and power control method for multi-UAV networking provided by the present invention can be applied to Figure 1 the application scenarios as shown. Among them, the multi-UAV system conducts end-to-end (D2D) communication through the U2U (UAV-to-UAV link) link to exchange key information and cooperate in searching for trapped persons, and at the same time transmits real-time video and sensor data to the ground station through the A2G (Air-to-Ground link) link. Assume that N A2G links (considering the uplink) have been pre-allocated orthogonal spectrum sub-bands with fixed transmission power, that is, the nth A2G link occupies the nth sub-band. Assume that all transceivers use a single antenna, then there are a total of N A2G links and J U2U links in the cluster. Denote the set of A2G links as {1,..., N}, and the set of U2U links as {1,..., J}.
[0024] In one embodiment, as Figure 2 shown, a spectrum resource allocation and power control method for multi-UAV networking is provided. Taking the operation scenario of multiple UAV systems in Figure 1 as an example, the method includes the following steps:
[0025] Step 202, construct a communication resource allocation model according to the U2U communication links between UAVs and the A2G communication links between UAVs and the ground station in the multi-UAV networking.
[0026] Step 204, use the dynamic observation value of the U2U communication link as an agent. Each agent conducts centralized learning through the local observation function generated by the Markov decision algorithm using the communication resource allocation model, obtains the channel state and data packet transmission action of the U2U link corresponding to the agent according to the observation value of the agent in the current iteration round, and obtains the global reward value based on reinforcement learning and the optimized reward function. Each agent generates the corresponding channel and power joint allocation strategy according to the global reward value.
[0027] Step 206, form a global resource allocation action set with the channel states and data packet transmission actions of all agents, and use the global resource allocation action set and the local observation value of the agent in the next iteration round as the state input to the communication resource allocation model for iterative learning of the channel and power joint allocation strategy until the optimal channel and power joint allocation strategy is output.
[0028] In the above method for spectrum resource allocation and power control in multi-UAV networking, first, a communication resource allocation model is constructed based on the U2U communication links between UAVs and the A2G communication links between UAVs and the ground station. This model abstracts the multi-UAV communication environment and provides a basic framework for subsequent optimization, avoiding the cooperation difficulties caused by the lack of a unified model in traditional methods. Next, the dynamic observation values of the U2U communication links are used as agents, and the agents perform centralized learning using the local observation functions generated by the Markov decision algorithm. This process does not require the intervention of complex expert knowledge. The agents can autonomously learn the change rules of the communication links from the dynamic observation values, reducing the dependence on manual parameter tuning and model assumptions compared with traditional methods and enhancing the scalability of the scheme. At the same time, based on reinforcement learning and the optimized reward function, a global reward value is obtained. The optimized design of the reward function can guide the agents to act in the direction of improving the overall communication efficiency. Each agent generates a joint channel and power allocation strategy according to the global reward value, achieving collaborative optimization among UAVs and enhancing the collaborative modeling ability. Finally, by forming a global resource allocation action set from the channel states and data packet transmission actions of the agents and using it as the state of the next iteration round to input into the communication resource allocation model for iteration, the joint channel and power allocation strategy is continuously optimized until the optimal strategy is output. This iterative optimization mechanism enables the scheme to adapt to the dynamic changes of the multi-UAV environment, effectively improving the operation flexibility and efficiency of multi-UAVs under limited communication resources and overcoming the limitations of traditional methods.
[0029] In one embodiment, the first path loss is obtained according to the U2U communication links between UAVs in multi-UAV networking:
[0030] L j,j′ (t) = 20log 10 d j,j′ (t) + 20log 10 f0 + 20log 10 (4π / c0)
[0031] Wherein, L j,j′ (t) is the first path loss between the j-th UAV and the j'-th UAV, d j,j′ is the three-dimensional spatial distance between the j -th UAV and the j '-th UAV, c0 is the speed of light, and f0 is the channel carrier frequency of the U2U communication link. Orthogonal frequency division multiplexing is used to occupy different subbands of the A2G communication link to obtain the second path loss between the UAV and the ground station on the A2G communication link:
[0032] L n,B (t) = P LoS ·L LoS + P NLoS·L NLoS
[0033] wherein, L n,B (t) is the second path loss between the UAV and ground station B on the nth A2U communication link for the transmission of the nth sub-band, P LoS and P NLoS are the probabilities of LoS and NLoS respectively, L LoS and L NLoS represent the path losses under LoS and NLoS conditions respectively; and the channel power gain is obtained according to the A2G communication link between the UAV and the ground station:
[0034]
[0035] wherein, g n,B [n] is the channel power gain of the nth A2G communication link for the transmission of the nth sub-band, H n [n] is the small-scale fading power component related to frequency in the nth A2G communication link, L n,B is the path loss, and the subscript B represents the base station. A spectrum resource sharing and power control strategy is constructed based on the first path loss, the second path loss, and the channel power gain:
[0036]
[0037]
[0038] wherein, is the channel capacity of the nth A2G communication link for transmission on the nth RB, N is the total number of A2G communication links, maxPr is the maximum load transmission success rate of the U2U communication link, T is the transmission time interval, β j [n] indicates whether the jth U2U communication link transmits on the nth RB, is the channel capacity of the UAV on the jth U2U communication link, Δ T is the channel coherence time, L is the size of the U2U data packet generated periodically, J is the total number of U2U communication links, β j′ [n] indicates whether the j'th U2U communication link transmits on the nth RB, is the channel power of the jth U2U communication link, is the channel power of the j'th U2U communication link, P max is the maximum channel power of the U2U communication link, and d is the ground station. A communication resource allocation model is constructed according to the spectrum resource sharing and power control strategy.
[0039] In one of the embodiments, the target network is obtained by optimizing the deep Q-network according to the spectrum resource sharing and power control strategy and the competitive network, and the deep Q-network is used as the main network and the target network to construct a communication resource allocation model:
[0040]
[0041] Among them, is the estimated Q value at the current moment in the communication resource allocation model, D3QN is the communication resource allocation model, L(θ) is the loss function of the communication resource allocation model, s t is the dynamic observation value of the U2U communication link at the current moment, a t is the data packet transmission action at the current moment, θ is the global network parameter of the communication resource allocation model, r t+1 is the reward at the next moment, γ is the discount factor, s t+1 is the channel state of the U2U communication link at the next moment.
[0042] In one of the embodiments, the dynamic observation values of different agents are input into the communication resource allocation model for distributed Markov decision operation to generate the local observation function corresponding to the agent:
[0043]
[0044] G j [n] = {g j [n], g j′,j [n], g j,B [n], g n,j [n]}
[0045] Among them, is the local observation value of the j-th U2U communication link at the current moment, O(S t , j) is, for the observation function of the j-th U2U communication link at the current moment, S t is the dynamic observation value at the current moment, L j is the remaining load of the j-th U2U communication link, T j is the remaining delay of the j-th U2U communication link, e is the number of model iteration trainings, ε is the probability of random action selection, I j is the received interference of the j-th U2U communication link, G j is the observation space of the j-th U2U communication link, N is the total number of sub-bands, g j [n] is the channel gain of the j-th U2U communication link, g j′,j [n] is the co-channel interference of the j'-th U2U communication link on the j-th U2U communication link, g j,B[n] is the co-channel interference of the j-th U2U communication link to the ground station, g n,j [n] is the co-channel interference of the j-th U2U communication link by the n-th A2G communication link.
[0046] In one embodiment, the dynamic observation value of the U2U communication link is used as an agent. Each agent performs centralized learning through a local observation function generated by a communication resource allocation model using the Markov decision algorithm. The observation value of the agent in the current iteration round is used as a training sample and input into the communication resource allocation model. According to the estimated Q value, the loss function is minimized using the gradient descent method, and the network parameters of the deep Q network of the training agent are copied according to a fixed iteration step to update the global network parameters of the communication resource allocation model. The channel state and packet transmission actions of the U2U link corresponding to the agent are obtained according to the global network parameters.
[0047] In one embodiment, based on the reward function of reinforcement learning, the packet delivery result of each U2U communication link is detected at the end of each transmission. If the remaining load value of the current U2U communication link is greater than zero and the packet delivery result is partial delivery, the link reward value is one times the transmission rate value; otherwise, if the packet delivery result is all successfully delivered, the link reward value is two times the transmission rate value. The training reward is generated using cumulative discounted return according to a preset discount rate:
[0048]
[0049] where, G t is the training reward, γ is the discount rate, R t+k+1 is the reward value, k is the number of interactions. The reward function is optimized according to the link reward values and training rewards of all U2U communication links to obtain the optimized reward function:
[0050]
[0051] where, R t+1 is the optimized reward function, δ d is the positive weight of the ground station for balancing the A2G communication link and the U2U communication link, is the channel capacity of the n-th A2G communication link transmitted on the n-th RB, δ u is the positive weight of the UAV for balancing the A2G communication link and the U2U communication link, B j is the reward value of the j'-th U2U communication link. Each agent generates a corresponding joint channel and power allocation strategy according to the global reward value.
[0052] In one embodiment, the channel states of all agents and the data packet transmission actions are combined to form a global resource allocation action set. After the global resource allocation action set and the local observations of the agents in the next iteration round are used as the state inputs of the next iteration round into the communication resource allocation model, the joint channel and power allocation strategy is iteratively trained according to the global reward function and the loss function until the optimal joint channel and power allocation strategy is output.
[0053] In one embodiment, the modeling steps for the U2U channel are as follows: The drones operate in an unobstructed open space. Since the signal transmission cycle time is short, the influence of the change in the drone position within each transmission task can be ignored. Considering mainly the large-scale fading, the U2U link uses the Line-of-Sight (LOS) transmission channel. The path loss between the j-th drone and the j'-th drone can be regarded as the free-space path loss, which can be expressed as:
[0054] L j,j′ (t) = 20log 10 d j,j′ (t) + 20log 10 f0 + 20log 10 (4π / c0)(1)
[0055] where f0 is the carrier frequency of the U2U channel, c0 is the speed of light, and d j,j′ (t) is the 3D spatial distance between the j-th drone and the j'-th drone, that is:
[0056]
[0057] where h is the flight altitude of the drone, and (x, y) are the horizontal coordinates of the drone. Within a coherence time, the channel power gain of the j-th U2U link transmitted through the n-th subband is g j [n], and its expression is:
[0058]
[0059] where, H j [n] represents the small-scale fading related to the frequency, and satisfies E(H j [n]) = 1.
[0060] In one of the embodiments, the A2G channel is modeled as follows: Different A2G links occupy different subbands in the form of orthogonal frequency division multiplexing (OFDM), and OFDM can transform a frequency-selective channel into a flat channel on different subcarriers. Several consecutive subcarriers are grouped into a spectral subband, and it is assumed that the fading within a subband is approximately the same, while different subbands are independent of each other. The A2G channel model between the UAV and the ground station is defined according to 3GPP specifications Release 15, and its path loss is jointly determined by the states of the line-of-sight (LoS) link and the non-line-of-sight (NLoS) link. The path loss between the nth UAV and the ground station is expressed as:
[0061] L n,B (t) = P LoS ·L LoS + P NLoS ·L NLoS (4)
[0062] Where, P LoS and P NLoS are the probabilities of LoS and NLoS respectively, and their expressions are:
[0063]
[0064] P NLoS = 1 - P LoS (6)
[0065] Where,
[0066] d o = max{294.05log 10 h n (t) - 432.98, 18}(7)
[0067] p1 = 233.98log 10 h n (t) - 0.95(8)
[0068]
[0069] d n,B (t) is the 3D spatial distance between the nth UAV and the ground station, that is:
[0070]
[0071] Where, (x b , y b ) are the horizontal coordinates of the ground station, and the expressions of L LoS and L NLoS are:
[0072] L LoS= 30.9+(22.25 - 0.5log 10 h n (t))log 10 d n,B (t)+20log 10 f c (11)
[0073] L NLoS = max{L LoS , 32.4+(43.2 - 7.6log 10 h n (t))log 10 d n,B (t)+20log 10 f c}
[0074] where f c is the carrier frequency of the A2G channel. Considering small-scale fading, the channel power gain of the nth A2G link transmitted through the nth subband is:
[0075]
[0076] where H n [n] is the small-scale fading power component related to frequency in the nth A2G link. Assuming it is a fading coefficient with a unit mean exponential distribution, satisfying E(H n [n]) = 1.
[0077] In one embodiment, the symbols used in the following text are first explained. The superscript represents the receiver type, u represents the receiver is a UAV, d represents the receiver is a ground station, the subscript represents the transmitter / receiver, and the number in the parentheses is the channel number. The channel gain symbols are sorted as shown in Table 1:
[0078] Table 1 Explanation of Channel Gain Symbols
[0079] <![CDATA[g j > Channel gain of link j - U2U <![CDATA[g n,B > Channel gain of link n - A2G <![CDATA[g n,j > Interference gain from link n - A2G to link j - U2U <![CDATA[g j,B > Interference gain from link j - U2U to the ground station <![CDATA[g j′,j > Interference gain from link j'- U2U to link j - U2U
[0080] The SINR (signal-to-interference-plus-noise ratio) of the nth A2G at the ground station is:
[0081]
[0082] where P is the transmit power, σ 2 is the noise power, β j [n] is a boolean quantity indicating whether the j th U2U link is transmitted on the nth RB. A value of 1 indicates that the two links share the spectrum, and a value of 0 indicates the opposite. It is restricted that each U2U link accesses only one orthogonal subband at the same time, that is:
[0083]
[0084] The SINR of the j-th U2U transmitted on the n-th sub-band at the receiving end is:
[0085]
[0086] where I j [n] contains interference generated by A2G and other U2U links on the same sub-band:
[0087]
[0088] According to Shannon's theorem, the channel capacity of the n-th A2G link transmitted on the n-th RB is:
[0089]
[0090] The channel capacity of the j-th U2U link transmitted on the n-th RB is:
[0091]
[0092] where B is the bandwidth of the spectral sub-band. The UAV generates data packets periodically at different frequencies. The U2U link is mainly used for the transmission of periodic cooperative information. Its original objective function is designed as the transmission rate of data packets within time T. Since the faster the transmission rate, the higher the transmission success rate of the U2U load, and this is exactly the manifestation of the transmission reliability of the U2U link. Therefore, in order to more directly express the optimization objective for the U2U link, here the second optimization objective is equivalently transformed into the transmission success rate of the U2U link load, that is, the success rate of transmitting a data packet of a specified size within a specified time, which is expressed as:
[0093]
[0094] L represents the size of the U2U data packets generated periodically, in bits, Δ T is the channel coherence time, and t is the index of the coherence time. The t in represents that the channel capacity of the U2U link is time-varying.
[0095] It should be noted that the model optimization objective can be expressed as:
[0096]
[0097] Due to the high mobility of the drones themselves, the center cannot instantaneously collect the global CSI. Therefore, using distributed U2U resource allocation is a better solution. However, how to coordinate multiple drone links so that they can achieve the system optimization goal in a collaborative manner instead of competing with each other for their own interests is also a key issue that needs to be considered. In addition, Equation (20) involves sequential decisions for multiple coherent time slots within the time constraint T, but it brings difficulties to traditional optimization methods due to its exponential complexity. To address these challenges, this paper converts the problem into a Markov decision process and proposes a joint optimization algorithm for U2U link channel allocation and power control based on centralized training and distributed execution multi-agent deep reinforcement learning (CTDE-MARL), namely the improved Multi-Agent Double Dueling Deep Q-Network with Value-Advantage Variant (MAD3QN-VAV).
[0098] In one embodiment, as Figure 3 shown, each U2U link is regarded as an agent that interacts with the operating environment to obtain experience and feedback, and continuously improves the spectrum allocation (Boolean quantity β j [n]) and power control (P j [n]) strategies. Multiple agents jointly explore an environment and optimize the resource allocation strategy according to the changes in the environmental state. To improve the global performance of the neural network and prevent competition among individual links for maximizing their own interests, a same reward is used for all agents, thus transforming the multiple-agent process from a game into a fully cooperative form for training to optimize the overall system performance.
[0099] It should be noted that the adopted CTDE-MARL algorithm is divided into two stages, namely centralized training and distributed execution. Each agent is assigned an improved D3QN network. In the training stage, each agent can obtain system-level rewards and train its own deep neural network in a centralized learning manner to continuously improve the existing strategy. In the execution stage, each agent selects actions on a time scale comparable to the small-scale channel fading according to the locally obtained environmental observation information and the existing policy output network to perform the resource allocation task.
[0100] In one embodiment, at each time step t, given an environmental value S t , considering that the observation ability of a single agent for the environment is limited, for the j-th U2U link agent, a local environmental observation value is obtained, which is determined by the observation function O(S t ,j). Each agent selects an action according to its local observation The actions of all agents together constitute the joint action A t . After that, the agent will receive the same global reward R t+1 , and the environmental state S t transitions to the state S t+1 at the next moment with the probability of P(S t |S t , A t+1 . At this time, each agent will obtain a new observation value . The environmental state S t includes all channel states and joint action A t , but a single agent cannot fully observe this information and can only understand the current environment through the observation function O. Its observation space includes: the channel gain g j [n] of the current U2U link, the co-channel interference g j′,j [n] from other U2U links, the co-channel interference g j,b [n] of itself to the ground station, and the co-channel interference g n,j [n] from the A2G link. Among them, g j,b [n] is obtained at the ground station and will be broadcast to all drones within the communication coverage; all other information can be obtained immediately at the receiving end of the U2U link agent. All received interferences I j [n] on the nth sub-band are measured by the U2U link receiver and added to its local observation space. In addition, the local observation space also includes the remaining U2U load L j , the remaining time delay T j , the number of training iterations e, and the exploration rate ε (the probability of random action selection). Therefore, the local observation space of each agent is represented as:
[0101]
[0102] where,
[0103] G j [n] = {g j [n], g j′,j [n], g j,B [n], g n,j [n]}
[0104] It should be noted that the resource allocation design of the drone is attributed to the spectrum sub-band selection and power control of the U2U link. Here, for the convenience of learning and control, the power control is set to M discrete quantities within the maximum power range. Therefore, each agent has M power selection and N channel selections, that is, the dimension of the action space is M*N, and each action corresponds to a specific combination of the spectrum sub-band and power.
[0105] In one of the embodiments, an optimization reward function is designed, that is, there are two optimization objectives, namely, maximizing the total capacity of the A2G link while maximizing the transmission rate of the U2U link. The first objective can be specifically the instantaneous value of the total capacity of the A2G link; the second objective can be specifically the load transmission capacity of the U2U link. By detecting the remaining load L of the U2U link at the end of each transmission j to determine whether the data packet is successfully delivered by judging whether it is greater than 0. If all data packets are successfully delivered, that is, L j ≤0, then the reward B j (t) is twice the value of the transmission rate; if only partial delivery is completed, that is, L j >0, then the reward is once the value of the transmission rate, which can be expressed as:
[0106]
[0107]
[0108] where μ is the reward coefficient, which is determined by whether the load is completely delivered. A higher reward can motivate the agent to be more motivated to complete the transmission of all data packets, thus achieving the second optimization objective. For all U2U links, the training objective is to find an optimal policy π * , such that under the condition of any initialized environment, the revenue can be maximized. The training return is defined as the cumulative discounted return with a discount rate γ, and its expression is:
[0109]
[0110] To express the relative importance of the two optimization objectives, the reward function at each time step is designed as the weighted sum of the above two objective functions, that is:
[0111]
[0112] where δ d and δ u are positive weights that balance the A2G and U2U objectives, and satisfy δ d +δ u = 0.1.
[0113] In one of the embodiments, the MAD3QN-VAV algorithm is designed and applied to the D3QN-VAV network structure. This network is as Figure 4 shown. Based on DQN and DDQN, a dueling network structure is introduced. By introducing the dueling network, the neural network no longer directly outputs the Q value, but outputs the state value function and the advantage function, which are specifically expressed as follows:
[0114]
[0115] Formula (16) shows that after adding the competitive network, the Q value is composed of the value function V(s) of the state and the advantage function of each action Add up to get θ as the shared network parameter, θ V and Represents a separate network parameter, where the state value function indicates the quality of the state, and the advantage function indicates the quality of an action relative to other actions in the state. The value Q of an action in the state is obtained by combining them. The role of introducing the advantage function is: unlike the DQN network that directly learns all Q values, the D3QN network can distinguish whether the current reward is caused by the state itself or by the selected action. However, there is also a problem with formula (16). Since V(s) is a scalar, if a Q value is given, it is impossible to determine the unique V(s) and Therefore, Equation (16) is improved into the following form:
[0116]
[0117] Where |A| is the action dimension. Equation (17) is based on Equation (16) by subtracting an average from the advantage function. This allows the average value of all advantage functions to be 0, thus obtaining a unique value function. Because the average value of the advantage function is defined as 0, when an action is sampled and updated, the Q values of other actions are also updated. Combined with the DDQN estimation method, the D3QN estimate is expressed as:
[0118]
[0119] Formula (18) represents the process of D3QN estimating Q value, θ and θ - are the neural network parameters of the main network and the target network respectively. After estimating the Q value, the neural network parameters are updated using the gradient descent method:
[0120]
[0121] Formula (19) represents the loss function. The gradient descent method is used to minimize the loss function L(θ) to update the neural network parameters. The target network parameter θ - The main network parameters are copied every fixed number of steps to complete the update of the target network.
[0122] It should be understood that although Figure 2 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 2At least some of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed and completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0123] In one embodiment, as Figure 5 shown, a spectrum resource allocation and power control system for multi-UAV networking is provided, including: a model construction module 502, an allocation strategy generation module 504, and a resource allocation module 506, where:
[0124] The model construction module 502 is configured to construct a communication resource allocation model according to the U2U communication links between UAVs and the A2G communication links between UAVs and ground stations in multi-UAV networking.
[0125] The allocation strategy generation module 504 is configured to use the dynamic observation values of the U2U communication links as agents. Each agent performs centralized learning through a local observation function generated by the Markov decision algorithm using the communication resource allocation model, obtains the channel state and data packet transmission actions of the U2U link corresponding to the agent according to the observation values of the agent in the current iteration round, and obtains a global reward value based on reinforcement learning and an optimized reward function. Each agent generates a corresponding channel and power joint allocation strategy according to the global reward value.
[0126] The resource allocation module 506 is configured to form a global resource allocation action set with the channel states and data packet transmission actions of all agents, and use the global resource allocation action set and the local observation values of the agents in the next iteration round as the state input to the communication resource allocation model for iterative learning of the channel and power joint allocation strategy until the optimal channel and power joint allocation strategy is output.
[0127] For the specific limitations on the spectrum resource allocation and power control system for multi-UAV networking, reference can be made to the limitations on the spectrum resource allocation and power control method for multi-UAV networking in the above text, which will not be elaborated here. Each module in the above spectrum resource allocation and power control system for multi-UAV networking can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0128] Those skilled in the art can understand that Figure 5The structure shown is only a block diagram of some of the structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0129] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as these combinations of technical features do not conflict, they should be considered as within the scope described in this specification.
[0131] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A method for spectrum resource allocation and power control in multi - UAV networking, characterized in that, The method includes: Constructing a communication resource allocation model based on the U2U communication links between the unmanned aerial vehicles (UAVs) in a multi-UAV network and the A2G communication links between the UAVs and the ground station; Taking the dynamic observation values of the U2U communication links as agents, each agent performs centralized learning through the local observation function generated by the Markov decision algorithm using the communication resource allocation model, obtains the channel state and packet transmission actions of the U2U link corresponding to the agent according to the observation values of the agent in the current iteration round, and obtains the global reward value based on reinforcement learning and the optimized reward function. Each agent generates the corresponding channel and power joint allocation strategy according to the global reward value; Forming the global resource allocation action set with the channel states and the packet transmission actions of all the agents, and using the global resource allocation action set and the local observation values of the agents in the next iteration round as the state input to the communication resource allocation model for iterative learning of the channel and power joint allocation strategy until the optimal channel and power joint allocation strategy is output.
2. The method according to claim 1, characterized in that, Constructing a communication resource allocation model based on the U2U communication links between the UAVs in a multi-UAV network and the A2G communication links between the UAVs and the ground station, including: Obtaining the first path loss according to the U2U communication links between the UAVs in a multi-UAV network; L j,j′ (t) = 20 log 10 d j,j′ (t) + 20 log 10 f0 + 20 log 10 (4π / c0) Among them, L j,j′ (t) is the first path loss between the j-th drone and the j'-th drone, and d j,j′ is the j distance in three-dimensional space between the j -th drone and the ′-th drone. c0 is the speed of light, and f0 is the channel carrier frequency of the U2U communication link; Using orthogonal frequency division multiplexing to occupy different subbands of the A2G communication links to obtain the second path loss between the UAVs and the ground station on the A2G communication links; L n,B (t) = P LoS ·L LoS +P NLoS ·L NLoS Among them, L n,B (t) is the second path loss between the UAV and ground station B on the nth A2U communication link for the transmission of the nth sub-band, P LoS and P NLoS are the probabilities of LoS and NLoS respectively, L LoS and L NLoS represent the path losses under LoS and NLoS conditions respectively; And obtaining the channel power gain according to the A2G communication links between the UAVs and the ground station; where, g n,B [n] is the channel power gain of the n-th A2G communication link transmitted in the n-th subband, H n [n] is the small-scale fading power component related to frequency in the n-th A2G communication link, and B is the reward function; Constructing a spectrum resource sharing and power control strategy according to the first path loss, the second path loss and the channel power gain; Among them, is the channel capacity of the nth A2G communication link transmitted on the nth RB, N is the total number of A2G communication links, maxPr is the maximum load transmission success rate of the U2U communication link, T is the transmission time interval, β j [n] indicates whether the jth U2U communication link is transmitted on the nth RB, is the channel capacity of the UAV on the jth U2U communication link, Δ T is the channel coherence time, L is the size of the U2U data packet periodically generated according to the first loss and the second loss, J is the total number of U2U communication links, β j′ [n] is the j indication of whether the j'th U2U communication link is transmitted on the nth RB, is the channel power of the jth U2U communication link, is the j channel power of the j'th U2U communication link, P max is the maximum channel power of the U2U communication link, d is the ground station; Constructing a communication resource allocation model according to the spectrum resource sharing and power control strategy.
3. The method according to claim 2, wherein Constructing a communication resource allocation model according to the spectrum resource sharing and power control strategy, including: Obtaining a target network according to the spectrum resource sharing and power control strategy and the competing network optimized deep Q network, and constructing a communication resource allocation model using the deep Q network as the main network and the target network; Among them, is the estimated Q value at the current moment in the communication resource allocation model, D3QN is the communication resource allocation model, L(θ) is the loss function of the communication resource allocation model, s t is the dynamic observation value of the U2U communication link at the current moment, a t is the data packet transmission action at the current moment, θ is the global network parameter of the communication resource allocation model, r t+1 is the reward at the next moment, γ is the discount factor, s t+1 is the channel state of the U2U communication link at the next moment.
4. The method according to claim 3, wherein The communication resource allocation model generates a local observation function using the Markov decision algorithm, including: Inputting the dynamic observation values of different agents into the communication resource allocation model for distributed Markov decision operation to generate the local observation function corresponding to the agent; Among them, is the local observation value of the j-th U2U communication link at the current moment, O(S t , j) is the observation function of the j-th U2U communication link at the current moment, S t is the dynamic observation value at the current moment, L j is the remaining load of the j-th U2U communication link, T j is the remaining delay of the j-th U2U communication link, e is the number of model iteration trainings, ε is the probability of random action selection, I j is the received interference of the j-th U2U communication link, G j is the observation space of the j-th U2U communication link, N is the total number of sub-bands, g j [n] is the channel gain of the j-th U2U communication link, g j′,j [n] is the co-channel interference of the j'-th U2U communication link on the j-th U2U communication link, g j,B [n] is the co-channel interference of the j-th U2U communication link on the ground station, g n,j [n] is the co-channel interference of the j-th U2U communication link by the n-th A2G communication link.
5. The method according to claim 4, wherein Taking the dynamic observation values of the U2U communication links as agents, each agent performs centralized learning through the local observation function generated by the Markov decision algorithm using the communication resource allocation model, and obtains the channel state and packet transmission actions of the U2U link corresponding to the agent according to the observation values of the agent in the current iteration round, including: Take the dynamic observations of the U2U communication links as agents. Each agent conducts centralized learning through the local observation function generated by the communication resource allocation model using the Markov decision algorithm. Input the observations of the agent in the current iteration round as training samples into the communication resource allocation model. According to the estimated Q value, use the gradient descent method to minimize the loss function and copy and train the network parameters of the deep Q network of the agent at a fixed iteration step to update the global network parameters of the communication resource allocation model; Obtain the channel state and packet transmission actions of the U2U link corresponding to the agent according to the global network parameters.
6. The method according to claim 5, characterized in that, Obtain the global reward value based on reinforcement learning and the optimized reward function. Each agent generates a corresponding channel and power joint allocation strategy according to the global reward value, including: Detect the packet delivery result of each U2U communication link at the end of each transmission based on the reward function of reinforcement learning. If the remaining load value of the current U2U communication link is greater than zero and the packet delivery result is partial delivery, the link reward value is one times the transmission rate value; otherwise, if the packet delivery result is all successfully delivered, the link reward value is two times the transmission rate value; Generate the training return using cumulative discounted return according to the preset discount rate: Among them, G t is the training return, γ is the cumulative discount rate, and R t+k+1 is the reward value, k and is the number of interactions; Optimize the reward function according to the link reward values of all U2U communication links and the training return to obtain the optimized reward function: Among them, R t+1 is the optimized reward function, and δ d is the positive weight of the ground station for balancing the A2G communication link and the U2U communication link. is the channel capacity of the nth A2G communication link transmitted on the nth RB, and δ u is the positive weight of the UAV for balancing the A2G communication link and the U2U communication link, and B j is the reward value of the j'th U2U communication link; Each agent generates a corresponding channel and power joint allocation strategy according to the global reward value.
7. The method according to claim 6, characterized in that, Form a global resource allocation action set with the channel states and the packet transmission actions of all the agents. Input the global resource allocation action set and the local observations of the agents in the next iteration round as the state of the next iteration round into the communication resource allocation model for iterative learning of the channel and power joint allocation strategy until the optimal channel and power joint allocation strategy is output, including: Form a global resource allocation action set with the channel states and the packet transmission actions of all the agents. After inputting the global resource allocation action set and the local observations of the agents in the next iteration round as the state of the next iteration round into the communication resource allocation model, iteratively train the channel and power joint allocation strategy according to the global reward function and the loss function until the optimal channel and power joint allocation strategy is output.
8. A spectrum resource allocation and power control system for multi - UAV networking, characterized in that, The system includes: A model construction module, which is used to construct a communication resource allocation model according to the U2U communication links between drones in a multi-UAV network and the A2G communication links between the drones and the ground station; The allocation strategy generation module is used to take the dynamic observations of the U2U communication link as agents. Each agent conducts centralized learning through a local observation function generated by a communication resource allocation model using the Markov decision algorithm, obtains the channel state and packet transmission actions of the U2U link corresponding to the agent according to the observations of the agent in the current iteration round, and obtains the global reward value based on reinforcement learning and the optimized reward function. Each agent generates corresponding channel and power joint allocation strategies respectively according to the global reward value; The resource allocation module is used to form a global resource allocation action set with the channel states and the packet transmission actions of all the agents, and input the global resource allocation action set and the local observations of the agents in the next iteration round as the state of the next iteration round into the communication resource allocation model for iterative learning of the channel and power joint allocation strategy until the optimal channel and power joint allocation strategy is output.
Citation Information
Cited By
Power grid control method, system and equipment based on space-time mapping function
CN120999602A
Power grid control method, system and device based on space-time mapping function
CN120999602B