Unmanned aerial vehicle resource allocation method, system and device based on multi-agent reinforcement learning, and medium
By employing a multi-agent reinforcement learning approach, combined with a conditional variational autoencoder and the MATD3 algorithm, the coupling relationship between discrete and continuous optimization variables in UAV-assisted wireless communication was resolved. This enabled resource allocation for UAVs in dynamic and stochastic scenarios, improving system performance and user data service quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIDIAN UNIV
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies in UAV-assisted wireless communication have failed to effectively handle the suddenness and randomness of ground user data demands, and have failed to effectively characterize the coupling relationship between discrete and continuous optimization variables, resulting in high difficulty in resource allocation optimization and poor model performance.
We design a multi-agent reinforcement learning method that combines discrete and continuous actor networks, uses a conditional variational autoencoder to represent the coupling relationship between discrete and continuous actions, and employs the MATD3 algorithm for training to achieve resource allocation for UAVs in dynamic random data service scenarios.
In dynamic random data service scenarios, the system enables real-time action mode selection and communication resource allocation for UAV base stations, improving system performance and scalability, and better meeting the data needs of ground users.
Smart Images

Figure CN122028075A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) wireless communication technology, and specifically relates to a UAV resource allocation method, system, device and medium based on multi-agent reinforcement learning. Background Technology
[0002] In recent years, due to the numerous advantages of drones, such as multifunctionality, high mobility, flexible deployment, low cost, and high probability of establishing line-of-sight links, research on drones as aerial base stations has attracted widespread attention worldwide. On the one hand, small drones play an important role in achieving airspace coverage in 6G integrated air-ground systems, serving as airspace-assisted communication platforms. For example, in high-density communication scenarios, drones can act as temporary base stations or relay stations to support wireless communication and enhance user capacity. On the other hand, in the field of emergency communication, although basic communication infrastructure can usually handle daily communication loads, drones can be used to support communication in emergency situations or unconventional temporary scenarios. For example, in network reconstruction after major natural disasters, drones can quickly adjust their positions to provide rapid restoration of post-disaster wireless services.
[0003] These characteristics of drones offer opportunities to establish efficient, flexible, and reliable drone-assisted communication networks in situations where fixed ground infrastructure is insufficient or faulty. However, resource allocation in current data service scenarios still faces many challenges. On the one hand, the suddenness and randomness of data demands from ground users pose significant challenges to the fairness of data services and improving the quality of service for ground users. On the other hand, the assumptions about the prior information of ground users are too idealistic. Although some studies have considered the dynamic changes of ground users, they only yielded drone downlink transmission strategies under the condition that ground user information (location distribution, etc.) is known in advance, without considering how to perceive the state of ground users. At the same time, the coexistence of discrete and continuous optimization variables in the resource allocation process often further increases the optimization difficulty. How to characterize and decouple the coupling relationship between discrete and continuous optimization variables is a major challenge.
[0004] Flight trajectory planning, mode selection, and transmit power allocation for drones are fundamental to ensuring energy-efficient drone-assisted wireless networks. Drone trajectories should be precisely designed to meet the dynamic data service needs of ground users in real time. Simultaneously, drone transmit power should be well-allocated to adapt to dynamic changes in channel state information between drones and ground users, while minimizing co-channel interference between drones. Furthermore, drones need the ability to select appropriate operational modes to perceive ground user information, thereby enhancing their ability to meet ground user data needs. In conclusion, developing more practical methods to effectively advance drone-assisted wireless networks is of urgent importance.
[0005] Patent application number 202411903606.6 discloses a UAV resource allocation method based on multi-agent reinforcement learning. It achieves hybrid action decision-making by designing discrete and continuous Actor networks separately, and introduces the MATD3 deep reinforcement algorithm for network training to realize mode selection, path planning, and power allocation for multiple UAVs. However, this method separates discrete and continuous actions for decision-making, without considering the coupling relationship between them, often resulting in poor model performance after training convergence. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, the present invention aims to provide a method, system, device and medium for allocating UAV resources based on multi-agent reinforcement learning. By designing discrete and continuous actor networks and combining them with a multi-agent centralized training and distributed execution training framework, the present invention enables UAV base stations to select action modes and allocate communication resources in dynamic and random data service scenarios. The present invention has good scalability and high system performance.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for allocating drone resources based on multi-agent reinforcement learning includes the following steps: Step 1: Establish an optimization problem for the allocation of resources among multiple drones in a data service scenario assisted by multimodal drones and design an evaluation index for drone data services based on the average peak information age. Step 2: The optimization problem of multi-UAV resource allocation in the data service scenario established in Step 1 is expressed as a Markov decision problem, and the state space, action space and corresponding instant reward function are designed respectively. Step 3: By designing a conditional variational autoencoder, the coupling relationship between discrete and continuous actions is characterized. Combined with the MATD3 multi-agent reinforcement learning algorithm, the MATD3 algorithm representing the coupling relationship is obtained. Based on the Markov decision model obtained in Step 2, the network parameters of the UAV are trained and updated using the MATD3 algorithm representing the coupling relationship. The trained UAV is then applied to the data service scenario for communication resource allocation.
[0008] The specific method of step 1 includes: Step 1.1: Make specific assumptions about the data service scenario; within a designated area, deploy multiple drones as airborne base stations to provide data service support to ground users; during drone operation, choose to use radar sensing mode to locate ground users, or communicate with multiple ground users based on an orthogonal frequency division multiplexing (OFDM) scheme; simultaneously, the locations of some ground users are constantly changing, and they will have different data needs at different times; in addition, drones provide data services by sending multiple data packets to ground users; as airborne base stations, drones, while deciding on their movement trajectories, select their own behavior mode, radar sensing mode, or data transmission mode, and simultaneously implement power control in different modes; Step 1.2: Establish a communication link channel model between the UAV and the ground user; the channel capacity of the communication link is:
[0009] in, For the spectrum subband bandwidth, Let LoS be the probability channel gain of the communication link between the drone m and the ground user n. This indicates the drone's transmission power on this communication link. This represents the power spectral density of additive white Gaussian noise. This refers to co-channel interference on the communication link between the UAV m and the ground user n, specifically:
[0010] Where M and N represent the number of drones and ground users, respectively. This represents the transmit power of the communication link between UAV i and ground user j. Use a binary value to represent whether drone i is serving ground user j, where 1 indicates service and 0 indicates no service; The binary value indicates whether ground user j and ground user n are using frequency band reuse, with 1 indicating reuse and 0 indicating no reuse. Step 1.3: Design an evaluation index for UAV data services based on the average peak information age; in time... The age of the update data packet with timestamp u is determined as When the updated timestamp corresponds to the current time t and the age is zero, the updated data packet is considered fresh; the average peak information age of ground user i in time slot j is expressed as:
[0011] in, This represents the cumulative peak count of ground user i in time slot j. Let represent the age of the 'a'th peak information; finally, the average peak information age of each ground user within each time slot is expressed as:
[0012] Where K represents the total number of time slots; Step 1.4, combining the scenario assumptions of Step 1.1, the communication link channel model between the UAV and the ground user established in Step 1.2, and the UAV data service evaluation index of average peak information age designed in Step 1.3, yields the optimization problem of multi-UAV resource allocation in the data service scenario, specifically described as follows:
[0013] Among them, (a), (b), and (c) ensure the maximum flight speed constraint of the UAV and its location constraint; (d) represents the formulaic description of the UAV behavior mode selection; (e) ensures that each ground user is served by only one UAV; (f) ensures that the number of ground users served by each UAV does not exceed the number of subcarriers; and (g) ensures that the power consumption of each UAV does not exceed the upper limit of the transmit power.
[0014] The specific method for step 2 includes: Step 2.1, design the state space, including the state changes of ground users and drones, namely the relative distance between drones and ground users and the data demand. The specific relative distance between drones and other drones is as follows:
[0015] in, and These represent the relative distances between the drone m and the ground user n on the x and y axes, respectively. This represents the data demand of ground user n. and Let m represent the relative distances between drone m and other drone i on the x-axis and y-axis, respectively. Step 2.2: Design the action space. The UAV's actions consist of discrete actions and continuous actions. Discrete actions are the selection of action modes, including radar perception mode or data transmission mode. Continuous actions are the continuous parameters across all action modes. When the UAV selects radar perception mode, the radar power is fixed, and continuous actions only vary along the x and y axes of the UAV's position coordinates. Furthermore, when the data transmission mode is selected, continuous actions also include the transmit power allocated on each subcarrier when the UAV selects data transmission. ; Step 2.3, design the reward function as follows:
[0016] in, The average peak information for time period t is represented by the age penalty factor.
[0017] The specific method for step 3 includes: Step 3.1: Design a conditional variational autoencoder to characterize the coupling relationship between discrete and continuous actions; Construct a decodable latent action space, represent dependencies in the hybrid action space, and build an embedding table. to indicate A discrete action, where each row It is Discrete actions of dimension Continuous vector, where, For row indexing, a latent representation space for continuous actions is then constructed using a conditional variational autoencoder (CDAE), implicitly modeling and embedding dependencies. The CDAE consists of an encoder and a decoder. For discrete actions... Continuous actions and state encoder by As parameters, in terms of state and embedding vector As a condition, continuous actions Encoding as latent actions For encoders Using Gaussian distribution ,in, and These are the mean and standard deviation of the encoder output, respectively, and the decoder's response to any potential action. Perform deterministic decoding, in addition, by Parameterized decoder Under the same conditions, from potential actions Decoding continuous actions Arbitrary embedding and any potential action Both are performed as a hybrid action through nearest neighbor lookup in the embedded table and decoding by the decoder. and Therefore, the encoding and decoding process can be represented as:
[0018] Subsequently, utilizing environmental dynamics, a squared error loss based on state change prediction is designed to further refine the representation of hybrid actions; therefore, the data service volume of each UAV after executing the hybrid actions is used as the state change prediction, i.e. ,in, Let m represent the amount of data service provided by drone m to ground user n, and let m be the data service provided by drone m to ground user n. These represent the parameters of the transformation network, decoder network, and state change prediction network, respectively; the continuous actions of VAE reconstruction. and prediction Using experience pool Batch status and mixed actions and By minimizing the following loss function Training embedding table Conditional VAE:
[0019] in, It is a continuous action of VAE reconstruction. The squared error, It is the encoder distribution The Kullback-Leibler (KL) divergence between the standard Gaussian distribution and the standard Gaussian distribution; furthermore, minimizing the state change prediction. Prediction Squared error:
[0020] Therefore, the total training loss of the conditional variational autoencoder is expressed as: ; In addition, two mechanisms are employed: latent action constraints and training experience correction, to handle unreliable latent actions and outdated off-policy training experience, respectively. Constrain potential actions within a reasonable range; that is, for each potential action in the experience pool... Through calculation Obtain the boundary of the center range Then, each dimension of the potential action is rescaled to a bounded range. ; A training experience correction mechanism is adopted to check the timeliness of potential actions in the experience pool and use the latest potential strategies to correct outdated potential actions. Step 3.2: Combine the MATD3 multi-agent reinforcement learning algorithm to obtain the MATD3 algorithm with coupling relationship representation, and use the MATD3 algorithm with coupling relationship representation to train and update the network parameters of the UAV based on the Markov decision model obtained in Step 2. The latent policy is trained using the MATD3 algorithm; the Actor network of the UAV uses a learned conditional variational autoencoder that represents the coupling relationship between discrete and continuous actions, i.e., the parameters are... Potential strategies To output the latent action vector ,in Simultaneously, a conditional variational autoencoder is trained for each agent, and each agent uses a conditional variational autoencoder to... Decode into mixed actions respectively and ; A dual-critic network for each agent , Using global information as input, an approximate state-action value function is used. From the experience pool Randomly select a batch of experience , As training samples, compute the Critic network. The mean squared error loss function is:
[0021] in, , The estimated target Q value is expressed as:
[0022] in, , and For the parameters of the target network, the Actor network for each agent is trained to maximize the state-action value of the potential action output by the network, which is updated through policy gradients; therefore, the loss function of the Actor network is:
[0023] in, ; Step 3.3: Apply the trained drones to the data service scenario for communication resource allocation; For the hybrid action space design, each UAV is equipped with an Actor network and a Conditional Variational Autoencoder (CVA). The Actor network outputs a latent policy, which is then input into the trained CVA for decoding, yielding discrete and continuous actions. Discrete actions determine the action mode, either radar sensing or data transmission, while continuous actions output continuous action parameters for each action mode. After observing the current state of the environment, the UAV selects either the radar sensing or data transmission mode based on the discrete actions, and then obtains the corresponding continuous action parameters from the continuous actions. When the UAV selects the radar sensing mode, it activates radar to sense the location of ground users and adjusts its own position. When the UAV selects the data transmission mode, it transmits data to ground users on each channel using the transmission power obtained from the continuous actions and adjusts its own position.
[0024] This invention also provides a UAV resource allocation system based on multi-agent reinforcement learning, comprising: The optimization problem establishment module is used to establish an optimization problem for the allocation of resources among multiple drones in a data service scenario assisted by multimodal drones and to design an evaluation index for drone data services based on the average peak information age. The Markov decision problem formulation module is used to formulate the optimization problem of multi-UAV resource allocation in data service scenarios as a Markov decision problem, and designs the state space, action space and corresponding instant reward function respectively. The UAV communication resource allocation module is used to characterize the coupling relationship between discrete and continuous actions by designing a conditional variational autoencoder. Combined with the MATD3 multi-agent reinforcement learning algorithm, a MATD3 algorithm representing the coupling relationship is obtained. Based on the Markov decision model, the network parameters of the UAV are trained and updated using the MATD3 algorithm representing the coupling relationship. The trained UAV is then applied to data service scenarios for communication resource allocation.
[0025] This invention also provides a drone resource allocation device based on multi-agent reinforcement learning, comprising: Memory: A computer-readable device that stores the computer program of the above-mentioned UAV resource allocation method based on multi-agent reinforcement learning; Processor: Used to implement the UAV resource allocation method based on multi-agent reinforcement learning when executing the computer program.
[0026] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the aforementioned UAV resource allocation method based on multi-agent reinforcement learning.
[0027] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Step 1 of this invention establishes a data service scenario that integrates sensing and communication with the assistance of multimodal drones. Compared with the scenario of drone trajectory optimization and resource allocation under prior known conditions, the scenario of this invention considers how drones can timely perceive user status under prior unknown conditions. The scenario built is more adaptable to the situation where the dynamic changes of user status are unknown in actual data service scenarios.
[0028] 2. Step 1 of this invention designs an average peak information age evaluation index. Compared with traditional evaluation indexes such as data transmission rate, the evaluation index of this invention is geared towards individual users rather than the overall system, and can better reflect the quality of data services for users.
[0029] 3. Step 3 of this invention designs a conditional variational autoencoder to characterize the coupling relationship between discrete and continuous actions, and combines it with the MATD3 multi-agent reinforcement learning algorithm with centralized training and distributed execution. This enables the proposed algorithm to be applicable to UAVs carrying various payloads, has good scalability, and is more in line with the needs of actual scenarios.
[0030] In summary, this invention establishes a data service scenario assisted by multimodal UAVs, sets up user-oriented evaluation indicators, and designs the MATD3 algorithm to represent coupling relationships. This enables UAVs to make real-time decisions on action modes and communication resource allocation strategies in scenarios where ground users experience random and dynamic changes and have unknown prior knowledge. In particular, it can better solve the optimization process of resource allocation where discrete and continuous optimization variables coexist, thereby maximizing the data service quality for ground users. It also exhibits high scalability and excellent system performance. Attached Figure Description
[0031] Figure 1 This is a scenario diagram of the integrated sensory data service assisted by the multimodal drone of the present invention.
[0032] Figure 2 This is a graph showing the change in the age of ground users' information in the context of a multimodal drone-assisted integrated data service scenario, as described in this invention.
[0033] Figure 3 This invention presents a conditional variational autoencoder for characterizing the coupling relationship between discrete and continuous variables.
[0034] Figure 4 This is a graph showing the convergence of rewards for all drones under different methods in this embodiment of the invention.
[0035] Figure 5 This is a graph showing the variation of the average peak age of ground users under different methods in the embodiments of the present invention. Detailed Implementation
[0036] To more clearly understand the technical features, objectives, and effects of this invention, its technical solution is described in detail below. Obviously, the described embodiments are only some examples of this invention, not all, and should not be considered as limitations on the scope of implementation of this invention; based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.
[0037] A method for allocating drone resources based on multi-agent reinforcement learning includes the following steps: Step 1: Establish an optimization problem for the allocation of resources among multiple drones in a data service scenario assisted by multimodal drones and design an evaluation index for drone data services based on the average peak information age. The specific method of step 1 includes: like Figure 1 As shown, step 1.1 makes specific assumptions about the data service scenario; within a designated area, multiple drones are deployed as airborne base stations to provide data service support to ground users; during drone operation, the drones may choose to use radar sensing mode to locate ground users, or communicate with multiple ground users based on an orthogonal frequency division multiplexing (OFDM) scheme; simultaneously, the locations of some ground users are constantly changing, and they will have different data needs at different times; in addition, the drones provide data services by sending multiple data packets to ground users; as airborne base stations, the drones, while deciding on their movement trajectories, select their own behavior mode, radar sensing mode, or data transmission mode, and simultaneously implement power control in different modes; Step 1.2: Establish a communication link channel model between the UAV and the ground user; the channel capacity of the communication link is:
[0038] in, For the spectrum subband bandwidth, Let LoS be the probability channel gain of the communication link between the drone m and the ground user n. This indicates the drone's transmission power on this communication link. This represents the power spectral density of additive white Gaussian noise. This refers to co-channel interference on the communication link between the UAV m and the ground user n, specifically:
[0039] Where M and N represent the number of drones and ground users, respectively. This represents the transmit power of the communication link between UAV i and ground user j. Use a binary value to represent whether drone i is serving ground user j, where 1 indicates service and 0 indicates no service; The binary value indicates whether ground user j and ground user n are using frequency band reuse, with 1 indicating reuse and 0 indicating no reuse. like Figure 2 As shown in step 1.3, design an evaluation index for UAV data services based on the average peak information age; in time... The age of the update data packet with timestamp u is determined as When the updated timestamp corresponds to the current time t and the age is zero, the updated data packet is considered fresh; the average peak information age of ground user i in time slot j is expressed as:
[0040] in, This represents the cumulative peak count of ground user i in time slot j. Let represent the age of the 'a'th peak information; finally, the average peak information age of each ground user within each time slot is expressed as:
[0041] Where K represents the total number of time slots; Step 1.4, combining the scenario assumptions of Step 1.1, the communication link channel model between the UAV and the ground user established in Step 1.2, and the UAV data service evaluation index of average peak information age designed in Step 1.3, yields the optimization problem of multi-UAV resource allocation in the data service scenario, specifically described as follows:
[0042] Among them, (a), (b), and (c) ensure the maximum flight speed constraint of the UAV and its location constraint; (d) represents the formulaic description of the UAV behavior mode selection; (e) ensures that each ground user is served by only one UAV; (f) ensures that the number of ground users served by each UAV does not exceed the number of subcarriers; and (g) ensures that the power consumption of each UAV does not exceed the upper limit of the transmit power.
[0043] Step 2: The optimization problem of multi-UAV resource allocation in the data service scenario established in Step 1 is expressed as a Markov decision problem, and the state space, action space and corresponding instant reward function are designed respectively. The specific method for step 2 includes: Step 2.1, design the state space, including the state changes of ground users and drones, namely the relative distance between drones and ground users and the data demand. The specific relative distance between drones and other drones is as follows:
[0044] in, and These represent the relative distances between the drone m and the ground user n on the x and y axes, respectively. This represents the data demand of ground user n. and Let m represent the relative distances between drone m and other drone i on the x-axis and y-axis, respectively. Step 2.2: Design the action space. The UAV's actions consist of discrete actions and continuous actions. Discrete actions are the selection of action modes, including radar perception mode or data transmission mode. Continuous actions are the continuous parameters across all action modes. When the UAV selects radar perception mode, since the radar power is fixed, the continuous actions only vary along the x and y axes of the UAV's position coordinates. Furthermore, when selecting data transmission mode, the continuous actions also include the transmit power allocated on each subcarrier when the UAV selects data transmission. ; Step 2.3, design the reward function as follows:
[0045] in, The average peak information for time period t is represented by the age penalty factor.
[0046] Step 3: By designing a conditional variational autoencoder, the coupling relationship between discrete and continuous actions is characterized. Combined with the MATD3 multi-agent reinforcement learning algorithm, the MATD3 algorithm representing the coupling relationship is obtained. Based on the Markov decision model obtained in Step 2, the network parameters of the UAV are trained and updated using the MATD3 algorithm representing the coupling relationship. The trained UAV is then applied to the data service scenario for communication resource allocation.
[0047] The specific method for step 3 includes: like Figure 3 As shown in step 3.1, design a conditional variational autoencoder to characterize the coupling relationship between discrete and continuous actions; Construct a decodable latent action space to represent dependencies in the hybrid action space, and build an embedding table. to indicate A discrete action, where each row It is Discrete actions of dimension Continuous vector, where, For row indexing, a latent representation space for continuous actions is then constructed using a conditional variational autoencoder (CDAE), implicitly modeling and embedding dependencies. The CDAE consists of an encoder and a decoder. For discrete actions... Continuous actions and state encoder by As parameters, in terms of state and embedding vector As a condition, continuous actions Encoding as latent actions For encoders Using Gaussian distribution ,in, and These are the mean and standard deviation of the encoder output, respectively, and the decoder's response to any potential action. Perform deterministic decoding, in addition, by Parameterized decoder Under the same conditions, from potential actions Decoding continuous actions Arbitrary embedding and any potential action All are conveniently decoded into hybrid actions through nearest neighbor lookup in the embedded table and decoder. and Therefore, the encoding and decoding process can be represented as:
[0048] Subsequently, utilizing environmental dynamics, a squared error loss based on state change prediction is designed to further refine the hybrid action representation; therefore, the data service volume of each UAV after executing the hybrid action is used as the state change prediction, i.e. ,in, Let m represent the amount of data service provided by drone m to ground user n, and let m be the data service provided by drone m to ground user n. These represent the parameters of the transformation network, decoder network, and state change prediction network, respectively; the continuous actions of VAE reconstruction. and prediction Using experience pool Batch status and mixed actions and By minimizing the following loss function Training embedding table Conditional VAE:
[0049] in, It is a continuous action of VAE reconstruction. The squared error, It is the encoder distribution The Kullback-Leibler (KL) divergence between the standard Gaussian distribution and the standard Gaussian distribution; furthermore, minimizing the state change prediction. Prediction Squared error:
[0050] Therefore, the total training loss of the conditional variational autoencoder is expressed as: ; Furthermore, two mechanisms are employed: latent action constraints and training experience correction, to handle unreliable latent actions and outdated off-policy training experience, respectively. In the early stages of training, the latent policy may output some anomalous latent actions, which can be highly unreliable when decoding and estimating Q-values; this can quickly damage the policy and lead to poor performance. Therefore, latent actions are adaptively constrained within a reasonable range; specifically, for each latent action in the experience pool... Through calculation Obtain the boundary of the center range Then, each dimension of the potential action is rescaled to a bounded range. ; Furthermore, the latent action space is continuously optimized during the reinforcement learning process. After a period of training, the representation distribution of mixed actions in the latent action space will change. To address this issue, this invention proposes a training experience correction mechanism that checks the timeliness of latent actions in the experience pool and uses the latest latent policy to correct outdated latent actions. In this way, the multi-agent reinforcement learning policy learning is always based on the latest representation distribution.
[0051] Step 3.2: Combine the MATD3 multi-agent reinforcement learning algorithm to obtain the MATD3 algorithm with coupling relationship representation, and use the MATD3 algorithm with coupling relationship representation to train and update the network parameters of the UAV based on the Markov decision model obtained in Step 2. The MATD3 algorithm is used to train the latent policy; each UAV determines its action mode selection, self-motion, and power allocation; therefore, the UAV's Actor network uses a learned conditional variational autoencoder representing the coupling relationship between discrete and continuous actions, i.e., parameters are... Potential strategies To output the latent action vector ,in Simultaneously, a conditional variational autoencoder is trained for each agent. Specifically, each agent uses a conditional variational autoencoder to... Decode into mixed actions respectively and ; In a centralized training, distributed execution framework, each agent makes decisions in a decentralized manner based on its local observations, while training is centralized, leveraging the collective experience of all agents. Therefore, the critic network for each agent requires global information—the actions and observations of all agents—while the actor network operates based on each agent's own observations. Thus, a bi-critic network for each agent is required. , Using global information as input, an approximate state-action value function is used. From the experience pool Randomly select a batch of experience , As training samples, compute the Critic network. The mean squared error loss function is:
[0052] in, , The estimated target Q value is expressed as:
[0053] in, , and For the parameters of the target network, the Actor network for each agent is trained to maximize the state-action value of the potential action output by the network, which is updated through policy gradients; therefore, the loss function of the Actor network is:
[0054] in, ; Step 3.3: Apply the trained drones to the data service scenario for communication resource allocation; For the hybrid action space design, each UAV is equipped with an Actor network and a Conditional Variational Autoencoder (CVA). The Actor network outputs a latent policy, which is then input into the trained CVA for decoding, yielding discrete and continuous actions. Discrete actions determine the action mode, either radar sensing or data transmission, while continuous actions output continuous action parameters for each action mode. After observing the current state of the environment, the UAV selects either the radar sensing or data transmission mode based on the discrete actions, and then obtains the corresponding continuous action parameters from the continuous actions. When the UAV selects the radar sensing mode, it activates radar to sense the location of ground users and adjusts its own position. When the UAV selects the data transmission mode, it transmits data to ground users on each channel using the transmission power obtained from the continuous actions and adjusts its own position.
[0055] This invention also provides a UAV resource allocation system based on multi-agent reinforcement learning, comprising: The optimization problem establishment module is used to establish the optimization problem of multi-drone resource allocation in the data service scenario for the integrated sensory data service scenario assisted by multimodal drones in step 1, and at the same time design the drone data service evaluation index of average peak information age. The Markov decision problem formulation module is used to formulate the optimization problem of multi-UAV resource allocation in the data service scenario established in step 1 as a Markov decision problem in step 2, and to design the state space, action space and corresponding instant reward function respectively. The UAV communication resource allocation module is used to implement step 3 by designing a conditional variational autoencoder to represent the coupling relationship between discrete and continuous actions, combining it with the MATD3 multi-agent reinforcement learning algorithm to obtain the MATD3 algorithm representing the coupling relationship, and using the MATD3 algorithm representing the coupling relationship to train and update the network parameters of the UAV based on the Markov decision model obtained in step 2, and applying the trained UAV to the data service scenario for communication resource allocation.
[0056] This invention also provides a drone resource allocation device based on multi-agent reinforcement learning, comprising: Memory: A computer-readable device that stores the computer program of the above-mentioned UAV resource allocation method based on multi-agent reinforcement learning; Processor: Used to implement the UAV resource allocation method based on multi-agent reinforcement learning when executing the computer program.
[0057] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the aforementioned UAV resource allocation method based on multi-agent reinforcement learning.
[0058] The discrete and continuous Actor networks proposed in this invention can select different action modes (radar sensing mode or data transmission mode) to choose appropriate time slots to perceive the location of ground users, enabling real-time data services even when their location is unknown prior, thus better reflecting real-world scenarios. This invention also proposes a hybrid action space multi-agent TD3 algorithm, which can be used for data service resource allocation in UAV-assisted sensor-integrated scenarios.
[0059] Example: The problem scenario studied in this embodiment is as follows: Figure 1 As shown, the framework diagram of the conditional variational autoencoder used for representing the coupling relationship between discrete and continuous variables is as follows: Figure 3 As shown, the specific implementation method is introduced in three parts: the first part is the description of the simulation scenario and the setting of parameters; the second part is the specific application process of the MATD3 algorithm for representing coupling relationship; and the third part is the simulation results and effect verification.
[0060] Part 1: Simulation of the scenario and parameter setting; The simulation experiment was conducted within a 500m x 500m square area, where three drones flew at a fixed altitude of 100m, providing data services to 25 randomly distributed ground users. The MAHTD3 algorithm's drone actor and critic networks consisted of three convolutional layers, two hidden layers, and a residual network. A total of 3000 training iterations were performed, each containing 40 steps, with a performance test conducted every 50 training iterations to evaluate the agent's effectiveness.
[0061] In this simulation, the initial positions of ground users are uniformly distributed, and to simulate the dynamic changes in user positions in a real-world scenario, ground users move back and forth along a certain trajectory. Furthermore, to simulate the randomness of user data demands in a real-world scenario, the simulation experiment sets the data demand of each ground user in different time slots to follow a normal distribution. (MB). Furthermore, according to the TCP / IP protocol, the drone, acting as a base station, will provide data services by sending multiple data packets to ground users. The drone's operating time is divided into K time slots, each time slot having a length of... The other important parameter settings are shown in Table 1 below.
[0062] Table 1: Simulation Experiment Parameter Settings
[0063] Part Two: The practical application of the MATD3 algorithm for representing coupling relationships in the scenario of integrated sensory data services assisted by multimodal UAVs, which is divided into the following 7 steps; Step 1: Initialize the Actor and Critic networks of the UAV, as well as their corresponding target network parameters, network update frequency, delayed update frequency, experience pool, and simulation environment parameters; Step 2: Determine if the maximum number of iterations has been reached. If yes, end the algorithm and output the optimal trajectory optimization and communication resource allocation strategy, the maximum total reward value, the maximum cumulative data service volume, the minimum UAV energy consumption, and the average peak information age. Otherwise, proceed to Step 3. Step 3: Determine if the maximum number of steps has been reached. If not, proceed to Step 4; otherwise, proceed to Step 2. Step 4: The agent obtains the current environmental state. The state is input into the Actor network of the agent to obtain the corresponding potential actions. Then The mixed motion is obtained by decoding the input conditional variational autoencoder. and Meanwhile, the environmental state changes according to the action, and the agent obtains the reward for that action. and the next environmental state And whether the next state is the final state. ; Step 5: Apply experience Store the data in the experience pool. Determine if the size of the experience pool is greater than the number of training batch samples. If yes, proceed to step 6; otherwise, proceed to step 4. Step 6: Randomly select a batch of samples from the experience pool, calculate the loss function values of the Actor network and Critic network of the UAV, and update the Critic network and Actor network of the UAV according to the delayed update strategy. Step 7: Determine whether the target network's update frequency has been reached. If so, copy the parameters of the main network to the target network according to the soft update strategy; otherwise, proceed to step 3.
[0064] Step 8: Randomly select a batch of samples from the experience pool, calculate the loss function value of the conditional variational autoencoder of the UAV, and update the conditional variational autoencoder.
[0065] Part Three: Simulation Experiment Results Display and Analysis; First, here are some comparison algorithms: MADDPG algorithm for coupling relationship representation: Compared with MATD3 algorithm for coupling relationship representation, MADDPG algorithm is used to train potential policies.
[0066] MAHPPO: Multi-Agent HPPO uses two policy heads, one for discrete actions and the other for continuous actions.
[0067] P-DQN: Parametric DQN, a policy-free reinforcement learning algorithm for handling mixed action spaces.
[0068] like Figure 4As shown, the total reward of the MATD3 algorithm, which represents coupling relationships, gradually increases with the number of training epochs and converges after approximately 500 epochs. Compared with MAHPPO, MADDPG, and P-DQN algorithms, the final convergence reward of the algorithm in this invention is 0.036, 0.057, and 0.112 higher, respectively, demonstrating the effectiveness of the MATD3 algorithm. Compared with the MAHPPO algorithm, the algorithm in this invention requires a longer training convergence time because it is a joint training of the conditional variational autoencoder and the MATD3 algorithm. However, the conditional variational autoencoder achieves superior performance by representing the coupling relationship between discrete and continuous actions. Compared with the MADDPG algorithm, the dual-Critic network architecture and delayed policy update mechanism of MATD3 help to train the latent policy more stably and efficiently, thereby improving performance. Furthermore, because the P-DQN algorithm lacks centralized training, each agent learns independently, lacking collaborative decision-making between discrete and continuous actions, resulting in unstable training and unsatisfactory overall performance.
[0069] like Figure 5 As shown, the algorithm proposed in this invention outperforms all three benchmark algorithms in terms of average peak information age performance. The proposed algorithm achieves the minimum average peak information (28.3 seconds), surpassing the coupling relationship characterization algorithms MADDPG (41.5 seconds), MAHPPO (34.9 seconds), and P-DQN (47.6 seconds), with improvements of 31.8%, 18.9%, and 40.5%, respectively. It can be seen that the proposed algorithm better meets the data requirements of ground users, thanks to the effective coordination between discrete and continuous action decisions. These results further demonstrate the superiority of the proposed algorithm in handling mixed action spaces.
[0070] Therefore, this invention employs the aforementioned MATD3 algorithm with a coupling relationship representation and applies it to a UAV-assisted integrated sensing data service scenario to solve the problems of UAV trajectory optimization and communication resource allocation. Simulation results show that the proposed method has better stability and convergence than comparable algorithms such as the MADDPG, MAHPPO, and P-DQN algorithms with coupling relationship representations, ultimately obtaining more rewards and achieving superior data service performance. This demonstrates the effectiveness of the proposed MATD3 algorithm with a coupling relationship representation in complex and dynamic UAV-assisted data service scenarios.
Claims
1. A method for allocating unmanned aerial vehicle (UAV) resources based on multi-agent reinforcement learning, characterized in that, Includes the following steps: Step 1: Establish an optimization problem for the allocation of resources among multiple drones in a data service scenario assisted by multimodal drones and design an evaluation index for drone data services based on the average peak information age. Step 2: The optimization problem of multi-UAV resource allocation in the data service scenario established in Step 1 is expressed as a Markov decision problem, and the state space, action space and corresponding instant reward function are designed respectively. Step 3: By designing a conditional variational autoencoder, the coupling relationship between discrete and continuous actions is characterized. Combined with the MATD3 multi-agent reinforcement learning algorithm, the MATD3 algorithm representing the coupling relationship is obtained. Based on the Markov decision model obtained in Step 2, the network parameters of the UAV are trained and updated using the MATD3 algorithm representing the coupling relationship. The trained UAV is then applied to the data service scenario for communication resource allocation.
2. The UAV resource allocation method based on multi-agent reinforcement learning according to claim 1, characterized in that, The specific method of step 1 includes: Step 1.1: Make specific assumptions about the data service scenario; within a designated area, deploy multiple drones as airborne base stations to provide data service support to ground users; during drone operation, choose to use radar sensing mode to locate ground users, or communicate with multiple ground users based on an orthogonal frequency division multiplexing (OFDM) scheme; simultaneously, the locations of some ground users are constantly changing, and they will have different data needs at different times; in addition, drones provide data services by sending multiple data packets to ground users; as airborne base stations, drones, while deciding on their movement trajectories, select their own behavior mode, radar sensing mode, or data transmission mode, and simultaneously implement power control in different modes; Step 1.2: Establish a communication link channel model between the UAV and the ground user; the channel capacity of the communication link is: in, For the spectrum subband bandwidth, Let LoS be the probability channel gain of the communication link between the drone m and the ground user n. This indicates the drone's transmission power on this communication link. This represents the power spectral density of additive white Gaussian noise. This refers to co-channel interference on the communication link between the UAV m and the ground user n, specifically: Where M and N represent the number of drones and ground users, respectively. This represents the transmit power of the communication link between UAV i and ground user j. Use a binary value to represent whether drone i is serving ground user j, where 1 indicates service and 0 indicates no service; The binary value indicates whether ground user j and ground user n are using frequency band reuse, with 1 indicating reuse and 0 indicating no reuse. Step 1.3: Design an evaluation index for UAV data services based on the average peak information age; In time The age of the update data packet with timestamp u is determined as When the updated timestamp corresponds to the current time t and the age is zero, the updated data packet is considered fresh; the average peak information age of ground user i in time slot j is expressed as: in, This represents the cumulative peak count of ground user i in time slot j. Let represent the age of the 'a'th peak information; finally, the average peak information age of each ground user within each time slot is expressed as: Where K represents the total number of time slots; Step 1.4, combining the scenario assumptions of Step 1.1, the communication link channel model between the UAV and the ground user established in Step 1.2, and the UAV data service evaluation index of average peak information age designed in Step 1.3, yields the optimization problem of multi-UAV resource allocation in the data service scenario, specifically described as follows: Among them, (a), (b), and (c) ensure the maximum flight speed constraint of the UAV and its location constraint; (d) represents the formulaic description of the UAV behavior mode selection; (e) ensures that each ground user is served by only one UAV; (f) ensures that the number of ground users served by each UAV does not exceed the number of subcarriers; and (g) ensures that the power consumption of each UAV does not exceed the upper limit of the transmit power.
3. The UAV resource allocation method based on multi-agent reinforcement learning according to claim 1, characterized in that, The specific method for step 2 includes: Step 2.1, design the state space, including the state changes of ground users and drones, namely the relative distance between drones and ground users and the data demand. The specific relative distance between drones and other drones is as follows: in, and These represent the relative distances between the drone m and the ground user n on the x and y axes, respectively. This represents the data demand of ground user n. and Let m represent the relative distances between drone m and other drone i on the x-axis and y-axis, respectively. Step 2.2: Design the action space. The UAV's actions consist of discrete actions and continuous actions. Discrete actions are the selection of action modes, including radar perception mode or data transmission mode. Continuous actions are the continuous parameters across all action modes. When the UAV selects radar perception mode, the radar power is fixed, and continuous actions only vary along the x and y axes of the UAV's position coordinates. Furthermore, when the data transmission mode is selected, continuous actions also include the transmit power allocated on each subcarrier when the UAV selects data transmission. ; Step 2.3, design the reward function as follows: in, The average peak information for time period t is represented by the age penalty factor.
4. The UAV resource allocation method based on multi-agent reinforcement learning according to claim 1, characterized in that, The specific method for step 3 includes: Step 3.1: Design a conditional variational autoencoder to characterize the coupling relationship between discrete and continuous actions; Construct a decodable latent action space, represent dependencies in the hybrid action space, and build an embedding table. to indicate A discrete action, where each row It is Discrete actions of dimension Continuous vector, where, For row indexing, a latent representation space for continuous actions is then constructed using a conditional variational autoencoder (CDAE), implicitly modeling and embedding dependencies. The CDAE consists of an encoder and a decoder. For discrete actions... Continuous actions and state encoder by As parameters, in terms of state and embedding vector As a condition, continuous actions Encoding as potential actions For encoders Using Gaussian distribution ,in, and These are the mean and standard deviation of the encoder output, respectively, and the decoder's response to any potential action. Perform deterministic decoding, in addition, by Parameterized decoder Under the same conditions, from potential actions Decoding continuous actions Arbitrary embedding and any potential action Both are performed as a hybrid action through nearest neighbor lookup in the embedded table and decoding by the decoder. and Therefore, the encoding and decoding process can be represented as: Subsequently, utilizing environmental dynamics, a squared error loss based on state change prediction is designed to further refine the representation of hybrid actions; therefore, the data service volume of each UAV after executing the hybrid actions is used as the state change prediction, i.e. ,in, Let m represent the amount of data service provided by drone m to ground user n, and let m be the data service provided by drone m to ground user n. These represent the parameters of the transformation network, decoder network, and state change prediction network, respectively; the continuous actions of VAE reconstruction. and prediction Using experience pool Batch status and mixed actions and By minimizing the following loss function Training embedding table Conditional VAE: in, It is a continuous action of VAE reconstruction. The squared error, It is the encoder distribution The Kullback-Leibler (KL) divergence between the standard Gaussian distribution and the standard Gaussian distribution; furthermore, minimizing the state change prediction. Prediction Squared error: Therefore, the total training loss of the conditional variational autoencoder is expressed as: ; In addition, two mechanisms are employed: latent action constraints and training experience correction, to handle unreliable latent actions and outdated off-policy training experience, respectively. Constrain potential actions within a reasonable range; that is, for each potential action in the experience pool... Through calculation Obtain the boundary of the center range and Then, each dimension of the potential action is rescaled to a bounded range. ; A training experience correction mechanism is adopted to check the timeliness of potential actions in the experience pool and use the latest potential strategies to correct outdated potential actions. Step 3.2: Combine the MATD3 multi-agent reinforcement learning algorithm to obtain the MATD3 algorithm with coupling relationship representation, and use the MATD3 algorithm with coupling relationship representation to train and update the network parameters of the UAV based on the Markov decision model obtained in Step 2. The latent policy is trained using the MATD3 algorithm; the Actor network of the UAV uses a learned conditional variational autoencoder that represents the coupling relationship between discrete and continuous actions, i.e., the parameters are... Potential strategies To output the latent action vector ,in Simultaneously, a conditional variational autoencoder is trained for each agent, and each agent uses a conditional variational autoencoder to... Decode into mixed actions respectively and ; A dual-critic network for each agent , Using global information as input, an approximate state-action value function is used. From the experience pool Randomly select a batch of experience , As training samples, compute the Critic network. The mean squared error loss function is: in, , The estimated target Q value is expressed as: in, , and For the parameters of the target network, the Actor network for each agent is trained to maximize the state-action value of the potential action output by the network, which is updated through policy gradients; therefore, the loss function of the Actor network is: in, ; Step 3.3: Apply the trained drones to the data service scenario for communication resource allocation; For the hybrid action space design, each UAV is equipped with an Actor network and a Conditional Variational Autoencoder (CVA). The Actor network outputs a latent policy, which is then input into the trained CVA for decoding, yielding discrete and continuous actions. Discrete actions determine the action mode, either radar sensing or data transmission, while continuous actions output continuous action parameters for each action mode. After observing the current state of the environment, the UAV selects either the radar sensing or data transmission mode based on the discrete actions, and then obtains the corresponding continuous action parameters from the continuous actions. When the UAV selects the radar sensing mode, it activates radar to sense the location of ground users and adjusts its own position. When the UAV selects the data transmission mode, it transmits data to ground users on each channel using the transmission power obtained from the continuous actions and adjusts its own position.
5. A UAV resource allocation system based on multi-agent reinforcement learning, based on the method of claim 1, characterized in that, include: The optimization problem establishment module is used to establish an optimization problem for the allocation of resources among multiple drones in a data service scenario assisted by multimodal drones and to design an evaluation index for drone data services based on the average peak information age. The Markov decision problem formulation module is used to formulate the optimization problem of multi-UAV resource allocation in data service scenarios as a Markov decision problem, and designs the state space, action space and corresponding instant reward function respectively. The UAV communication resource allocation module is used to characterize the coupling relationship between discrete and continuous actions by designing a conditional variational autoencoder. Combined with the MATD3 multi-agent reinforcement learning algorithm, a MATD3 algorithm representing the coupling relationship is obtained. Based on the Markov decision model, the network parameters of the UAV are trained and updated using the MATD3 algorithm representing the coupling relationship. The trained UAV is then applied to data service scenarios for communication resource allocation.
6. A resource allocation device for unmanned aerial vehicles (UAVs) based on multi-agent reinforcement learning, characterized in that, include: Memory: A computer program for a UAV resource allocation method based on multi-agent reinforcement learning as described in any one of claims 1-4, which is a computer-readable device; Processor: Used to implement the UAV resource allocation method based on multi-agent reinforcement learning as described in any one of claims 1-4 when executing the computer program.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the implementation of the UAV resource allocation method based on multi-agent reinforcement learning as described in any one of claims 1-4.