An energy efficiency unmanned aerial vehicle resource scheduling method based on deep reinforcement learning
Patent Information
- Application Number
- CN202410310272.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-03-19
AI Technical Summary
[0005]1.现在技术大多只考虑到无人机部署完毕后的资源调度,未综合考虑无人机前期部署导致的功耗问题;
[0012]本发明提出一种基于注意力机制与多模态融合离散-连续混合动作空间的无人机三维部署联合资源调度算法,加入对无人机飞行与悬停功耗的建模估计,一方面通过机载设备采集风速与空气密度信息等多模态数据,结合注意力机制,赋予无人机代理观测状态中各个分量不同的权重来突出重要且关键的信息,实现对周边飞行环境的预测,另一方面通过深度强化学习代理和历史数据的交互来学习数据中的关键信息,在每一个时间步中,代理能够更加关注有价值的环境状态,从而简化动作空间,将无人机飞行的连续动作与功率、频谱分配的离散动作合并作为混合动作空间,联合选取最佳动作,在保证服务质量的前提下最小化无人机功耗。
Smart Images

Figure CN117993475B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of communication technology, and specifically relates to an energy-efficient UAV resource scheduling method based on deep reinforcement learning. Background Technology
[0002] Unmanned Aerial Vehicles (UAVs), also known as drones, are unmanned aircraft that execute commands using radio remote control equipment and self-written programs. Compared to manned aircraft, UAVs are less maneuverable, cheaper, and suitable for performing low-difficulty, high-risk, and low-tolerance missions.
[0003] In emergency disaster relief, drones are primarily used for tasks such as detection, search, and support. In search and detection, their effectiveness is limited by the drones' low intelligence and computing power. They mostly rely on computer vision-based detection methods, transmitting images or audio / video data from onboard sensors. In support roles, their low energy and intelligence limit their ability to perform largely pre-defined tasks, making them ill-equipped to handle complex disaster environments and unable to adapt to dynamically changing conditions.
[0004] In summary, significant progress has been made in exploring resource scheduling methods for UAV-supported communication, but problems still remain:
[0005] 1. Current technologies mostly only consider resource scheduling after drone deployment is complete, without comprehensively considering the power consumption issues caused by the initial deployment of drones;
[0006] 2. The power consumption of the drone was not modeled when scheduling resources, and there was a lack of accurate consideration of the power consumption of the drone, making it difficult to estimate the power consumption in actual applications;
[0007] 3. Existing resource scheduling methods based on deep reinforcement learning mainly target purely discrete action spaces and purely continuous action spaces, or simplify the continuous action space into a discrete action space. This results in unavoidable biases and does not consider the relationship between the two parts. Summary of the Invention
[0008] To address the above problems, this invention proposes an energy-efficient UAV resource scheduling method based on deep reinforcement learning, which specifically includes the following steps:
[0009] Sensors mounted on the drone collect wind speed and air density data in real time, and the collected data is preprocessed.
[0010] The preprocessed data is input into a pre-trained GRU network for feature extraction. The extracted features are then concatenated with UAV observation information to form the UAV state.
[0011] The drone's status is input into the agent, which then selects the optimal resource scheduling strategy for the drone.
[0012] This invention proposes a joint resource scheduling algorithm for UAV 3D deployment based on an attention mechanism and a multimodal fusion discrete-continuous hybrid action space. It incorporates modeling and estimation of UAV flight and hovering power consumption. On one hand, it collects multimodal data such as wind speed and air density information through onboard equipment. Combined with the attention mechanism, it assigns different weights to each component in the UAV agent's observation state to highlight important and critical information, enabling prediction of the surrounding flight environment. On the other hand, it learns key information from the data through deep reinforcement learning agent interaction with historical data. At each time step, the agent can focus more on valuable environmental states, thus simplifying the action space. It merges the continuous actions of UAV flight with the discrete actions of power and spectrum allocation into a hybrid action space, jointly selecting the optimal action to minimize UAV power consumption while ensuring service quality. Attached Figure Description
[0013] Figure 1 This is a schematic diagram illustrating an application scenario of the energy-efficient UAV resource scheduling method based on deep reinforcement learning according to the present invention.
[0014] Figure 2 This is a schematic diagram illustrating line-of-sight transmission and non-line-of-sight transmission of the present invention;
[0015] Figure 3 This is a schematic diagram of the GRU network structure of the present invention;
[0016] Figure 4 This is a schematic diagram of the hybrid action space model of the present invention;
[0017] Figure 5 This is a schematic diagram of the deep reinforcement learning model based on a hybrid action space of attention mechanism and multimodal data. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] This invention proposes an energy-efficient UAV resource scheduling method based on deep reinforcement learning, which specifically includes the following steps:
[0020] Sensors mounted on the drone collect wind speed and air density data in real time, and the collected data is preprocessed.
[0021] The preprocessed data is input into a pre-trained GRU network for feature extraction. The extracted features are then concatenated with UAV observation information to form the UAV state.
[0022] The drone's status is input into the agent, which then selects the optimal resource scheduling strategy for the drone.
[0023] In this embodiment, the UAV resource scheduling method can be roughly divided into two stages: the first stage is UAV deployment, in which it is necessary to model and estimate the flight power consumption of UAVs, divide the target area according to the user distribution, and carry out three-dimensional deployment of UAVs; the second stage uses deep reinforcement learning to make allocation decisions on UAV spectrum, transmission power, and user association.
[0024] Assuming the target area is a square of length Lm × Lm, a two-dimensional mesh is used to decompose the target area into cells. The target area is divided into M squares with side length lm, and a temporary communication base station is placed at the center of the target area I = {λ}. x ,λ y ,0},λ x ,λ y , 0 represent the coordinates of the center of the target area in the three-dimensional coordinate system, and static obstacles (such as buildings, trees, etc.) are randomly distributed within the area; The position of drone i at time t is represented as in U represents the horizontal position of drone i at time t. z The altitude of drone i is represented by [symbol missing]. The position of the ground user at time t is represented by [symbol missing]. This represents the location of the k-th ground user. This embodiment assumes that all drones have sufficient information about the target area before takeoff. To avoid wireless interference, FDMA is used. All drones start from the same location at a constant speed and choose one of eight discrete directions that satisfy the turning angle constraint to fly in. During flight, collision-free operation is ensured, and the drones eventually reach their destination. Application scenarios include... Figure 1 .
[0025] The motion of each UAV i in each unit time interval depends on its flight speed V i (t) and azimuth θ i Since (t)∈(0,2π), the drone position can be updated after each unit time interval, specifically including the following steps:
[0026] First, this embodiment uses a widely used quadcopter drone as the research object to construct a power consumption model during flight. The drone operates at a speed V.i The power consumption model during flight (t) is constructed as follows:
[0027]
[0028] Where S and C represent the wing density and rotor area of the rotary-wing UAV, It is the fuselage drag ratio, S FP The equivalent flat panel area of the fuselage; P0 and P i The blade power and inductive load power are expressed as:
[0029]
[0030]
[0031] Where ρ, ω, R, and W are the air density, blade angular velocity, rotor radius, and UAV weight, respectively; δ is the drag coefficient of the UAV's equivalent flat surface area; v0 is the average induced velocity of the UAV rotor; and U... ti The tip velocity of the rotor blade is represented as:
[0032]
[0033] Therefore, the current flight speed of the drone can be calculated:
[0034]
[0035] Among them, V air The wind speed at the current location of the drone, as collected by sensors.
[0036] When the drone hovers, its speed is V = 0, and its power consumption is calculated as P = P0 + P i ;
[0037] When a drone flies at a speed V, its power consumption is calculated as P(V). Therefore, the energy required for a quadcopter drone to fly and hover can be calculated using the following formula:
[0038] E V =P(V)·T 飞行
[0039] E 悬停 =P(0)·T 悬停
[0040] Among them, E V T is the energy required for the drone to fly at speed V. 飞行 E is the flight time of the drone flying at speed V. 悬停 T is the energy required for the drone to hover. 悬停 The time the drone hovers.
[0041] Based on the above formula, the energy consumed by UAV i during flight within the time interval Δt can be calculated:
[0042]
[0043] Next, the remaining energy of drone i can be updated:
[0044]
[0045] The communication channel model for unmanned aerial vehicles (UAVs), i.e., air-to-ground (A2G), is divided into two types: line-of-sight (LoS) and non-line-of-sight (NLos). When the propagation channel between the transmitter and receiver is unobstructed, the channel model is LoS. In this case, the electromagnetic wave only attenuates during propagation, and this attenuation is positively correlated with distance. When there are obstacles obstructing the propagation channel between the transmitter and receiver, in addition to attenuation, the electromagnetic wave will also be reflected, diffracted, and penetrate between obstacles. These phenomena will cause frequency shift and distortion of the waveform, further increasing the loss. Therefore, it can be determined that the communication channel models for UAV-to-UAV (U2U), UAV-to-base station (U2I), and UAV-to-ground user (U2E) are LoS and NLoS. The probability of these occurrences depends on the propagation environment and the elevation angle of the UAV. Figure 2 The link between the drone and the ground on the left is a LoS link, and the links between the drone and the other two ground users are NLoS links.
[0046] The probability of a Loss of Space (LoS) between drone i and user k is:
[0047]
[0048] in, Let α be the probability of a Loss of Space (LoS) between UAV i and user k at time t; α represents the ratio of building area to total area in the target region; and β represents the number of buildings per unit area (buildings / square kilometer). U represents the elevation angle between the UAV i and the ground user at time t. i (t) represents the three-dimensional position coordinates of UAV i at time t; e k (t) represents the ground user's location coordinates at time t; u z (t) represents the altitude of UAV i at time t; ‖·‖ represents finding the Euclidean distance between two coordinate points.
[0049] The probability of NLoS occurring between drone i and user k at time t is:
[0050]
[0051] Therefore, the path loss between drone i and ground user k can be expressed as:
[0052]
[0053] in, This represents the free space path loss, f c where c is the carrier frequency, η is the speed of light, and c is the carrier frequency. LoS and η NLoS These represent the average additional path loss for LoS and NLoS links, respectively.
[0054] The channel gain between UAV i and user k at time t is:
[0055]
[0056] Assume the total transmit power of the UAV is defined as P max The power allocated to user k by the drone is represented as p. k (t), then we have:
[0057]
[0058] in, This represents the set of ground users.
[0059] The formula for drone energy consumption at time t is updated as follows:
[0060]
[0061] in,
[0062] Assumption The threshold is the reference received power for the signal received by the user, therefore the power received by user k at time t is... Less than This means that the link will not affect the system's throughput, if the received power... Greater than The information transmission rate between the drone network and user k can be expressed as:
[0063]
[0064] Where, σ 2 The sub-band Gaussian white noise is represented by B, the network bandwidth of the drone is B, and the total number of users is K.
[0065] Assuming each user is served by at most one drone, then the following constraints apply:
[0066]
[0067]
[0068] Where ψ i,k (t) = 1 indicates that at time t, ground user k is provided with service by drone i, and vice versa. i,k (t) = 0; K represents the total number of users, and it is stipulated that each user can only be served by one drone at the same time.
[0069] Based on the aforementioned flight power consumption model, the establishment of the channel model requires joint optimization of the user-UAV correlation matrix. drone trajectory and drone power The optimization can be summarized as follows:
[0070]
[0071] Constraints:
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079] Where N is the deadline; γ is the maximum displacement distance of the UAV per unit time; r is the safe radius of the UAV; t The report discount rate is represented by r(t+1); the reward obtained from the environment is represented by r(t+1); T represents the total time; b i,k (t) represents the spectrum resources allocated by UAV i to ground user k at time t; τ represents the set of experiences extracted from the experience pool. A set of experiences is represented as (s,a,r,s′), which means that after selecting action a according to the current state s, the reward value r is obtained and the user transitions from state s to state s′.
[0080] Constraint (1) To guarantee QoS requirements, η is the minimum rate required for each ground user. Constraints (2) and (3) ensure that a user can only be served by one UAV within a single time unit. Constraint (4) represents the maximum range of movement for each UAV within a single time unit. Constraint (5) ensures that all UAVs do not collide with each other. Constraint (6) represents the maximum transmission power constraint for UAVs. Constraint (7) ensures that the bandwidth allocated to each user by a UAV does not exceed the spectrum resources it possesses. Constraint (8) ensures that the remaining energy of the UAV is sufficient for landing.
[0081] In this embodiment, the observation information of the UAV comes from a ground-based edge server and the onboard meteorological sensor carried by the UAV itself. During UAV flight, the current state of the UAV is inevitably affected by the environmental information of the previous moment. If the wind force and air density in the airspace where the UAV is located change abruptly due to a disaster, the UAV's flight state and energy estimation will also inevitably change. Therefore, this invention uses a GRU network to process the multimodal observation data proxies by the UAV and to discover the correlation between time series data in order to make reasonable predictions about the UAV's flight environment.
[0082] like Figure 3 , where h t-1 This represents the hidden state at the previous time step. The hidden state acts as the memory of the neural network, containing information about the data that the previous node had seen. t This represents the hidden state that is passed on to the next time step. Represents the candidate hidden state, r t To reset the door, z t This is the update gate. u(·) represents the sigmoid function. From this structure, the network hidden state h at time t can be derived. t The expression is as follows:
[0083]
[0084] The expression is as follows:
[0085]
[0086] The expressions for resetting the gate and updating the gate are as follows:
[0087] r t =u(W r x t +K r h t-1 +b r )
[0088] z t =u(W z x t +Kz h t-1 +b z )
[0089] Among them, W h K h b h W r K r b r W z K z b z x represents the learnable matrices of each layer in the GRU network. t This represents the input to the GRU network at time t, which includes UAV observation information and sensor observation information.
[0090] Attention mechanisms are mechanisms that simulate how humans selectively focus on important parts when processing large amounts of information. In deep learning, attention mechanisms allow models to learn which parts are important at a given time, thus making information processing more efficient. Essentially, it involves learning a weight distribution for the input features and then applying this weight distribution to the original features, causing the task to focus primarily on key features and ignore less important ones, thereby improving task efficiency. Let the input sequence vector be:
[0091] Γ=[Γ1(t),…,Γ Π (t)]
[0092] The formula for calculating the attention mechanism is as follows:
[0093]
[0094] in: The weight matrix is used to perform matrix operations with the input sequence Γ, followed by a Softmax activation function, and finally multiplied by the input sequence to obtain a new sequence Γ′. Γ1(t) represents the first component of the input sequence at time t, and Π represents the number of components in the input sequence. The attention mechanism can highlight important features and reduce the influence of useless features, enabling the model to make better choices and improve prediction accuracy. In this embodiment, the attention network consists of a fully connected layer and a Softmax activation function. First, the input vector passes through the first fully connected layer, then the Softmax activation function is used to obtain the weights of each component in the input vector, and finally multiplied by the input vector to obtain a new vector.
[0095] Drones interact with ground users via wireless channels. The drone acts as an agent, while the user and the wireless channel represent the environment. The drone's flight causes changes in its position, which in turn leads to changes in the wireless channel. Therefore, this process can be modeled as a Markov Decision Process (MDP). It is defined as a quadruple. This involves the state space, action space, state transition probabilities, and reward. In each unit time interval, the drone transitions from the current state to the next state based on the action and transition probability, and receives a reward. The drone stores the iterative process as experience in a buffer and randomly samples data to train the neural network. The state space, action space, and reward are defined as follows:
[0096] State Space: In the model established by the invention, the environmental state s(t) is mainly constructed using the three-dimensional coordinates of all drones, ground users, and drone base stations, the remaining energy of the drones, and the wind speed and air density data collected by sensors. The coordinates of the drones and ground users are perceived by the ground base stations and the association between the drones and users, while the information collected by the sensors is perceived through airborne equipment. Therefore, a multimodal data fusion process exists, requiring processing to form the input sequence of the deep neural network. Here, we tentatively represent the processed wind speed and air density information collected by the sensors as a vector V. air The state-space representation of ρ(t) and ρ(t) is as follows:
[0097]
[0098] in, Let K be the total number of drones, and K be the total number of ground users.
[0099] Action Space: A reinforcement learning agent determines a joint action at the current time interval t by observing state information. In the application scenario described in this invention, the UAV's action space comprises four parts: user association strategy A, UAV trajectory design U, transmit power P, and spectrum selection B. However, since the UAV's flight state requires adjustments to multiple dimensions such as tilt angle, acceleration, and velocity, and these actions are continuous, they are difficult to convert into a discrete action space. Therefore, the UAV's action space is a hybrid of discrete and continuous actions. The discrete action space a is defined as follows: d =[A,P,B], then define the parameterized continuous action space a c = [dip_angle, acceleration], where dip_angle is the drone's tilt angle and acceleration is the drone's acceleration, such as Figure 4In this invention, each set of discrete action selections by the UAV agent corresponds to a set of continuously parameterized action selections; that is, each discrete action corresponds to a set of continuous action parameters. Therefore, the action space can be represented as:
[0100] a(t)={a d (t)|a c (t)}
[0101] Each discrete action a d Each (t) has a corresponding continuous action parameter a. c (t).
[0102] Reward Function: Based on the optimization objective proposed in this invention, the aim is to maximize the cumulative expected reward, thereby obtaining the maximum information and speed of the UAV support network system. Furthermore, considering the energy-saving issue of the UAV, the proposed reward function is as follows:
[0103]
[0104]
[0105] r(t) = r1(t) + r2(t)
[0106] Where ζ and μ represent the weights of the maximum information and rate rewards and the energy consumption reward, respectively, t complete This indicates the initially preset target time for the drone network to maintain its operation.
[0107] Generally, the higher the UAV signal transmission power, the greater its channel gain and the higher the information rate obtained. However, with the increase in transmission power, its energy consumption also increases, and the UAV network maintenance time becomes shorter. Therefore, the two weight parameters should be set reasonably. Generally, to ensure basic service quality, ζ should be set greater than μ. When the UAV's loitering time is greater than or equal to the preset time and the UAV's remaining energy is greater than or equal to the energy consumed during the return (assuming that the energy consumed by the UAV to fly to the destination and return to the starting point is equal), it receives the maximum reward for completing the preset goal; otherwise, it does not receive the reward. For existing deep reinforcement learning algorithms, most of them require the action space to be discrete or continuous. For example, Deep Q-Learning (DQL) and its variants are suitable for discrete action spaces; while Deep Deterministic Policy Gradient (DDPG) is widely used in continuous action spaces. However, the UAV's flight actions, power, and spectrum selection need to handle the case of a mixed discrete-continuous action space. In addition, in order to save UAV flight power consumption to the greatest extent, attention mechanisms are combined to specifically learn the changes in ground user distribution, channel environment, and high-altitude wind speed and air density during UAV loitering, and timely predict state changes to ensure optimal signal coverage.
[0108] In the DQN concept, Q(s,a) represents the quality of performing action a given state s. A similar Q value needs to be defined for the mixed action space when dealing with discrete-continuous mixed action space problems; therefore, Q(s,a) is defined as follows: d ,a c ), indicating that discrete action a is selected at time t. d and its associated continuous action a c .
[0109] By combining DQN and DDPG to operate directly in the parameterized action space (PAS), we first obtain the parameterized actions of the continuous actions corresponding to all discrete actions. A Q-network is used to output the Q-values of the discrete actions and provide gradients for the policy network. Therefore, the Bellman equation can be rewritten as:
[0110]
[0111] in, Represents the Bellman expectation equation; r t This indicates that at time t, state s t Take action a d ,a c The rewards received; Indicates with a c The upper bound of the Q-function is the variable.
[0112] The problem is that DQN can easily use the action with the highest Q value in a finite number of actions, but in a continuous action space a c Finding the maximum value is tricky because DQN requires iterating through all possible values in a continuous space to find the maximum Q(s). t ,a d ,a c The solution to this is to use a deterministic policy network χ(s). t ;θ) to approximate a c For discrete actions, the parameters used are Deep neural network representation Therefore, when When fixed, we want to determine θ such that:
[0113]
[0114] The state observed by the UAV is first input into the first policy network χ(θ), which uses the Deep Deterministic Policy Gradient (DDPG) algorithm to determine the action parameters in the continuous action space. After passing through several fully connected layers and activation functions, the output is the optimal value of the continuous parameters. Then, the state and the optimal continuous parameter values are concatenated, processed through a normalization layer, and input into the attention layer. In the attention layer, the attention mechanism distributes attention with different weights to different components of the input sequence, selectively focusing on the intrinsic relationship between the optimal action parameters and the state, ignoring less important states. Finally, the output is used as a discrete network. The input can be viewed as a DQN network, which, after passing through some fully connected layers and activation functions, outputs the optimal discrete action in the DQN and selects the corresponding continuous parameters.
[0115] Similarly, the loss function is also divided into two parts: the Q-network part for discrete actions is optimized using TD-error, similar to DQN.
[0116]
[0117] Where s, a∈mini_batch are mini-batch data sampled from the empirical replay memory, and y is... Given, where a′∈a d (t), where γ is the future reward discount factor. The policy network for continuous actions is designed to deterministically provide optimal continuous parameter values. With the discrete network parameters fixed, the discrete network can act as a "critic," inputting the continuous parameters into the discrete network to obtain multiple corresponding Q-values. These Q-values are then summed to maximize this result, thus optimizing the continuous Q-network. This can be represented as:
[0118]
[0119] The following is a flowchart of the actual operation of this invention, such as... Figure 5 Specifically, it includes the following steps:
[0120] 1. Initialize map and environment information (α,β,ρ,σ) 2 User distribution (K,e) k (t)), set the training batch size batch_size, ε-greedy policy constant.
[0121] 2. Initialize the channel model, UAV power consumption model, UAV energy E, and number of UAVs. Initial distribution location of the drone u i A. The relationship between drones and users.
[0122] 3. Construct the agent's state s, action a, action parameters, and reward r, initialize the reward function and loss function, and obtain the configured agent model;
[0123] The proxy model includes an online network and a target network. The online network consists of an online deterministic policy network χ(s;θ) with network parameters θ and a target network with network parameters θ. Online deep Q network The target networks also contain a target deterministic policy network χ(s; θ′) with network parameters θ′ and a network with parameters θ′. Target Depth Q-Network
[0124] The agent's state space vector is s = {s(t) = {u i (t),E i (t),v i (t),V air (t),ρ(t)}}, where u i (t),E i (t),v i (t) represents the observation information obtained by the UAV through data exchange with ground base stations and users, and its built-in battery module. air ρ(t) and ρ(t) are the wind speed and air density at the next moment, which are obtained by the information data collected by the UAV's onboard sensors and predicted by the GRU network. In this invention, the information collected by the sensor at the current moment is input into the GRU network to predict the wind speed and air density at the next moment, and then the predicted values are used as the wind speed and air density features in the current state of the UAV.
[0125] The GRU is pre-trained using a dataset of historical weather information at different altitudes provided by the meteorological bureau, and the model is saved. Information collected by the drone through sensors is integrated into the data through a data preprocessing module and used as the input sequence. During feature extraction, the GRU first converts the input sequence into a vector representation step by step over time, and then models it through the GRU network. In the GRU network, the hidden state at each time step is a weighted sum of the previous information, and the flow of information is controlled by a gating mechanism. Over time, the GRU can gradually capture the long-term dependencies in the sequence and learn useful features, namely air wind speed and air density.
[0126] After processing by the GRU, the information collected by the sensor is used as the input sequence, which is then used as the input sequence for the subsequent attention layer.
[0127] 4. Input the initial state s at the current time into the policy network χ(s; θ) of the online network, and select the continuous action parameters according to the policy network. To explore more rationally, noise N is added to the continuous action parameters. The continuous action parameters and the current state are normalized and then input into the attention layer. The attention layer assigns different weights to each state to capture key information and make the best prediction of the environment state at the next moment.
[0128] The output of the attention layer is then input into the online network. Output the Q-values of each discrete-continuous mixed action group, and use the ε-greedy strategy to explore and utilize the data to determine the next action and obtain the state at the next moment.
[0129] The reward function is used to evaluate the performance of the mixed action a under the current state s. The objectives are as follows: (1) avoid drone collisions; (2) adjust the drone attitude and altitude to maximize the information rate of the drone network system; (3) save drone energy as much as possible to ensure the lifespan of the drone network.
[0130] The obtained [s,a d ,a c ,r t ,s'](s' represents the action a to be performed in state s. d ,a c The final state is stored in the experience pool for subsequent neural network training. When the amount of data in the experience pool reaches batch_size, the experience is extracted from the experience pool. The extracted experience is represented as... Will Inputting into the online Q network, we obtain the state s. memory Next action group Q value And reward r′, will s′ memory Input the target policy network to obtain the set of parameters for the target's continuous actions. Combine the target action parameter set with the state s′ memory The Q-value of the target discrete-continuous hybrid action is obtained by inputting the target Q-network. Next, the loss function is calculated based on the TD-error, and the parameters are updated using gradient descent. The loss function is expressed as:
[0131]
[0132]
[0133] Update Loss(θ) = ∑χ(s,a) based on gradient ascent. d ,a c ;θ).
[0134] 5. Target network parameters θ′ and Obtained through a soft update.
[0135] 6. Repeat the above steps until the maximum number of training steps is reached, and finally obtain the trained model.
[0136] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A resource scheduling method for energy-efficient unmanned aerial vehicles (UAVs) based on deep reinforcement learning, characterized in that, Specifically, the following steps are included: Sensors mounted on the drone collect wind speed and air density data in real time, and the collected data is preprocessed. The preprocessed data is input into a pre-trained GRU network for feature extraction. The extracted features are then concatenated with UAV observation information to form the UAV state. The drone's status is input into the agent, which then selects the optimal resource scheduling strategy for the drone, specifically including: The agent inputs the UAV state into an online deterministic policy network to obtain action parameters in the continuous action space. The agent's action space includes a hybrid action space consisting of discrete action selection and continuous parameterized action selection. Discrete action selection includes user association policy A, transmit power P, and spectrum selection B. Continuous parameterized action selection includes UAV trajectory U. Each set of discrete action selection corresponds to a set of continuous parameterized action selection. The agent first obtains the action parameters in the continuous action space through the online deterministic policy network, and then determines the discrete action selection through an online deep Q-network. The continuous parameterized action selection includes UAV tilt angle and UAV acceleration. The motion parameters of the continuous motion space are concatenated with the UAV state and then normalized. Each component of the normalized data is weighted using an attention mechanism. The weighted data is input into an online deep Q-network, which outputs the Q-values of each discrete-parameterized action pair. intelligent agents according to - A greedy strategy randomly selects a discrete-parameterized action pair and follows... The agent's reward function includes: Probabilistically selecting discrete-parameterized action pairs that maximize Q-values. in, The reward value for the action performed at time t; A reward function that maximizes the total information rate and remaining energy consumption; As an additional bonus, a one-time bonus value will be awarded when the drone service time meets the predetermined time. Otherwise, the additional reward value is 0. It is a constant greater than 0; , These represent the weights of the maximum information reward and the rate reward, and the energy consumption reward, respectively. Let t be the information transmission rate between the drone and user k at time t; K is the total number of users. Let i be the remaining energy of the drone at time t; This represents the total number of drones; The target time for maintaining the pre-defined drone network; The energy required for the drone to return along the same route is ideally the same as the energy consumed when the drone arrives at the deployment point.
2. The energy-efficient UAV resource scheduling method based on deep reinforcement learning according to claim 1, characterized in that, The drone status is represented as follows: in, This represents the initial state of the UAV at time t. Let i be the position of UAV i at time t; Let be the coordinates of the ground user k at time t; Let i be the remaining energy of the drone at time t; Let be the velocity of UAV i at time t; The wind speed feature extracted by the UAV through the GRU network at time t; Let t be the air density feature extracted by the UAV through the GRU network.
3. The energy-efficient UAV resource scheduling method based on deep reinforcement learning according to claim 1, characterized in that, In the process of optimizing the parameters of the online deterministic policy network, the sum of the Q-values of the continuous parameters of the input online deep Q-network is used as the loss function of the online deterministic policy network. The parameters of the online deterministic policy network are optimized by gradient ascent. The loss function of the online deterministic policy network is expressed as: in, For online deterministic policy networks; This represents the Q-value of each discrete-parameterized action pair output by the online deep Q-network; The drone is in a certain state. Discrete actions selected for the drone; Selected continuous actions for the drone; These are the network parameters for an online deterministic policy network.
4. The energy-efficient UAV resource scheduling method based on deep reinforcement learning according to claim 1, characterized in that, Online deep Q-networks update network parameters using gradient descent with a loss function. The loss function of an online deep Q-network is expressed as: in, Let be the loss function for the online deep Q-network; Indicates an online deep Q network. Indicates state, Indicates discrete actions. This represents parameterized continuous actions; This indicates a demand for expectation; Indicates the execution of an action The reward received later; Indicates a time-related discount factor; This represents the output of the target Q-network. Indicates the next state after s. Indicates in Discrete actions performed, This represents the parameters of the target Q-network.
5. The energy-efficient UAV resource scheduling method based on deep reinforcement learning according to claim 1, characterized in that, The feature extraction process of the GRU network includes: in, Let t be the hidden state at time t, i.e., the feature extracted by the GRU network; The output of the GRU network update gate at time t; The output of the GRU network reset gate at time t; Let be the candidate hidden state of the GRU network at time t. , and These are the learnable parameters corresponding to the candidate hidden states. The data represents the wind speed and air density collected by the sensor at time t. For activation functions; , and Reset the learnable parameters of the gates in the GRU network; , and Update the learnable parameters of the gates for the GRU network.
Citation Information
Patent Citations
Three-dimensional deployment and power distribution joint optimization method for flight base station of unmanned aerial vehicle
CN113206701A
Unmanned aerial vehicle data acquisition trajectory and user association joint optimization method based on reinforcement learning in wireless network
CN115616906A