Energy efficiency optimization method for unmanned aerial vehicle assisted communication system based on deep reinforcement learning

By constructing a UAV-RIS assisted communication system model and adopting the OAN-SD3 algorithm, the problems of low energy efficiency and poor dynamic adaptability in UAV assisted communication systems are solved, achieving efficient resource allocation and improved communication quality.

CN121508617BActive Publication Date: 2026-05-12DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DALIAN UNIV OF TECH
Filing Date
2025-11-14
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Unmanned aerial vehicle (UAV) assisted communication systems suffer from low energy efficiency, poor communication quality, and poor dynamic adaptability. Traditional resource allocation methods struggle to respond to dynamic environmental changes in real time, and deep reinforcement learning algorithms have limitations when dealing with high-dimensional, mixed-integer non-convex optimization problems.

Method used

The OAN-SD3 algorithm based on deep reinforcement learning is adopted. By constructing a UAV-RIS assisted communication system model, the resource allocation strategy is optimized by combining time allocation, power allocation, RIS phase configuration and switching state. The Softmax Q-value aggregation mechanism and dual Actor network structure are used to solve the high-dimensional, mixed integer non-convex optimization problem.

Benefits of technology

It significantly improves the system's energy efficiency and communication quality, enables intelligent resource allocation in dynamic environments, optimizes the coordinated utilization of multi-dimensional resources such as power, time, and phase, and meets the service quality requirements of multiple users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121508617B_ABST
    Figure CN121508617B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of wireless communication, and particularly relates to an energy efficiency optimization method of a UAV-assisted communication system based on deep reinforcement learning. The present application models a UAV-assisted communication system with an RIS, the UAV-RIS provides channel enhancement function for BS to communicate with UE through downlink, and the interference factor of the jammer and the energy collection characteristics of the RIS are considered. Through accurate system modeling, the complex physical layer communication and energy transmission process are converted into an optimizable mathematical model. For the optimization problem of maximizing energy efficiency, the OAN-SD3 algorithm is innovatively proposed, which converts the original high-dimensional, mixed and non-convex problem into a Markov decision process, and then uses a deep reinforcement learning algorithm to solve it, effectively solving the high-dimensional mixed integer non-convex optimization problem that is difficult to handle by traditional optimization methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication technology, specifically relating to an energy efficiency optimization method for UAV-assisted communication systems based on deep reinforcement learning, which is particularly applicable to dynamic resource allocation and energy efficiency optimization in UAV communication systems equipped with reconfigurable smart metasurfaces. Background Technology

[0002] With the rapid development of Unmanned Aerial Vehicle (UAV) and Reconfigurable Intelligent Surfaces (RIS) technologies, UAV-assisted wireless communication systems are showing great application potential in emergency communication, the Internet of Things (IoT), and future 6G networks. UAVs, with their high mobility and flexible deployment capabilities, can effectively extend communication coverage, especially in scenarios where traditional ground base stations cannot reach or are damaged. Reconfigurable intelligent metasurfaces, as a novel wireless communication technology, can significantly enhance channel quality and improve communication performance by intelligently controlling the propagation characteristics of electromagnetic waves.

[0003] However, unmanned aerial vehicle (UAV)-assisted communication systems face many challenges: First, the onboard energy of UAVs is limited, and energy supply becomes a key constraint on the continuous operation of the system; second, the wireless channel environment is highly dynamic, and factors such as UAV position, user movement, and environmental obstruction cause the channel state to change rapidly; third, there are multiple sources of interference in the system, including active interference from malicious jammers, which further increases the uncertainty of communication.

[0004] Traditional resource allocation methods are mainly based on mathematical methods such as convex optimization and alternation optimization. These methods perform well in static or small-scale networks, but they face serious limitations in dynamic and large-scale networks: first, they are difficult to adapt to changes in the environment in real time; second, they have high computational complexity when dealing with high-dimensional state spaces and complex decision spaces; and third, they cannot effectively handle mixed-integer nonlinear programming problems.

[0005] Deep Reinforcement Learning (DRL) technology learns optimal policies through continuous interaction between agents and the environment, providing a new approach to solving dynamic resource allocation problems. Among commonly used DRL algorithms, the Deep Q-Network (DQN) algorithm, while automatically extracting state features by introducing a deep neural network, is similar to Q-learning in that it is only applicable to discrete action spaces. It requires selecting the maximum value operation to solve for the optimal action, cannot directly handle continuous actions, and suffers from the risk of Q-value overestimation. While the Deep Deterministic Policy Gradient (DDPG) algorithm achieves efficient policy optimization in continuous action spaces, a major drawback is the overestimation problem. The Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm effectively alleviates the overestimation problem in DDPG by introducing a twin-delayed deep deterministic policy gradient technique. However, when facing the complex non-convex problem of energy efficiency optimization in the RIS-assisted UAV communication system of this invention, the TD3 algorithm still has several limitations. For example, the TD3 algorithm uses a minimum Q-value estimation strategy to suppress overestimation bias, but this pessimistic estimation may be overly conservative in non-convex optimization problems, causing the strategy to get stuck in local optima and unable to escape.

[0006] Therefore, there is an urgent need for an optimization method that can effectively solve the above problems in order to improve the energy efficiency, communication quality and dynamic adaptability of UAV-assisted communication systems. Summary of the Invention

[0007] The purpose of this invention is to overcome the technical problems of low energy efficiency, poor communication quality, and poor adaptability to dynamic environments in existing UAV-assisted communication systems for resource allocation scenarios.

[0008] 1. Low energy efficiency: Traditional UAV communication systems use energy inefficiently, and the limited onboard energy of UAVs severely restricts their endurance and service quality.

[0009] 2. Poor dynamic adaptability: In dynamic environments with user movement and malicious interference, traditional optimization methods based on static assumptions (such as convex optimization) need to be recalculated every time resources are allocated, making it difficult to respond in real time and causing a sharp drop in performance.

[0010] 3. Uneven resource allocation makes it difficult to achieve coordinated optimization among multiple resources such as power, time, and phase, resulting in low resource utilization efficiency and an inability to meet the Quality of Service (QoS) requirements of multiple users.

[0011] 4. Limitations of the algorithm: Existing deep reinforcement learning algorithms (such as DDPG and TD3) have problems such as Q-value estimation bias, insufficient policy exploration, and reward convergence fluctuation when dealing with the high-dimensional, mixed-integer non-convex optimization problem involved in this invention, and are prone to getting trapped in local optima.

[0012] Therefore, this invention specifically addresses the scenario of UAV-assisted communication equipped with RIS, and provides an energy efficiency optimization method based on deep reinforcement learning, aiming to achieve intelligent resource allocation in dynamic environments and significantly improve the overall energy efficiency of the system.

[0013] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:

[0014] An energy efficiency optimization method for a drone-assisted communication system based on deep reinforcement learning includes the following steps:

[0015] S1: Establish a UAV-RIS assisted communication system model

[0016] Construct a complete system model that includes downlink communication channels, interference signal transmission channels, energy harvesting, and energy consumption.

[0017] In a communication scenario, ground-based fixed base stations (BS) and ground-based user equipment (UE) cannot communicate via a direct transmission link. UAVs equipped with a Resonant Array (RIS) have phase modulation capabilities, providing channel enhancement services. The BS has Z antennas, providing communication services to M ground-based UEs with single antennas. The RIS on the UAV has L elements, and it possesses both phase modulation and energy harvesting capabilities. Furthermore, an airborne jammer exists. The BS provides services to ground users through the RIS on the UAV, while the ground users are affected by interference signals reflected from the RIS by the jammer. In a complete mission, the total time... Divided into K time slots, with index as To facilitate the description of the segmented constant positional behavior of UAVs within a time slot, K+1 time slot boundary times are introduced. And let the three-dimensional positions of the UAV at the boundary time be respectively Therefore, in the k-th time slot Inside, the position of the UAV is considered to remain constant and equal to The communication channel state is considered unchanged.

[0018] (1) Downlink communication channel

[0019] The downlink communication channels include the BS-UAV link, the UAV-RIS signal enhancement, and the UAV-UE link.

[0020] Channel matrix of BS-UAV link in the kth time slot The model is as follows:

[0021]

[0022] in, express The complex set of dimensionless numbers, i.e., the channel matrix It is A complex matrix of dimension 1 This is large-scale fading. When modeling, free-space path loss and the line-of-sight / non-line-of-sight (LoS / NLoS) probability are considered. The calculation formula is as follows:

[0023]

[0024] in, The probability of a Loss of Service (LoS) transmission occurring. The probability of NLoS transmission occurring. Additional attenuation factor for NLoS, The air-to-ground path attenuation index. and These are the positions of BS and UAV, respectively. For small-scale Rayleigh fading, its elements follow a complex Gaussian distribution. .

[0025] UAV-RIS signal enhancement is achieved by the RIS reflection coefficient matrix. It is indicated that it is a diagonal matrix. The characteristics of each reflective unit are determined by... Description, in which, Represents the imaginary unit. It is the continuously adjustable phase shift of the l-th unit. It is the switch state that controls whether the l-th unit is turned on or off, where 1 indicates on and 0 indicates off.

[0026] For a UAV-UE link, the channel vector from UAV-RIS to the m-th user Using the Rice fading model:

[0027]

[0028] in, express The set of complex numbers of dimension 1 The Rician K factor, This represents large-scale path loss. Let the distance between the UAV in the k-th time slot and the m-th user be denoted as . This is the path loss index. Let be the path loss constant. for A 1-dimensional column vector of all ones. This represents the small-scale fading component, which follows a complex-valued Gaussian distribution with a mean vector that is an L-dimensional zero vector. Covariance matrix for An identity matrix of dimension 1.

[0029] Based on the above, an equivalent cascaded communication channel can be constructed for the downlink communication link from BS to UAV, then enhanced by UAV-RIS, and transmitted to UE. Considering that the BS has Z antennas, and employing an equal-power omnidirectional precoding strategy, the effective scalar channel coefficients from the BS to the m-th user can be obtained. for:

[0030] .

[0031] (2) Interference signal transmission channel

[0032] jammer to UAV channel Also using the Rice model This indicates that the channel model is a Using a complex vector of dimension m, combined with the downlink communication model, we can obtain the equivalent interference channel coefficients from the jammer to the m-th user. for:

[0033] .

[0034] (3) Energy harvesting model

[0035] Each duration is The time slots are divided into an Energy Harvesting (EH) phase and an Information Transmission (IT) phase, with the EH phase lasting for [duration missing]. The IT phase lasts for [duration]. , The time allocation factor is optimizable. The total energy collected by RIS in time slot k. for:

[0036]

[0037] in, and These are the RF power received during the EH and IT phases, respectively. It refers to energy conversion efficiency.

[0038] (4) Energy consumption model

[0039] The system's energy consumption model takes into account the UAV propulsion energy consumption and the BS launch energy consumption.

[0040] UAV propulsion energy consumption Adopting flight speed A precise physical model:

[0041]

[0042] in, and These represent the constant airfoil power and induced power during hovering, respectively. The tip velocity of the rotor blades. The average rotor induced velocity during hovering. s and s represent the fuselage drag ratio and rotor solidity, respectively. air density, R is the rotor disk area, and R is the rotor radius.

[0043] BS launch energy consumption The calculation is as follows:

[0044]

[0045] in, Let be the transmit power that BS allocates to user m in the k-th time slot.

[0046] Combine energy harvesting models to construct the system's net energy consumption. for:

[0047] .

[0048] S2: Constructing an optimization problem with energy efficiency as the core.

[0049] The optimization objective is to maximize the average energy efficiency over a single task cycle.

[0050]

[0051] in, To determine the system energy efficiency in the k-th time slot, The effective communication capacity of the m-th user in the k-th time slot, taking into account the time allocation factor, is calculated as follows:

[0052]

[0053] in, Let be the channel bandwidth for the m-th user to communicate via the downlink. The jammer's transmission power (considered constant) is given. Let H be the noise power, where the superscript H of the vector / matrix indicates the conjugate transpose of the corresponding original vector / matrix.

[0054] Optimization variables include four categories: time allocation factors. Power allocation vector RIS phase configuration vector RIS switch state vector ,in, , Let M and L represent real vector spaces of dimensions respectively.

[0055] The optimization problem is subject to several constraints, namely C1-C7, regarding the duration, power, phase, switching state, quality of service, and net energy consumption of the EH phase.

[0056]

[0057] The time allocation constraint ensures that the time allocation factor is a reasonable proportion;

[0058]

[0059] Power constraints ensure that there is an upper limit to the power allocation for each user; among them, The minimum power allocated to a single user, The maximum power allocated to a single user;

[0060]

[0061] Total power constraints ensure that the total power of the BS's RF power amplifier does not exceed the maximum limit; among which, This represents the maximum total power for all users within a single time slot.

[0062]

[0063] RIS constraints ensure the range of phase offset;

[0064]

[0065] RIS constraints ensure that integer variables are used to enable and disable RIS units;

[0066]

[0067] Quality of Service (QoS) constraints require that each user's communication capacity must meet minimum requirements; among which, The minimum communication capacity required to satisfy QoS constraints for a single user;

[0068]

[0069] Net energy consumption positive constraint, among which, This is the minimum power limit for the system within a single time slot.

[0070] S3: Transform the optimization problem into a Markov decision process.

[0071] The optimization problem constructed in S2 is a complex nonconvex mixed-integer nonlinear programming problem, which is modeled as a Markov decision process. To distinguish the time scale of the communication system from that of the Markov decision process, this invention uses different symbols to represent time variables: the time slot for establishing the communication model. Used to describe changes in UAV location, channel status, and other system states over time; The time step, used to describe the Markov decision-making process, represents the sequence in which the agent (i.e., the subject controlling the decision-making, the Actor network in deep reinforcement learning algorithms) interacts with the environment. In this invention, the agent performs an action at the beginning of each communication time slot; therefore, each time slot is considered the smallest time unit for one decision by the agent, and the two satisfy the following correspondence:

[0072]

[0073] In other words, a communication time slot corresponds to a time step in reinforcement learning. Therefore, all variables indexed by time slot k in the communication model (such as UAV location, channel coefficients, etc.) can be represented by time step t in the Markov decision process. Furthermore, an episode represents a complete task cycle, with a fixed number of time steps, K. After performing an action and receiving a reward, the turn ends and the environment is reset.

[0074] state space The design primarily considers the state information required to establish a communication channel at time step t, specifically:

[0075]

[0076] in, This indicates the distance between the base station (BS) and the user aerial vehicle (UAV). Indicates the distance between the jammer and the UAV. This represents the distance between the UAV and the m-th user. This represents the end-to-end equivalent channel coefficients from the BS to the m-th user. This represents the equivalent interference channel coefficient of the signal transmitted by the jammer after being reflected by the RIS and reaching the m-th user. The state space is represented as A dimensional real vector space.

[0077] The motion space at time step t is designed as follows: That is, all the optimization variables in the original optimization problem, where, As a time allocation factor, For power allocation vector, Configure vectors for RIS phases. This is the RIS switch state vector.

[0078] Considering the complexity of the optimization objective, a reconstructed reward function is designed. This is a combination of communication capacity reward and constraint violation penalty, compressed to the (-1,1) interval using the tanh function to ensure training stability.

[0079]

[0080] in, Let t be the communication capacity of the m-th user. Let be the capacity normalization constant. , and As a penalty weight, Penalties for violating QoS Penalty for exceeding total power limit. Penalties for excessive net energy consumption.

[0081] S4: Training using the OAN-SD3 algorithm

[0082] The proposed Softmax Deep Double Deterministic Policy Gradient (OAN-SD3) algorithm for Optimized Actor Networks is an optimization of the traditional TD3 algorithm, used to solve the problem established in S3. First, based on the TD3 algorithm, a Softmax Q-value aggregation mechanism is introduced, and a dual-Actor network structure is maintained, forming the SD3 algorithm. Then, the structure of the Actor network is optimized to obtain the complete OAN-SD3 algorithm, the main parts of which are as follows:

[0083] (1) Softmax Q-value aggregation and dual-actor network

[0084] Building upon the TD3 algorithm, the SD3 algorithm maintains two pairs of independently trained Actor-Critic networks: and Each pair of networks has its own master network and target network, and the corresponding target Actor-Critic networks are respectively... and ,in Both refer to Actor networks. Both refer to Critic networks. These are network parameters.

[0085] For the next state N noisy action samples are generated using a target actor network. First, for each sample, a bi-objective Critic network is used for evaluation, and the minimum value among them is taken. Then use the Softmax function on N We obtain the soft Q-value objective by performing a weighted average:

[0086]

[0087] in, This refers to the temperature parameter.

[0088] (2) Optimized Actor Network Structure

[0089] Based on the SD3 algorithm, the OAN-SD3 algorithm proposed in this invention introduces a new algorithm between the first and second fully connected layers of the traditional Actor network architecture. The convolutional feature reorganization layer refactors the output features of the first fully connected layer. The channel dimension is linearly transformed. Specifically, the convolution operation expression is:

[0090]

[0091] in, Indicates the kernel size as One-dimensional convolutional layer, Convolution can linearly map the "channel dimension" without changing the dimension of the feature space, thereby enhancing the coupling relationship between features of different channels.

[0092] Then, the convolutional features With original features The residuals are summed and input into the feature fusion layer, using the weight matrix. With bias vector Perform a linear transformation and apply the ReLU activation function to obtain the fused features:

[0093]

[0094] in, This is the weight matrix of the feature fusion layer, used to perform a linear transformation on the fusion result of the original features and the convolutional features. This is the bias vector of the feature fusion layer. The output features of the feature fusion layer.

[0095] In addition, random perturbations are added to the actions output by the Actor network during network training. To promote exploration of the environment, it follows a mean of 0 and a variance of . It follows a normal distribution. During training, an exponential decay strategy is employed, adjusting the decay rate according to the training progress. Dynamically adjust the standard deviation of exploration noise To adapt to the needs of different training stages:

[0096]

[0097] in, It is the maximum standard deviation. It is the minimum standard deviation. It is the attenuation rate.

[0098] S5: Verify the online resource allocation strategy

[0099] First, set the environment and training parameters. Then, continuously interact with the environment through the agent (i.e., the Actor network of the OAN-SD3 algorithm in S4) to collect experience data. ,in The environmental state at time step t, Actions performed by the intelligent agent The immediate reward obtained by the intelligent agent. The state of the environment at time step t+1. This indicates whether the task has ended. The OAN-SD3 algorithm is used for training. After training, the optimized policy network is deployed for verification testing, enabling it to generate optimal resource allocation decisions online based on real-time perceived system states. By dynamically adjusting time allocation, transmit power, RIS phase, and switching states, the system's energy efficiency is continuously maximized.

[0100] The beneficial effects of this invention are:

[0101] (1) A UAV-assisted communication system equipped with RIS is modeled. The UAV-RIS provides channel enhancement for the BS to communicate with the UE via the downlink, and takes into account the interference factors of the jammer and the energy harvesting characteristics of the RIS. Through accurate system modeling, the complex physical layer communication and energy transmission process is transformed into an optimizable mathematical model.

[0102] (2) For the optimization problem of maximizing energy efficiency, the OAN-SD3 algorithm is innovatively proposed. By transforming the original high-dimensional, mixed, non-convex problem into a Markov decision process, and then using a deep reinforcement learning algorithm to solve it, the high-dimensional mixed integer non-convex optimization problem that traditional optimization methods cannot handle is effectively solved.

[0103] (3) The effectiveness of the innovative mechanism in the proposed OAN-SD3 algorithm was demonstrated through simulation, showing its superiority over other methods. Attached Figure Description

[0104] Figure 1 This is a diagram illustrating a UAV-assisted communication scenario equipped with RIS.

[0105] Figure 2 The optimized Actor network structure diagram;

[0106] Figure 3 This is a traditional Actor network structure diagram;

[0107] Figure 4 Here is a diagram of the Critic network structure;

[0108] Figure 5 A preset motion trajectory diagram of the user, UAV, and jammer within a mission cycle;

[0109] Figure 6 A graph showing the change in reward for each round during training;

[0110] Figure 7 This is a graph showing the average energy efficiency change for each round during the test.

[0111] Figure 8 This is a comparison chart of the average energy efficiency over the entire test process;

[0112] Figure 9 This is a graph showing the average communication capacity change for each round during the test.

[0113] Figure 10 This is a graph showing the change in average net energy consumption for each round during the test.

[0114] Figure 11 This is a graph showing the QoS satisfaction rate changes for each round during the test. Detailed Implementation

[0115] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings.

[0116] This invention provides an energy efficiency optimization method for a UAV-assisted communication system based on deep reinforcement learning, comprising:

[0117] (1) Model a base station with multiple antennas, which uses a UAV equipped with RIS to enhance the channel and provide services to multiple ground users. Also, consider the existence of an air jammer that will emit interference signals when modeling.

[0118] (2) The OAN-SD3 algorithm is proposed, which learns resource allocation strategies by interacting and training with the environment through an agent, and solves the problem of energy harvesting efficiency optimization of UAV-RIS assisted communication system.

[0119] The specific steps of this invention are as follows:

[0120] 1. System Modeling

[0121] 1.1 Scene Construction

[0122] In situations where natural disasters damage existing communication infrastructure, current communication resources may be insufficient to meet the communication needs of users in certain areas. Ground-based fixed base stations and ground user equipment may be unable to communicate via direct transmission links. Drones, with their high mobility, can amplify signals, thereby satisfying the communication needs of ground users. Furthermore, with the development of 6G technology, Reflective Radiation Components (RIS) possess the ability to actively control the channel environment and reshape wireless propagation channels, effectively improving the energy efficiency of communication. RIS is a planar structure containing numerous low-cost passive reflective elements, each of which can independently adjust the incident signal, thus providing new degrees of freedom and further enhancing wireless communication performance. Integrating RIS onto drones allows for the full combination of their advantages, achieving a significant performance improvement in the communication system.

[0123] like Figure 1 As shown, in this scenario, BS has multiple antennas, and the number of antennas is Z. The three-dimensional position coordinates of BS are... The role of the BS is to provide communication services to M single-antenna UEs on the ground, represented as... Due to obstacles and other reasons, the BS and UE cannot directly transmit signals. Instead, the phase modulation capability of the UAV equipped with RIS is used to optimize the signal transmission path to provide temporary communication coverage. In addition, RIS also has the ability to harvest energy, which can convert radio frequency signals into electrical energy.

[0124] The time spent on a complete task Divided into K time slots, with index as When the time slot is sufficiently small, it can be assumed in engineering that the communication channel state will not change within that time slot. To facilitate the description of the segmented constant position behavior of the UAV within the time slot, K+1 time slot boundary times are introduced. And let the three-dimensional positions of the UAV at the boundary time be respectively Therefore, in the k-th time slot Inside, the position of the UAV is considered to remain constant and equal to The position of the m-th UE in the k-th time slot is... The UE starts moving from its initial position and reaches the mission endpoint after a fixed number of time slots. The UAV also moves from its starting point to the endpoint in a complete mission; its position in the k-th time slot is... The RIS system on the UAV has a total of L units, each represented as... Since the RIS is mounted on a small UAV and each unit on the RIS has a relatively small area, it can be assumed that the position of each unit is the same as that of the UAV.

[0125] UAVs typically fly at low altitudes (10-100m), making them easy targets for malicious jamming. Furthermore, the Reflection Signal Controller (RIS) relies on the reflection and manipulation of incident signals. If the jamming signal is mistakenly reflected to the user end by the RIS, or if it directly suppresses the useful signal link between the RIS and the base station / user, communication quality will be significantly degraded. Here, the jammer is a single-antenna UAV equipped with a jamming signal transmission module, achieving interference with the user through the intelligent reflection of the RIS. The RIS-based jamming channel also constitutes a typical 3-node cascaded system. The path of the jamming signal from the jammer to the UE includes: air-to-air wireless propagation, RIS reflection, and air-to-ground wireless propagation.

[0126] The scenario studied in this invention is a downlink communication scenario where users receive video streams and download data. It is necessary to establish a composite channel model from the BS to the UE through the RIS, and a composite channel model from the jammer to the UE through the RIS. In addition, it is also necessary to establish a system energy consumption model.

[0127] 1.2 Establishing a Communication Model

[0128] The signal transmitted by the BS is enhanced by the UAV equipped with RIS before being transmitted to the UE. This communication link involves three stages: the transmission link from the BS to the UAV, the signal enhancement by the UAV equipped with RIS, and finally the transmission link from the UAV to the UE.

[0129] 1.2.1 BS-UAV Link

[0130] When establishing a channel transmission model from BS to UAV, the channel matrix can be decomposed into a large-scale fading component. and small-scale fading part Independent modeling of large-scale fading and small-scale fading:

[0131]

[0132] in, It is the channel matrix. This indicates that the matrix is ​​a A complex matrix of dimension 1, which describes the amplitude scaling and phase rotation of the signal by the channel. The matrix's 1st dimension is... element This represents the channel coefficient from the z-th antenna of the BS to the l-th unit of the UAV-RIS.

[0133] Large-scale fading, considering free-space loss and line-of-sight / non-line-of-sight (LoS / NLoS) attenuation, is modeled as follows:

[0134]

[0135]

[0136]

[0137] in, The probability of Loss propagation occurring. The probability of NLoS propagation. Additional attenuation factor for NLoS, The air-to-air path attenuation exponent is given by e, which is a natural constant, and a and b are constants that depend on the carrier frequency and environment type (such as rural or urban areas). The elevation angle between the BS and the UAV is calculated as follows:

[0138]

[0139] A complex Gaussian stochastic process is used to model small-scale fading. Considering the mobility of UAVs and the rich scattering characteristics of urban environments, a Rayleigh fading model is adopted:

[0140]

[0141] in, Represents the imaginary unit, satisfying , These are the real and imaginary parts of the channel complex gain, respectively. A real matrix of dimension 1. and Each element in and Independently and identically distributed ,set up for The element in the i-th row and j-th column:

[0142]

[0143] 1.2.2 Signal Enhancement of UAV-RIS

[0144] RIS reflection coefficient matrix For diagonal matrices:

[0145]

[0146] This indicates that the reflection coefficient matrix is A complex matrix of dimension 1, the values ​​in the matrix With phase configuration matrix and switch state matrix Relevant, specifically The l-th diagonal element can be represented as:

[0147]

[0148] in, For a superscript in a complex form, satisfying , Responsible for phase control, Responsible for the switching status of the RIS unit.

[0149] 1.2.3 UAV-UE Link

[0150] The Rician fading model is used to characterize the downlink channel characteristics from UAV-RIS to the m-th ground user. This model assumes that a strong LoS is always maintained between UAV-RIS and the user, and uses a fixed... A factor is used to characterize the LoS-dominated propagation characteristics. Specifically, within the k-th time slot... The l-th RIS unit of the UAV and located at Channel coefficients between the m-th ground users The model is as follows:

[0151]

[0152]

[0153] in, This represents the large-scale path loss coefficient. Let be independent and identically distributed small-scale fading random variables, and follow a complex circular symmetric Gaussian distribution with zero mean and unit variance. Here, K is the Rician K-factor, representing the ratio of the power of the LosS component to the power of the NLoS component. Let be the path loss constant. This is the path loss index. The distance between the UAV and the m-th user in the k-th time slot is calculated as follows:

[0154]

[0155] Combine the channel coefficients from all RIS units to the m-th user into a vector. , It indicates that it is a If we consider a complex column vector of dimension 1, then the channel model of the UAV-UE link can be compactly represented as:

[0156]

[0157] in, Let L be an L-dimensional column vector of all 1s. Let L be a small-scale fading vector, and its elements are... .

[0158] 1.2.4 Equivalent Cascaded Communication Channel Model

[0159] In this embodiment of the invention, the RIS has L adjustable phase degrees of freedom, which is much greater than that of the antenna, sufficient to achieve beamforming. Therefore, to reduce system complexity, an equal-power omnidirectional precoding strategy is adopted, and the precoding vector of the communication signal transmitted to the m-th user is... for:

[0160]

[0161] in It is a column vector of length Z consisting of all 1s, meaning all antennas transmit with equal amplitude and in phase, and it can ensure the total transmit power of BS in the k-th time slot. for:

[0162]

[0163] in As the expectation operator, BS simultaneously provides services to M users, therefore BS transmits signals... Superposition of signals from various users:

[0164]

[0165] in, For user m, the normalized data symbol. This is the transmit power allocated to user m.

[0166] End-to-end equivalent channel from BS to the m-th user Defined as a cascade of three communication stages:

[0167]

[0168] in, This indicates that the equivalent channel from the BS to the m-th user is a If the signal received by user m is a complex vector of dimension m, then it can be represented as:

[0169]

[0170] Here, the vector / matrix with the superscript H represents the conjugate transpose of the corresponding original vector / matrix. Let be the additive noise received by the m-th user in the k-th time slot, with a mean of 0 and a variance of . In the above formula Let m be the expected signal of the m-th user. Interference signals from other users Let be the precoding vector of the communication signal transmitted by BS to the i-th user. In this invention, orthogonal frequency division multiple access technology is used, and there is no interference between signals from different users. Therefore, this interference signal is not considered in the following.

[0171] 1.3 Interference Model

[0172] The jammer is in the air and almost always maintains a Loss of Space (LoS) propagation state with the UAV equipped with the Rician fading model. The air-to-air channel from the jammer to the UAV-RIS is described using the Rician fading model, and the channel vector from the jammer to the UAV-RIS is defined as follows: , This indicates that the vector is a A complex vector of dimension l, whose l-th element for:

[0173]

[0174]

[0175] in This represents the large-scale path loss coefficient. For independent and identically distributed small-scale fading random variables, Here, K is the Rician K-factor, representing the ratio of the power of the Loss component to the power of the scattered component. Let be the path loss constant. This is the path loss index. To be located in the k-th time slot The distance between the jammer and the UAV is calculated as follows:

[0176]

[0177] Combining 1.2.2 and 1.2.3, the equivalent interference channel coefficient of the signal transmitted by the jammer reaching the m-th user after being reflected by RIS is:

[0178]

[0179] Let the jammer's transmitted signal be:

[0180]

[0181] in, This refers to the jammer's transmission power. The interference waveform, after being reflected by the RIS, reaches the m-th user as follows:

[0182]

[0183] In this context, a vector / matrix with the superscript H represents the conjugate transpose of the corresponding original vector / matrix.

[0184] 1.4 Energy Harvesting Model

[0185] This invention employs wireless information and energy simultaneous transmission technology, which allows the system to transmit information and energy simultaneously via the same radio frequency signal, a key enabling technology for achieving energy self-sufficiency RIS. A time-division multiplexing strategy is used, with each time slot lasting for a specific duration. This is sufficient to complete both wireless power and information transmission. Each time slot is time-division duplex, divided into two sub-phases: the energy harvesting (EH) phase and the information transmission (IT) phase.

[0186]

[0187] in, This is an optimizable time allocation factor.

[0188] The duration of the EH phase is The radio frequency signal transmitted by the BS propagates to the RIS via a wireless channel. All L reflecting units of the RIS then switch to power receiving mode, converting the received RF power into DC power through a built-in rectifier circuit. The total received power of the RIS during the EH phase in the k-th time slot is:

[0189]

[0190] Considering conversion efficiency The energy collected during the EH phase is:

[0191]

[0192] The duration of the IT phase is RIS dynamically adjusts the phase offset of each reflector unit. and switch status This enables intelligent reflection of incident signals. A layered resource reuse strategy is adopted in the IT phase, with on / off states... The RIS unit participates in signal reflection to support communication, while Although the cells do not reflect signals, they can still receive energy, enabling parallel communication and energy harvesting. The total received power of the RIS in the IT phase of the k-th time slot is:

[0193]

[0194] Considering conversion efficiency The energy collected during the IT phase is:

[0195]

[0196] The total energy collected by RIS in the k-th time slot is:

[0197]

[0198] 1.5 Energy Consumption Model

[0199] In this embodiment of the invention, the energy consumed by the system in time slot k is considered to be two parts: the UAV carrying the RIS moves in the air, and its hovering and flight require a lot of electrical energy to drive the rotor to generate lift and thrust. Therefore, the UAV propulsion energy consumption is the main component of the system's energy efficiency; the other part is the transmission power consumed by the BS in the IT phase.

[0200] UAV at boundary moment The three-dimensional positions are respectively In the k-th time slot In order to establish the communication model, the position of the UAV is considered to remain constant and equal to... However, when calculating UAV energy consumption, the change in the UAV's position within the time slot needs to be considered. Therefore, in the k-th time slot... Within a two-dimensional plane, the flight speed of a UAV can be expressed as:

[0201]

[0202] in, This is the maximum flight speed limit for UAVs.

[0203] The propulsion energy consumption of the UAV in the kth time slot is calculated as follows:

[0204]

[0205] in, and These represent the constant airfoil power and induced power during hovering, respectively. The tip velocity of the rotor blades. The average rotor induced velocity during hovering. s and s represent the fuselage drag ratio and rotor solidity, respectively. air density, The area of ​​the rotor disk, Where is the rotor radius. and The calculation is as follows:

[0206]

[0207]

[0208] in, The blade profile drag coefficient is... The blade angular velocity, This is the incremental correction coefficient for induced power, used to correct the deviation between the theoretical model and the actual value. For the drone's gravity.

[0209] The transmit energy consumed by the BS during the information transmission phase of the k-th time slot. for:

[0210]

[0211] Taking into account the energy consumption of UAV propulsion, BS launch, and RIS harvesting, the net energy consumption of the system in the k-th time slot is:

[0212]

[0213] 1.6 Energy Efficiency Model

[0214] During the information transmission phase, the signal received by the m-th user includes the desired received signal, the interference signal generated by the jammer, and the noise signal, as shown in the following formula:

[0215]

[0216] From the above formula, the expected signal power received by the m-th user can be obtained as:

[0217]

[0218] The jamming power of the jammer is:

[0219]

[0220] The noise power is:

[0221]

[0222] Therefore, the signal-to-interference-plus-noise ratio (SIR) of the m-th user can be obtained as follows:

[0223]

[0224] According to Shannon's information theory, considering the time allocation factor... The effective communication capacity of the m-th user in the IT phase during the k-th time slot is:

[0225]

[0226] in, Based on the above analysis of this embodiment of the invention, the energy efficiency of the system in the k-th time slot is as follows: (The channel bandwidth allocated to the m-th user is given by the following formula.)

[0227]

[0228] 2. Establish the optimization problem

[0229] This invention takes maximizing the average energy efficiency of K time slots over the entire mission cycle as the optimization objective, that is...

[0230]

[0231] st

[0232]

[0233]

[0234]

[0235]

[0236]

[0237]

[0238]

[0239] The optimization variables include four types of variables: time allocation factors. This represents the proportion of time slots occupied by the energy harvesting phase; power allocation vector. , is an M-dimensional vector representing the downlink transmit power allocated to each user by the BS during the IT phase; RIS phase configuration vector is an L-dimensional vector representing the continuous phase shift of each unit in the RIS; RIS switch state vector is an L-dimensional vector representing the on or off state of each RIS unit during the IT phase. Therefore, for each time slot, the dimension of the optimization variables is 1+M+2L, and it includes both integer and continuous variables.

[0240] The constraints of the optimization problem include:

[0241] (1) C1: Time allocation constraint, which ensures that the time allocation factor is a reasonable ratio, which determines the proportion of energy collection and information transmission in each time slot.

[0242] (2) C2 and C3: power constraints. C2 ensures that the power allocation for each user has a minimum power requirement. and maximum power Due to limitations, C3 ensures that the total power of the BS's RF power amplifier does not exceed the maximum limit. .

[0243] (3) C4 and C5: RIS constraints. C4 ensures the range of phase offset, and C5 is an integer variable constraint for turning the RIS unit on and off.

[0244] (4) C6: Quality of Service (QoS) constraint, the communication capacity of each user must meet the minimum communication capacity requirement. To ensure communication quality.

[0245] (5) C7: Positive net energy consumption constraint, This is the minimum net energy consumption constraint.

[0246] 3. Solution process based on deep reinforcement learning algorithm

[0247] This optimization problem is a non-convex mixed-integer nonlinear programming problem, and the OAN-SD3 algorithm is used to solve it in this embodiment of the invention.

[0248] 3.1 Establishing a reinforcement learning model

[0249] First, the energy efficiency optimization problem constructed in the embodiments of this invention needs to be transformed into a Markov decision process, which requires the reasonable design of three core elements: state space, action space, and reward function. The state space needs to contain enough information to satisfy the Markov property, so that the optimal decision depends only on the current state; the action space needs to encode all optimization variables and handle complex constraints; and the reward function needs to reflect the constraints through a penalty mechanism while maintaining consistency with the original optimization objective.

[0250] Considering that in a UAV-RIS system, channel quality is mainly determined by geometric topology, this embodiment of the invention defines the state as: the Euclidean distance vector between nodes in the system, the end-to-end equivalent channel coefficient from the BS to the user, and the equivalent interference channel coefficient of the signal transmitted by the jammer after being reflected by the RIS to the user.

[0251] In communication scenarios, indexes are used. Represents each physical time slot, The time step, used to describe the Markov decision-making process, represents the sequence of interactions between the agent (the controlling entity) and the environment. The agent performs one action at the beginning of each communication time slot; that is, a time slot is the smallest unit of time for one decision by the agent. Therefore, the two satisfy... Specifically, for a system containing M users, the state at time step t is represented by a dimension... Real vectors:

[0252]

[0253] in, This indicates the distance from the base station (BS) to the user entrance (UAV). This indicates the distance from the jammer to the UAV. This represents the distance from the UAV to the m-th user. Indicates the location of the UAV. Indicates the position of BS. Indicates the location of the jammer. This represents the position of the m-th user. Let represent the end-to-end equivalent channel coefficients from the BS to the m-th user, where , This represents the equivalent interference channel coefficient of the jamming signal from the jammer reaching the m-th user after being reflected by the RIS, where... .

[0254] The action space needs to encode all decision variables in the original optimization problem, including energy harvesting time allocation. User power allocation RIS phase configuration and RIS switch status .

[0255] The design of the reward function needs to strike a balance among three objectives: consistency with the original optimization objective, representation of constraints, and numerical stability. Directly using the relatively complex energy efficiency as the reward may lead to convergence difficulties during training. Therefore, this invention proposes a reconstructed reward function that decomposes the original EE objective into maximizing communication capacity and energy consumption penalty, while introducing explicit penalties for constraint violations. The reward function is defined as follows:

[0256]

[0257] in, Let m be the communication capacity of the m-th user at time step t. Let be the capacity normalization constant. , and These are the penalty weights. The reward function contains three types of penalty terms, each targeting different constraints:

[0258] (1) QoS penalty The proportion of users who violate the minimum rate (communication capacity) requirement is quantified, with a value range of [0,1].

[0259]

[0260] in Let m be the communication capacity of the m-th user at time step t. This is a minimum communication capacity constraint.

[0261] (2) Power penalty This measures the relative extent to which total power exceeds the budget and only takes effect when constraints are violated.

[0262]

[0263] in Let m be the transmit power allocated to the m-th user at time step. The maximum limit for the total power of all users.

[0264] (3) Energy consumption penalty Normalizing net energy consumption to the [0,1] range encourages minimizing energy consumption while meeting capacity requirements, which is consistent with the essence of energy efficiency.

[0265]

[0266] in and The preset energy consumption reference value, The net energy consumption of the system at time step t.

[0267] The outermost layer of the reward function uses the hyperbolic tangent function to compress the output to the (-1,1) interval. The bounded reward ensures the boundedness of the value function, making the iterative solution of the Bellman equation more stable during model training.

[0268] 3.2 Implementation process of OAN-SD3 algorithm

[0269] The OAN-SD3 algorithm proposed in this invention is an optimized version of the traditional TD3 algorithm. It mainly includes two core innovations: (1) introducing a Softmax Q-value aggregation mechanism to replace the direct minimum Q-value operation, and extending the single-actor to a dual-actor network structure; (2) fusing The Actor network structure optimization mechanism of the convolutional feature recombination module. Based on the TD3 algorithm, innovation point (1) is added to obtain the SD3 algorithm, and further innovation point (2) is added to obtain the OAN-SD3 algorithm proposed in this invention. The technical implementation is described in detail below.

[0270] The empirical data generated by the interaction between the intelligent agent and the environment is ,in Let t be the environmental state at time step t. Actions performed by the intelligent agent The immediate reward obtained by the intelligent agent. Let the environment be in state at time step t+1. This indicates whether the task has ended. The number of samples in the experience pool is... The number of experience batches extracted in each training session is .

[0271] 3.2.1 Multi-sample two-stage Softmax Q-value aggregation and dual-actor network architecture

[0272] The first phase of this mechanism is multi-sample evaluation, for a given next state. and target Actor network Instead of generating a single noisy action, it generates N action samples with independent noise added, where These are the parameters of the target Actor network. Specifically, for each sample... Independently sample noise vectors from the standard normal distribution Then construct the noise action:

[0273]

[0274] in, This is the action boundary constraint function. This is the lower bound of the action vector. This represents the upper limit of the action vector.

[0275] The second stage is two-level Q-value aggregation, for each noise action. First, a dual-critic target network is used. and Two Q-value estimates were obtained by evaluating them separately. and ,in and The parameters of the two target Critic networks are given. The aggregation process is divided into two levels: first, aggregation is performed at the Critic level, and the aggregation Q-value is calculated for each sample i. Then, aggregation is performed at the sample level for N samples. Perform a weighted average.

[0276] In Critic-level aggregation, using minimum aggregation can avoid overestimation issues:

[0277]

[0278] This yielded N candidate actions and their corresponding Q-value estimates. Inspired by the successful application of the Softmax function in classification problems, the SD3 algorithm proposes to introduce the Softmax mechanism into Q-value aggregation at the sample level. This provides a "soft" aggregation method that can assign different weights based on the magnitude of the Q-value. The soft Q-value is defined as follows: The exponentially weighted average of the Q values ​​for each sample is:

[0279]

[0280] in, It is an exponential function. For temperature parameters, The sharpness of the weight distribution can be controlled: when When all weights tend to be equal, it has high smoothness and small variance; when At that time, the maximum Q value will dominate, approaching the max operation in the DDPG algorithm.

[0281] Therefore, the objective Q value when updating the Critic network is:

[0282]

[0283] To further enhance the strategy exploration capability and global search performance, the SD3 algorithm extends the dual-Critic architecture of the TD3 algorithm to the Actor side, maintaining two pairs of independently trained Actor-Critic networks: and Each network pair has its own master network and target network, therefore the complete architecture contains 8 networks: ,in Both refer to Actor networks. Both refer to Critic networks. These are network parameters.

[0284] When updating the Critic k network to calculate the target Q value, all N noisy actions are processed by their paired Actor k target network. When Actor k is generated and updated via gradients using a deterministic policy, it is updated only based on feedback from Critic k. The loss function for updating Actor k is defined as follows:

[0285]

[0286] Here, summing over t represents summing the empirical samples extracted during each training iteration. The parameter is The Actor network in a given state is Actions generated in real time.

[0287] When generating experience, a dynamic action selection mechanism based on Q-value comparison is employed. Given a state... Two Actors generate candidate actions respectively. and ,in, The parameter is The Actor network in state Actions generated in time The parameter is The Actor network in state The actions generated at that time are then evaluated by the corresponding Critic network, yielding two corresponding Q values: and Select the action with the higher Q value to execute.

[0288] 3.2.2 Actor Network Structure Optimization

[0289] like Figure 2 The diagram shows the optimized Actor network structure of this invention, compared to the traditional Actor network structure as follows: Figure 3 As shown. The Actor proposed in this invention adopts a four-layer architecture, including:

[0290] (1) Input layer and first fully connected layer

[0291] The input layer receives dimensions as follows: state vector via the first fully connected layer Perform a linear transformation on it, and then apply the ReLU activation function:

[0292]

[0293] in, It is the weight matrix of the first fully connected layer. This indicates that the weight matrix is ​​a real number matrix, and from... Dimensional input mapping to Dimensional output, This is the bias vector of the first fully connected layer. The activation function is the first fully connected layer. The output dimension is .

[0294] (2) Innovative Convolutional Feature Reconstruction Layer

[0295] In traditional fully connected structures, the linear combination of features is global. To enhance the coupling and reorganization capabilities between feature channels, this invention introduces... Convolutional layers, whose convolution kernels are The step size is also 1. The features output by the first fully connected layer... Shape Reshape it into Convolution is performed after tensor formatting:

[0296]

[0297] in, The convolution kernel is One-dimensional convolutional layer, These are the features after convolution.

[0298] Convolution operations can linearly transform the "channel dimension" without changing the dimension of the feature space, which is equivalent to performing a fully connected operation on the "channel feature vector" at each spatial location. This transformation can enhance the coupling between "originally unrelated channel features" and extract local high-order feature combination relationships.

[0299] (3) Feature fusion layer

[0300] Restore the feature dimensions after convolution to Add the residuals to the original features, and then transform the result using a linear transformation matrix. The optimal weight allocation between "original global features" and "local features after convolutional reconstruction" is learned to determine the contribution ratio of the two types of features in the final representation. Furthermore, non-linearity is introduced to allow the network to fit more complex feature relationships.

[0301]

[0302] in, This is the weight matrix of the feature fusion layer. This is the bias vector of the feature fusion layer. The output features of the feature fusion layer.

[0303] (4) Second fully connected layer and output layer

[0304] The output of the feature fusion module passes through a second fully connected layer. Further transformations compress the features, thus refining them:

[0305]

[0306] in, This indicates that the weight matrix is ​​a real number matrix, and from... Dimensional input mapping to Therefore, the output dimension of the second fully connected layer is 1. , This is the bias vector for the second fully connected layer.

[0307] The final output layer uses the tanh activation function to limit the range of actions.

[0308]

[0309] in, , To output the dimension of the action vector, As the bias vector of the output layer, the tanh activation function can restrict the output to the interval [-1, 1], naturally satisfying the requirement of normalizing the action space. Finally, it is multiplied by... Mapping normalized actions to a physically meaningful action space. This represents the maximum physical possible value for each action dimension in the action space.

[0310] In addition, such as Figure 4 As shown, the Critic network in this invention is a single-output value function estimation network based on a fully connected layer architecture. The network input contains two types of features: state features and action features, which need to be concatenated to form a unified input feature vector. The hidden layer employs a two-layer fully connected neural network structure. The first hidden layer uses ReLU activation function to perform nonlinear transformation and dimensional mapping on the concatenated input features; the second hidden layer also uses ReLU activation function to further extract high-dimensional abstract features, enhancing the network's ability to fit complex state-action spaces. The output layer is a single-node fully connected layer with no activation function, directly outputting the Q-value of the corresponding input state-action pair.

[0311] Figure 2 , Figure 3 and Figure 4In this context, state_dim is the state vector dimension, action_dim is the action vector dimension, and max_action represents the maximum physical possible value for each action dimension in the action space. Figure 2 In the formula (batch, 400, 1), `batch` represents the empirical batch size selected for each training iteration, `400` represents the number of channels, and `1` represents the spatial dimension. `Conv1d(400, 400, kernel_size=1)` represents a single kernel. A one-dimensional convolutional layer with 400 input and 400 output channels and a kernel size of 1.

[0312] 3.2.3 Dynamic Adjustment of Hyperparameters During Training

[0313] The performance of deep reinforcement learning algorithms is highly dependent on the choice of hyperparameters. However, fixed hyperparameters often fail to meet different learning needs at different stages of training: high exploratory power and low variance are needed in the early stages of training to quickly accumulate experience, while low bias and precise policies are needed in the later stages of training to converge to the optimal solution.

[0314] The training progress is defined as follows:

[0315]

[0316] Where t is the number of steps in the current time step. These are the steps taken during the warm-up period (the agent only explores the environment and collects experience, without performing network training). This represents the total number of training steps. Training progress. A normalized time scale is provided for parameterizing the scheduling scale.

[0317] Exploratory noise is used to add random perturbations to the Actor's output actions during training to facilitate environmental exploration; noise standard deviation. Employing an exponential decay strategy:

[0318]

[0319] in and These are the maximum and minimum values ​​of the noise standard deviation, respectively. It is a noise parameter.

[0320] The complete execution flow of the OAN-SD3 algorithm is shown in Table 1.

[0321] Table 1 Execution flow of the OAN-SD3 algorithm

[0322]

[0323] 4. Experimental Verification

[0324] The energy efficiency optimization method of the UAV-assisted communication system based on deep reinforcement learning is verified. First, the simulated trajectories of the user, UAV and jammer in a complete task are pre-generated and relevant environmental parameters are set. Then, the convergence of the reward function of the algorithm in the training phase is verified. Finally, the trained model is tested in the test environment to verify the effect of its energy efficiency optimization.

[0325] This invention employs a weighted random walk model to generate realistic user trajectories. This model captures both the global trend of group movement and retains the randomness of individual behavior. In designing the user trajectory model, the trend aspect ensures that users do not engage in completely random walks, but rather move in a general direction. Secondly, the trajectory needs to maintain local randomness, with each step containing a certain degree of random perturbation.

[0326] In multi-user communication scenarios, the location of the UAV directly affects the channel quality and coverage fairness between it and the users. This embodiment of the invention adopts a centroid following algorithm to minimize the "average distance" of the UAV to all users in the horizontal position.

[0327] In real-world adversarial scenarios, jammers often cannot accurately determine the real-time location of a UAV, but can only estimate its approximate area. This embodiment of the invention employs a region-constrained random walk jamming strategy. This strategy assumes that the jammer can estimate that the UAV is located within a bounded region, but cannot know its exact coordinates. The maximum search radius is set to 1m, and the jammer randomly selects and moves within this region to probabilistically jam the UAV.

[0328] In this embodiment of the invention, the map size is selected as follows: m, the drone's flight altitude is fixed at 20m, the jammer's flight altitude is fixed at 25m, and the specified number of users. And the number of RIS units is The trajectory changes of UAVs, jammers, and ground users during a complete mission are as follows: Figure 5 As shown, the total number of time slots required to complete a task. The actual duration of the time slot is set to 0.45s. In reinforcement learning training, the time step corresponds to the smallest decision granularity of the interaction between the environment and the agent. Therefore, each communication time slot k is a decision opportunity, which means that the length of one episode is 40 time steps.

[0329] The total number of training steps is set at 10,000, with the first 1,200 steps serving as a warm-up phase where randomly generated actions allow the agent to fully explore the environment. Therefore, the total number of training rounds is 2,500. Other relevant parameter settings are shown in Table 2, and training parameters are shown in Table 3.

[0330] Table 2 Environmental Parameters

[0331]

[0332] Table 3 Training Parameters

[0333]

[0334] To compare and verify the effectiveness of the algorithm proposed in this invention, it was compared with other algorithms, where (1) and (2) are the comparison algorithms, and (3) is the algorithm proposed in this invention:

[0335] (1) TD3 algorithm

[0336] (2) SD3 algorithm: Based on the TD3 algorithm, a Softmax Q-value aggregation mechanism was added and extended to a dual-Actor network structure;

[0337] (3) OAN-SD3 algorithm: The algorithm proposed in this invention adds a Softmax Q-value aggregation mechanism to the TD3 algorithm and extends it to a dual Actor network structure. In addition, the Actor network structure has also been optimized.

[0338] This invention is implemented using Python 3.6 and PyTorch 1.10.1, and trained using an Nvidia 4060 GPU. The reward changes during the training process are recorded as follows: Figure 6 As shown, the reward values ​​of each algorithm generally show an upward trend as the number of training rounds increases, indicating that they are all continuously optimizing performance through training. However, there are differences in the convergence speed, final reward level, and stability of different algorithms.

[0339] The OAN-SD3 algorithm proposed in this invention performs best among all algorithms. Its reward increases rapidly and the growth process is relatively stable. It enters the high reward range earlier and the reward value stabilizes at a high level in the later stage of training (after about 1600 rounds) with relatively small fluctuations. The TD3 algorithm has a very slow reward increase in the early stage of training and is in a low reward range until about 1300 rounds, when it achieves rapid reward growth. Compared with the TD3 algorithm, the SD3 algorithm has improved the reward increase trend and reaches a relatively stable reward value more quickly, but its reward fluctuations during training are larger.

[0340] To verify the resource allocation effectiveness of the trained model, simulated trajectories of users, drones, and jammers were regenerated for testing. To obtain the average performance of the policy in a random channel environment, 100 rounds of testing were conducted in the test environment, with each round consisting of 40 steps. Relevant metrics are as follows: Figure 7-11 As shown.

[0341] Figure 7 The average energy efficiency during each round of the test was measured. Overall, the OAN-SD3 algorithm proposed in this invention showed the best performance, with an average energy efficiency generally at [value missing]. - In the bps / J range, although there are some fluctuations due to the randomness of channel variations, the efficiency remains at a high level; the average energy efficiency of the SD3 algorithm remains basically within a certain range. - In the bps / J range, slightly higher than the TD3 algorithm, it is basically in the range of bps / J. - The average energy efficiency in the bps / J range is lower than that of the TD3 algorithm, but the average energy efficiency of the SD3 algorithm is better than that of the TD3 algorithm across all test rounds.

[0342] Figure 8 The average energy efficiency of all rounds during the test is the average value, which shows that the OAN-SD3 algorithm proposed in this invention performs best.

[0343] Figure 9 This represents the average communication capacity for each round during the test. Figure 10 The average net energy consumption for each round during the test shows that, under the goal of maximizing energy efficiency, the OAN-SD3 algorithm proposed in this invention mainly learns strategies to improve communication capacity when optimizing energy efficiency. In other words, it stems more from the "coverage effect of increased communication capacity on energy consumption" rather than simply reducing energy consumption. By sacrificing relatively reasonable energy consumption, it achieves a higher communication capacity than the comparison algorithms, ultimately realizing the optimal energy efficiency effect of "maximizing the amount of data transmitted per unit of energy consumption."

[0344] Figure 11 This represents the QoS satisfaction rate for each round during the test. The QoS satisfaction rate is defined as follows: in the current round, after each action step is executed, if all users' needs meet the minimum QoS requirement, then that step is considered to satisfy the QoS requirement. The QoS satisfaction rate is the proportion of time steps in the current round that satisfy the QoS requirement out of the total number of time steps. It can be seen that the OAN-SD3 algorithm proposed in this invention has the highest QoS satisfaction rate during the test, generally maintaining above 90%, and exhibiting good stability with only minor fluctuations. The QoS satisfaction rates of the SD3 and TD3 algorithms are relatively low, mainly fluctuating within the 80%-90% range, and the fluctuation range is relatively large, indicating that the performance of these algorithms in ensuring QoS needs improvement.

[0345] The downlink communication model of the UAV-assisted communication system with RIS proposed in this invention improves channel quality by incorporating RIS into the UAV due to the UAV's mobility and portability. It also takes into account the presence of jammers transmitting interference signals, and requires reasonable allocation of resources such as time allocation factor, RIS unit switching and phase configuration, and power to maximize energy efficiency.

[0346] The OAN-SD3 algorithm proposed in this invention combines the Actor network structure optimization strategy with the SD3 algorithm based on TD3 optimization, which can solve the problem of energy harvesting efficiency optimization in UAV-RIS assisted communication systems.

Claims

1. An energy efficiency optimization method for an unmanned aerial vehicle (UAV) assisted communication system based on deep reinforcement learning, characterized in that, Includes the following steps: S1: Establish a UAV-RIS assisted communication system model Construct a complete system model that includes downlink communication channels, interference signal transmission channels, energy harvesting, and energy consumption; The ground-based fixed base station (BS) has Z antennas, providing communication services to M ground user equipment (UEs) with single antennas. The UAV (Unmanned Aerial Vehicle) carries a reconfigurable intelligent metasurface (RIS) with L elements, possessing both phase modulation and energy harvesting capabilities. Additionally, an airborne jammer exists; the BS provides services to ground users through the RIS on the UAV, while the ground users are affected by interference signals reflected from the RIS. In a complete mission, the total time... Divided into K time slots, with index as To describe the segmented constant positional behavior of UAVs within a time slot, K+1 time slot boundary times are introduced. And let the three-dimensional positions of the UAV at the boundary time be respectively Therefore, in the k-th time slot Inside, the position of the UAV is considered to remain constant and equal to The communication channel state is considered unchanged; (1) Downlink communication channel The downlink communication channels include the BS-UAV link, the UAV-RIS signal enhancement, and the UAV-UE link; Channel matrix of BS-UAV link in the kth time slot The model is as follows: ; in, express The complex set of dimensionless numbers, i.e., the channel matrix It is A complex matrix of dimension 1 This is large-scale fading. When modeling, free-space path loss and line-of-sight / non-line-of-sight probabilities are considered. The calculation formula is: ; in, The probability of Loss of Service (LoS) transmission occurring. The probability of NLoS transmission occurring. This is an additional attenuation factor for NLoS. The air-to-ground path attenuation index. and These are the locations of the BS and UAV, respectively. For small-scale Rayleigh fading, its elements follow a complex Gaussian distribution. ; UAV-RIS signal enhancement is achieved by the RIS reflection coefficient matrix. It is indicated that it is a diagonal matrix. The characteristics of each reflective unit are determined by... Description, in which, Represents the imaginary unit. It is the continuously adjustable phase shift of the l-th unit. It is the switch state that controls whether the l-th unit is turned on or off, where 1 indicates on and 0 indicates off; For a UAV-UE link, the channel vector from UAV-RIS to the m-th user Using the Rice fading model: ; in, express The set of complex numbers of dimension 1 The Rician K factor, This represents large-scale path loss. Let the distance between the UAV in the k-th time slot and the m-th user be denoted as . This is the path loss index. The path loss constant is... for A 1-dimensional column vector of all ones. This represents the small-scale fading component, which follows a complex-valued Gaussian distribution with a mean vector that is an L-dimensional zero vector. Covariance matrix for An identity matrix of 3D; Based on the above, the equivalent concatenated communication channel for the downlink communication link from BS to UAV, then enhanced by UAV-RIS, and transmitted to UE is as follows: Considering that the BS has Z antennas, and employing an equal-power omnidirectional precoding strategy, the effective scalar channel coefficients from the BS to the m-th user are obtained. for: ; (2) Interference signal transmission channel jammer to UAV channel Also using the Rice model This indicates that the channel model is a Using a complex vector of dimension m and combining it with the downlink communication model, we obtain the equivalent interference channel coefficients from the jammer to the m-th user. for: ; (3) Energy harvesting model Each duration is The time slots are divided into an Energy Harvesting (EH) phase and an Information Transmission (IT) phase, with the EH phase lasting for [duration missing]. The IT phase lasts for [duration]. , The optimizable time allocation factor; the total energy collected by RIS in time slot k. for: ; in, and These are the RF power received during the EH and IT phases, respectively. It refers to energy conversion efficiency; (4) Energy consumption model The system's energy consumption model considers the UAV propulsion energy consumption and the BS launch energy consumption; UAV propulsion energy consumption Adopting flight speed Physical model: ; in, and These represent the constant airfoil power and induced power during hovering, respectively. The tip velocity of the rotor blades. The average rotor induced velocity during hovering. s and s represent the fuselage drag ratio and rotor solidity, respectively. air density, R is the rotor disk area, and R is the rotor radius; BS launch energy consumption The calculation is as follows: ; in, The transmit power that BS allocates to user m in the k-th time slot; Combine energy harvesting models to construct the system's net energy consumption. for: ; S2: Constructing an optimization problem with energy efficiency as the core. The optimization objective is to maximize the average energy efficiency over a single task cycle. ; in, To determine the system energy efficiency in the k-th time slot, The effective communication capacity of the m-th user in the k-th time slot, taking into account the time allocation factor, is calculated as follows: ; in, Let be the channel bandwidth for the m-th user to communicate via the downlink. The jammer's transmission power (considered constant) is given. For noise power, the superscript H of the vector / matrix in the formula represents the conjugate transpose of the corresponding original vector / matrix; Optimization variables include four categories: time allocation factors. Power allocation vector RIS phase configuration vector RIS switch state vector ,in, , Let M and L represent real vector spaces of dimensions respectively; S3: Transform the optimization problem into a Markov decision process. The optimization problem constructed in S2 is a complex nonconvex mixed-integer nonlinear programming problem, which is modeled as a Markov decision process. To distinguish the time scale of the communication system from the time scale of the Markov decision process, different symbols are used to represent the time variables: the time slot for establishing the communication model. Used to describe changes in UAV location, channel status, and other system states over time; The time step used to describe the Markov decision-making process represents the order in which the agent interacts with the environment. The agent is the main body controlling the decision-making, i.e., the Actor network in the OAN-SD3 algorithm in step S4. The agent performs an action once at the beginning of each communication time slot, so each time slot is regarded as the smallest time unit for one decision by the agent. The two satisfy the following correspondence: ; In other words, a communication slot corresponds to a time step in reinforcement learning. Therefore, all variables indexed by slot k in the communication model are represented by time step t in the Markov decision process. Furthermore, a round represents a complete task cycle, with a fixed number of time steps, K. After performing an action and receiving a reward, the turn ends and the environment is reset; state space The design primarily considers the state information required to establish a communication channel at time step t, specifically: ; in, This indicates the distance between the base station (BS) and the user aerial vehicle (UAV). Indicates the distance between the jammer and the UAV. This represents the distance between the UAV and the m-th user. This represents the end-to-end equivalent channel coefficients from the BS to the m-th user. This represents the equivalent interference channel coefficient of the signal transmitted by the jammer after being reflected by the RIS and reaching the m-th user. The state space is represented as 3D real vector space; The motion space at time step t is designed as follows: That is, all the optimization variables in the original optimization problem, where, As a time allocation factor, For power allocation vector, Configure vectors for RIS phases. This is the RIS switch state vector; Considering the complexity of the optimization objective, a reconstructed reward function is designed. This is a combination of communication capacity reward and constraint violation penalty, compressed to the (-1,1) interval using the tanh function to ensure training stability. ; in, Let t be the communication capacity of the m-th user. Let be the capacity normalization constant. , and As a penalty weight, Penalties for violating QoS Penalty for exceeding total power limit. Penalty for excessive net energy consumption; S4: Training using the OAN-SD3 algorithm The Softmax deep bideterministic policy gradient OAN-SD3 algorithm for optimizing Actor networks is an optimization algorithm based on the dual-delay deep deterministic policy gradient TD3 algorithm, used to solve the problem established in S3. First, based on the TD3 algorithm, a Softmax Q-value aggregation mechanism is introduced and the dual Actor network structure is maintained to form the SD3 algorithm. Then, the structure of the Actor network is optimized to obtain the complete OAN-SD3 algorithm. S5: Verify the online resource allocation strategy First, set the environmental and training parameters. Then, through the agent (the Actor network in the OAN-SD3 algorithm of S4), continuously interact with the environment to collect experience data. ,in The environmental state at time step t. Actions performed by the intelligent agent The immediate reward obtained by the intelligent agent. The state of the environment at time step t+1. This indicates whether the task has ended; the OAN-SD3 algorithm is used for training, and after training, the optimized policy network is deployed for verification testing, enabling it to generate optimal resource allocation decisions online based on the real-time perceived system status. By dynamically adjusting time allocation, transmit power, RIS phase, and switching status, the system's energy efficiency is ultimately maximized continuously.

2. The energy efficiency optimization method for a UAV-assisted communication system based on deep reinforcement learning according to claim 1, characterized in that, In step S2, the constraints of the optimization problem are as follows: ; Time allocation constraints ensure that the time allocation factor is a reasonable proportion; ; Power constraints ensure that there is an upper limit to the power allocation for each user; among which, The minimum power allocated to a single user, The maximum power allocated to a single user; ; Total power constraint ensures that the total power of the BS's RF power amplifier does not exceed the maximum limit; among which, This represents the maximum total power for all users within a single time slot. ; RIS constraints ensure the range of phase offset; ; RIS constraints ensure that integer variable constraints are used to enable and disable RIS units; ; Quality of Service (QoS) constraints require that each user's communication capacity must meet minimum requirements; among which, The minimum communication capacity required to satisfy QoS constraints for a single user; ; Net energy consumption positive constraint, among which, This is the minimum power limit for the system within a single time slot.

3. The energy efficiency optimization method for a UAV-assisted communication system based on deep reinforcement learning according to claim 1, characterized in that, In step S4, the OAN-SD3 algorithm includes: (1) Softmax Q-value aggregation and dual-actor network Building upon the TD3 algorithm, the SD3 algorithm maintains two pairs of independently trained Actor-Critic networks: and Each pair of networks has its own master network and target network, and the corresponding target Actor-Critic networks are respectively... and ,in Both refer to Actor networks. Both refer to Critic networks. For network parameters; For the next state N noisy action samples are generated using a target actor network. First, for each sample, a bi-objective Critic network is used for evaluation, and the minimum value among them is taken. Then use the Softmax function on N We obtain the soft Q-value objective by performing a weighted average: ; in, For temperature parameters; (2) Optimized Actor Network Structure Based on the SD3 algorithm, a method is introduced between the first and second fully connected layers of the Actor network architecture. The convolutional feature reorganization layer refactors the output features of the first fully connected layer. The channel dimension is linearly transformed; specifically, the convolution operation expression is: ; in, Indicates the kernel size as A one-dimensional convolutional layer; Then, the convolutional features With original features The residuals are summed and input into the feature fusion layer, using the weight matrix. With bias vector Perform a linear transformation and apply the ReLU activation function to obtain the fused features: ; in, This is the weight matrix of the feature fusion layer, used to perform a linear transformation on the fusion result of the original features and the convolutional features. This is the bias vector of the feature fusion layer. The output features of the feature fusion layer; In addition, random perturbations are added to the actions output by the Actor network during training. To promote exploration of the environment, it follows a mean of 0 and a variance of . The normal distribution, and according to the training progress Dynamically adjust the standard deviation of exploration noise To adapt to the needs of different training stages.