Anti-interference UAV-IOS assisted wireless communication energy harvesting method

By using a joint optimization and energy harvesting model of the UAV-IOS system, the problems of UAV energy limitation and malicious interference were solved, achieving high-efficiency and high-reliability wireless communication and extending the system's service time.

CN119743782BActive Publication Date: 2025-10-28XIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411938700.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-10-28
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

The energy limitations of drones restrict communication coverage, increase mobile frequency, reduce system energy efficiency, and malicious interference in dynamic environments severely affect communication quality.

Method used

An anti-jamming UAV-IOS assisted wireless communication method is adopted. By jointly optimizing the trajectory, transmission power and IOS component calls of the UAV-IOS system through reinforcement learning and alternating optimization algorithms, and combining the soft Actor-Critic algorithm to optimize phase shift, an energy harvesting model is constructed to improve system energy efficiency and resist interference.

Benefits of technology

Improve system energy efficiency, ensure communication quality, effectively combat malicious interference, and extend UAV-assisted IOS communication time in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119743782B_ABST
    Figure CN119743782B_ABST
Patent Text Reader

Abstract

This invention discloses an anti-interference UAV-IOS assisted wireless communication energy harvesting method, comprising the following steps: 1) modeling the total transmit power constraint of the BS; 2) modeling the BS-UAV link communication; 3) modeling the UAV-IoTD link communication, obtaining an expression for the data rate of IoTD i within a time slot in the UAV-IOS network; 4) constructing a UAV-IOS energy harvesting model to obtain an expression for the energy consumed by the UAV-IOS assisted communication system within time slot t; 5) optimizing the objective function formula, using reinforcement learning to solve the joint problem of continuous optimization, and using the AO algorithm to obtain the final phase shift. This invention belongs to the field of communication system energy control technology, and solves the problems in the prior art where the energy limitation of UAVs restricts communication coverage and the increased frequency of UAV movement leads to low system energy efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of energy control technology for communication systems, and relates to an anti-interference UAV-IOS-assisted wireless communication energy harvesting method. Background Technology

[0002] Intelligent Reflecting Surfaces (RIS) are considered a promising technology for future wireless networks. Specifically, each element of a RIS is made of electromagnetic materials and metal patches that can reflect incident signals by intelligently changing its phase shift. Therefore, the wireless propagation environment can be reshaped by collaboratively adjusting the phase shifts of a large number of reflective elements. RIS has been widely studied by industry and academia due to its ease of deployment, low power consumption, and low hardware cost, improving accessibility, security, and reliability. However, despite the significant progress made in RIS technology, most systems are statically deployed (e.g., installed on buildings), limiting its effective application in dynamic scenarios.

[0003] Combining unmanned aerial vehicles (UAVs) with Reliable Information Systems (RIS) can provide on-demand services in dynamic situations. Due to the controllability and flexibility of UAVs, they have numerous applications in blind spots of fixed communication infrastructure, such as acting as temporary base stations (BS), assisting the Internet of Things (IoT), vehicle-to-everything (V2X) networks, and enhancing hotspot coverage. It is worth noting that deploying RIS can only reflect signals to networked devices located on the same side, which limits coverage and increases the frequency of UAV movement.

[0004] Compared to traditional RIS (Radio Line of Attachment), the Simultaneously Transmitting and Reflecting RIS (STAR-RIS), or IOS (Inductively coupled Internet of Things), allows networked devices on the back to also be served, achieving full-space coverage. Furthermore, by adjusting the phase shift and power ratio of the reflected and transmitted signals, IOS can create a favorable propagation environment for both networked devices. Therefore, combining IOS with UAVs to reshape virtual line-of-sight lines in dynamic environments and obtain a better wireless network environment is of great significance. However, the limited onboard battery capacity of UAVs restricts the performance and durability of UAV-assisted IOS communication.

[0005] Energy harvesting (EH) ensures longer duration of UAV-assisted IOS communication, while the Simultaneous Wireless Power and Information Transfer (SWIPT) system harvests energy from impact radio frequency (RF), alleviating the onboard power problem of the UAV-IOS system. One of the most efficient SWIPT modes is the Harvest-Transmit-Store (HTS) model, which divides each time block into two time slots for EH and information transmission. However, resource allocation in the HTS model of the UAV-IOS system involves joint optimization of transmit power, reflector phase shift, transmit timing scheduling, and IOS scheduling under UAV trajectory design and communication quality requirements, which is difficult to achieve efficiently. Furthermore, various interferences exist in dynamic environments, especially malicious jammers, which can severely interfere with UAV trajectory design and communication between networked devices and the network, causing the entire system to consume more energy and reduce the overall service quality.

[0006] Therefore, it is of practical significance to develop an effective wireless communication system with high reliability, high energy efficiency and low latency to combat malicious jammers. Summary of the Invention

[0007] The purpose of this invention is to provide an anti-interference UAV-IOS-assisted wireless communication energy harvesting method, which solves the problems in the prior art where the energy limitation of the UAV restricts the communication coverage and the increased UAV movement frequency leads to low system energy efficiency.

[0008] The technical solution adopted in this invention is an anti-interference UAV-IOS assisted wireless communication energy harvesting method, implemented according to the following steps:

[0009] Step 1: Model the total transmit power constraint of BS, and determine the constraint conditions that the total transmit power generated by all IoTDs and BS must satisfy;

[0010] Step 2: Model the BS-UAV link communication to obtain the channel gain from the jammer to the UAV-IOS link;

[0011] Step 3: Model the UAV-IoTD link communication and obtain the expression for the data rate of IoTDi within the time slot in the UAV-IOS network;

[0012] Step 4: Construct a UAV-IOS energy harvesting model to obtain an expression for the energy consumed by the UAV-IOS-assisted communication system in time slot t;

[0013] Step 5: Optimize the objective function formula definition, use reinforcement learning methods to solve the joint problem of continuous optimization, and use the AO algorithm to obtain the final phase shift.

[0014] The beneficial effects of this invention are: 1) This invention is based on a UAV-IOS-assisted communication system under malicious interference, aiming to maximize system energy efficiency (EE) and ensure service quality while meeting the requirements of resisting interference attacks. 2) Considering the mixture of continuous and discrete high-dimensional variables, this invention first proposes a scheme based on the Alternating Optimization (AO) algorithm to obtain the optimal phase shift and thus solve the discrete variable problem; subsequently, it proposes an algorithm based on Soft Actor-Critic (SAC) to jointly optimize the trajectory, transmission power, and IOS component calls of the UAV-IOS, thereby solving the continuous variable problem. 3) Simulation results show that, compared with existing conventional methods, the method of this invention achieves better system energy efficiency performance under different experimental environments. Attached Figure Description

[0015] Figure 1 This is a system model diagram for the method of this invention;

[0016] Figure 2 The present invention is used to solve the design structure diagram;

[0017] Figure 3 This is a comparison of the reward values ​​of the three algorithms used in the method of this invention;

[0018] Figure 4 This invention compares the energy efficiency of four schemes with the same number of training episodes under different IOS component counts.

[0019] Figure 5 This invention compares the energy gain of four schemes with the same number of training episodes under different IOS component conditions.

[0020] Figure 6 This invention compares the energy efficiency of four schemes under different maximum transmission powers for the same training episodes.

[0021] Figure 7 This invention compares the energy gains of four schemes under different maximum transmission powers in the same training episodes.

[0022] Figure 8 This is a comparison chart of the energy efficiency of four schemes under different interference power conditions for the same training episodes.

[0023] Figure 9 This invention compares the energy gains of four schemes under different interference power conditions on the same training episodes.

[0024] Figure 10 This invention compares the average throughput of four schemes under the same training episodes under different two-stage durations.

[0025] Figure 11 This invention compares the energy gains of four schemes under different two-stage durations for the same training episodes. Detailed Implementation

[0026] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0027] like Figure 1 As shown, the method of this invention considers a typical UAV-IOS assisted communication system. In a typical scenario, the base station (BS) is built on one side of a building to transmit signals to Internet of Things (IoT) devices, often applicable to maritime communication systems and war zones. Since the line-of-sight (LoS) link between the BS and IoTD is blocked by obstacles (i.e., buildings and hills), a RIS mounted on the UAV is used as a wireless relay to assist communication from the BS to the IoTD. Furthermore, a malicious jammer J equipped with a single antenna exists in this environment to interfere with legitimate transmissions from the BS to the IoTD. The BS is equipped with Z antennas, where I = {1,2,…,i,…,I} represents the set of all IoTDs, and each IoTD is equipped with a single antenna.

[0028] The present invention provides an anti-interference UAV-IOS-assisted wireless communication energy harvesting method, comprising the following steps:

[0029] Preliminary steps: Set model parameters,

[0030] First, the entire UAV-IOS assisted communication system is discretized into T time slots of equal length τ, then we have Without loss of generality, the locations of IoTD, UAV-IOS, and BS are represented using a 3D Cartesian coordinate system. The BS is deployed at the origin and denoted by the symbol B, and the total number of antennas it can provide is Z = {1, 2, ..., z}. The UAV-IOS flies at low altitude and is denoted by the symbol K. That is, B and K refer to the base station and UAV-IOS, respectively. Each IoTD in the system is considered to be randomly roaming and is represented by I = {1, 2, ..., i}. Let qi [t] = (x i [t],y i [t],z i [t]) represents the position of the i-th IoTD at time t; the position of the UAV-IOS in time slot t is determined by q. k [t] = (x k [t],y k [t],z k [t]) represents the position of the jammer determined by q. J [t] = (x J [t],y J [t],z J [t]) indicates that the maximum flight speed of the UAV is v max The IOS has a uniform planar array of passively collecting or reflecting elements, where the total number of elements is represented as (M×N), and the IOS element in the nth row and mth column is... express;

[0031] Secondly, without loss of generality, the entire iOS element array is represented as Assuming all IOS components operate in T&R mode, the signal energy incident on each component is typically divided into transmitted signal energy and reflected signal energy. Energy harvesting can reduce energy consumption and increase the service time of the UAV-IOS assisted communication system. Each time slot is divided into two phases: Information Transmission (IT) and Energy Harvesting (EH).

[0032] Step 1: Model the system based on the total transmit power constraint of the BS.

[0033] The UAV-IOS assisted communication system considers passive beamforming and has two types of links in its communication modes: one is the link between the BS and the UAV-IOS, called the BS-UAV link, and the other is the link between the UAV-IOS and the IoTD, called the UAV-IoTD link; each IOS element A time slot includes an EH phase and an IT phase. During the EH and IT phases, the transmission signal expression generated by the BS is as follows:

[0034] X=∑ i∈I v i s i (1)

[0035] Among them, v i ∈C Z×1 Represents the beamforming vector of IoTDi; s iThis represents the signal transmitted from the BS to the i-th IoTD, and it is a random variable, E(x). i ) = 0 and E(|x i | 2 ) = 1, that is Therefore, the expression for the constraint condition satisfied by the total transmit power of BS is:

[0036]

[0037] Where, ‖·‖ represents the Euclidean norm of the vector. X corresponds to the upper limit of the total transmit power of the BS. H It is a self-conjugate matrix;

[0038] Equation (2) is the modeling result obtained in step 1 for the transmission power problem, that is, the constraint condition that is satisfied for the total transmission power generated by all IoTDs and BSs.

[0039] Step 2: Model the BS-UAV link communication.

[0040] Assume the channel matrix of the BS-UAV link Following the Rayleigh fading distribution, where g n,m This represents the channel vector, from the Z antenna at BS to the IOS reflector. Represented as

[0041] For the BS-UAV link, assuming the IoTD can only receive signals reflected from the UAV-IOS, the BS-UAV link is considered as a Loss of Speed ​​(LoS) transmission. Therefore, the link communication between the BS and IOS components employs a composite communication channel that simultaneously considers both large-scale and small-scale fading. The expression for the channel gain between them is:

[0042]

[0043] in, Represents iOS elements Path loss of BS-UAV links, This indicates the decline of Rice, leading to Representing BS and IOS elements Given the communication distance between them, the path loss expression is:

[0044]

[0045] in, An exponent representing path loss; A K This indicates the path loss when the distance of the BS-UAV link equals the reference distance d0, i.e. The deviation, in the fitting process, is described as having zero mean and standard deviation. Gaussian random variable, Rice fading parameter The expression is:

[0046]

[0047] in, and K K The Rice factor determines the ratio between scattering and line-of-sight transmission. Convert to The expression is:

[0048]

[0049] The transmitter's antenna gain is expressed as G. BS The receiver's antenna gain is expressed as G. r Then connect the z-th antenna of BS with the IOS element. The simplified expression for the channel gain between them is:

[0050]

[0051] Among them, the channel gain of the jammer to UAV-IOS link is used express.

[0052] Step 2 of this paper addresses the BS-UAV link communication problem, which involves modeling a communication mode.

[0053] Step 3: Model the UAV-IoTD link communication.

[0054] For the UAV-IoTD link, since the UAV-IOS is at a much higher altitude than the IoTD, the UAV-IoTD link may experience NLoS communication due to other factors. Therefore, the modeling of low-altitude flight platforms considering both LoS and NLoS factors in the UAV-IoTD link communication is used to predict the probability of geometric LoS between the UAV-IOS and IoTD. Thus, from the IOS element... The expression for the path loss to IoTDi is:

[0055]

[0056] in,

[0057]

[0058] Among them, h k ξ is the altitude of the UAV-IOS, c is the speed of light, and λ represents the carrier frequency; Losξ NLos a and b are environment-dependent Channel State Information (CSI) values; similarly, a and b are... Transform into The expression is:

[0059]

[0060] Therefore, the expression for the channel gain of the UAV-IoTD link is:

[0061]

[0062] in, Representing small-scale fading, similarly, the diagonal reflection phase matrix of IOS is expressed as:

[0063]

[0064] Similarly, the diagonal transmission phase matrix of IOS is expressed as:

[0065]

[0066] in, Represents RIS elements Phase shift;

[0067] After reflection and adjustment by the IOS element, the expression of the signal received by the UAV-IOS system at IoTDi is:

[0068]

[0069] Similarly, after transmission and adjustment by the IOS element, the signal expression received by the UAV-IOS system at IoTDi is:

[0070]

[0071] Among them, P J It is the transmission power of the jammer. It is additive white Gaussian noise (AWGN) at the i-th IoTD. This is the channel vector between the jammer and the UAV-IOS. The channel gain matrix h of the UAV-IoTD link and the IOS scheduling matrix r,i The expression is:

[0072]

[0073] in, It is an iOS element. The length of the IT phase; similarly, the scheduling matrix at the transmission point is obtained. This step mainly considers severe interference effects, assuming IoTD i It can completely cancel out signals from IOS to other IoTDs, therefore, IoTD i The expression for the signal-to-interference-plus-noise ratio (SINR) in time slot t is:

[0074]

[0075] Therefore, in the UAV-IOS network, the expression for the data rate of IoTDi within a time slot is:

[0076]

[0077] Equation (19) is the modeling result obtained in step 3 for the UAV-IoTD link communication problem.

[0078] Step 4: Construct the UAV-IOS energy harvesting model.

[0079] The Adaptive Energy Harvesting (AEH) model operates in two distinct phases within a specific time slot t: the Energy Harvesting (EH) phase and the Information Processing (IT) phase. This step considers that in the AEH model, the length of the IT phase is an element individually allocated to the operation of each individual IOS. This allocation method facilitates independent control over the duration of the EH phase, thereby enhancing system energy efficiency. By adopting this strategy, the overall energy efficiency of the UAV-IOS assisted communication system is improved, effectively addressing the energy shortage and consumption issues of UAVs.

[0080] make Represents iOS elements The length of the IT phase in the t-th time slot will be determined by the IOS elements within time slot t. The length of the EH stage is expressed as Therefore, the total energy collected by UAV-IOS in each time slot t is expressed as:

[0081]

[0082] Where μ∈(0,1) represents EH efficiency, the total energy consumption of the UAV-IOS assisted communication system in each time slot is divided into two main parts: the energy consumed during UAV-IOS flight and the energy consumed by the communication link between BS and IoTD. The energy consumed by the communication link between BS and IoTD is significantly lower than the energy consumed during flight and is essentially negligible. Therefore, it is inferred that the total energy consumption of the UAV-IOS assisted communication system consists only of the energy consumed during flight. For simplicity, E... k Treating it as a constant related to the UAV, the expression for the energy consumed by the UAV-IOS assisted communication system in time slot t is:

[0083]

[0084] Step 5: Optimize the objective function by formulating its definition (this problem is essentially an optimization problem, so an optimization function needs to be formulated).

[0085] The objective of this invention (since it is an optimization problem, an optimization objective must ultimately be formed) is to maximize the energy efficiency (EE) of the UAV-IOS assisted communication system while considering the total BS transmit power constraint, RIS component configuration constraints, and all IoTD requirements. Specifically, by jointly optimizing the UAV-IOS layout, BS transmit power, phase shift vector, and IOS component scheduling, the system EE is maximized under practical constraints. The expression for maximizing the system EE under practical constraints is:

[0086]

[0087] Therefore, the function expression after the optimization problem is transformed is:

[0088]

[0089] Where q=(x k ,y k ) indicates the location of the drone, γ max This represents the minimum signal-to-interference-plus-noise ratio (SINR) threshold. It is the transmit power vector from BS to IoTD, θ=(θ 1,1 ,…,θ M,N ) is the phase shift vector of the IOS element. This is the iOS scheduling matrix, expressed as:

[0090]

[0091] In the constraints, C1 is the maximum power control of the total transmit power from BS to IoTD; C2 is the length range of the IT phase of the IOS element; C3 is the phase shift range of the IT phase of the IOS element; C4 is the range of EH efficiency; C5 is the upper bound of the signal transmitted by the jammer in a time slot; C6 is the constraint to ensure the quality of service of each IoTD, where the received SINR should be greater than a minimum threshold; C7 limits the flight range of the UAV; and C8 limits the speed range of the UAV to prevent the UAV from exceeding the maximum service area.

[0092] Since equation (23) is a non-convex mixed integer programming problem and also an NP-hard problem, it is challenging to solve, and the multiple optimization variables are strongly coupled with each other. In addition, the characteristics of UAV, IoTD and local environment change dynamically, and the communication strategy of UAV-IOS needs to be adjusted in real time according to the CSI observed in the environment. Therefore, in order to solve the optimization problem, this invention develops a UAV-IOS assisted wireless communication energy harvesting framework based on DRL. For equation (23), the optimization problem is solved according to the following design scheme:

[0093] Product Solution: First, the joint optimization problem of UAV-IOS trajectory, transmit power, IOS scheduling, and IOS phase shift matrix is ​​transformed into a Markov decision process. The optimization problem is then decomposed into a continuous optimization joint problem and a discrete phase shift problem. Two optimization schemes are proposed to solve these two sub-problems. Specifically, reinforcement learning methods are used to solve the continuous optimization joint problem, and the AO algorithm is used to obtain the final phase shift.

[0094] 5.1) A Markov Decision Process (MDP) is used to simulate and solve the probability distribution of random changes in the problem's environmental state.

[0095] MDP models are typically defined as tuples<S,A,R,P,γ> S represents the global state space model of the environment, treating UAV-IOS as agents exploring the unknown environment; A represents the action space of each UAV-IOS; R represents the reward function; P represents the state transition function P:S×A→S′, which means that the MDP transitions to the new state S′ after performing joint actions; γ is the discount factor, 0<γ<1, representing the importance between future rewards and current rewards; in the time-slot environment, the goal of UAV-IOS is to minimize its own energy consumption. The detailed description of each component in the MDP model is as follows:

[0096] a) State space:

[0097] Within time slot t, the state space of the system consists of three parts: the CSI of the UAV-IOS network G, the interference power P of the jammer on the UAV-IOS network, and the interference space P of the jammer on the UAV-IOS network.J The estimated CSI of the jammer is called G. J Therefore, state s t Represented as: s t ={G,P J G J};

[0098] b) Action space:

[0099] Within time slot t, the UAV-IOS is in the currently observed state s t Then choose a t The action at point A, where a t It consists of four parts: the transmit power p from BS to all IoTDi i iOS scheduling variable τ i,j ∈(0,1),τ i,j ∈(0,1), i∈[1,M],j∈[1,N]; each IOS element phase shift The horizontal position of UAV-IOS is q = (x k ,y k );

[0100] c) Reward function:

[0101] The reward function is a fundamental component for evaluating the effectiveness and quality of learning decision-making strategies. It determines the expected reward the agent receives after performing an operation. Within time slot t, a positive reward is set to maximize the energy harvesting efficiency of the UAV-IOS. Before constructing the reward function, the proposed reward function also considers the minimum SINR requirement; therefore, r t =ρ×∈(t), where ρ is the symbol for the number of users that meet the minimum SINR requirement; for ρ, we have the following definition: ρ=∏ i∈I ρ i Similarly, when considering whether time slot t meets the minimum SINR requirement, the following definition applies:

[0102]

[0103] Otherwise, penalties are imposed for situations such as failing to meet the minimum speed condition, collisions, and exceeding the service area. In the proposed MDP model, the UAV-IOS actions are hybrid, involving both continuous UAV flight control and discrete phase shift control problems. This typically performs poorly in a hybrid action space, and the introduction of a transmission scheme further exacerbates the problem, making the action space quite large and difficult to converge. To improve this, the AO algorithm is introduced to handle the discrete phase shift problem, determining the phase shift control based on the data rate to ensure basic user requirements are met. Simultaneously, for UAV flight control, maximum diffusion learning is employed to optimize UAV movement. This approach improves sample efficiency, enhances the agent's exploration capabilities, provides robustness against randomness and environmental changes, and offers the ability to successfully learn in a single deployment.

[0104] 5.2) When communication between the IoTD and the BS is obstructed using the IOS-AO algorithm, RIS-assisted communication requires determining the phase shift angle of the reflector to provide better communication services.

[0105] In the t-th time slot, given the current position of the UAV and the predetermined time of IoTDi, the optimal phase shift is obtained by the following expression:

[0106] P2:Findθ(t) (26)

[0107]

[0108] To address this issue, this step employs a simple yet robust AO algorithm to handle the discrete phase shifts in IOS.

[0109] First, the initial phase shift of all IOS elements is randomly configured based on all available angle values;

[0110] Then, in one iteration, each reflecting element is optimized alternately while the others remain fixed; for a given element, its phase shift is set to the achievable rate R by examining all its possible phase shift values. i (t) Maximize the phase shift; repeat the above operation in each time slot until R is satisfied. i (t)≥γ min Or reach the maximum number of iterations I max ;

[0111] Finally, the final phase shift of all RIS elements θ(t) in time slot t is obtained.

[0112] It is worth noting that this step will use this method to simultaneously obtain the final phase shift of the transmission, which saves a lot of time compared to directly using neural networks for optimization. At the same time, it allows reinforcement learning to focus on the trajectory of the moving UAV, power allocation, and IOS scheduling, greatly improving system performance and thus achieving the maximum system energy gain.

[0113] 5.3) Solve using the SAC algorithm.

[0114] This step proposes applying the SAC algorithm to the joint optimization of UAVs. Compared with traditional RL algorithms, the SAC algorithm introduces the concepts of entropy regularization and softening policy updates during the optimization process, enabling the agent to better explore unknown states and improve learning efficiency. The reasons for adopting the SAC algorithm in this step are as follows: 1) This algorithm is suitable for tasks involving continuous action spaces, such as robot control and autonomous driving; 2) By introducing entropy adjustment terms and softening value functions, the employability and stability of the policy are improved without reducing efficiency; 3) Compared with traditional RL algorithms, the SAC algorithm typically performs better in continuous control tasks; 4) Parameters such as the entropy adjustment coefficient and softening value function can be flexibly adjusted according to task requirements to obtain optimal performance. Figure 2 The specific structure diagram of the solution design using the SAC algorithm is shown.

[0115] Low sampling efficiency and fragile convergence make reinforcement learning frameworks slow in practical applications. Therefore, the SAC framework, which uses the maximum entropy principle for effective sample training, is adopted. An entropy term is added to the objective function to enhance exploration, and the objective function is defined as follows:

[0116]

[0117] Here, α is a temperature factor representing the stochastic nature of the entropy weights and the optimal policy π, which depends on the different tasks and the magnitude of the reward during training. By using the average entropy as a bound, the entropy weights are flexibly adjusted to obtain a new objective function, expressed as:

[0118]

[0119] in, Assuming minimum entropy constraints, the optimal paired variables for each time slot are obtained using the recursive expression of the objective function and strong pairwise properties. The expression is as follows:

[0120]

[0121] in, Indicates the relationship with temperature α t The corresponding optimal strategy for problem P1 is to be solved using the even gradient descent method, where the objective is defined as:

[0122]

[0123] The optimal temperature and optimal policy are interdependent. SAC consists of two phases: policy evaluation and policy improvement. In the policy evaluation phase, the focus is on using Bellman's fair expectation to evaluate the action value of a specific policy, also known as the Q-function. Therefore, within time slot t, the Q-function expression corresponding to policy π is:

[0124]

[0125] in, Indicates action a within time slot t t and state s t The corresponding reward function value, Indicates the intrinsic action a in time slot t. t Next, the state changes from s t Convert to s′ t+1 The transition probability,

[0126] Derivation of soft state value function by combining entropy The expression is:

[0127]

[0128] Q-networks aim to approximate the values ​​of the true state. To minimize the soft Bellman residuals, the following equation is obtained:

[0129]

[0130] Where ω is a weight parameter used to update the Q-function value. The Q-function without the weight parameter ω is defined as:

[0131]

[0132] The goal of this step is to achieve a continuous action setup, and since the action space of the optimization problem is discrete, the expected value is calculated using the dispersed action probability, and equation (34) is rewritten as:

[0133]

[0134] The goal of the policy improvement phase is to improve the policy, specifically the neo-Q value. For discrete action settings, a Boltzmann policy is adopted, and the improved policy expression is as follows:

[0135]

[0136] in, This represents the optimal strategy obtained from the preceding time slot;

[0137] Based on the principle of policy refinement, the expression for the loss function of a policy network is:

[0138]

[0139] Among them, Q ω (s t ,·) applies only to the Q-function value obtained in the given state. Let D represent the sum of policy probabilities. KL To quantify the pairwise similarity of Kullback-Leibler (KL) scatter plots, considering If it depends only on the state, then the loss function should be simplified to:

[0140]

[0141] For a discrete action space, the action expectation based on a specific action probability is calculated using the following formula, and then equation (38) is rewritten as:

[0142]

[0143] Thus, the SAC algorithm not only improves sample efficiency and enhances agent exploration, but also provides robustness to randomness and environmental changes, as well as the ability to learn successfully in a single deployment.

[0144] Experimental verification

[0145] In the simulated scenario, six IoTDs (uniformly or non-uniformly distributed on both sides of the UAV-IOS), one UAV-IOS, one BS, and one jammer were set up, acting on a 200×200×100 (m) area. 3 The system is initialized with the BS at the origin and the UAV-IOS at (100, 100, 100). The initial positions of the IoTDs on both sides of the IOS are randomly generated, and all IoTDs move randomly. Furthermore, the horizontal movement range of the UAV-RIS along the x or y axis is [-30, 30] centered on its initial position. ξ is set... Los =1dB, ξ NLos=20dB, a=9.61, b=0.16. It should be noted that the horizontal movement range represents the range that the UAV-IOS can move along the x-axis and y-axis respectively. The jammer moves according to a pre-planned trajectory in the communication system. As the IoTD and jammer move, the UAV-IOS selects the optimal position at each step to enhance the EH and SINR values. Furthermore, the simulation was conducted on an Ubuntu 20.04 server with four NVIDIA RTX3090 graphics cards and 14 cores / GPU, a Xeon(R) Gold 6330, a single-precision floating-point operation capability of 35.58 TFLOPS, and a half-precision tensor operation capability of 71 TFLOPS.

[0146] Four benchmark algorithms are designed to evaluate the performance of each method and its performance in terms of energy efficiency, throughput, and energy harvesting. These four benchmark algorithms are described below:

[0147] (Benchmark 1) SAC-AO (proposed schema): refers to the method of this invention. Communication between the BS and IoTD is enhanced using UAV-IOS, and the optimal phase shift matrix is ​​obtained using the AO algorithm. The more robust and faster convergent SAC algorithm of this invention is used to achieve multi-dimensional optimization of UAV, so as to achieve efficient exploration.

[0148] Benchmark 2) SAC-RIS: Unlike Benchmark 1, Benchmark 2 utilizes RIS to provide a virtual line-of-sight link between BS and IoTD. It also uses the AO algorithm to obtain the optimal phase shift matrix. Benchmark 2 aims to compare the gap between RIS and IOS to verify the superiority of IOS.

[0149] Benchmark 3 (SAC): Benchmark 3 is mainly used to verify the practicality and superiority of the AO algorithm. Other conditions remain unchanged from Benchmark 1. The random phase shift matrix should be inferior to the AO algorithm in obtaining system benefits.

[0150] Benchmark 4) SD3-AO: SD3 also has good performance and outperforms other algorithms on real machines. However, its exploration capabilities are worth exploring in this environment. Under the same conditions, this benchmark 4 uses the SD3 algorithm to optimize UAV-IOS.

[0151] Example 1

[0152] Design the SAC algorithm, MaxDiff algorithm, and SD3 algorithm to compare their convergence performance.

[0153] Using the parameter settings verified in the aforementioned experiment, the conclusions are as follows: Figure 3 As shown. Figure 3 The rewards of three reinforcement learning algorithms in this environment are compared. Figure 3 As can be seen, since sufficient data needs to be obtained from the environment to begin training, the three algorithms initially have almost identical reward outcomes in their episodes. When the data reaches the training set, during the training process, an interesting phenomenon emerges: due to SAC's inherent good performance, all three algorithms achieve good rewards in a relatively short training time. However, SAC still demonstrates its strong capabilities in human-computer interaction with the environment. Starting from 1200 episodes, the proposed algorithm begins to stabilize and exhibits considerable stability, while MaxDiff and SD3 begin to converge after 1300 episodes, and their reward outcomes are far inferior to the scheme proposed in this invention. With an increase in the training set, the training results of the three algorithms show significant fluctuations. However, the scheme of this invention can normalize policy improvement within a reasonable range and possesses corresponding robustness and strong exploratory capabilities.

[0154] Example 2

[0155] The energy utility performance of the SAC-AO method of this invention was compared with that of three other benchmark algorithms (SAC-RIS, SAC, and SD3-AO) under different numbers of IOS components; using the parameter settings verified in the aforementioned experiments, the conclusions are as follows. Figure 4 As shown.

[0156] Figure 4 This paper compares the energy efficiency of four schemes with different numbers of IOS components on the same training episodes. The method utilizes the alternating AO optimization algorithm to first obtain the phase shift configuration matrix, and then iteratively optimizes the phase shifts while meeting the minimum rate requirement until convergence. Because the expression is closed-form, AO has lower complexity, significantly reducing the action space of reinforcement learning and accelerating training. Due to the powerful exploration capability and excellent robustness of SAC-AO, although the method initially has almost the same energy efficiency as the SAC algorithm, as the number of IOS components increases, the number of matrices increases, leading to an increased action space and a difference of approximately 10 Mbps / joule in the later stages, demonstrating the strong performance of the method in the later stages. The application of IOS reduces UAV movement, which is a significant advantage compared to RIS. The previously proposed SD3 algorithm also has good performance, continuously improving with the number of IOS components, but its results are inferior to the method of this invention.

[0157] Example 3

[0158] The energy harvesting performance of the SAC-AO method of this invention was compared with that of three other benchmark algorithms (SAC-RIS, SAC, and SD3-AO) under different numbers of IOS elements; using the parameter settings verified in the aforementioned experiments, the conclusions are as follows. Figure 5 As shown.

[0159] Figure 5 This paper compares the energy gains of four schemes under different numbers of IOS components on the same training episodes. As shown in the figure, interestingly, the energy gains of all algorithms steadily increase. However, the energy gained by the method of this invention is significantly better than the other schemes. Furthermore, despite malicious interference, the energy gained by the method of this invention shines in the later stages as the number of IOS components increases. In the later stages, the energy gain of the method of this invention is close to 20 joules, 30 joules, and 40 joules compared to the others, which is quite outstanding. Utilizing IOS reduces the movement of the UAV, which implicitly increases the system's performance difference. It is not difficult to imagine that the method of this invention will continue to have even stronger performance as the number of IOS components increases.

[0160] Example 4

[0161] The energy utility performance of the SAC-AO method of this invention was compared with that of three other benchmark algorithms (SAC-RIS, SAC, and SD3-AO) under different maximum transmit powers; using the parameter settings verified in the aforementioned experiments, the conclusions are as follows. Figure 6 As shown.

[0162] Figure 6 This paper compares the energy efficiency of four schemes under different maximum transmit powers for the same number of training episodes. With the number of IOS components set to 40 and the IT phase duration set to 0.6, the paper observes that since maximum transmit power is positively correlated with data rate and energy consumption, the overall system throughput increases with increasing maximum transmit power. Furthermore, the efficiency of RF energy harvesting improves due to the increased transmit power. Fortunately, the method of this invention utilizes IOS to provide a virtual line-of-sight link for IoTDs. The energy used by UAV-IOS is less than that of schemes utilizing RIS, resulting in a later energy efficiency that is nearly 6 Mbps / Joule higher than SAC-RIS. This significant improvement demonstrates the superiority of the method.

[0163] Example 5

[0164] The energy harvesting of the SAC-AO method of this invention was compared with that of three other benchmark algorithms (SAC-RIS, SAC, and SD3-AO) under different numbers of IOS devices; using the parameter settings verified in the aforementioned experiments, the conclusions are as follows. Figure 7 As shown.

[0165] Figure 7 This paper describes a comparison of energy harvesting for four schemes under different maximum transmit powers on the same training episodes. Figure 7 Interestingly, all algorithm schemes show an increase in energy gain, and the scheme using SAC to obtain the phase shift configuration has a similar energy gain to the method of this invention. A reasonable explanation is that the SAC algorithm itself has strong exploration capabilities and the advantages of the maximum entropy algorithm, which enable the final result to remain relatively good. At the same time, the use of IOS also brings benefits to the system performance.

[0166] Example 6

[0167] The energy utility performance of the SAC-AO method of this invention was compared with that of three other benchmark algorithms (SAC-RIS, SAC, and SD3-AO) under different interference powers. Using the parameter settings verified in the aforementioned experiments, the conclusions are as follows: Figure 8 As shown.

[0168] Figure 8 This paper compares the energy efficiency of four schemes under different interference powers on the same training episodes. As the interference power P_j increases, the energy efficiency of almost all schemes decreases, with the initial decrease being quite pronounced. Since interference power primarily affects the data rate in the system, a decrease in the data rate forces the drone to expend more energy exploring to meet the basic requirements of IoTDs, resulting in significant energy consumption. When the interference power increases from 18dBm to 38dBm, the overall system energy efficiency gradually declines, but the method described in this invention still maintains the best performance advantage and stabilizes in the later stages.

[0169] Example 7

[0170] The energy harvesting performance of the SAC-AO method of this invention was compared with that of three other benchmark algorithms (SAC-RIS, SAC, and SD3-AO) under different interference powers. Using the parameter settings verified in the aforementioned experiments, the conclusions are as follows: Figure 9 As shown.

[0171] Figure 9 The paper compares the energy gains of four schemes under different interference powers on the same training episodes. Interestingly, as shown in the figure, the energy gain capability of all schemes does not show a significant decreasing trend with increasing interference power. This is because the reward function setting in the entire UAV-IOS trajectory planning causes the UAV-IOS to focus on the data rate requirements of IoTDs, thus the energy gain capability of the UAV-IOS is not significantly affected even with malicious interference.

[0172] Example 8

[0173] The energy utility performance of the SAC-AO method of this invention was compared with that of three other benchmark algorithms (SAC-RIS, SAC, and SD3-AO) under different information transmission ratios. Using the parameter settings verified in the aforementioned experiments, the conclusions are as follows: Figure 10 As shown.

[0174] Figure 10 This describes a comparison of the average throughput of four schemes for the same training episodes under different two-stage durations. Figure 10 As shown, over a period of time, as the IT phase ratio increases, all solutions show stable growth in terms of average throughput. The method of this invention still performs well, but the difference from other solutions is not significant. This is mainly because the number of IoTDs is insufficient, and the services provided by a single UAV-IOS to a user are always limited.

[0175] Example 9

[0176] The energy harvesting performance of the SAC-AO method of this invention was compared with that of three other benchmark algorithms (SAC-RIS, SAC, and SD3-AO) under different information transfer ratios. Using the parameter settings verified in the aforementioned experiments, the conclusions are as follows: Figure 11 As shown.

[0177] Figure 11 This paper compares the energy gains of four schemes under different two-stage durations for the same training episodes. As the proportion of the IT stage increases, the EE decreases significantly. This is because the total received SINR per IoTD increases over a period of time when the proportion of the IT stage increases. However, the decrease in the EH stage length simultaneously leads to a decrease in the sum of EH over a period of time. The allocation of the two-stage duration within a time period is a critical IOS resource and has a significant impact on the trade-off between IT and EH. Therefore, the goal of the method in this invention is to maximize the system's EE, but due to communication limitations and interference, maximizing EE does not equate to a trade-off between IT and EH.

Claims

1. An anti-interference UAV-IOS-assisted wireless communication energy harvesting method with intelligent metasurface capable of simultaneous transmission and reflection, characterized in that, Follow these steps: Step 1: Model the total transmit power constraint of the base station (BS) to determine the constraint conditions that the total transmit power generated by all IoT devices (IoTD) and the BS must satisfy. Step 2: Model the BS-UAV link communication to obtain the channel gain from the jammer to the UAV-IOS link; Step 3: Model the UAV-IoTD link communication and obtain the expression for the data rate of the i-th IoT device IoTDi in the time slot in the UAV-IOS network; Step 4: Construct a UAV-IOS adaptive energy harvesting model to obtain an expression for the energy consumed by the UAV-IOS assisted communication system in time slot t; Step 5: Optimize the objective function formula, solve the joint problem of continuous optimization using reinforcement learning methods, and use the alternating optimization algorithm (AO algorithm) to obtain the final phase shift. The specific process is as follows: By jointly optimizing the UAV-IOS layout, BS transmit power, phase shift vector, and IOS element scheduling, the system energy efficiency (EE) is maximized under practical constraints. The joint optimization problem of UAV-IOS trajectory, transmit power, IOS scheduling, and IOS phase shift matrix is ​​then transformed into a Markov decision process. This joint optimization problem is further decomposed into a continuous optimization joint problem and a discrete phase shift problem. Two optimization schemes are proposed to solve these two sub-problems. Specifically, reinforcement learning methods are used to solve the continuous optimization joint problem, and the AO algorithm is used to obtain the final phase shift. The process includes the following: 5.1) The Markov Decision Process (MDP) model is used to simulate and solve the probability distribution of random changes in the problem environment state; 5.2) When the communication between IoTD and BS is blocked, the IOS-AO algorithm is used to solve the problem. When the communication is assisted by the reconfigurable smart surface RIS, the phase shift angle of the reflective element on it needs to be determined to obtain the final phase shift of all RIS elements θ(t) in time slot t. 5.3) Solve using the soft actor-critic algorithm (SAC algorithm).

2. The anti-interference UAV-assisted wireless communication energy harvesting method with simultaneous transmission and reflection capabilities using intelligent metasurfaces, as described in claim 1, is characterized in that... It also includes preparatory steps, the specific process of which is as follows: Set model parameters, First, the entire UAV-IOS assisted communication system is discretized into T time slots of equal length τ, then we have Without loss of generality, the locations of the IoTD, UAV-IOS, and BS are represented using a 3D Cartesian coordinate system. The BS is deployed at the origin and denoted by the symbol B, and the total number of antennas it can provide is... UAV-IOS flying at low altitude is represented by the symbol K, where B and K refer to the base station and UAV-IOS, respectively; each IoTD in the system is considered to be a random roaming, represented as I = {1,2,…,i}, let q i [t] = (x i [t],y i [t],z i [t]) represents the position of the i-th IoTD at time t; the position of the UAV-IOS in time slot t is determined by q. k [t] = (x k [t],y k [t],z k [t]) represents the position of the jammer determined by q. J [t] = (x J [t],y J [t],z J [t]) indicates that the maximum flight speed of the UAV is v max The IOS has a uniform planar array of passively collecting or reflecting elements, where the total number of elements is represented as (M×N), and the IOS element in the nth row and mth column is... express; Secondly, without loss of generality, the entire iOS element array is represented as All components of IOS operate in T&R mode, where the signal energy incident on each component is divided into the energy of the transmitted signal and the energy of the reflected signal; energy harvesting can reduce energy consumption and increase the service time of the UAV-IOS auxiliary communication system by dividing each time slot into two phases: information transmission and energy harvesting.

3. The anti-interference UAV-assisted wireless communication energy harvesting method with simultaneous transmission and reflection capabilities using intelligent metasurfaces, as described in claim 2, is characterized in that... In step 1, the specific process is as follows: The UAV-IOS assisted communication system takes into account passive reflection beamforming. Its communication mode has two types of links: one is the link between BS and UAV-IOS, called the BS-UAV link, and the other is the link between UAV-IOS and IoTD, called the UAV-IoTD link. Each iOS element A time slot includes an energy harvesting (EH) phase and an information transmission (IT) phase. During the EH and IT phases, the transmission signal expression generated by the BS is as follows: X=∑ i∈I v i s i (1) Among them, v i ∈C Z×1 Represents the beamforming vector of IoTDi; s i The signal transmitted from the BS to the i-th IoTD is a random variable, E(x). i ) = 0 and E(|x i | 2 ) = 1, that is Therefore, the expression for the constraint condition satisfied by the total transmit power of BS is: Where, ‖·‖ represents the Euclidean norm of the vector. X corresponds to the upper limit of the total transmit power of the BS. H It is a self-conjugate matrix; Equation (2) is the modeling result obtained in step 1 for the transmission power problem, that is, the constraint condition that is satisfied for the total transmission power generated by all IoTDs and BSs.

4. The anti-interference UAV-assisted wireless communication energy harvesting method with simultaneous transmission and reflection capabilities using intelligent metasurfaces, as described in claim 3, is characterized in that... Step 2, the specific process is as follows: Channel matrix of BS-UAV link Following a Rayleigh fading distribution, where g n,m This represents the channel vector, from the Z antenna at BS to the IOS reflector. Represented as For the BS-UAV link, the IoTD can only receive signals reflected from the UAV-IOS; therefore, the BS-UAV link is considered as a Loss of Speed ​​(LoS) transmission. Consequently, the link communication between the BS and IOS components employs a composite communication channel that simultaneously considers both large-scale and small-scale fading. The expression for the channel gain between them is: in, Represents iOS elements Path loss of BS-UAV links, This indicates the decline of Rice, leading to Representing BS and IOS elements Given the communication distance between them, the path loss expression is: in, An exponent representing path loss; A K This indicates the path loss when the distance of the BS-UAV link equals the reference distance d0, i.e. The deviation, in the fitting process, is described as having zero mean and standard deviation. Gaussian random variable, Rice fading parameter The expression is: in, and K K The Rice factor determines the ratio between scattering and line-of-sight transmission. Convert to The expression is: The transmitter's antenna gain is expressed as G. BS The receiver's antenna gain is expressed as G. r Then connect the z-th antenna of BS with the IOS element. The simplified expression for the channel gain between them is: Among them, the channel gain of the jammer to UAV-IOS link is used express.

5. The anti-interference UAV-assisted wireless communication energy harvesting method with simultaneous transmission and reflection capabilities using intelligent metasurfaces, as described in claim 4, is characterized in that... Step 3, the specific process is as follows: For the UAV-IoTD link, the low-altitude flight platform is considered in the UAV-IoTD link communication model, taking into account the LoS and NLoS factors, to predict the probability of geometric LoS between UAV-IOS and IoTD. Therefore, from the IOS element... The expression for the path loss to IoTDi is: in, Among them, h k ξ is the altitude of the UAV-IOS, c is the speed of light, and λ represents the carrier frequency; Los ξ NLos a and b are channel state information values ​​that depend on the environment; Transform into The expression is: Therefore, the expression for the channel gain of the UAV-IoTD link is: in, To represent small-scale fading, the diagonal reflection phase matrix of IOS is expressed as: The diagonal transmission phase matrix of IOS is represented as: in, Represents RIS elements Phase shift; After reflection and adjustment by the IOS element, the expression of the signal received by the UAV-IOS system at IoTDi is: After transmission and adjustment by the IOS element, the signal expression received by the UAV-IOS system at IoTDi is: Among them, P J It is the transmission power of the jammer. It is additive white Gaussian noise (AWGN) at the i-th IoTD. This is the channel vector between the jammer and the UAV-IOS. The channel gain matrix h of the UAV-IoTD link and the IOS scheduling matrix r,i The expression is: in, It is an iOS element. The length of the IT phase; similarly, the scheduling matrix at the transmission point is obtained. This step considers severe interference effects. IoTDi can completely cancel the signal from IOS to other IoTDs. Therefore, the expression for the signal interference plus noise ratio of IoTDi in time slot t is: Therefore, in the UAV-IOS network, the expression for the data rate of IoTDi within a time slot is: Equation (19) is the modeling result obtained in step 3 for the UAV-IoTD link communication problem.

6. The anti-interference UAV-assisted wireless communication energy harvesting method with simultaneous transmission and reflection capabilities using intelligent metasurfaces, as described in claim 5, is characterized in that... Step 4, the specific process is as follows: The Adaptive Energy Harvest (AEH) model, or simply the IOS model, involves two distinct phases within a specific time slot t: the Energy Harvest (EH) phase and the Information Transmission (IT) phase. This step considers that, in the AEH model, the length of the IT phase is an element individually deployed to the operation of each individual IOS. make Represents iOS elements The length of the IT phase in the t-th time slot will be determined by the IOS elements within time slot t. The length of the EH stage is expressed as Therefore, the total energy collected by UAV-IOS in each time slot t is expressed as: Where μ∈(0,1) represents EH efficiency, the total energy consumption of the UAV-IOS assisted communication system in each time slot is divided into two parts: the energy consumed during UAV-IOS flight and the energy consumed by the communication link between BS and IoTD; for the energy consumed by the communication link between BS and IoTD, it is inferred that the total energy consumption of the UAV-IOS assisted communication system consists only of the energy consumption during flight. For simplicity, E k Treating it as a constant related to the UAV, the expression for the energy consumed by the UAV-IOS assisted communication system in time slot t is:

7. The anti-interference UAV-assisted wireless communication energy harvesting method with simultaneous transmission and reflection capabilities using intelligent metasurfaces, as described in claim 6, is characterized in that... Step 5, the specific process is as follows: The expression for EE is: Therefore, the function expression after the optimization problem is transformed is: Where q=(x k ,y k ) indicates the location of the drone, γ max This represents the minimum signal-to-interference-plus-noise ratio (SINR) threshold. It is the transmit power vector from BS to IoTD, θ=(θ 1,1 ,…,θ M,N ) is the phase shift vector of the IOS element. This is the iOS scheduling matrix, expressed as: In the constraints, C1 is the maximum power control of the total transmit power from BS to IoTD; C2 is the length range of the IT phase of the IOS element; C3 is the phase shift range of the IT phase of the IOS element; C4 is the range of EH efficiency; C5 is the upper bound of the signal transmitted by the jammer in a time slot; C6 is the constraint to ensure the quality of service of each IoTD, where the received SINR should be greater than a minimum threshold; C7 limits the flight range of the UAV; and C8 limits the speed range of the UAV to prevent the UAV from exceeding the maximum service area. Equation (23) is a non-convex mixed integer programming problem, and it is also an NP-hard problem. Therefore, in order to solve the optimization problem, a UAV-IOS-assisted wireless communication energy harvesting framework based on deep reinforcement learning (DRL) is adopted. For equation (23), the optimization problem is solved according to the following design scheme: 5.1) The Markov Decision Process (MDP) model is adopted. The MDP model is defined as a tuple.<S,A,R,P,γ> S represents the global state space model of the environment, treating UAV-IOS as agents exploring the unknown environment; A represents the action space of each UAV-IOS; R represents the reward function; P represents the state transition function P:S×A→S′, which means that the MDP model transitions to a new state S′ after performing joint actions; γ is a discount factor, 0<γ<1, representing the importance between future rewards and current rewards; in a time-slot environment, the goal of UAV-IOS is to minimize its own energy consumption. A detailed description of each component in the MDP model is as follows: a) State space: Within time slot t, the state space of the system consists of three parts: the CSI of the UAV-IOS network G, the interference power P of the jammer on the UAV-IOS network, and the interference space P of the jammer on the UAV-IOS network. J The estimated CSI of the jammer is called G. J Therefore, state s t Represented as: s t ={G,P J G J }; b) Action space: Within time slot t, the UAV-IOS is in the currently observed state s t Then choose a t The action at point A, where a t It consists of four parts: the transmit power p from BS to all IoTDi i iOS scheduling variable τ i,j ∈(0,1),τ i,j ∈(0,1), i∈[1,M],j∈[1,N]; each IOS element phase shift The horizontal position of UAV-IOS is q = (x k ,y k ); c) Reward function: The reward function is a fundamental component for evaluating the effectiveness and quality of learning decision-making strategies. It determines the expected reward the agent receives after performing an operation. Within time slot t, a positive reward is set to maximize the energy harvesting efficiency of the UAV-IOS. Before constructing the reward function, the proposed reward function also considers the minimum SINR requirement; therefore, r t =ρ×∈(t), where ρ is the symbol for the number of users that meet the minimum SINR requirement; for ρ, we have the following definition: ρ=∏ i∈I ρ i (t) is defined as follows, depending on whether time slot t meets the minimum SINR requirement: Otherwise, penalties will be imposed for failing to meet the minimum speed requirement, causing a collision, or exceeding the service area. 5.2) When using the IOS-AO algorithm to address communication obstruction between the IoTD and BS, the phase shift angle of the reflective element on the reconfigurable smart surface RIS needs to be determined to assist communication. In the t-th time slot, given the current position of the UAV and the predetermined time of IoTDi, the optimal phase shift is obtained by the following expression: This step uses a simple yet robust AO algorithm to handle the discrete phase shift of IOS. First, the initial phase shift of all IOS elements is randomly configured based on all available angle values; Then, in one iteration, each reflecting element is optimized alternately while the others remain fixed; for an element, its phase shift is set to the achievable rate R by examining all its possible phase shift values. i (t) Maximize the phase shift; repeat the above operation in each time slot until R is satisfied. i (t)≥γ min Or reach the maximum number of iterations I max ; Finally, the final phase shift of all RIS elements θ(t) in time slot t is obtained; 5.3) Solve using the soft actor-critic algorithm (SAC algorithm). The SAC framework, based on the maximum entropy principle for effective sample training, is adopted. An entropy term is added to the objective equation to enhance exploration, and the objective function is defined as follows: Here, α is a temperature factor representing the stochastic nature of the entropy weights and the optimal policy π, which depends on the different tasks and the magnitude of the reward during training. By using the average entropy as a bound, the entropy weights are flexibly adjusted to obtain a new objective function, expressed as: in, Assuming minimum entropy constraints, the optimal paired variables for each time slot are obtained using the recursive expression of the objective function and strong pairwise properties. The expression is as follows: in, Indicates the relationship with temperature α t The corresponding optimal strategy for problem P1 is to be solved using the even gradient descent method, where the objective is defined as: The optimal temperature and the optimal policy are interdependent. SAC consists of two phases: policy evaluation and policy improvement. In the policy evaluation phase, the focus is on using Bellman's fair expectation to evaluate the action value of a specific policy, also known as the Q function. Therefore, within time slot t, the Q-function expression corresponding to the obtained strategy π is: in, Indicates action a within time slot t t and state s t The corresponding reward function value, Indicates the intrinsic action a in time slot t. t Next, the state changes from s t Convert to s′ t+1 The transition probability, Derivation of soft state value function by combining entropy The expression is: Q-networks aim to approximate the values ​​of the true state. To minimize the soft Bellman residuals, the following equation is obtained: Where ω is a weight parameter used to update the Q-function value. The Q-function without the weight parameter ω is defined as: The goal of this step is to achieve a continuous action setup, and since the action space of the optimization problem is discrete, the expected value is calculated using the dispersed action probability, and equation (34) is rewritten as: The goal of the policy improvement phase is to improve the policy, specifically the neo-Q value. For discrete action settings, a Boltzmann policy is adopted, and the improved policy expression is as follows: in, This represents the optimal strategy obtained from the preceding time slot; Based on the principle of policy refinement, the expression for the loss function of a policy network is: Among them, Q ω (s t ,·) applies only to the Q-function value obtained in the given state. Let D represent the sum of policy probabilities. KL To quantify the pairwise similarity of KL scatter plots, considering If it depends only on the state, then the loss function should be simplified to: For a discrete action space, the action expectation based on a specific action probability is calculated using the following formula, and then equation (38) is rewritten as: Thus, the SAC algorithm provides robustness to randomness and environmental changes, as well as the ability to learn successfully in a single deployment.

Citation Information

Patent Citations

  • Real scene three-dimensional modeling method based on power grid GIS platform

    CN116977586A

  • Unmanned aerial vehicle emergency communication method, system and equipment assisted by intelligent reflecting surface, and medium

    CN119052830A