A control method and related device for RIS-assisted UAV communication
By optimizing the UAV's trajectory and RIS phase shift angle using a Markov decision process model and the TD3-PER-AO algorithm, the problems of signal attenuation and insufficient energy in complex environments are solved, achieving efficient communication and safe flight.
Patent Information
- Application Number
- CN202411125152.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-08-16
AI Technical Summary
In densely populated and complex areas with tall buildings, signal propagation between drones and users is attenuated due to building obstruction, and drones have insufficient power supply. Their flight paths pose safety risks, and the random movement of target users causes dynamic changes in the channel, affecting communication quality.
A Markov decision process model is used to jointly optimize the UAV's trajectory, transmit power, and RIS phase shift angle. The UAV communication is optimized through the TD3 and AO algorithms. The improved TD3-PER-AO algorithm is used to adjust the probability of sampling empirical data and optimize the RIS phase shift angle to improve the communication rate.
Under time and energy constraints, the goal is to achieve a reasonable arrangement of UAV flight paths, maximize communication throughput and energy utilization efficiency, and ensure that UAVs fly safely and maintain high-quality communication in complex environments.
Smart Images

Figure CN119095057B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a control method and related apparatus for RIS-assisted unmanned aerial vehicle (UAV) communication, belonging to the field of UAV-assisted wireless communication. Background Technology
[0002] In drone-assisted wireless communication, drones mainly serve as mobile aerial communication platforms, providing wireless access to ground users from high altitudes. However, in dense and complex areas with tall buildings, the signal propagation between drones and users is still attenuated due to the obstruction of buildings, which poses a certain challenge to drone deployment.
[0003] Reconfigurable smart metasurfaces (RIS) technology has proven to be one of the most promising technologies for improving data transmission efficiency. As an important auxiliary means for 5G and 6G wireless communication networks, signal propagation can be optimized by adjusting the phase shift angle and other parameters of each reflective element in the RIS. RIS are typically fixedly deployed on building walls to enhance signal transmission quality and supplement wireless communication networks. In recent years, the integration of RIS with drone technology has attracted increasing attention, raising numerous challenging issues such as drone trajectory design, resource allocation, and system capacity, and has been widely applied in various scenarios.
[0004] In urban environments, RIS-assisted drones are crucial for information dissemination and data collection, but several problems and challenges remain. First, drones primarily rely on their onboard batteries for power, and advancements in battery technology still cannot meet the demands of prolonged, high-intensity communication. Therefore, improving the energy efficiency of drones is a pressing issue in wireless communication. Second, certain buildings or environmental areas prohibit drones from crossing for security and privacy reasons, posing safety risks during flight and necessitating careful planning of flight paths. Finally, the random movement of target users within the environment can cause dynamic changes in the communication channel, impacting communication quality. Summary of the Invention
[0005] This invention provides a control method and related apparatus for RIS-assisted unmanned aerial vehicle (UAV) communication, which solves the problems disclosed in the background art.
[0006] According to one aspect of this disclosure, a control method for RIS-assisted unmanned aerial vehicle (UAV) communication is provided, comprising:
[0007] The elements of the Markov decision process model are updated using the UAV's 3D coordinates, the user's 3D coordinates, and the RIS phase shift matrix at the current time step. The Markov decision process model is transformed into an optimization problem, which is a problem of jointly optimizing the UAV's trajectory, UAV's transmit power, and RIS phase shift deflection, with the goal of improving the UAV's communication capabilities and energy efficiency in data acquisition.
[0008] The TD3 algorithm is used to solve the updated Markov decision process model to obtain the UAV movement control command and transmit power control command, as well as the phase shift angle control command of each component of the RIS at the current time step. During training, the TD3 algorithm adjusts the probability of sampling experience data by assigning priority weights to the data in the experience pool. When solving the problem using the TD3 algorithm, the AO algorithm is used to find the optimal phase shift angle at the current time step. The optimal phase shift angle is the phase shift angle that maintains the maximum communication rate between the UAV and the user.
[0009] In some embodiments of this disclosure, the Markov decision process model is pre-built, and the building process includes:
[0010] Based on the environmental parameters of RIS-assisted UAV communication, a communication model and an energy consumption model are constructed.
[0011] Based on the communication model and energy consumption model, an optimization problem is constructed and then transformed into a Markov decision process model.
[0012] In some embodiments of this disclosure, the communication model is as follows:
[0013] ;
[0014] In the formula, r u,k,m Let B be the communication rate between the drone and the m-th user in the k-th user cluster, and let B be the total bandwidth. Let P be the average bandwidth from the drone to each user, P be the drone's transmit power, and σ be additive white Gaussian noise. Let p be the average channel gain between the UAV and the m-th user in the k-th user cluster. u,k,m Let g be the blocking probability between the drone and the m-th user in the k-th user cluster. u,k,m Let g be the channel gain when the m-th user in the k-th user cluster communicates directly with the drone. u,r,k,m = g r,k,m Θg r,u Let g be the channel gain when the m-th user in the k-th user cluster communicates indirectly with the UAV. r,k,m Let g be the channel gain from the m-th user in the k-th user cluster to the RIS. r,u Let Θ be the channel gain from the UAV to the RIS, and Θ be the RIS phase shift matrix;
[0015] The energy consumption model is as follows:
[0016] E u =E mov +E t ;
[0017] In the formula, E u E represents the total energy consumption of the drone. t For the launch energy consumption of the drone, E mov For the mobile energy consumption of drones, P0 and P1 are the blade power and derived power of the UAV in hovering state, respectively. tip denoted as the tip velocity of the UAV's rotor blades, v0 as the average rotor-induced velocity of the UAV in hovering state, v as the UAV's moving speed, and d0, ρ, s0, and A as parameters related to the UAV's fuselage drag ratio, air density, rotor solidity, and rotor disk area, respectively.
[0018] In some embodiments of this disclosure, the optimization problem is:
[0019] ;
[0020] st C1:0≤x u (t) ≤L;
[0021] C2:0≤y u (t) ≤L;
[0022] C3:L u (0)={(x u ,y u )|(x u ,y u )∈A ini};
[0023] C4:L u (N-1)={(x u ,y u )|(x u ,y u )∈A des};
[0024] C5:|| L u (t-1)- L u (t)||2≤V max δ t ;
[0025] C6: 0≤P(t) ≤P max ;
[0026] C7: ;
[0027] C8: ;
[0028] In the formula, L u Here, K represents the three-dimensional coordinates of the drone, M represents the number of users in the user cluster, and x represents the number of users in the user cluster. u (t) represents the three-dimensional coordinates of the UAV at time step t. u The x-axis coordinate and y-axis coordinate in (t) u (t) represents the three-dimensional coordinates of the UAV at time step t. u In (t), the Y-axis coordinate is L, which is the maximum boundary value of the UAV's flight area, and A is... ini A des These represent the start and end point ranges of the drone, L. u (0) and L u (N-1) represent the three-dimensional coordinates of the UAV at the initial and final time steps, respectively, and L u (t-1) represent the three-dimensional coordinates of the UAV at time step t-1, V max δ is the maximum speed of the drone's flight. t Let P(t) be the duration of time step t, and P(t) be the transmit power of the UAV at time step t. max E is the upper limit of the drone's transmit power. u (t) represents the total energy consumption of the UAV at time step t. For the first The phase shift angle of each element, where n is the number of quantization bits.
[0029] In some embodiments of this disclosure, in the Markov decision process model:
[0030] The state space is represented as S ={s t =( S 1, S 2 )}, where s t The state at time step t. S 1 represents the two-dimensional coordinates of the UAV at time step t. S 2 This is the set of distances between the drone and the user, and between the drone and the destination at time step t;
[0031] Action space is represented as A ={a t =(θ(t),l(t),P(t)}, where, a t Let θ(t) be the action at time step t, θ(t) be the UAV's flight direction at time step t, l(t) be the UAV's flight distance at time step t, and P(t) be the UAV's transmit power at time step t.
[0032] The reward function is expressed as:
[0033] ;
[0034] In the formula, R(t) is the reward value at time step t, and δ c1 ~δ c4 r is a constant u,k,m Let E be the communication rate between the drone and the m-th user in the k-th user cluster. u R represents the total energy consumption of the drone. des These are the guidance parameters for the drone to reach its destination. d des (t) represents the horizontal distance between the UAV and the destination center at time step t, and d des (t-1) represents the horizontal distance between the UAV and the destination center at time step t-1, R ser R provides an additional reward for drones completing user data collection. pun This is a penalty for abnormal situations.
[0035] In some embodiments of this disclosure, the priority of data allocation in the experience pool is the absolute value of the TD error δ plus ɛ, and the probability of the j-th sample data being sampled is... ;in, p represents the total number of sample data points in each sampling. i and p j The priorities of the i-th and j-th sample data are respectively. The priority level is indicated by ɛ, which is a relatively small constant.
[0036] According to another aspect of this disclosure, a control device for RIS-assisted unmanned aerial vehicle (UAV) communication is provided, comprising:
[0037] The update module uses the UAV's 3D coordinates, the user's 3D coordinates, and the RIS phase shift matrix at the current time step to update the elements of the Markov decision process model. The Markov decision process model is transformed from an optimization problem, which is a problem of jointly optimizing the UAV's trajectory, UAV's transmit power, and RIS phase shift angle deflection, with the goal of improving the UAV's communication capabilities and energy efficiency in data acquisition.
[0038] The solution module uses the TD3 algorithm to solve the updated Markov decision process model, obtaining the UAV movement control command and transmit power control command for the current time step, as well as the phase shift angle control command for each component of the RIS. During training, the TD3 algorithm adjusts the probability of sampling experience data by assigning priority weights to the data in the experience pool. When solving using the TD3 algorithm, the AO algorithm is used to find the optimal phase shift angle for the current time step. The optimal phase shift angle is the phase shift angle that maintains the maximum communication rate between the UAV and the user.
[0039] In some embodiments of this disclosure, the Markov decision process model in the update module is pre-built, and the building process includes:
[0040] Based on the environmental parameters of RIS-assisted UAV communication, a communication model and an energy consumption model are constructed.
[0041] Based on the communication model and energy consumption model, an optimization problem is constructed and then transformed into a Markov decision process model.
[0042] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to perform a control method for RIS-assisted unmanned aerial vehicle communication.
[0043] According to another aspect of this disclosure, a computer device is provided, including one or more processors and one or more memories, wherein one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for performing a control method for RIS-assisted unmanned aerial vehicle communication.
[0044] The beneficial effects achieved by this invention are as follows: 1. This invention transforms the joint optimization problem of UAV movement trajectory, UAV transmission power, and RIS phase shift deflection into a Markov decision process model. At each time step, the elements of the Markov decision process model are updated, and an improved TD3 algorithm is used to solve the updated model. This yields the UAV movement control commands and transmission power control commands, as well as the phase shift angle control commands for each RIS element at each time step. This ensures that, under strict time and energy constraints, the UAV flight path is rationally arranged, maximizing communication throughput (i.e., providing communication quality) and energy utilization efficiency. 2. This invention uses an improved TD3 algorithm to solve the model. It introduces PER (Preferred Experience Replay) technology on top of the traditional TD3 algorithm, which can improve the UAV's exploration of the environment and utilization of effective data. The AO algorithm is used to optimize the discrete phase shift angle of the RIS, thereby improving the quality of the wireless channel environment. Attached Figure Description
[0045] Figure 1 A flowchart of a control method for RIS-assisted UAV communication;
[0046] Figure 2 A schematic diagram of a RIS-assisted unmanned aerial vehicle (UAV) communication system;
[0047] Figure 3 Here is a structural diagram of the TD3-PER-AO algorithm;
[0048] Figure 4A two-dimensional diagram of the first type of user cluster distribution in a simple and accessible environment;
[0049] Figure 5 A two-dimensional diagram of the second type of user cluster distribution in a simple and accessible environment;
[0050] Figure 6 A two-dimensional diagram of the third type of user cluster distribution in a simple, accessible environment;
[0051] Figure 7 A 3D graph of the first scenario, showing the distribution of user clusters in an environment full of obstacles;
[0052] Figure 8 A 3D graph of a second scenario where user clusters are distributed in an environment full of obstacles;
[0053] Figure 9 A 3D graph of a third scenario where user clusters are distributed in an environment full of obstacles;
[0054] Figure 10 A comparison chart of the TD3 algorithm with and without PER in an accessible environment;
[0055] Figure 11 Reward graphs for RIS models with and without TD3-PER algorithm in an accessible environment;
[0056] Figure 12 A comparison chart of the TD3 algorithm with and without PER in obstacle environments;
[0057] Figure 13 Reward graphs for RIS models with and without TD3-PER algorithm in obstacle environments;
[0058] Figure 14 A comparison chart of drone energy consumption under different environments and user distribution scenarios;
[0059] Figure 15 A comparison chart of drone energy efficiency under different environments and user distribution scenarios;
[0060] Figure 16 A comparison chart of the total speed of drones under different environments and user distribution scenarios;
[0061] Figure 17 Block diagram of the control device for RIS-assisted drone communication. Detailed Implementation
[0062] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0063] Unless otherwise stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0064] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0065] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0066] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0067] It should be noted that similar symbols and letters in the following figures represent similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0068] To achieve rational flight path planning for UAVs, maximize communication throughput (i.e., provide communication quality), and improve energy efficiency, this disclosure proposes a control method and related apparatus for RIS-assisted UAV communication. Specifically, the joint optimization problem of UAV trajectory, UAV transmission power, and RIS phase shift angle deflection is transformed into a Markov decision process model. By solving the Markov decision process model, RIS-assisted UAV communication is controlled.
[0069] Figure 1 This is a schematic diagram of one embodiment of the RIS-assisted UAV communication control method disclosed herein. Figure 1 The implementation can be executed by the server in the control center.
[0070] like Figure 1As shown, in step 1 of the embodiment, the elements of the Markov decision process model are updated using the UAV's three-dimensional coordinates, the user's three-dimensional coordinates, and the RIS phase shift matrix at the current time step. The Markov decision process model is transformed from an optimization problem, which is a problem of jointly optimizing the UAV's movement trajectory, UAV's transmission power, and RIS phase shift deflection, with the goal of improving the UAV's communication capabilities and energy efficiency in data acquisition.
[0071] It should be noted that Markov decision process models need to be pre-built, and the building process may include:
[0072] 1) Construct a communication model and an energy consumption model based on the environmental parameters of RIS-assisted UAV communication.
[0073] The environment for RIS-assisted drone communication can be as follows:
[0074] like Figure 2 The communication system is defined as having one drone and several users (i.e., mobile user terminals). The drone's flight altitude is fixed at H1, and the users are divided into K clusters, with M users in each cluster, used for communication via v. ue The speed of the random movement is different in different communication areas. Assume that the RIS is a uniform linear array composed of S elements, fixed in a certain position, and the environment is set with several no-fly zones / obstacles.
[0075] It should be noted that communication between drones and users is divided into two types: direct communication and indirect communication via RIS. Therefore, the communication model can be defined as follows:
[0076] Define the three-dimensional coordinates of the UAV at time step t as L u (t)=[x u (t),y u [(t),H1],x u (t) and y u (t) are respectively L u The X-axis and Y-axis coordinates in (t) are given, and the three-dimensional coordinates of the m-th user in the k-th user cluster at time step t are given by L. k,m (t)=[x k,m (t),y k,m [(t),0],x k,m (t) is L k,m The x-axis coordinate and y-axis coordinate in (t) k,m (t) is L k,m Given the Y-axis coordinate in (t), the distance between the UAV and the m-th user in the k-th user cluster at time step t is d. u,k,m (t)=||L k,m (t)- L u(t)||2.
[0077] The channel gain g when the m-th user in the k-th user cluster communicates directly with the drone u,k,m It can be represented as:
[0078] ;
[0079] In the formula, ρ0 is the channel gain at a reference distance of 1m, α1 is the path loss factor, and d u,k,m Let be the distance between the drone and the m-th user in the k-th user cluster.
[0080] Define the three-dimensional coordinates of the fixed position of RIS as L r =[x r ,y r [,H2],x r and y r L respectively r The X and Y coordinates are shown in the figure, H2 is the fixed height of the RIS, and the RIS phase shift matrix is represented. , for The phase shift of the first element, the first The phase shift of each element is expressed as a complex number. Without losing generality, , For the first The phase shift angle of each element, .
[0081] The link from the drone to the user can be indirectly represented by two links: the user-to-RIS and the RIS-to-drone. Using a Ricean channel fading model to model these two links, the channel gain g from the drone to the RIS is then calculated. r,u And the channel gain g from the m-th user in the k-th user cluster to the RIS. r,k,m It can be represented as:
[0082] ;
[0083] ;
[0084] ;
[0085] ;
[0086] In the formula, d r,u d represents the distance from the drone to the RIS. r,k,m Let α2 and α3 be the distance between the user and the RIS, α3 be the path loss factors, and K1 and K2 be the Rice factors. The components are random NLoS, conforming to a cyclic complex Gaussian distribution. The Loss of Service (LoS) component in the link can be specifically represented as:
[0087] ;
[0088] ;
[0089] In the formula, λ and d are the carrier wavelength and the spacing between RIS elements, respectively, and the cosine of the launch angle. The cosine of the angle reached .
[0090] Therefore, the channel gain when the m-th user in the k-th user cluster communicates indirectly with the UAV is g. u,r,k,m =g r,k,m Θg r,u .
[0091] The probability of direct link congestion is evaluated using an air-to-ground channel model in an urban environment. a and b are two constants determined by environmental factors. The congestion probability p between the UAV and the m-th user in the k-th user cluster is then calculated. u,k,m It can be represented as:
[0092] .
[0093] Let P be the drone's transmit power, B be the total bandwidth, and B be the average bandwidth from the drone to each user. Then the average channel gain h between the UAV and the m-th user in the k-th user cluster is... u,k,m It can be represented as:
[0094] ;
[0095] The communication rate r between the drone and the m-th user in the k-th user cluster u,k,m It can be represented as:
[0096] ;
[0097] In the formula, σ represents additive white Gaussian noise.
[0098] It should be noted that there are two main types of drone energy consumption: one is the energy consumption for drone movement, and the other is the energy consumption for launch. Therefore, the energy consumption model can be defined as follows:
[0099] Define the blade power and derived power of the UAV in hovering state as P0 and P1, respectively, and the tip velocity of the UAV blade as U. tip If the average rotor-induced velocity of the drone while hovering is v0, and the drone's moving speed is v, then the drone's kinetic energy consumption E mov It can be represented as:
[0100] ;
[0101] In the formula, d0, ρ, s0 and A are parameters related to the drag ratio of the UAV fuselage, air density, rotor solidity and rotor disk area, respectively.
[0102] Total energy consumption can be expressed as:
[0103] E u =E mov +E t ;
[0104] In the formula, E u E represents the total energy consumption of the drone. t For the launch energy consumption of the drone, E t =Pδ t P is the transmit power of the UAV, δ t Let t be the duration of time step t (assuming the entire communication period T is divided into N time steps, and the duration of each step is δ). t =T / N).
[0105] 2) Based on the communication model and energy consumption model, construct the optimization problem and transform the optimization problem into a Markov decision process model.
[0106] It should be noted that, based on the analysis of the above model, a joint optimization problem is constructed for the UAV's movement trajectory, UAV's transmission power, and RIS phase shift deflection, with the goal of improving the UAV's communication capability and energy efficiency in data acquisition. Specifically, the aim is to improve energy efficiency while optimizing the overall system communication rate to meet the needs of all users in the cluster during the data acquisition operation, and at the same time minimize the total propulsion energy consumption of the UAV at all time steps.
[0107] The optimization problem can be represented as:
[0108] ;
[0109] st C1:0≤x u (t) ≤L;
[0110] C2:0≤y u (t) ≤L;
[0111] C3:L u (0)={(x u ,y u )|(x u ,y u )∈A ini};
[0112] C4:L u (N-1)={(xu ,y u )|(x u ,y u )∈A des};
[0113] C5:|| L u (t-1)- L u (t)||2≤V max δ t ;
[0114] C6: 0≤P(t) ≤P max ;
[0115] C7: ;
[0116] C8: ;
[0117] In the formula, L u Let L be the three-dimensional coordinates of the UAV, and A be the maximum boundary value of the UAV's flight area. ini A des These represent the start and end point ranges of the drone, L. u (0) and L u (N-1) represent the three-dimensional coordinates of the UAV at the initial and final time steps, respectively, and L u (t-1) represent the three-dimensional coordinates of the UAV at time step t-1, V max P is the maximum speed of the UAV, and P(t) is the transmit power of the UAV at time step t. max E is the upper limit of the drone's transmit power. u (t) represents the total energy consumption of the UAV at time step t, and n represents the quantization bit depth (limited by current hardware capabilities, the phase shift angle is discretized into 2^n bits). n (each corner).
[0118] Of the constraints mentioned above, C1 and C2 define the flight of the UAV within the defined boundaries, C3 and C4 restrict its initial and final positions, C5 sets the maximum speed of the UAV, C6 restricts the power allocation of the user cluster, C7 ensures that the UAV's energy consumption remains within the battery capacity range, and C8 represents the range of values for the phase shift angle.
[0119] It should be noted that the above optimization problem is transformed into a Markov Decision Process (MDP) model, where state transitions depend entirely on the previous state and are unaffected by the entire historical sequence. MDP uses quintuples (...). S , A , R , P Let ,γ) represent, where Ѕ Representing the state space,A Represents the action space. R Represents the reward function, P This represents the state transition probability, with γ serving as a discount factor, ranging from 0 to 1.
[0120] Based on the characteristics of the problem, the state space, action space, and reward function are summarized as follows:
[0121] The state space is represented as S ={s t =( S 1, S 2 )}, used to characterize a series of observations of the wireless environment, where s t The state at time step t. S 1 represents the two-dimensional coordinates of the UAV at time step t. S 1={x u (t),y u (t)}, S 2 Let be the set of distances between the drone and the user, and between the drone and the destination at time step t. S 2 ={d 1,1 (t),…, d K,M (t), d des (t)},d K,M (t) represents the distance between the UAV and the Mth user in the Kth user cluster at time step t, and d des (t) represents the distance between the UAV and the destination at time step t.
[0122] Action space is represented as A ={a t =(θ(t),l(t),P(t)}, where, a t Let θ(t) ∈ (0, π) be the action at time step t, and l(t) ∈ [0, π]. max δ t Let P(t) represent the distance the UAV flies at time step t, where P(t) ∈ [0, P]. max [ ] represents the UAV's transmit power at time step t. Furthermore, each time the UAV performs a maneuver, the RIS controller manipulates the RIS phase shift component, thereby adjusting the phase shift angle accordingly.
[0123] The reward function consists of four parts: the communication rate between the drone and the user, the energy consumed by the drone during flight (i.e., E), and so on. u The distance-based destination reward and anomaly penalty are described in detail below:
[0124] ;
[0125] In the formula, R(t) is the reward value at time step t, and δ c1 ~δ c4 r is a constant u,k,m With E u The ratio of energy efficiency to the efficiency of drone flight is crucial. Furthermore, once the drone has collected sufficient information from the m-th user in the k-th user cluster, such as r... u,k,m ≥r min r min If the communication rate is set to the lower limit, the user is marked as complete, thus excluding further access from the drone.
[0126] R des The guidance parameters for the UAV to reach its destination are defined as follows:
[0127] ;
[0128] In the formula, d des (t) represents the horizontal distance between the UAV and the destination center at time step t, and d des (t-1) represents the horizontal distance between the UAV and the destination center at time step t-1, R ser The additional reward for drones completing user data collection is primarily intended to accelerate information gathering and set up R... ser =1, to prevent the drone from staying within a certain location range, R pun The penalty for abnormal situations is mainly to prevent drones from engaging in abnormal behaviors such as exceeding time limits, crossing boundaries, or entering no-fly zones during flight. R is set up... pun =-1, and simultaneously set δ c4 Set to a large value, and use a termination flag to end the exploration of the current scene if any of the above-mentioned exceptions are encountered.
[0129] return Figure 1 In step 2 of the embodiment, the TD3 algorithm is used to solve the updated Markov decision process model to obtain the UAV movement control command and transmit power control command, as well as the phase shift angle control command of each component of the RIS at the current time step. During training, the TD3 algorithm adjusts the probability of sampling experience data by assigning priority weights to the data in the experience pool. When solving the problem using the TD3 algorithm, the AO algorithm is used to find the optimal phase shift angle at the current time step. The optimal phase shift angle is the phase shift angle that maintains the maximum communication rate between the UAV and the user.
[0130] It should be noted that, in order to address the mixed action space issues of continuous UAV transmission power and trajectory movement, as well as discrete phase shifts in RIS, an improved TD3 algorithm was pre-designed, specifically the TD3-PER-AO algorithm, see [link to relevant documentation]. Figure 3Specifically as follows:
[0131] For the UAV control problem, the TD3-PER algorithm is used to optimize the transmission power and trajectory during the process. Specifically, the Priority Experience Playback (PER) technique is introduced on the traditional dual-delay deep deterministic policy gradient (TD3).
[0132] In the traditional TD3 algorithm training process, after receiving a state, the UAV executes the algorithm-specified action, and then receives feedback rewards, new states, and mission states, storing this data in an experience replay pool. However, the traditional TD3 algorithm uses random sampling, which, while reducing the correlation between experience data, reduces the utilization rate of important experience data. Data utilization can only be guaranteed through a sufficient number of samples. Here, the PER method is introduced, which adjusts the probability of sampling experience data by assigning priority to data in the experience pool; that is, data with higher values are extracted first. Specifically, the priority p is determined by the TD error δ, which is the absolute value of the TD error plus ɛ, with the formula p = |δ| + ɛ. The probability of the j-th sample data being sampled is... ;in, p represents the total number of sample data points in each sampling. i and p j The priorities of the i-th and j-th sample data are respectively. The priority level is indicated by ɛ, which is a small constant used to prevent a probability of 0.
[0133] During the training process of this algorithm, data containing priorities is stored in PER. At each network update, sampling is performed based on P(j), and importance sampling (IS) weights w are applied. j To resolve the deviation.
[0134] In the TD3-PER algorithm, the main network of the actor... and its target network Critics Network and its target network The principles for updating parameters are as follows:
[0135] ;
[0136] ;
[0137] ;
[0138] ;
[0139] Among them, the critics' main network, through the The loss function is updated using gradient descent. This indicates the addition of a small amount of random noise. The action of the target network aims to achieve a smoother target policy. Every κ time steps, the actor network and all target networks are updated according to the above formula. This indicates that the critic network estimates the value of state-action pairs. For the actor network strategy, represent the mapping relationship between states and consecutive actions. Random noise. It follows a Gaussian distribution pruned to a small range (-c, c), with a mean of 0 and a variance of . , The objective function of the actor network is to maximize... The value is used to update the actor network, where Rj is the reward value for the j-th sample data. Parameters for soft updates of the target network. This indicates that the gradient of action a is being calculated. Calculate the gradient of the actor network parameters.
[0140] When using the TD3 algorithm, the phase shift angles used in the calculation process are all optimal phase shift angles. Specifically, the AO algorithm is used to find the current optimal phase shift angle. The optimal phase shift angle is the one that maintains the maximum communication rate between the UAV and the user, that is, it maintains the optimal communication rate between the UAV and the user at each time step. , To achieve the optimal phase shift angle, To use RIS at time step t Phase is the communication rate between the drone and the user, where Constraint C8 is satisfied. In each iteration of the AO algorithm, the communication rate is maximized iteratively by optimizing the phase shift of each component. This method ensures that the overall performance of the system is enhanced by fine-tuning the phase shift, thereby achieving the highest communication rate.
[0141] In each round of the algorithm, the UAV executes the actions specified by the TD3 algorithm, namely, flight direction, flight distance, and transmission power. At this time, the RIS controller, in conjunction with the UAV, optimizes the RIS phase shifter to obtain environmental feedback rewards, new states, and task states. This data is stored in the PER, and the algorithm network is updated. Through continuous iteration of the TD3-PER-AO algorithm, the optimal strategy for the UAV is finally obtained, thereby maximizing long-term rewards.
[0142] The above method uses an improved TD3 algorithm to solve the model. It introduces PER (Preferred Experience Playback) technology on the traditional TD3 algorithm, which can improve the UAV's exploration of the environment and the utilization of effective data. The AO algorithm is used to optimize the discrete phase shift angle of RIS, thereby improving the quality of the wireless channel environment.
[0143] To verify the above method, the following simulation experiment was conducted:
[0144] A 1000m × 1000m urban area was designed, with user clusters randomly distributed throughout the designated area, and users moving randomly. To measure the effectiveness of the method, three different scenarios with different user cluster distributions were designed: a simple scenario with a single cluster accommodating 4 users, and two more complex scenarios where two clusters were located nearby or far away, each accommodating 7 users. Furthermore, two environments were configured to increase the complexity of the UAV mission: an obstacle-free environment allowing UAVs to fly freely on buildings, and a no-fly zone established on specific high-rise buildings, requiring UAVs to adjust their course when encountering obstacles, thus increasing the mission difficulty. Results are shown below. Figures 4-9 As shown, Figures 4-6 The diagram shows a two-dimensional representation of the drone's flight trajectory in an obstacle-free environment. The orange area represents the start and end positions, the gray area represents the user clustering area, and the green area represents the location of the RIS (Real-Indicator System). The trajectory in the diagram is the flight trajectory generated by guiding the drone using this invention. Figures 7-9 This provides a more detailed 3D description of the UAV flight trajectory in obstacle-prone environments. The red path corresponds to the TD3-PER-AO algorithm, the blue path corresponds to the traditional TD3 algorithm, and the green path corresponds to the TD3-PER algorithm excluding the RIS model.
[0145] The TD3-AO algorithm and the TD3-PER algorithm were established to verify the performance of the TD3-PER-AO algorithm. Figure 10 and Figure 11 The experiment showcases the reward curves of the proposed optimization scheme and two benchmark algorithms in an obstacle-free environment. Clearly, the TD3-PER-AO algorithm demonstrates superior convergence speed and task performance across different scenarios, validating the effectiveness of the state space and reward formula. However, in a simplified obstacle-free environment, the performance differences between the UAV and the algorithm are not particularly pronounced. To address this issue, no-fly zones / obstacles were introduced into the experiment, increasing environmental complexity and enhancing user mobility, thus exacerbating the challenge of UAV exploration. Figure 12 and Figure 13The reward curves of the proposed optimization scheme and two benchmark algorithms are shown in an obstacle-filled environment. Clearly, the UAV, guided by the benchmark algorithm, fails to fully explore the environment and gets trapped in a local optimum; while the proposed scheme avoids obstacles, completes the data collection task, and quickly reaches the destination. Comparative analysis of two different environments shows that all three algorithms perform well in simple, obstacle-free environments, but the proposed algorithm performs slightly better than the benchmark algorithms in training. In more complex environments, the limitations of the benchmark algorithms hinder their ability to fully complete the assigned tasks, while the proposed algorithm demonstrates adaptability by effectively handling environmental complexity, thus achieving the expected goals in the face of these challenges.
[0146] To comprehensively evaluate the effectiveness of the TD3-PER-AO algorithm in different environments and user scenarios Figures 14-16 The study focuses on a comparative analysis of three key indicators for UAVs: energy consumption, energy efficiency, and total speed. Compared to an unobstructed environment, the proposed algorithm achieves significantly higher values for all indicators when operating in an environment with obstacles. This indicates that UAVs operating in obstacle-filled environments exhibit longer flight distances, higher energy consumption levels, greater environmental complexity, and higher mission difficulty.
[0147] The above method transforms the joint optimization problem of UAV trajectory, UAV transmit power, and RIS phase shift deflection into a Markov decision process model. At each time step, the elements of the Markov decision process model are updated, and the updated model is solved using the TD3 algorithm. This yields the UAV movement control commands and transmit power control commands, as well as the phase shift angle control commands for each RIS element at each time step. This ensures that, under strict time and energy constraints, the UAV flight path is rationally arranged, communication throughput (i.e., communication quality) is maximized, and energy utilization efficiency is achieved.
[0148] Figure 17 This is a schematic diagram of one embodiment of the control device for RIS-assisted unmanned aerial vehicle communication disclosed herein. Figure 3 An embodiment of this is a virtual device that can be loaded and executed by a server in the control center, including an update module and a solution module.
[0149] The update module in this embodiment is configured to update the elements of the Markov decision process model using the UAV's three-dimensional coordinates, the user's three-dimensional coordinates, and the RIS phase shift matrix at the current time step. The Markov decision process model is transformed from an optimization problem, which is a problem of jointly optimizing the UAV's trajectory, UAV's transmit power, and RIS phase shift angle deflection, with the goal of improving the UAV's communication capabilities and energy efficiency in data acquisition.
[0150] It should be noted that the Markov decision process model in the update module is pre-built. The construction process may include: building a communication model and an energy consumption model based on the environmental parameters of RIS-assisted UAV communication; building an optimization problem based on the communication model and the energy consumption model; and transforming the optimization problem into a Markov decision process model.
[0151] The solution module in this embodiment is configured to use the TD3 algorithm to solve the updated Markov decision process model, and obtain the UAV movement control command and transmit power control command, as well as the phase shift angle control command of each component of the RIS at the current time step. During training, the TD3 algorithm adjusts the probability of sampling experience data by assigning priority weights to the data in the experience pool. When solving with the TD3 algorithm, the AO algorithm is used to find the optimal phase shift angle at the current time step. The optimal phase shift angle is the phase shift angle that maintains the maximum communication rate between the UAV and the user.
[0152] Similar to the above method, the above device transforms the joint optimization problem of UAV movement trajectory, UAV transmission power, and RIS phase shift deflection into a Markov decision process model. At each time step, the elements of the Markov decision process model are updated, and the updated model is solved using the TD3 algorithm to obtain the UAV movement control command and transmission power control command, as well as the phase shift angle control command of each RIS element at each time step. This ensures that under strict time and energy constraints, the UAV flight path is rationally arranged, communication throughput (i.e., communication quality) is maximized, and energy utilization efficiency is achieved.
[0153] Based on the same technical solution, this disclosure also relates to a computer-readable storage medium that stores one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to perform a control method for RIS-assisted unmanned aerial vehicle communication.
[0154] Based on the same technical solution, this disclosure also relates to a computer device, including one or more processors and one or more memories, wherein one or more programs are stored in one or more memories and configured to be executed by one or more processors, and the one or more programs include instructions for performing a control method for RIS-assisted unmanned aerial vehicle communication.
[0155] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0156] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0157] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0158] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0159] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of the claims of the present invention pending approval.
Claims
1. A control method for RIS-assisted unmanned aerial vehicle (UAV) communication, characterized in that, include: The elements of the Markov decision process model are updated using the UAV's 3D coordinates, the user's 3D coordinates, and the RIS phase shift matrix at the current time step. The Markov decision process model is transformed into an optimization problem, which is a problem of jointly optimizing the UAV's trajectory, UAV's transmit power, and RIS phase shift deflection, with the goal of improving the UAV's communication capabilities and energy efficiency in data acquisition. The TD3 algorithm is used to solve the updated Markov decision process model to obtain the UAV movement control command and transmit power control command, as well as the phase shift angle control command of each component of the RIS at the current time step. During training, the TD3 algorithm adjusts the probability of sampling experience data by assigning priority weights to the data in the experience pool. When solving the problem using the TD3 algorithm, the AO algorithm is used to find the optimal phase shift angle at the current time step. The optimal phase shift angle is the phase shift angle that maintains the maximum communication rate between the UAV and the user. In the above Markov decision process model: The state space is represented as S = {s} t = (S1, S2)}, where s t S1 represents the state at time step t, S2 represents the two-dimensional coordinates of the UAV at time step t, and S3 represents the set of distances between the UAV and the user, and between the UAV and the destination at time step t. The action space is represented as A = {a t =(θ(t),l(t),P(t)}, where a t Let θ(t) be the action at time step t, θ(t) be the UAV's flight direction at time step t, l(t) be the UAV's flight distance at time step t, and P(t) be the UAV's transmit power at time step t. The reward function is expressed as: In the formula, R(t) is the total reward value at time step t, and δ c1 ~δ c4 r is a constant u,k,m Let E be the communication rate between the drone and the m-th user in the k-th user cluster. u R represents the total energy consumption of the drone. des These are the guidance parameters for the drone to reach its destination. d des (t) represents the horizontal distance between the UAV and the destination center at time step t, and d des (t-1) represents the horizontal distance between the UAV and the destination center at time step t-1, R ser R provides an additional reward for drones completing user data collection. pun This is a penalty for abnormal situations.
2. The control method for RIS-assisted UAV communication according to claim 1, characterized in that, The Markov decision process model is pre-built, and the building process includes: Based on the environmental parameters of RIS-assisted UAV communication, a communication model and an energy consumption model are constructed. Based on the communication model and energy consumption model, an optimization problem is constructed and then transformed into a Markov decision process model.
3. The control method for RIS-assisted UAV communication according to claim 2, characterized in that, The communication model is as follows: In the formula, B is the total bandwidth. Let P be the average bandwidth from the drone to each user, P be the drone's transmit power, σ be additive white Gaussian noise, and h be the noise level. u,k,m =(1+p) u,k,m )g u,k,m +p u,k,m g u,r,k,m Let p be the average channel gain between the UAV and the m-th user in the k-th user cluster. u,k,m Let g be the blocking probability between the drone and the m-th user in the k-th user cluster. u,k,m Let g be the channel gain when the m-th user in the k-th user cluster communicates directly with the drone. u,r,k,m =g r,k,m Θg r,u Let g be the channel gain when the m-th user in the k-th user cluster communicates indirectly with the UAV. r,k,m Let g be the channel gain from the m-th user in the k-th user cluster to the RIS. r,u Let Θ be the channel gain from the UAV to the RIS, and Θ be the RIS phase shift matrix; The energy consumption model is as follows: AND u =And mov +E t ; In the formula, E t For the launch energy consumption of the drone, E mov For the mobile energy consumption of drones, P0 and P1 are the blade power and derived power of the UAV in hovering state, respectively. tip denoted as the tip velocity of the UAV rotor blades, v0 as the average rotor-induced velocity of the UAV in hovering state, v as the moving speed of the UAV, and d0, ρ, s0 and A as the relevant parameters related to the UAV fuselage drag ratio, air density, rotor solidity and rotor disk area, respectively.
4. The control method for RIS-assisted UAV communication according to claim 3, characterized in that, The optimization problem is: s.t.C1:0≤x u (t)≤L; C2:0≤y u (t)≤L; C3:L u (0)={(x u ,y u )|(x u ,y u )∈A ini }; C4:L u (N-1)={(x u ,y u )|(x u ,y u )∈A des }; C5:||L u (t-1)-L u (t)||2≤V max δ t ; C6:0≤P(t)≤P max ; C7: C8: In the formula, L u Here, K represents the three-dimensional coordinates of the drone, M represents the number of users in the user cluster, and x represents the number of users in the user cluster. u (t) represents the three-dimensional coordinates of the UAV at time step t. u The x-axis coordinate and y-axis coordinate in (t) u (t) represents the three-dimensional coordinates of the UAV at time step t. u In (t), the Y-axis coordinate is L, which is the maximum boundary value of the UAV's flight area, and A is... ini A des These represent the start and end point ranges of the drone, L. u (0) and L u (N-1) represent the three-dimensional coordinates of the UAV at the initial and final time steps, respectively, and L u (t-1) represent the three-dimensional coordinates of the UAV at time step t-1, V max δ is the maximum speed of the drone's flight. t P is the duration of time step t. max E is the upper limit of the drone's transmit power. u (t) represents the total energy consumption of the UAV at time step t. For the first The phase shift angle of each element, where n is the number of quantization bits.
5. The control method for RIS-assisted UAV communication according to claim 1, characterized in that, The priority of data allocation in the experience pool is the absolute value of the TD error δ plus ε, and the probability of the j-th sample data being sampled is... in, p represents the total number of sample data points in each sampling. i and p j Let α and ε be the priorities of the i-th and j-th sample data, respectively, where α∈[0,1] represents the importance of the priority, and ε is a small constant.
6. A control device for RIS-assisted unmanned aerial vehicle communication, characterized in that, include: The update module uses the UAV's 3D coordinates, the user's 3D coordinates, and the RIS phase shift matrix at the current time step to update the elements of the Markov decision process model. The Markov decision process model is transformed from an optimization problem, which is a problem of jointly optimizing the UAV's trajectory, UAV's transmit power, and RIS phase shift angle deflection, with the goal of improving the UAV's communication capabilities and energy efficiency in data acquisition. The solution module uses the TD3 algorithm to solve the updated Markov decision process model, obtaining the UAV movement control commands and transmit power control commands for the current time step, as well as the phase shift angle control commands for each component of the RIS. During training, the TD3 algorithm adjusts the probability of sampling experience data by assigning priority weights to the data in the experience pool. When solving using the TD3 algorithm, the AO algorithm is used to find the optimal phase shift angle for the current time step, which is the phase shift angle that maintains the maximum communication rate between the UAV and the user. In the above Markov decision process model: The state space is represented as S = {s} t = (S1, S2)}, where s t S1 represents the state at time step t, S2 represents the two-dimensional coordinates of the UAV at time step t, and S3 represents the set of distances between the UAV and the user, and between the UAV and the destination at time step t. The action space is represented as A = {a t =(θ(t),l(t),P(t)}, where, a t Let θ(t) be the action at time step t, θ(t) be the UAV's flight direction at time step t, l(t) be the UAV's flight distance at time step t, and P(t) be the UAV's transmit power at time step t. The reward function is expressed as: In the formula, R(t) is the total reward value at time step t, and δ c1 ~δ c4 r is a constant u,k,m Let E be the communication rate between the drone and the m-th user in the k-th user cluster. u For the total energy consumption of the drone, R des These are the guidance parameters for the drone to reach its destination. d des (t) represents the horizontal distance between the UAV and the destination center at time step t, and d des (t-1) represents the horizontal distance between the UAV and the destination center at time step t-1, R ser R provides an additional reward for drones when they complete user data collection. pun This is a penalty for abnormal situations.
7. The control device for RIS-assisted UAV communication according to claim 6, characterized in that, The Markov decision process model in the update module is pre-built, and the building process includes: Based on the environmental parameters of RIS-assisted UAV communication, a communication model and an energy consumption model are constructed. Based on the communication model and energy consumption model, an optimization problem is constructed and then transformed into a Markov decision process model.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to perform the method of any one of claims 1 to 5.
9. A computer device, characterized in that, include: One or more processors and one or more memories, one or more programs stored in one or more memories and configured to be executed by one or more processors, the one or more programs including instructions for performing the method of any one of claims 1 to 5.
Citation Information
Patent Citations
6G RIS assisted full duplex unmanned aerial vehicle system energy efficiency optimization method and device
CN115694583A
Energy efficiency optimization method of unmanned aerial vehicle and IRS auxiliary wireless charging edge network
CN118233926A