A communication network optimization method for the attitude change of an aerial RIS
Through Euler angle control and SAC deep reinforcement learning, the attitude and phase shift of UAV-RIS is optimized, which solves the beam offset problem caused by the fuselage tilt during acceleration and deceleration, and improves the performance of RIS-assisted communication in the air.
Patent Information
- Application Number
- CN202510475283.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The prior art has failed to effectively deal with the fuselage tilt caused by inertial resistance and aerodynamic effects during acceleration and deceleration, resulting in beam offset and channel changes, affecting the performance of RIS assisted communication in the air.
The flight control paradigm based on Euler angle is adopted, combined with the Soft Actor Critic (SAC) deep reinforcement learning method, optimize the posture of UAV and the phase shift of RIS, and ensure that the signal is accurately aligned with the target user through real-time phase regulation and beam adjustment.
Real-time optimization of the air RIS attitude is achieved, the system communication rate and channel gain are improved, and the communication performance stability and efficiency are improved.
Smart Images

Figure CN120018159B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of wireless communication, and particularly relates to a method for optimizing a communication network assisted by a reconfigurable intelligent surface carried by an unmanned aerial vehicle. Background Art
[0002] With the increasing tension of airspace resources, as an important part of the low-altitude economy, unmanned aerial vehicles (UAVs) have been widely used in scenarios such as emergency communication, signal coverage, and intelligent transportation due to their high mobility and rapid deployment capabilities. At the same time, as a key technology, reconfigurable intelligent surfaces (RISs) can effectively improve spectral efficiency and reduce energy consumption by dynamically controlling the phase of reflected signals. However, the complexity of the low-altitude airspace often seriously affects communication quality, and traditional independent deployment of UAVs or RISs is difficult to effectively address this challenge. By integrating UAVs with RISs, a UAV-RIS assisted communication system can enhance the reliability and adaptability of transmission, thereby ensuring stable and efficient communication performance in key tasks of the low-altitude economy.
[0003] However, in the actual aerial RIS deployment, due to the influence of inertial resistance and aerodynamic effects during the acceleration and deceleration of UAVs, the fuselage inevitably tilts, resulting in beam offset and channel variation, thus reducing the performance of aerial RIS assisted communication. In addition, existing research has shown that the actual gain of RISs is highly sensitive to the incident angle and reflection angle of signals. Despite these physical limitations, most current research ignores the impact of the attitude change of aerial RISs, resulting in the system performance not reaching the theoretical upper bound of the aerial RIS gain. This persistent modeling defect severely limits the effectiveness of aerial RISs in practical applications. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to propose a flight control paradigm based on Euler angles for the deficiencies in the background art, which realizes the attitude optimization of aerial RISs; through the Euler angle control strategy, this framework can perform real-time phase offset compensation while optimizing attitude adjustment, thereby maintaining the optimal beam alignment.
[0005] The present invention adopts the following technical solutions to solve the above technical problems:
[0006] A method for optimizing a communication network for the attitude change of aerial RISs, including a communication network scenario assisted by aerial RISs, where the communication network scenario assisted by aerial RISs includes a communication environment between a base station and a ground user and a control environment of UAV-RIS;
[0007] The communication environment between the base station and the ground users includes the set of ground users and the base station; in this environment, the base station can provide communication services for multiple ground users at the same time; among them, the communication links between the base station and the users are divided into two types: direct link and reflection link; the direct link is that the signal is directly transmitted to the ground user through free space propagation; the reflection link is that the signal is sent from the base station and transmitted to the ground user after being reflected by the aerial RIS.
[0008] The control environment of the UAV-RIS includes the UAV and the RIS array; in this control environment, the UAV platform serves as a carrier and provides an aerial signal reflection relay between the base station and the ground users through its flexible flight ability; during flight and hovering, affected by inertia and air resistance, the body attitude of the UAV, including the roll angle, pitch angle, and yaw angle, will change dynamically, which will lead to the deviation of the RIS reflection angle and the signal incident angle, thus causing beam alignment error and channel gain fluctuation, directly affecting the communication performance of the reflection link.
[0009] Specifically, it includes the following steps:
[0010] Step 1, the UAV carrying the RIS hovers or flies over the ground users; the base station sends communication signals to the aerial RIS through the downlink, and the RIS reflects and enhances the signals and adjusts the beam through real-time phase control to ensure that the reflected signals are accurately aligned with the direction of the target users; the optimized signals are transmitted to the ground users through the reflection link.
[0011] Step 2, in order to cope with the influence of the UAV attitude changes, including the roll angle, pitch angle, and yaw angle, on the angles of the received and reflected signals, that is, on the system communication performance, the Soft Actor-Critic (SAC) deep reinforcement learning method is used to jointly optimize the UAV attitude and the RIS phase shift, so as to maximize the system communication rate.
[0012] Step 3, in each time slot l, through the state analysis of the aerial RIS attitude and the UAV trajectory, the optimal strategy is determined, and the UAV performs the next action according to the obtained optimal strategy.
[0013] As a further preferred solution of a communication network optimization method for the aerial RIS attitude change in the present invention, in Step 1, the environment adopts a three-dimensional Cartesian coordinate system, the UAV flies horizontally at a fixed height, and the ground users are randomly distributed.
[0014] As a further preferred solution of the communication network optimization method for the attitude change of the airborne RIS in the present invention, in step 2, the angles of the received and reflected signals are represented by Euler angles, so as to represent the received and reflected beams, and the system communication rate sum is calculated; the SAC deep reinforcement learning method improves the exploration efficiency through the maximum entropy strategy, so that the exploration and exploitation can be balanced in the policy evaluation and policy improvement stages.
[0015] As a further preferred solution of the communication network optimization method for the attitude change of the airborne RIS in the present invention, in step 3, the optimal policy includes: adjusting the flight trajectory of the UAV and the attitude of the airborne RIS according to the current environment, so as to ensure the optimal signal incident angle and reflection angle, and then maximize the channel gain of the RIS and improve the communication performance of the system.
[0016] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0017] 1. The method for controlling the airborne RIS based on Euler angles in the present invention jointly optimizes the UAV trajectory and the attitude of the airborne RIS, and adopts a deep reinforcement learning algorithm to maximize the system communication rate sum; the present invention considers the coupling relationship between the attitude of the airborne RIS and the UAV trajectory, as well as the relationship between the RIS channel gain and the orientation, making the application scenario more in line with the actual situation, and good results can be obtained when applied to the actual scenario;
[0018] 2. The SAC algorithm proposed by the present invention not only focuses on maximizing the cumulative reward, but also takes the policy entropy as one of the key optimization objectives; the algorithm aims to maximize the weighted sum of the cumulative reward and the policy entropy, and by introducing policy randomness, encourages the agent to have higher exploration ability when selecting actions. Brief Description of the Drawings
[0019] Figure 1 is the application scenario of the present invention, a communication network environment assisted by a UAV carrying an RIS;
[0020] Figure 2 is the flowchart of the SAC algorithm of the present invention;
[0021] Figure 3 is the comparison of the convergence performance of the SAC algorithm used in the present invention and other deep reinforcement learning algorithms (PPO, DDPG) under the system model of the present invention;
[0022] Figure 4 is the flowchart of the present invention. Detailed Embodiments
[0023] The present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments. The objectives and effects of the present invention will become more clear. The preferred embodiments described herein are only for explaining the present invention and are not used to limit the present invention:
[0024] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. The present invention will be described in detail below with reference to the accompanying drawings and preferred embodiments. The objectives and effects of the present invention will become more clear. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0025] A communication network optimization method for the attitude change of an aerial RIS. The described aerial RIS-assisted communication network scenario includes the communication environment between the base station and the ground user and the control environment of the UAV-RIS.
[0026] The described communication environment between the base station and the ground user includes the set of ground users and the base station. In this environment, the base station can provide communication services for multiple ground users at the same time. The communication links between the base station and the users can be divided into two types: direct link and reflected link. The described direct link is that the signal is directly transmitted to the ground user through free space propagation; the described reflected link is that the signal is emitted from the base station and transmitted to the ground user after being reflected by the aerial RIS.
[0027] The described control environment of the UAV-RIS includes the UAV and the RIS array. In this control environment, the UAV platform serves as a carrier and provides an aerial signal reflection relay between the base station and the ground user through its flexible flight ability. During flight and hovering, the UAV is affected by inertia and air resistance, and its body attitude (including roll angle, pitch angle, and yaw angle) will change dynamically, which will further cause the deviation of the RIS reflection angle and the signal incident angle, resulting in beam alignment error and channel gain fluctuation, directly affecting the communication performance of the reflected link.
[0028] The optimization of the attitude change of the aerial RIS is achieved by constructing a dynamic model based on Euler angles, establishing the relationship between the incident signal angle and the reflected signal angle in the reflected link and the Euler angles, and further obtaining the relationship between the RIS gain and the Euler angles. Under this modeling, in order to maximize the RIS gain and thus improve the communication performance, the present invention proposes a SAC deep reinforcement learning method, which transforms the communication performance optimization problem into a sequential decision-making problem to find the weight of the adjusted policy entropy, a feasible and optimal policy.
[0029] The described communication network optimization method based on the attitude change of the aerial RIS specifically includes the following steps:
[0030] (1) The UAV carries the RIS and hovers or flies above the ground users. The base station sends communication signals to the aerial RIS through the downlink. The RIS reflects and enhances the signals and adjusts the beam through real-time phase regulation to ensure that the reflected signals are accurately aligned with the direction of the target users. The optimized signals are transmitted to the ground users through the reflection link.
[0031] (2) To address the impact of UAV attitude changes, including roll angle, pitch angle, and yaw angle, on the angles of the received and reflected signals, i.e., on the system communication performance, a SAC-based deep reinforcement learning method is proposed to jointly optimize the UAV's attitude and the RIS phase shift, thereby maximizing the system communication rate sum.
[0032] (3) At each time slot l, by analyzing the states of the aerial RIS attitude and the UAV trajectory, the optimal policy is determined, and the UAV performs the next action according to the obtained optimal policy.
[0033] Furthermore, the environmental system in step (1) has the following characteristics: The environment uses a three-dimensional Cartesian coordinate system. The UAV flies horizontally at a fixed height, and the ground users are randomly distributed. Considering that each time slot interval is very short, it can be assumed that the terminal users are static within each time slot. At the same time, the base station position is fixed.
[0034] Furthermore, the angles of the received and reflected signals in step (2) will be represented by Euler angles, thereby representing the received and reflected beams, and finally calculating the system communication rate sum. In addition, the core idea of the proposed SAC algorithm is to improve the exploration efficiency through the maximum entropy policy, enabling a balance between exploration and exploitation in the policy evaluation and policy improvement phases.
[0035] Furthermore, the optimal policy in step (3) includes adjusting the UAV's flight trajectory and the attitude of the aerial RIS according to the current environment to ensure the optimal signal incident angle and reflection angle, thereby maximizing the channel gain of the RIS and ultimately improving the system's communication performance. The optimization objective is the system communication rate sum.
[0036] In view of the deficiencies of existing research, the present invention proposes a flight control paradigm based on Euler angles to achieve the attitude optimization of the aerial RIS. Through the Euler angle control strategy, this framework can perform real-time phase offset compensation while optimizing the attitude adjustment, thereby maintaining the optimal beam alignment. In addition, the present invention uses the system communication rate sum as a performance indicator to reflect the effectiveness of the aerial RIS attitude optimization. Furthermore, the present invention reconstructs the communication rate maximization problem into a model based on the Markov decision process and proposes a SAC-based deep reinforcement learning method.
[0037] The present invention designs an aerial RIS control scheme based on Euler angles to jointly optimize the UAV trajectory and the aerial RIS attitude. A deep reinforcement learning algorithm is adopted to maximize the system communication rate and...
[0038] Compared with other inventions, the present invention takes into account the coupling relationship between the aerial RIS attitude and the UAV trajectory, as well as the connection between the RIS channel gain and the orientation, making the application scenario more in line with the actual situation and achieving better results when applied to the actual scenario.
[0039] Some other inventions use traditional optimization methods, such as the block coordinate descent method, the successive convex approximation method, and the semidefinite relaxation method. However, these traditional methods usually have a high computational complexity and often generate static solutions in complex environments, unable to adapt to dynamically changing communication scenarios, resulting in the optimization results may become suboptimal or outdated in actual deployment. Another part often adopts algorithms such as DQN, DDPG, or PPO. These traditional deep reinforcement learning algorithms often exhibit problems such as slow convergence speed and low training efficiency, and are prone to falling into local optimal solutions, making it difficult to obtain a globally optimal UAV autonomous maneuver decision. The SAC algorithm proposed by the present invention not only focuses on maximizing the cumulative reward but also takes the policy entropy as one of the key optimization objectives. The algorithm aims to maximize the weighted sum of the cumulative reward and the policy entropy, and by introducing policy randomness, encourages the agent to have a higher exploration ability when choosing actions. Specific embodiments:
[0041] System model: As Figure 1 shown, the described aerial RIS-assisted wireless communication scenario includes K single-antenna user equipments, a multi-antenna base station, and a UAV equipped with an RIS at the bottom. The entire flight period of the UAV is T. For ease of processing, the total flight time T is equally divided into L time slots, and the length of each time slot is δ = T / L. The position of the UAV changes between adjacent time slots. To accurately describe the flight trajectory of the UAV, a three-dimensional Cartesian coordinate system is used to model the flight state of the UAV, assuming that the flight height of the UAV is a fixed value H. In the l-th time slot, the position coordinates of the UAV are expressed as q[l] = (x[l], y[l], H).
[0042] RIS Angle Calculation and Attitude Transformation: In the local coordinate system, the initial unit normal vector of the airborne RIS is represented as e. To achieve the spatial coordinate transformation from the global coordinate system to the local coordinate system, a two-stage coordinate transformation method is adopted. First, an initial translation transformation is performed to translate the coordinate origin (0, 0, 0) to the instantaneous position (x[l], y[l], H) of the UAV, thereby ensuring the accurate spatial positioning of the RIS in the global reference system. Next, the attitude is adjusted by means of a rotation transformation, and attitude correction is achieved through a series of rotation operations parameterized by Euler angles. The rotation transformation includes three consecutive operations: the roll angle of rotation around the x-axis, the pitch angle of rotation around the y-axis, and the yaw angle of rotation around the z-axis, denoted as φ[l], θ[l], The corresponding transformation matrices for controlling these rotations are represented as:
[0043]
[0044] where R φ is the transformation matrix for controlling the roll angle, R θ is the transformation matrix for controlling the pitch angle, is the transformation matrix for controlling the yaw angle.
[0045] Then the complete attitude transformation matrix is Furthermore, after applying this coordinate transformation, the unit normal vector of the airborne RIS can be derived in the global coordinate system as e ⊥ = R BE e, where e represents the initial normal vector of the airborne RIS, and the obtained e ⊥ after multiplication is the transformed unit normal vector. In addition, for the incident signal from the base station to the airborne RIS and the reflected signal from the airborne RIS to the ground user, their unit direction vectors can be expressed as:
[0046]
[0047] where and respectively represent the azimuth angle and elevation angle from the base station or user equipment to the airborne RIS at time slot l. By using these direction vectors, the angle between the plane normal vector of the airborne RIS and the incident / reflected signal can be derived as:
[0048]
[0049] It can be seen that the change in the azimuth angle of the airborne RIS significantly affects the directions of the incident signal and the reflected signal, thereby changing the equivalent receiving aperture. This change directly affects the gain of the airborne RIS and thus changes the performance of the entire system because the change in the signal reflection characteristics directly determines the link quality and communication efficiency.
[0050] Communication model: At any time slot, the channel gains between the base station and the RIS in the air, and between the RIS in the air and the ground user are respectively denoted as H B,R [l] and h R,k [l].
[0051] Since the actual gain of the RIS in the air is significantly affected by the signal incident angle and reflection angle, the actual gain model introduces an expression considering the azimuth angle and elevation angle. Specifically, the actual gain of the RIS in the air can be modeled as the product of the maximum directivity coefficient of the RIS in the air and the reflection angle characteristic function:
[0052]
[0053] where, and respectively represent the receiving gain from the base station to the RIS in the air and the transmitting gain from the RIS in the air to the user equipment. and are respectively the normalized directional radiation functions in the direction from the base station to the RIS in the air and in the direction from the RIS in the air to the user equipment.
[0054] Meanwhile, in order to more accurately describe the gain characteristics of the RIS, the maximum directivity coefficient D max and the normalized directional radiation function F(ζ,η) are introduced. This function reflects the directional radiation characteristics of the RIS in the air and is modeled as an exponential-Lambert radiation model, and its expression is
[0055]
[0056] where ζ and η respectively represent the azimuth angle and elevation angle, which are used to define the spatial relationship between the ground user (or base station) and the RIS in the air.
[0057] Based on these mathematical expressions, the RIS gain expression can be further deduced as:
[0058]
[0059] where Φ[l] is the RIS phase shift matrix, and this expression reflects the angle selection characteristic of the RIS for the reflected signal, that is, when the reflection angle meets certain conditions, the RIS can effectively reflect the signal, otherwise the signal cannot be reflected.
[0060] Furthermore, in the RIS-assisted wireless communication system in the air, the signal received by the ground user at time slot l can be expressed as:
[0061]
[0062] where v k[l] is the composite channel from the base station to the ground user, including the direct link and the reflection link from the base station to the ground user. w k [l] is the beamforming vector; x k [l] is the transmitted signal; n k is additive white Gaussian noise and follows a complex Gaussian distribution, and the noise variance is σ 2 . Then, the communication rate of the ground user within time slot l is:
[0063]
[0064] Finally, the total rate of the entire system over all time slots and all users can be expressed as:
[0065]
[0066] Markov Decision Process: In the application of deep reinforcement learning, we first define the Markov decision process, which is the basic framework for solving sequential decision-making problems in a stochastic environment. A Markov decision process usually consists of five parts: state space, action space, state transition probability function, reward function, and discount factor. At each time slot l, the agent selects an action based on the policy π, which defines the probability distribution of selecting actions given the state where S l and A l are the state space set and the action space set at time slot l; s and a represent the state and action under the current policy.
[0067] After each action is executed, the system transitions to the next state according to the state transition probability function and gives the agent an immediate reward. This state transition and reward feedback process guides policy optimization, and the agent adjusts the policy through iterative interactions with the environment to gradually approach the optimal policy.
[0068] SAC Algorithm: As Figure 2 shown. Under the SAC algorithm framework, the distribution of the state-action trajectory under the policy π is denoted as τ π . Different from traditional deep reinforcement learning algorithms, SAC introduces an entropy regularization term and integrates it into the objective function to improve the exploration efficiency of the policy. Its optimization goal is to maximize the sum of the cumulative reward and the policy entropy:
[0069]
[0070] where is the entropy of the policy distribution; represents the distribution τ l , a l ) over all state-action pairs (s πSampling is performed and the expectation is calculated. The entropy regularization term controls the weight of entropy in the optimization process by introducing a temperature parameter α, balancing the relationship between exploration and exploitation. The SAC algorithm adopts a policy iteration framework, alternating between policy evaluation and policy improvement. In the policy evaluation phase, the soft state-value function is calculated based on the Bellman expectation equation, and its expression is:
[0071]
[0072] where r(s l , a l ) is the immediate reward function for the state-action pair (s l , a l ); γ is the discount factor; p s is the state transition probability distribution of the environment; V π (s l+1 ) is the expected cumulative return when continuing to act according to the policy π in state s l+1 .
[0073] To improve the sampling efficiency, the SAC algorithm adopts a dual-network structure, including an actor network and two critic networks. The critic network is used to estimate the Q value by minimizing the temporal difference loss, and the loss function is expressed as:
[0074]
[0075] where is the target Q value, defined as:
[0076]
[0077] where R(s l , a l ) is the immediate reward function; Q ω (s l+1 , a l+1 ) is the Q value in the next state estimated by the current Q network.
[0078] In the policy improvement phase, SAC uses the policy gradient method to optimize the policy, and the update objective is to minimize the policy network loss function:
[0079]
[0080] where, is the state s sampled from the experience replay l ; π φ (a l | s l ) is the probability distribution of the previous policy taking action a l in state s l ; Q ω (sl , a l ) is the currently estimated state - action value function.
[0081] In the SAC algorithm, in each training iteration, a batch of samples is extracted from the experience replay buffer, the critic network is optimized to reduce the temporal difference error, and at the same time, the policy performance is improved by optimizing the actor network. In policy optimization, entropy regularization effectively enhances the randomness of action selection, improving the exploration efficiency and the stability of the algorithm.
[0082] Under the communication network model with the aerial RIS attitude change proposed in the present invention, the proposed SAC algorithm exhibits good performance. Compared with the traditional deep reinforcement learning algorithms PPO and DDPG, the SAC algorithm has a faster convergence speed and a larger reward value. In addition, due to its deterministic characteristics, DDPG usually has a faster and more stable convergence speed, but it lacks an inherent exploration mechanism, which will limit its performance in complex environments and thus is prone to falling into local optimal solutions. As Figure 3 shown.
[0083] Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not used to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, for those skilled in the art, they can still modify the technical solutions described in the foregoing examples, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the invention shall be included within the protection scope of the invention. All technical features in this embodiment can be freely combined according to actual needs.
[0084] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A communication network optimization method for the attitude change of an aerial RIS, characterized in that: It includes an aerial RIS-assisted communication network scenario, and the aerial RIS-assisted communication network scenario includes the communication environment between the base station and the ground users and the control environment of the UAV-RIS; The communication environment between the base station and the ground users includes a set of ground users and the base station; in this environment, the base station can provide communication services for multiple ground users at the same time; among them, the communication links between the base station and the users are divided into two types: direct link and reflected link; the direct link is that the signal is directly transmitted to the ground user through free space propagation; the reflected link is that the signal is sent from the base station and transmitted to the ground user after being reflected by the aerial RIS; The control environment of the UAV-RIS includes the UAV and the RIS array; in this control environment, the UAV platform serves as a carrier and provides an aerial signal reflection relay between the base station and the ground users through its flexible flight ability; during flight and hovering, affected by inertia and air resistance, the body attitude of the UAV, including the roll angle, pitch angle and yaw angle, will change dynamically, which will lead to the deviation of the RIS reflection angle and the signal incident angle, thus causing beam alignment error and channel gain fluctuation, directly affecting the communication performance of the reflected link; Specifically, it includes the following steps: Step 1, the UAV carrying the RIS hovers or flies over the ground users; the base station sends a communication signal to the aerial RIS through the downlink, and the RIS reflects and enhances the signal and adjusts the beam through real-time phase regulation to ensure that the reflected signal is accurately aligned with the direction of the target user; the optimized signal is transmitted to the ground user through the reflected link; Step 2, in order to cope with the influence of the UAV attitude change, including the roll angle, pitch angle and yaw angle, on the angles of the received and reflected signals, that is, on the system communication performance, the SAC deep reinforcement learning method is used to jointly optimize the UAV attitude and the RIS phase shift, so as to maximize the system communication rate; Step 3, in each time slot l, through the state analysis of the aerial RIS attitude and the UAV trajectory, the optimal strategy is determined, and the UAV performs the next action according to the obtained optimal strategy; The described aerial RIS-assisted wireless communication scenario includes K single-antenna user equipments, a multi-antenna base station and a UAV with an RIS mounted on its bottom; the entire flight period of the UAV is T, and for the convenience of processing, the total flight time T is equally divided into L time slots, and the length of each time slot is δ = T / L; the position of the UAV will change between adjacent time slots; a three-dimensional Cartesian coordinate system is used to model the flight state of the UAV, and the flight height of the UAV is set as a fixed value H; in the l-th time slot, the position coordinate of the UAV is expressed as q[l] = (x[l], y[l], H); RIS Angle Calculation and Attitude Transformation: In the local coordinate system, the initial unit normal vector of the airborne RIS is represented as e; First, an initial translation transformation is performed to translate the coordinate origin (0, 0, 0) to the instantaneous position of the UAV (x[l], y[l], H), thereby ensuring the accurate spatial positioning of the RIS in the global reference system; The attitude is adjusted by means of rotation transformation, and attitude correction is achieved through a series of rotation operations parameterized by Euler angles; The rotation transformation includes three consecutive operations: the roll angle of rotation around the x-axis, the pitch angle of rotation around the y-axis, and the yaw angle of rotation around the z-axis, denoted as φ[l], θ[l], The corresponding transformation matrices that control these rotations are represented as: Among them, R φ is the transformation matrix for controlling the roll angle, R θ is the transformation matrix for controlling the pitch angle, is the transformation matrix for controlling the yaw angle; The complete attitude transformation matrix is After applying this coordinate transformation, the unit normal vector of the airborne RIS can be derived in the global coordinate system as e ⊥ = R BE e, where e represents the initial normal vector of the airborne RIS, and the resulting e after multiplication ⊥ is the transformed unit normal vector; the unit direction vectors of the incident signal from the base station to the airborne RIS and the reflected signal from the airborne RIS to the ground user are expressed as: where and respectively represent the azimuth and elevation angles of the base station or user equipment to the airborne RIS at time slot l; by adopting these direction vectors, the angle between the normal vector of the airborne RIS plane and the incident / reflected signal can be deduced as: Communication model: At any time slot, the channel gains between the base station and the aerial RIS, and between the aerial RIS and the ground user are respectively denoted as H B,R [l] and h R,k [l]; Since the actual gain of the aerial RIS is significantly affected by the signal incident angle and the reflection angle, the actual gain model introduces an expression considering the azimuth angle and the pitch angle; the actual gain of the aerial RIS can be modeled as the product of the maximum directivity coefficient of the aerial RIS and the reflection angle characteristic function Among them, and respectively represent the receiving gain from the base station to the RIS in the air and the transmitting gain from the RIS in the air to the user equipment; and are respectively the normalized directional radiation functions in the direction from the base station to the RIS in the air and in the direction from the RIS in the air to the user equipment; To more accurately describe the gain characteristics of the RIS, the maximum directivity coefficient D max and the normalized directional radiation function F(ζ,η) are introduced; this function reflects the directional radiation characteristics of the RIS in the air and is modeled as an exponential-Lambert radiation model, and its expression is where ζ and η represent the azimuth angle and the pitch angle respectively, and are used to define the spatial relationship between the ground user and the aerial RIS; Based on these mathematical expressions, the RIS gain expression is further derived as follows: where Φ[l] is the RIS phase shift matrix, and this expression reflects the angle selection characteristic of the RIS for the reflected signal, that is, when the reflection angle meets certain conditions, the RIS can effectively reflect the signal, otherwise the signal cannot be reflected; Furthermore, in the air RIS-assisted wireless communication system, the signal received by the ground user in time slot l can be expressed as: where v k [l] is the composite channel from the base station to the ground user, including the direct link and the reflected link from the base station to the ground user; w k [l] is the beamforming vector; x k [l] is the transmitted signal; n k is additive white Gaussian noise and follows a complex Gaussian distribution, with a noise variance of σ 2 ; The communication rate of the ground user within time slot l is:
2. The communication network optimization method for the attitude change of the aerial RIS according to claim 1, wherein: In step 1, a three-dimensional Cartesian coordinate system is adopted for the environment, the UAV flies horizontally at a fixed altitude, and the ground users are randomly distributed.
3. The communication network optimization method for the attitude change of the aerial RIS according to claim 1, wherein: In step 2, the angles of the received and reflected signals will be represented by Euler angles, so as to represent the received and reflected beams, and the system communication rate is calculated; based on the SAC deep reinforcement learning method, the exploration efficiency is improved through the maximum entropy strategy, so that the exploration and exploitation can be balanced in the policy evaluation and policy improvement stages.
4. A communication network optimization method for the attitude change of an aerial RIS according to claim 1, characterized in that: In step 3, the optimal policy includes: adjusting the flight trajectory of the UAV and the attitude of the air RIS according to the current environment, so as to ensure the optimal signal incident angle and reflection angle, and then maximize the channel gain of the RIS and improve the communication performance of the system.
Citation Information
Patent Citations
Methods for near-field detection and beam optimization including reconfigurable intelligent surfaces
WO2025014817A1