An intelligent reflector assisted unmanned aerial vehicle emergency communication method, system, device and medium

By optimizing the UAV trajectory and reflector phase using the H-PPO algorithm, the problem of efficient emergency communication for UAVs in extreme environments is solved, ensuring service quality and system energy efficiency for ground users and adapting to complex and dynamic scenarios.

CN119052830BActive Publication Date: 2025-12-05XIAN MOLI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411272370.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2025-12-05
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

Existing technologies for equipping UAVs with intelligent reflective surface systems struggle to achieve efficient emergency communication in extreme environments, fail to guarantee service quality for all ground users, and traditional optimization methods cannot adapt to complex and dynamic scenarios, resulting in high energy consumption and insufficient communication performance of UAVs.

Method used

The Hierarchical Near-End Policy Optimization (H-PPO) algorithm, combined with deep reinforcement learning, is used to optimize the flight trajectory of the UAV and the phase of the intelligent reflector. By constructing a channel model and an energy consumption model, a reward function is designed to maximize system energy efficiency and ensure QoS for ground users.

Benefits of technology

It achieved efficient and persistent communication services in extreme environments, ensuring service quality for all ground users, while reducing the energy consumption of drones and adapting to complex and dynamic scenario requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119052830B_ABST
    Figure CN119052830B_ABST
Patent Text Reader

Abstract

An unmanned aerial vehicle emergency communication method, system, device and medium assisted by a smart reflecting surface, the method comprising: constructing an unmanned aerial vehicle carrying a smart reflecting surface assisted energy saving emergency communication scene; modeling the communication process of the UAV-RIS assisted multiple ground devices in the unmanned aerial vehicle carrying a smart reflecting surface assisted energy saving emergency communication scene; modeling the propulsion energy consumption and trajectory constraint of the unmanned aerial vehicle; modeling the optimization target of the final energy consumption ratio and the objective function when solving the model; constructing a deep reinforcement learning algorithm model according to the optimization target; constructing a deep reinforcement learning training model and setting a state space, an action space and a reward function; training the deep reinforcement learning training model to obtain an unmanned aerial vehicle carrying a smart reflecting surface assisted energy saving emergency communication decision model and obtain the optimal solution of the optimization problem; the system, device and medium are used to realize the method; the present application has the advantages of being able to provide more flexible and more persistent communication services in a harsh emergency scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) communication technology, and specifically relates to an emergency communication method, system, device and medium for UAVs assisted by an intelligent reflective surface. Background Technology

[0002] The increasing frequency and intensity of natural disasters pose significant challenges to effective communication in extreme and dynamic environments. Timely and reliable data transmission from ground devices (GDs) to central processing units is crucial for effective disaster detection and response. Leveraging the flight flexibility of unmanned aerial vehicles (UAVs), UAV communication systems can generally provide line-of-sight (LAS) communication over wide terrains, thus enabling UAVs to be used as aerial communication platforms to enhance the performance of ground communication systems. Innovations in UAV technology have been effectively utilized in disaster relief. UAVs can assist ground devices in disaster perception, offering advantages such as flexibility, wide coverage, and LAS communication.

[0003] The combination of reconfigurable smart surfaces (RIS) and unmanned aerial vehicles (UAVs) has recently attracted widespread attention. When a RIS is mounted on a UAV, it can act as a mobile relay. This innovative combination allows for precise manipulation of electromagnetic wave propagation and provides a new dimension for improving the performance of communication systems. For example, existing work has simultaneously adjusted the position of the UAV in a smart reflector system and the reflectivity of the smart reflector, ultimately achieving maximum downlink transmission capacity. Furthermore, in UAV communication, the propulsion energy consumption of the UAV is crucial for maintaining a sustainable emergency communication system. Most existing work employs traditional optimization methods to optimize the UAV's trajectory; however, traditional optimization methods face significant challenges in real-time decision-making regarding the UAV's flight trajectory and the phase of the RIS. Deep reinforcement learning (DRL) offers a feasible solution to these challenges. Existing work using reinforcement learning for optimization mostly employs discrete DRL algorithms to optimize the UAV's trajectory and the phase shift of the RIS. However, the discrete action algorithms used to process trajectory and phase limit the design space, negatively impacting the full utilization of its potential performance.

[0004] Furthermore, current research has not addressed the issues of simultaneously optimizing energy efficiency and ensuring quality of service (QoS) for individual users in scenarios requiring long decision sequences when integrating UAVs with intelligent reflectors. Resolving competition among users for reconfigurable intelligent surface (RIS) resources is crucial for guaranteeing superior service to individual users, a problem often overlooked in research. For scenarios where UAV-integrated intelligent reflector systems assist multiple users in emergency communication, how can the need for rapid decision-making be met? Considering the mobility of ground equipment, how can the UAV trajectory be designed to maximize system energy efficiency? Considering the competition among all ground equipment for reconfigurable intelligent surface (RIS) resources, how can the phase of the intelligent reflector be designed to guarantee QoS for all ground users?

[0005] In summary, the shortcomings of existing technologies are as follows:

[0006] (1) Existing technologies only consider the communication problem of UAVs equipped with intelligent reflective surface networks. However, the propulsion energy consumption of UAVs is very important for UAVs. UAV systems that only consider communication performance are difficult to adapt to emergency communication in extreme environments.

[0007] (2) Existing technologies do not take into account the competition for reconfigurable smart surface (RIS) resources among all ground devices, and cannot guarantee the individual service quality of each ground user.

[0008] (3) Existing optimization methods are difficult to adapt to complex scenarios and cannot solve the proposed joint trajectory and phase design optimization problem.

[0009] Patent application CN113873575B discloses an energy-saving optimization method for non-orthogonal multiple access (NOMA) UAV air-to-ground communication networks assisted by a smart reflector. This method effectively reduces the UAV's transmission power, improves its endurance, and saves energy. However, its reconfigurable smart surface (RIS) is placed next to distant NOMA users, resulting in poor mobility of the RIS relay, which cannot meet the needs of mobile scenarios.

[0010] Patent application CN117596616A discloses a method for optimizing the energy efficiency of an uplink NOMA UAV network with intelligent reflector assistance. This method aims to maximize the energy efficiency of the intelligent reflector-assisted uplink NOMA UAV network. However, its optimization method is traditional, employing alternating optimization to address two sub-problems. Due to the high complexity of the optimization problem and the dynamic nature of the scenario, it is impossible to derive a low-complexity optimal solution using mathematical methods, thus hindering the implementation of dynamic and rapid service decisions for multiple users. Summary of the Invention

[0011] To overcome the shortcomings of the prior art, the present invention aims to provide an emergency communication method for unmanned aerial vehicles (UAVs) assisted by a smart reflector system (UAV-RIS). This method utilizes a smart reflector system carried by the UAV to assist uplink communication in emergency scenarios. Furthermore, a hierarchical near-end policy optimization (H-PPO) algorithm is proposed to simultaneously optimize the UAV's flight trajectory and the phase of the smart reflector. This solves the complex joint optimization problem of UAV trajectory and smart reflector phase shift, achieving the goal of improving system energy efficiency while effectively ensuring the Quality of Service (QoS) requirements of all ground equipment. The present invention also provides a system, equipment, and medium for implementing the above method.

[0012] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0013] A smart reflective surface-assisted emergency communication method for unmanned aerial vehicles (UAVs) includes the following steps:

[0014] Step 1: Construct a scenario for energy-saving emergency communication using a drone equipped with an intelligent reflective surface;

[0015] Step 2: Model the communication process of multiple ground devices assisted by the UAV-Intelligent Reflector System (UAV-RIS) in the UAV-equipped intelligent reflector-assisted energy-saving emergency communication scenario constructed in Step 1.

[0016] Step 3: Model the propulsion energy consumption of the UAV;

[0017] Step 4: Combining the UAV-carrying intelligent reflector-assisted energy-saving emergency communication scenario in Step 2, which describes the communication process of multiple ground devices assisted by the UAV-carrying intelligent reflector system (UAV-RIS), and the UAV propulsion energy consumption model in Step 3, we model the optimization target of the final energy consumption ratio and construct the objective function for solving the model.

[0018] Step 5: Construct a deep reinforcement learning algorithm model based on the optimization objective proposed in Step 4;

[0019] Step 6: Construct a deep reinforcement learning training model based on the deep reinforcement learning algorithm model in Step 5. Combine the UAV equipped with intelligent reflective surface to assist energy-saving emergency communication scenario in Step 1 and the objective function in Step 4 to set the state space, action space and reward function of the deep reinforcement learning training model.

[0020] Step 7: Train the deep reinforcement learning training model obtained in Step 6 to obtain the UAV-equipped intelligent reflective surface-assisted energy-saving emergency communication decision model, and obtain the optimal solution to the optimization problem.

[0021] The process of step 1 is as follows:

[0022] In an energy-saving emergency communication scenario, a drone equipped with a smart reflector system (UAV-RIS) is used. The UAV carries an active reconfigurable smart surface (RIS), and an airborne base station (ABS) processes information transmitted from ground equipment. The ABS acts as a control center, interacting with the UAV-RIS and issuing commands. The UAV-RIS assists in uplink transmission between multiple ground devices and the ABS. The set of ground devices is denoted as g = {G}. m Let m = 1, 2, ..., M, where M is the number of ground devices. It is assumed that there is no direct communication link between the ground devices and the airborne base station (ABS), and the airborne base station (ABS) is located at a high position.

[0023] The process of step 2 is as follows:

[0024] Step 2.1, construct the channel model;

[0025] In the channel model, each ground device has an omnidirectional antenna, while the airborne base station (ABS) is equipped with a Q antenna array. The active reconfigurable smart surface (RIS) consists of N reflective elements. Assume the reflection coefficient of the nth element of the RIS is... Where, φ n ∈[0,2π), β represents an amplification factor greater than 1, and the reflection coefficient matrix of an active reconfigurable smart surface (RIS) is defined as Θ=diag([θ1,θ2,…,θ)). N ]);

[0026] H m,r H represents the channel of the link between the m-th ground device and the UAV-borne Intelligent Reflector System (UAV-RIS). r,b Let represent the channel between the UAV-carrying Intelligent Reflective Surface System (UAV-RIS) and the Airborne Base Station (ABS). Assuming the channel gain from the UAV-carrying Intelligent Reflective Surface System (UAV-RIS) to the Airborne Base Station (ABS) follows a Ricean distribution, then the channel is represented as:

[0027]

[0028] Among them, κ r,b Here, d is the Rice coefficient, l is the path loss at the reference distance D = 1m, and d is the path loss at the reference distance D = 1m. r,b The distance between the UAV carrying the Intelligent Reflective Surface System (UAV-RIS) and the Airborne Base Station (ABS), α r,b The path loss index for the UAV-RIS-ABS link; line-of-sight component. For the first level, it is represented as:

[0029]

[0030] in, The departure angle (AoD) and arrival angle (AoA) of the link from the UAV carrying the Intelligent Reflector System (UAV-RIS) to the Airborne Base Station (ABS), where ω is the antenna spacing;

[0031] Non-line-of-sight components Each element follows an independent and identically distributed complex Gaussian distribution with a mean of zero and a variance of 1; similarly, the channel H m,r It also follows a similar distribution as described above, with the Loss component of the m-th link... Represented as:

[0032]

[0033] in, Let AoD be the link departure angle (AoD) from the m-th ground device to the UAV-RIS. The arrival angle (AoA) and departure angle (AoD) are determined by their relative positional relationship.

[0034] Finally, it is assumed that all channels follow block fading, global channel information is known at the airborne base station (ABS) and remains constant in each time slot, but changes from one time slot to another;

[0035] Step 2.2: Based on the channel model constructed in Step 2.1, construct the transmission model;

[0036] In the uplink transmission model from ground equipment to the airborne base station (ABS), in-band interference is considered. The ABS uses maximum proportional combining (MRC) for each link, and the MRC beamforming matrix is ​​F = [f1, ..., f M ] indicates that, among which, f m The unit norm beamforming vector of the m-th link is represented as: The signal received by the airborne base station (ABS) from the m-th ground device is represented as:

[0037]

[0038] Among them, P m Let s be the transmission power of the m-th ground device. m A unit energy signal sample associated with monitoring data; noise vector n generated at the airborne base station (ABS). m , represented as n m =[n1,…,n M ] T ,in

[0039] Active reconfigurable smart surfaces (RIS) require additional power, and each component is equipped with an amplifier, while considering the thermal noise generated at the UAV-borne smart reflective surface system (UAV-RIS). i ,in, The uplink signal-to-noise ratio (SINR) of the m-th link of the airborne base station (ABS) is expressed as:

[0040]

[0041] Therefore, the rate of the m-th link is C. m =Blog(1+η) m ), where B is the bandwidth.

[0042] The process of step 3 is as follows:

[0043] The unmanned aerial vehicle (UAV) carrying a Smart Reflector System (UAV-RIS) and an Airborne Base Station (ABS) uses [the following technology] at a horizontal position in time slot t. and It means that, among them, The vertical altitude of the UAV carrying the Intelligent Reflective Surface System (UAV-RIS) and the Airborne Base Station (ABS) is set to h respectively. r and h b At a fixed altitude, the UAV carrying the Intelligent Reflective Surface System (UAV-RIS) will be positioned horizontally in the next time slot t by a flight distance D. t and azimuth variable ξ t If this is determined, then the horizontal coordinates of the UAV carrying the Intelligent Reflector System (UAV-RIS) at time t+1 are represented as follows:

[0044] Based on the given horizontal position setting of the UAV-RIS (Unmanned Aerial Vehicle-Reflector System), the horizontal flight speed of the UAV-RIS at time t is... For the maximum horizontal velocity, Δt is the length of each time interval; if This indicates that at time t, the drone is in a hovering state; based on the drone's horizontal speed, the propulsion energy consumption of the drone in each time interval is calculated as follows:

[0045]

[0046] Where P0 and P1 are the constant power and induced power of the blade in the hovering state, respectively, and U tip Δt is the tip velocity of the rotor blade, v0 is the average rotor induced velocity during hovering, d0 and s are the fuselage drag ratio and rotor firmness, respectively, and ρ and G represent the air density and rotor disk area, respectively.

[0047] The process of step 4 is as follows:

[0048] The reflection coefficient matrix of the intelligent reflector at the t-th time slot is denoted as Θ. t C t,m The communication rate of the m-th link at the t-th time slot obtained in step 2 is... The objective function for maximizing energy efficiency over a time span of length T, given the UAV propulsion energy consumption in the t-th time slot obtained in step 3, is expressed as follows:

[0049]

[0050] Wherein, constraint (1) represents the transmission rate requirement that each ground device needs to meet, and Υ represents the rate limit value; constraint (2) limits the range that the phase of the reconfigurable intelligent surface (RIS) can be adjusted; constraint (3) represents the limit on the maximum flight speed of the UAV carrying the intelligent reflective surface system (UAV-RIS).

[0051] The process of step 5 is as follows:

[0052] The deep reinforcement learning algorithm model is constructed by using the H-PPO algorithm to control the flight trajectory and reconfigurable smart surface (RIS) phase of the UAV carrying a smart reflector system (UAV-RIS). The H-PPO algorithm includes two independent proximal policy optimization algorithms (PPO). The first PPO is used to optimize flight maneuvers for trajectory control. The second PPO is used to optimize the reconfigurable smart surface (RIS) phase value to enhance channel gain. The independent PPO algorithms are used to optimize the control of the flight trajectory and the reconfigurable smart surface (RIS) phase.

[0053] The Proximal Policy Optimization (PPO) algorithm is a deep reinforcement learning (DRL) optimization algorithm that learns the optimal policy by iteratively updating the policy to maximize the expected cumulative reward. The policy of PPO is parameterized and represented by two networks: the actor network and the inverse policy network. and the network of critics Where, θ A and θ C This represents the corresponding parameter instance; the Proximal Policy Optimization (PPO) algorithm uses a surrogate objective function and a shearing mechanism to constrain policy updates; the objective function in the Proximal Policy Optimization (PPO) algorithm is:

[0054]

[0055] in, Representation strategy Compared to the old strategy Important sampling weights between them; for The corresponding advantage function, where γ is the discount factor; the pruning mechanism restricts strategy updates to a predefined range, where clip() represents the clipping function, and λ is a hyperparameter used to adjust ρ. t (θ A The range is controlled within the interval [1-ε, 1+ε].

[0056] During the training phase, the network is updated using the experience of the first batch of interactions with the environment; for the actor network... The parameter update method is as follows:

[0057]

[0058] At the same time, in order to update θ C Using the mean squared error function as the loss function, it is expressed as:

[0059]

[0060] Where, α C Indicates the learning rate. This represents the target state value function derived in a time-difference manner.

[0061] Step 6 sets the state space, action space, and reward function as follows:

[0062] 1) State Space: At time slot t, the state of the first near-end policy optimization algorithm (PPO) contains global position information, represented as follows: This represents the horizontal coordinate of the m-th ground device at time slot t; furthermore, the state of the second near-end policy optimization algorithm (PPO) includes global channel information, represented as... Before inputting this complex channel state into the network, separate the real and imaginary parts of the channel and input them into the neural network simultaneously.

[0063] 2) Action Space: At time slot t, the action of the first proximal policy optimization algorithm PPO (upper layer) is the flight distance and azimuth of the UAV carrying the Intelligent Reflector System (UAV-RIS), represented as... The action of the second proximal policy optimization algorithm PPO (lower layer) in time slot t is the phase value of the active RIS, denoted as Θ. t ;

[0064] 3) Reward Function: The reward structure of the first proximal policy optimization algorithm, PPO, consists of two parts. First, the real-time reward includes the ratio of communication traffic to energy consumption in the current time slot, and the difference between the horizontal distance between the UAV-carrying Intelligent Reflector System (UAV-RIS) and all ground equipment in the previous time slot, expressed as: Among them, P t,m This represents the additional positive reward value for the m-th ground device when the rate threshold is met in time slot t. Secondly, a termination reward reflecting energy efficiency throughout the entire service cycle is set, represented as... λ represents a positive fixed value used to adjust the ratio of the termination reward to the final reward; furthermore, the reward function of the first proximal policy optimization algorithm, PPO, is expressed as:

[0065]

[0066] Finally, in time slot t, the reward function of the second proximal policy optimization algorithm PPO is expressed as:

[0067]

[0068] The process of step 7 is as follows:

[0069] Step 7.1, Initialize the network

[0070] Randomly initialize the actor network θ for each Proximal Policy Optimization (PPO) algorithm. A and critics' network θ C Parameters;

[0071] Step 7.2, Training the deep reinforcement learning model

[0072] Randomly initialize the location of each ground device, the location of the UAV-RIS (Unmanned Aerial Vehicle-Reflective Surface System), and the location of the airborne base station (ABS);

[0073] In each time slot t, the first proximal policy optimization algorithm (PPO) interacts with the dynamic environment to obtain the state. Based on the current state, the first proximal policy optimization algorithm, PPO, obtains the drone's flight actions from its Actor network. To control the trajectory of the drone;

[0074] The second proximal policy optimization algorithm, PPO, interacts with the dynamic environment to obtain the state. Based on the current state, the second proximal policy optimization algorithm, PPO, obtains the action of the RIS phase from its Actor network. To control the intelligent reflective surface; and finally take action. Interact with the environment and calculate the reward r obtained from the environment. t ;

[0075] State state action action and reward rt Stored in the experience pool;

[0076] When the experience pool accumulates to a certain amount, N is taken from it. d The parameters of the Critic and Actor networks are updated using samples of varying sizes. After each training session, the experience pool is cleared. This process is repeated until the model training converges, resulting in a UAV-equipped intelligent reflector-assisted energy-saving emergency communication decision-making model.

[0077] Step 7.3, Decision-making stage

[0078] The decision model for energy-saving emergency communication with UAVs equipped with intelligent reflectors is applied to stochastic dynamic UAV-equipped intelligent reflector-assisted energy-saving emergency communication. In each time slot, the optimal reconfigurable intelligent surface (RIS) reflectivity matrix and UAV flight actions are determined to minimize the energy consumption ratio of the communication system throughout the process, while ensuring the communication quality of all individual ground users.

[0079] The present invention also provides an intelligent reflective surface-assisted UAV emergency communication system, comprising:

[0080] The UAV emergency communication scenario construction module is used to construct emergency communication scenarios for UAVs equipped with intelligent reflective surfaces to assist in energy saving.

[0081] The communication process establishment module is used to model the communication process of multiple ground devices carried by UAVs with intelligent reflective surface assisted energy-saving emergency communication scenarios.

[0082] The UAV propulsion energy consumption modeling module is used to model the propulsion energy consumption of UAVs.

[0083] The optimization objective and objective function construction module is used to model the optimization objective of the final energy consumption ratio by combining the communication process of UAV-carrying intelligent reflector system (UAV-RIS) assisting multiple ground devices in the UAV-equipped intelligent reflector system-assisted energy-saving emergency communication scenario and the UAV propulsion energy consumption model, and to construct the objective function for solving the model.

[0084] The deep reinforcement learning algorithm model building module is used to build a deep reinforcement learning algorithm model based on the proposed optimization objective.

[0085] The module for setting the state space, action space, and reward function is used to construct a deep reinforcement learning training model based on the deep reinforcement learning algorithm model. It combines the scenario of intelligent reflective surface-assisted energy-saving emergency communication on UAVs and the objective function to set the state space, action space, and reward function of the deep reinforcement learning training model.

[0086] The deep reinforcement learning algorithm model training module is used to train the deep reinforcement learning training model to obtain the UAV-equipped intelligent reflective surface-assisted energy-saving emergency communication decision model, and obtain the optimal solution to the optimization problem.

[0087] The present invention also provides an intelligent reflective surface-assisted emergency communication device for unmanned aerial vehicles, comprising:

[0088] Memory: A computer program that stores the above-mentioned intelligent reflective surface-assisted UAV emergency communication method, and is a computer-readable device;

[0089] Processor: Used to implement the intelligent reflective surface-assisted UAV emergency communication method when executing the computer program.

[0090] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the aforementioned intelligent reflective surface-assisted UAV emergency communication method.

[0091] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0092] 1. Existing traditional optimization methods cannot mathematically deduce low-complexity optimal solutions for highly complex and dynamic scenarios. Furthermore, most studies employ discrete DRL algorithms to optimize the trajectory of UAVs and the phase shift of reconfigurable intelligent surfaces (RIS). However, discrete motion algorithms used to handle trajectory and phase limit the design space and negatively impact performance. This invention employs a deep reinforcement learning algorithm to solve continuous optimization problems, simultaneously optimizing the UAV's flight trajectory and the phase of the intelligent reflector. This approach adapts to real-time decision-making requirements in dynamic scenarios and solves the complex optimization problems of trajectory and phase shift of joint unmanned aerial vehicles (UAVs) and reconfigurable intelligent surfaces (RIS).

[0093] 2. Currently, there is no research on UAVs equipped with intelligent reflectors that simultaneously solves the problems of optimizing energy efficiency and ensuring the quality of service for individual users in scenarios requiring long-term decision sequences. The H-PPO algorithm proposed in this invention effectively solves the joint trajectory and phase design, and through the effective design of the reward function of the H-PPO algorithm, it can maximize the energy efficiency of the system while ensuring the quality of service for each ground user.

[0094] 3. This invention utilizes a drone carrying a smart reflector for auxiliary emergency communication. Leveraging the high mobility of drones, the smart reflector, acting as a mobile relay, can be deployed over a wider range. This enables more flexible and longer-lasting communication services in harsh emergency scenarios.

[0095] In summary, this invention utilizes a UAV carrying a smart reflector for auxiliary emergency communication and proposes the H-PPO algorithm to simultaneously optimize the UAV's flight trajectory and the phase of the smart reflector. By considering the service quality of individual ground users in the reward function, it solves the complex optimization problem of the trajectory of the joint UAV (UAV) and the phase shift of the reconfigurable smart surface (RIS). This achieves the goal of maximizing the system's energy efficiency while ensuring the individual service quality of each ground user, and has the advantage of providing more flexible and durable communication services in harsh emergency scenarios. Attached Figure Description

[0096] Figure 1 This is a flowchart of the implementation method of the present invention.

[0097] Figure 2 This is a schematic diagram of a scenario where a drone is equipped with an intelligent reflective surface to assist in energy-saving emergency communication network, as provided in an embodiment of the present invention.

[0098] Figures 3(a)-3(d) Figure 3(a) is a schematic diagram of the convergence speed based on the H-PPO algorithm in an embodiment of the present invention. Figure 3(b) is a graph showing the trend of the reward value of the first proximal policy optimization algorithm (PPO) in the H-PPO algorithm as training progresses. Figure 3(c) is a graph showing the trend of the energy consumption ratio of the system during the training process of the H-PPO algorithm. Figure 3(d) is a graph showing the trend of the QoS completion rate of the ground equipment during the training process of the H-PPO algorithm.

[0099] Figure 4 The diagram illustrates the energy efficiency of different schemes provided in the embodiments of the present invention under different QoS thresholds.

[0100] Figure 5 This diagram illustrates the average QoS completion rate of different schemes provided in this invention under different QoS thresholds. Detailed Implementation

[0101] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0102] This invention studies the joint trajectory and phase design problem in UAV-equipped intelligent reflector-assisted energy-saving emergency communication. A corresponding non-convex optimization problem is proposed and solved using a deep reinforcement learning-based algorithm. This approach allows for simultaneous trajectory planning for the UAV and phase design of the intelligent reflector, enabling emergency communication within a UAV-equipped intelligent reflector communication network. The model considers UAV motion characteristics, Quality of Service (QoS) for all ground users, energy efficiency of the UAV-equipped intelligent reflector system, and the need for rapid decision-making. Simulation results demonstrate that the proposed method maximizes the energy efficiency of the UAV-equipped intelligent reflector system while ensuring QoS for all ground users.

[0103] This invention is implemented as follows: an emergency communication method for unmanned aerial vehicles (UAVs) assisted by an intelligent reflective surface. This method plans the flight trajectory of the UAV and designs the phase of the intelligent reflective surface, enabling the intelligent reflective surface mounted on the UAV to act as a mobile relay to provide emergency communication services to ground users.

[0104] First, a scenario model is created for an emergency network assisted by a UAV equipped with an intelligent reflector, using a wildfire scenario as an example. This includes a UAV carrying an intelligent reflector system (UAV-RIS) and M ground devices (GDs) for fire detection, and an airborne base station (ABS) responsible for processing information transmitted from the ground devices. Due to the adverse conditions of a wildfire environment, this invention assumes no direct connection between the ground devices and the airborne base station (ABS), and the ABS is positioned at a high location to ensure line-of-sight connectivity via the cascaded links of the RIS. The communication channel is considered to be Ricean distribution, with both line-of-sight and non-line-of-sight components. Each GD has an omnidirectional antenna, while the ABS is equipped with a linear array of Q antennas. The active RIS consists of N reflective elements. The altitude of the UAV-RIS and the ABS is considered constant, and the horizontal speed of the UAV-RIS is V. h The entire flight process is divided into T time slots.

[0105] Secondly, based on the cascaded channel of active RIS-assisted ground equipment communication, the uplink transmission rate of each user is derived. Then, based on the UAV propulsion energy consumption model, the energy consumption of the UAV-carrying intelligent reflector system (UAV-RIS) in each time slot is derived. Finally, these are integrated into an optimization problem of joint trajectory and reflector phase in the UAV-RIS communication network. The optimization objective is to maximize energy efficiency after a long decision period while ensuring the service quality of individual ground equipment. This is a multivariate non-convex optimization problem, and a Hierarchical Proximal Policy Optimization (H-PPO) algorithm based on deep reinforcement learning is proposed for solving it. This algorithm adopts a two-layer proximal policy optimization algorithm (PPO) structure. The upper and lower layers of H-PPO optimize the UAV trajectory and reflector phase control, respectively. It selects the optimal motion route for the UAV and simultaneously designs the phase of the intelligent reflector to improve the service quality. This invention maximizes the energy efficiency of the UAV-mounted intelligent reflector system while ensuring the Quality of Service (QoS) for all ground users. The proposed optimization algorithm has advantages such as high reliability and high real-time performance, meeting the requirements of emergency networks for latency and accuracy.

[0106] A smart reflective surface-assisted emergency communication method for unmanned aerial vehicles (UAVs) includes the following steps:

[0107] Step 1: Construct a scenario for energy-saving emergency communication using a drone equipped with an intelligent reflective surface;

[0108] The process of step 1 is as follows:

[0109] This invention considers a key scenario of UAV-equipped intelligent reflector-assisted energy-saving emergency communication, taking a wildfire monitoring scenario as an example. The specific scenario includes a UAV equipped with an active reconfigurable intelligent surface (RIS), i.e., a UAV-carrying intelligent reflector system (UAV-RIS), multiple ground devices for fire detection, and an airborne base station (ABS) for processing information transmitted from the ground devices. The airborne base station (ABS) acts as a control center, responsible for interacting with the UAV-carrying intelligent reflector system (UAV-RIS) and issuing commands. The set of ground devices is denoted as g = {G}. m Let m = 1, 2, ..., M, where M is the number of ground devices. Due to the adverse conditions of the wildfire environment, it is assumed that there is no direct communication link between the ground devices and the airborne base station (ABS), and the airborne base station (ABS) is located at a high position to ensure line-of-sight connection of the cascaded links.

[0110] Step 2: Model the communication process of multiple ground devices assisted by the UAV-Intelligent Reflector System (UAV-RIS) in the UAV-equipped intelligent reflector-assisted energy-saving emergency communication scenario constructed in Step 1.

[0111] Furthermore, the process of step 2 is as follows:

[0112] Step 2.1, construct the channel model;

[0113] In the channel model, each ground device has an omnidirectional antenna, while the airborne base station (ABS) is equipped with a Q antenna array. The active reconfigurable smart surface (RIS) consists of N reflective elements. Assume the reflection coefficient of the nth element of the RIS is... Where, φ n ∈[0,2π), β represents an amplification factor greater than 1, and the reflection coefficient matrix of an active reconfigurable smart surface (RIS) is defined as Θ=diag([θ1,θ2,…,θ)). N ]);

[0114] H m,r H represents the channel of the link between the m-th ground device and the UAV-borne Intelligent Reflector System (UAV-RIS). r,b Let represent the channel between the UAV-carrying Intelligent Reflective Surface System (UAV-RIS) and the Airborne Base Station (ABS). Assuming the channel gain from the UAV-carrying Intelligent Reflective Surface System (UAV-RIS) to the Airborne Base Station (ABS) follows a Ricean distribution, then the channel is represented as:

[0115]

[0116] Among them, κ r,b Here, d is the Rice coefficient, l is the path loss at the reference distance D = 1m, and d is the path loss at the reference distance D = 1m. r,b The distance between the UAV carrying the Intelligent Reflective Surface System (UAV-RIS) and the Airborne Base Station (ABS), α r,b The path loss index for the UAV-RIS-ABS link; line-of-sight component. For the first level, it is represented as:

[0117]

[0118] in, The departure angle (AoD) and arrival angle (AoA) of the link from the UAV carrying the Intelligent Reflector System (UAV-RIS) to the Airborne Base Station (ABS), where ω is the antenna spacing;

[0119] Non-line-of-sight components Each element follows an independent and identically distributed complex Gaussian distribution with a mean of zero and a variance of 1; similarly, the channel Hm,r It also follows a similar distribution as described above, with the Loss component of the m-th link... Represented as:

[0120]

[0121] in, Let AoD be the link departure angle (AoD) from the m-th ground device to the UAV-RIS. The arrival angle (AoA) and departure angle (AoD) are determined by their relative positional relationship.

[0122] Finally, it is assumed that all channels follow block fading, global channel information is known at the airborne base station (ABS) and remains constant in each time slot, but changes from one time slot to another.

[0123] Step 2.2: Based on the channel model constructed in Step 2.1, construct the transmission model;

[0124] In the uplink transmission model from ground equipment to the airborne base station (ABS), in-band interference is considered. The ABS uses maximum proportional combining (MRC) for each link, and the MRC beamforming matrix is ​​F = [f1, ..., f M ] indicates that, among which, f m The unit norm beamforming vector of the m-th link is represented as: The signal received by the airborne base station (ABS) from the m-th ground device is represented as:

[0125]

[0126] Among them, P m Let s be the transmission power of the m-th ground device. m A unit energy signal sample associated with monitoring data; noise vector n generated at the airborne base station (ABS). m , represented as n m =[n1,…,n M ] T ,in

[0127] Furthermore, unlike passive reconfigurable smart surfaces (RIS), active reconfigurable smart surfaces (RIS) require additional power, and because each component is equipped with an amplifier, the thermal noise generated at the UAV-borne smart reflector system (UAV-RIS) cannot be ignored. i ,in, Therefore, the uplink signal-to-noise ratio (SINR) of the m-th link of the airborne base station (ABS) is expressed as:

[0128]

[0129] Therefore, the rate of the m-th link is C. m =Blog(1+η) m ), where B is the bandwidth.

[0130] Step 3: Model the propulsion energy consumption of the UAV; this lays the foundation for the joint optimization problem proposed in the subsequent part of this invention.

[0131] The process of step 3 is as follows:

[0132] The UAV carrying a Smart Reflector System (UAV-RIS) and an Airborne Base Station (ABS) is available at a horizontal position in time slot t. and It means that, among them, The vertical altitude of the UAV carrying the Intelligent Reflective Surface System (UAV-RIS) and the Airborne Base Station (ABS) is set to h respectively. r and h b At a fixed altitude, the UAV carrying the Intelligent Reflective Surface System (UAV-RIS) will be positioned horizontally in the next time slot t by a flight distance D. t and azimuth variable ξ t If this is determined, then the horizontal coordinates of the UAV carrying the Intelligent Reflector System (UAV-RIS) at time t+1 are represented as follows:

[0133] Based on the horizontal position settings given above for the UAV-RIS (Unmanned Aerial Vehicle) carrying the Intelligent Reflective Surface System, the horizontal flight speed of the UAV-RIS at time t is... For the maximum horizontal velocity, Δt is the length of each time interval; if This indicates that at time t, the drone is in a hovering state; based on the drone's horizontal speed, the propulsion energy consumption of the drone in each time interval is calculated as follows:

[0134]

[0135] Where P0 and P1 are the constant power and induced power of the blade in the hovering state, respectively, and U tip Δt is the tip velocity of the rotor blade, v0 is the average rotor induced velocity during hovering, d0 and s are the fuselage drag ratio and rotor firmness, respectively, and ρ and G represent the air density and rotor disk area, respectively.

[0136] Step 4: Combining the UAV-equipped intelligent reflector-assisted energy-saving emergency communication scenario in Step 2, where the UAV carries an intelligent reflector system (UAV-RIS) to assist communication with multiple ground devices, and the UAV propulsion energy consumption model in Step 3, the optimization objective of the final energy consumption ratio of the present invention is modeled, and the objective function for solving the model is constructed; this lays the foundation for the subsequent use of deep reinforcement learning to solve the model in the present invention.

[0137] The process of step 4 is as follows:

[0138] The objective of this invention is to maximize the energy efficiency of an UAV-borne Intelligent Reflector System (UAV-RIS)-assisted emergency communication system by optimizing the UAV's trajectory and the phase of the intelligent reflector in each service cycle. Simultaneously, it aims to ensure the minimum communication rate requirements of all ground equipment in dynamic scenarios. The reflection coefficient matrix of the intelligent reflector at the t-th time slot is denoted as Θ. t C t,m The communication rate of the m-th link at the t-th time slot obtained in step 2 is... Given the UAV propulsion energy consumption in the t-th time slot obtained in step 3, the objective function to maximize energy efficiency over a time range of length T can be expressed as:

[0139]

[0140] Wherein, constraint (1) represents the transmission rate requirement that each ground device needs to meet, and Υ represents the rate limit value; constraint (2) limits the range that the phase of the reconfigurable intelligent surface (RIS) can be adjusted; constraint (3) represents the limit on the maximum flight speed of the UAV carrying the intelligent reflective surface system (UAV-RIS).

[0141] Step 5: Construct a deep reinforcement learning algorithm model based on the optimization objective proposed in Step 4; lay a theoretical foundation for the practical problem to be solved and reduce the difficulty of solving the optimization problem;

[0142] The process of step 5 is as follows:

[0143] To address the challenges of optimizing the flight trajectory and RIS phase of a UAV-RIS (Unmanned Aerial Vehicle-RIS) system, the H-PPO algorithm is proposed. The deep reinforcement learning algorithm model utilizes the H-PPO algorithm to control the flight trajectory and reconfigurable intelligent surface (RIS) phase of the UAV-RIS system. The H-PPO algorithm comprises two independent proximal policy optimization algorithms (PPOs): the first PPO optimizes flight maneuvers for trajectory control; the second PPO optimizes the RIS phase value to enhance channel gain. Since the state information required for flight trajectory and RIS phase control is independent, the independent PPOs are used to optimize the control of both the flight trajectory and the RIS phase.

[0144] The first proximal policy optimization algorithm (PPO) collects data including flight state, actions, and rewards to improve its policy in pursuit of the optimal trajectory. The second PPO algorithm specifically optimizes the reconfigurable smart surface (RIS) phase, utilizing channel state, phase actions, and reward functions to continuously update its policy and improve performance. Through decoupled optimization, it effectively utilizes the state information of each action, avoiding instability caused by interactions. This method aims to simplify coordination and optimization challenges and improve the overall performance of the UAV-RIS system.

[0145] The Proximal Policy Optimization (PPO) algorithm is a deep reinforcement learning (DRL) optimization algorithm that learns the optimal policy by iteratively updating the policy to maximize the expected cumulative reward. The policy of PPO is parameterized and represented by two networks: the actor network and the inverse policy network. and the network of critics Actor Network The role of the network is to output predicted actions based on different state information, while the role of the value network is to evaluate the quality of the current state, and thus evaluate the quality of the decisions made by the actor network. Wherein, θ A and θ C This represents the corresponding parameter instance; the core idea of ​​the Proximal Policy Optimization (PPO) algorithm is to maximize the objective function and use a shearing mechanism to constrain policy updates; the objective function in the Proximal Policy Optimization (PPO) algorithm is:

[0146]

[0147] in, Representation strategy Compared to the old strategy Important sampling weights between them; for The corresponding advantage function, where γ is the discount factor; the pruning mechanism restricts strategy updates to a predefined range, where clip() represents the clipping function, and λ is a hyperparameter used to adjust ρ. t (θ A The range is controlled within the interval [1-ε, 1+ε].

[0148] During the training phase, the network is updated using the experience of the first batch of interactions with the environment; for the actor network... The parameter update method is as follows:

[0149]

[0150] At the same time, in order to update θ C Using the mean squared error function as the loss function, it is expressed as:

[0151]

[0152] Where, α C Indicates the learning rate. This represents the target state value function derived in a time-difference manner.

[0153] Step 6: Construct a deep reinforcement learning training model based on the deep reinforcement learning algorithm model in Step 5. Combine the UAV equipped with intelligent reflective surface assisted energy-saving emergency communication scenario in Step 1 and the objective function in Step 4, set the state space, action space and reward function of the deep reinforcement learning training model to lay the foundation for obtaining the decision model in the future.

[0154] Simultaneous control of the flight trajectory and RIS phase of the UAV-carrying intelligent reflector system (UAV-RIS) is required. A deep reinforcement learning training model is constructed based on the deep reinforcement learning algorithm model from step 5. State spaces are set for the PPOs in the H-PPO deep reinforcement learning training model that optimize flight trajectory and RIS phase control. The action spaces of the two PPOs in the H-PPO deep reinforcement learning training model are set in conjunction with the UAV-carrying intelligent reflector-assisted energy-saving emergency communication scenario from step 1. The reward functions of the two PPOs in the H-PPO deep reinforcement learning training model are set in conjunction with the objective function and constraint (1) from step 4.

[0155] Step 6 sets the state space, action space, and reward function as follows:

[0156] 1) State Space: At time slot t, the state of the first near-end policy optimization algorithm (PPO) contains global position information, represented as follows: This represents the horizontal coordinate of the m-th ground device at time slot t; furthermore, the state of the second near-end policy optimization algorithm (PPO) includes global channel information, represented as... Since neural networks receive real numbers as input, before inputting this complex channel state into the network, the real and imaginary parts of the channel need to be separated and input into the neural network simultaneously.

[0157] 2) Action Space: At time slot t, the action of the first proximal policy optimization algorithm PPO (upper layer) is the flight distance and azimuth of the UAV carrying the Intelligent Reflector System (UAV-RIS), represented as... The action of the second proximal policy optimization algorithm PPO (lower layer) in time slot t is the phase value of the active RIS, denoted as Θ. t ;

[0158] 3) Reward Function: The reward structure of the first proximal policy optimization algorithm, PPO, mainly consists of two parts. First, the real-time reward includes the ratio of communication traffic to energy consumption in the current time slot, and the difference between the horizontal distance between the UAV-carrying Intelligent Reflector System (UAV-RIS) and all ground equipment in the previous time slot, expressed as: Among them, P t,m This represents the additional positive reward value for the m-th ground device when the rate threshold is met in time slot t. Secondly, a termination reward reflecting energy efficiency throughout the entire service cycle is set, represented as... λ represents a positive fixed value used to adjust the ratio of the termination reward to the final reward; furthermore, the reward function of the first proximal policy optimization algorithm, PPO, is expressed as:

[0159]

[0160] Finally, in time slot t, the reward function of the second proximal policy optimization algorithm PPO is expressed as:

[0161]

[0162] Step 7: Train the deep reinforcement learning training model obtained in Step 6 to obtain the UAV-equipped intelligent reflective surface-assisted energy-saving emergency communication decision model, and obtain the optimal solution to the optimization problem.

[0163] The process of step 7 is as follows:

[0164] Step 7.1, Initialize the network

[0165] Randomly initialize the actor network θ for each Proximal Policy Optimization (PPO) algorithm. A and critics' network θ C Parameters;

[0166] Step 7.2, Training the deep reinforcement learning model

[0167] Randomly initialize the location of each ground device, the location of the UAV-RIS (Unmanned Aerial Vehicle-Reflective Surface System), and the location of the airborne base station (ABS);

[0168] In each time slot t, the first proximal policy optimization algorithm (PPO) interacts with the dynamic environment to obtain the state. Based on the current state, the first proximal policy optimization algorithm, PPO, obtains the drone's flight actions from its Actor network. To control the trajectory of the drone;

[0169] The second proximal policy optimization algorithm, PPO, interacts with the dynamic environment to obtain the state. Based on the current state, the second proximal policy optimization algorithm, PPO, obtains the action of the RIS phase from its Actor network. To control the intelligent reflective surface; and finally take action. Interact with the environment and calculate the reward r obtained from the environment. t ;

[0170] State state action action and reward r t Stored in the experience pool;

[0171] When the experience pool accumulates to a certain amount, N is taken from it. d The parameters of the Critic and Actor networks are updated using samples of varying sizes. After each training session, the experience pool is cleared. This process is repeated until the model training converges, resulting in a UAV-equipped intelligent reflector-assisted energy-saving emergency communication decision-making model.

[0172] Step 7.3, Decision-making stage

[0173] The decision model for energy-saving emergency communication with UAVs equipped with intelligent reflectors is applied to stochastic dynamic UAV-equipped intelligent reflector-assisted energy-saving emergency communication. In each time slot, the optimal reconfigurable intelligent surface (RIS) reflectivity matrix and UAV flight actions are determined to minimize the energy consumption ratio of the communication system throughout the process, while ensuring the communication quality of all individual ground users.

[0174] The key points and protection points of this invention are as follows:

[0175] 1. A solution for realizing energy-saving emergency communication by equipping UAVs with intelligent reflective surfaces.

[0176] 2. An optimization problem for jointly designing trajectories and RIS phases is proposed in an energy-saving emergency communication network assisted by an intelligent reflector mounted on an unmanned aerial vehicle.

[0177] 3. Based on the Hierarchical Proximal Policy Optimization (H-PPO) algorithm based on deep reinforcement learning.

[0178] In scenarios where UAVs carry intelligent reflector-assisted communication, current research focuses on joint optimization of UAV trajectory planning and RIS phase to maximize the total capacity of the target user link, but it does not consider the Quality of Service (QoS) of individual users. Furthermore, in multi-user networks, resolving the competition for RIS resources among users is crucial for ensuring excellent service for individual users—a problem largely overlooked in many studies. This invention employs DRL joint design of UAV trajectories and RIS phase to simultaneously address the issues of optimizing energy efficiency and ensuring QoS for individual users in scenarios requiring long decision sequences. Therefore, currently, no other alternative can fully achieve the objectives of this invention.

[0179] like Figure 2 The diagram shows a scenario where a drone equipped with a smart reflective surface assists in an energy-saving emergency communication network. The drone, equipped with a smart reflective surface, provides uplink communication services to ground equipment.

[0180] like Figures 3(a)-3(d) The convergence of the H-PPO algorithm is shown in Figures 3(a) and 3(b). Figures 3(c) and 3(d) respectively show the trends of the reward values ​​of the two PPOs in the H-PPO algorithm as training progresses. Figures 3(a)-3(d) It can be seen that as training progresses, the system's energy efficiency and the individual service quality of all ground equipment have significantly improved, indicating that after sufficient training, it has a strong adaptability to high-complexity and high-dynamic scenarios.

[0181] like Figure 4 As shown, the system energy efficiency ratio under different schemes and different rate threshold requirements is illustrated. Figure 5The figures show the QoS completion rates of ground equipment under different schemes and different rate threshold requirements. Analysis of the data in the two figures clearly shows that although the H-PPO scheme of this invention lags slightly behind the DT-PPO scheme which uses heuristic control of the UAV trajectory in terms of QoS completion rate, the H-PPO scheme of this invention is the superior choice when considering energy efficiency. The H-PPO scheme of this invention innovatively coordinates the control of the UAV trajectory and the phase of the reconfigurable intelligent reflector (RIS), significantly optimizing energy efficiency, ensuring the quality of service for individual devices, and effectively solving the problem of optimizing energy efficiency and addressing competition for reconfigurable intelligent surface (RIS) resources among users in scenarios requiring long decision sequences when using UAVs equipped with intelligent reflectors. This is also true for cascaded channel information related to time-varying conditions.

[0182] The present invention also provides an intelligent reflective surface-assisted UAV emergency communication system, comprising:

[0183] The UAV emergency communication scenario construction module is used to realize the construction of the UAV equipped with intelligent reflective surface assisted energy-saving emergency communication scenario in step 1;

[0184] The communication process establishment module is used to model the communication process of multiple ground devices carried by the UAV-Intelligent Reflector System (UAV-RIS) in the UAV-equipped intelligent reflector-assisted energy-saving emergency communication scenario constructed in step 1 in step 2.

[0185] The UAV propulsion energy consumption modeling module is used to model the UAV propulsion energy consumption in step 3.

[0186] The optimization objective and objective function construction module is used to implement the communication process of multiple ground devices carried by the UAV-Intelligent Reflector System (UAV-RIS) in the UAV-equipped intelligent reflector-assisted energy-saving emergency communication scenario in step 4, combined with the UAV propulsion energy consumption model in step 2, and to model the optimization objective of the final energy consumption ratio and construct the objective function when solving the model.

[0187] The deep reinforcement learning algorithm model building module is used to build a deep reinforcement learning algorithm model based on the optimization objective proposed in step 4 in step 5.

[0188] The state space, action space, and reward function setting module is used to implement the deep reinforcement learning training model construction based on the deep reinforcement learning algorithm model in step 5 in step 6, and to set the state space, action space, and reward function of the deep reinforcement learning training model in combination with the UAV equipped with intelligent reflective surface assisted energy-saving emergency communication scenario in step 1 and the objective function in step 4.

[0189] The deep reinforcement learning algorithm model training module is used to train the deep reinforcement learning training model obtained in step 6 in step 7, so as to obtain the UAV-equipped intelligent reflective surface-assisted energy-saving emergency communication decision model and obtain the optimal solution of the optimization problem.

[0190] The present invention also provides an intelligent reflective surface-assisted emergency communication device for unmanned aerial vehicles, comprising:

[0191] Memory: A computer program that stores the above-mentioned intelligent reflective surface-assisted UAV emergency communication method, and is a computer-readable device;

[0192] Processor: Used to implement the intelligent reflective surface-assisted UAV emergency communication method when executing the computer program.

[0193] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the aforementioned intelligent reflective surface-assisted UAV emergency communication method.

Claims

1. A method for intelligent reflecting surface-assisted emergency communication of unmanned aerial vehicles, characterized in that, The method comprises the following steps: Step 1, constructing an unmanned aerial vehicle (UAV) carrying intelligent reflecting surface (RIS) assisted energy-saving emergency communication scenario; Step 2, modeling the communication process of the unmanned aerial vehicle (UAV) carrying intelligent reflecting surface (RIS) assisted multiple ground devices in the unmanned aerial vehicle (UAV) carrying intelligent reflecting surface (RIS) assisted energy-saving emergency communication scenario constructed in step 1; Step 3, modeling the propulsion energy consumption of the unmanned aerial vehicle (UAV); Step 4, modeling the optimization target of the final energy consumption ratio in combination with the communication process of the unmanned aerial vehicle (UAV) carrying intelligent reflecting surface (RIS) assisted multiple ground devices in the unmanned aerial vehicle (UAV) carrying intelligent reflecting surface (RIS) assisted energy-saving emergency communication scenario of step 2 and the unmanned aerial vehicle (UAV) propulsion energy consumption model of step 3, and constructing the objective function when solving the model; Step 5, constructing a deep reinforcement learning algorithm model according to the optimization target proposed in step 4; The deep reinforcement learning algorithm model is constructed by using an H-PPO algorithm to control the flight trajectory of the unmanned aerial vehicle (UAV) carrying intelligent reflecting surface (RIS) and the RIS phase, and the H-PPO algorithm comprises two independent proximal policy optimization (PPO) algorithms; the first proximal policy optimization (PPO) algorithm is used for optimizing flight actions for trajectory control; and the second proximal policy optimization (PPO) algorithm is used for RIS phase values to enhance channel gain; the flight trajectory and RIS phase control are optimized by using the independent proximal policy optimization (PPO) algorithms; Step 6, constructing a deep reinforcement learning training model according to the deep reinforcement learning algorithm model of step 5, combining the unmanned aerial vehicle (UAV) carrying intelligent reflecting surface (RIS) assisted energy-saving emergency communication scenario of step 1 and the objective function of step 4, and setting a state space, an action space and a reward function of the deep reinforcement learning training model; Step 7, training the deep reinforcement learning training model obtained in step 6 to obtain an unmanned aerial vehicle (UAV) carrying intelligent reflecting surface (RIS) assisted energy-saving emergency communication decision model, and obtaining an optimal solution of the optimization problem. 2.The smart reflector assisted emergency communication method for UAVs according to claim 1, wherein, The process of step 1 is as follows: The unmanned aerial vehicle carries an intelligent reflecting surface to assist energy-saving emergency communication scenarios, including an unmanned aerial vehicle, an unmanned aerial vehicle installed with an active reconfigurable intelligent surface (RIS), that is, an unmanned aerial vehicle carrying an intelligent reflecting surface system (UAV-RIS), and an air base station (ABS) for processing information transmitted from ground devices, the air base station (ABS) serving as a control center, responsible for interacting with and issuing commands to the unmanned aerial vehicle carrying the intelligent reflecting surface system (UAV-RIS), and the unmanned aerial vehicle carrying the intelligent reflecting surface system (UAV-RIS) assisting uplink transmission of multiple ground devices to the air base station (ABS), and the ground device set is denoted as g={G m ,m=1,2,…,M} where M is the number of ground devices, it is assumed that there is no direct communication link between the ground devices and the air base station (ABS), and the air base station (ABS) is located at a higher position. 3.The smart reflector assisted emergency communication method for UAVs of claim 1, wherein, The process of step 2 is as follows: Step 2.1, constructing a channel model; In the channel model, each ground device has an omnidirectional antenna, while the air base station (ABS) is equipped with Q antenna arrays, and the active reconfigurable intelligent surface (RIS) is composed of N reflecting elements, assuming that the reflection coefficient of the nth element of the RIS is where φ n ∈ [0, 2π), β represents an amplification factor greater than 1, and the reflection coefficient matrix of the active reconfigurable intelligent surface (RIS) is defined as Θ = diag ([θ1, θ2, …, θ N ]) H m,r Hm represents the channel of the link between the mth ground device and the unmanned aerial vehicle-carrying intelligent reflecting surface (UAV-RIS), H r,b H represents the channel between the unmanned aerial vehicle-carrying intelligent reflecting surface (UAV-RIS) and the air base station (ABS), assuming that the channel gain of the unmanned aerial vehicle-carrying intelligent reflecting surface (UAV-RIS) to the air base station (ABS) obeys the Rician distribution, the channel is represented as: wherein, κ r,b is the path loss at the reference distance D = 1 m, d r,b is the distance between the UAV-carried intelligent reflecting surface system (UAV-RIS) and the air base station (ABS), α r,b is the path loss exponent of the UAV-RIS-ABS link; the line-of-sight component is the first order, expressed as: wherein, for a UAV-carried intelligent reflecting surface system (UAV-RIS) to the angle of departure (AoD) and the angle of arrival (AoA) of an air base station (ABS) link, and ω is the antenna spacing; Non-line-of-sight component Each element of H follows a complex Gaussian distribution with zero mean and unit variance, independently of the other elements; similarly, the LoS component H m,r of the m-th link follows a distribution similar to the one described above, and is denoted by H . wherein, is the angle of departure (AoD) for the mth ground device to UAV-RIS link, the angle of arrival (AoA) and the angle of departure (AoD) are determined by their relative position relationship; Finally, it is assumed that all channels follow block fading, the global channel information is known at the air base station (ABS), and remains unchanged at each time slot but changes from one time slot to another; Step 2.2, constructing a transmission model based on the channel model constructed in step 2.1; The uplink transmission model from ground devices to aerial base stations (ABS) considers in-band interference, the aerial base stations (ABS) employ maximal ratio combining (MRC) for each link, the maximal ratio combining (MRC) beamforming matrix is denoted by F = [f1,..., fK], where f M represents the unit norm beamforming vector for the mth link; the signal received at the aerial base stations (ABS) from the mth ground device is denoted by ym = hTmFxm + n, where h m is the channel vector from the mth ground device to the aerial base stations (ABS), xmis the transmitted signal from the mth ground device, and n is the noise vector. where P m is the transmit power of the mth ground equipment, s m is a unit energy signal sample associated with the monitoring data; a noise vector n m produced at an air base station (ABS) is represented as n m = [n1,..., n M ] T where An active reconfigurable intelligent surface (RIS) requires additional power and each element is equipped with an amplifier, while considering the thermal noise n generated at the unmanned aerial vehicle carrying intelligent reflecting surface system (UAV-RIS) i wherein, The uplink signal-to-noise ratio (SINR) of the mth link of the aerial base station (ABS) is represented as: Thus, the rate of the mth link is C m = Blog(1 + η m ), where B is the bandwidth. 4.The smart reflector assisted emergency communication method for UAVs of claim 1, wherein, The process of step 3 is as follows: The horizontal position of the UAV-RIS and the ABS at time slot t is denoted by and where, The vertical height of the UAV-RIS and the ABS is set to h r and h b respectively, the horizontal position of the UAV-RIS at the next time slot t is determined by the flight distance D t and the azimuth variable ξ t The horizontal coordinates of the UAV-RIS at time t+1 are denoted by According to the setting given to the horizontal position of the unmanned aerial vehicle carrying intelligent reflecting surface system (UAV-RIS), the horizontal flight speed of the unmanned aerial vehicle carrying intelligent reflecting surface system (UAV-RIS) at time t The maximum horizontal speed, Δt is the length of each time period; if It indicates that at time t, the unmanned aerial vehicle is in a hovering state; according to the horizontal speed of the unmanned aerial vehicle, the propulsion energy consumption of the unmanned aerial vehicle in each time period is obtained: where P0 and P1 are the constant power and induced power of the airfoil in hover, U tip is the tip speed of the rotor blade, Δt is the duration of each time slot, v0 is the average rotor induced velocity in hover, d0 and s are the fuselage drag ratio and rotor solidity, respectively, and p and G represent the air density and rotor disc area, respectively; The process of step 4 is as follows: The reflection coefficient matrix of the smart reflecting surface at the t-th time slot is denoted as Θ t , C t,m is the communication rate of the m-th link at the t-th time slot obtained in step 2, The target function of maximizing energy efficiency in the time range of length T is expressed as: In the formula, constraint (1) represents the transmission rate requirement that each ground device needs to meet, and Y represents the rate limit value; constraint (2) limits the range that the RIS phase can be adjusted; and constraint (3) represents the limitation of the maximum flight speed of the unmanned aerial vehicle (UAV) carrying intelligent reflecting surface (RIS). 5.The smart reflector assisted emergency communication method for UAVs of claim 1, wherein, In step 5, the proximal policy optimization (PPO) algorithm is an optimization algorithm of deep reinforcement learning (DRL), and the proximal policy optimization (PPO) algorithm learns the optimal policy by iteratively updating the policy to maximize the expected cumulative reward; The policy of the proximal policy optimization algorithm PPO is parameterized and represented by two networks, an actor network and a critic network wherein θ A and θ C represent respective parameter instances; the proximal policy optimization algorithm PPO uses a surrogate objective function and a clipping mechanism to constrain the policy update; the objective function in the proximal policy optimization algorithm PPO is: wherein, representing a policy between the old policy and the importance sampling weight; is the corresponding advantage function, where γ is a discount factor; a clipping mechanism limits the policy update to a predefined range, where clip() denotes a clipping function, and λ is a hyperparameter that controls the trade-off between exploration and exploitation. t (θ A ) is controlled to be in the interval [1-ε, 1+ε]; In the training phase, the experience of interacting with the environment using I batches is used to update the network; for the parameter update mode of the actor network is: At the same time, in order to update θ C , the mean square error function is used as the loss function, which is expressed as: where α C denotes a learning rate, denotes a target state value function derived in a temporal difference manner. 6.The smart reflector assisted emergency communication method for UAVs of claim 1, wherein, The settings of the state space, the action space and the reward function in step 6 are as follows: 1) State space: At time slot t, the state of the first proximal policy optimization algorithm PPO contains global position information, denoted as denotes the horizontal coordinate of the mth ground device at time slot t; in addition, the state of the second proximal policy optimization algorithm PPO contains global channel information, denoted as Before inputting this complex channel state into the network, the real part and the imaginary part of the channel are separated and input into the neural network at the same time; 2) Action space: At time slot t, the action of the first proximal policy optimization algorithm PPO (upper layer) is the flight distance and flight direction of the unmanned aerial vehicle carrying the intelligent reflecting surface system (UAV-RIS), denoted as The action of the second proximal policy optimization algorithm PPO (lower layer) at time slot t is the phase value of the active RIS, denoted as Θ t ; 3) Reward function: The reward structure of the first proximal policy optimization algorithm PPO consists of two parts; first, the real-time reward, which includes the ratio of traffic to energy consumption in the current time slot and the difference between the horizontal distance between the UAV-RIS and all ground devices and the last time slot, denoted as where P t,m represents the additional positive reward value of the mth ground device when the rate threshold is met at time slot t, Secondly, the terminal reward reflecting the energy efficiency of the entire service period is set, denoted as λ represents a positive fixed value for adjusting the proportion of the terminal reward value and the final reward; in addition, the reward function of the first proximal policy optimization algorithm PPO is represented as: Finally, at time slot t, the reward function of the second proximal policy optimization (PPO) algorithm is represented as:

7. The smart reflector assisted drone emergency communication method of claim 1, wherein, The process of step 7 is as follows: Step 7.1, initializing the network randomly initialize parameters of an actor network θ for each proximal policy optimization algorithm PPO A and critic network θ C . Step 7.2, training the deep reinforcement learning training model Randomly initialize each ground device, the position of the unmanned aerial vehicle carrying the intelligent reflecting surface system (UAV-RIS), and the position of the air base station (ABS); At each time slot t, the first proximal policy optimization algorithm PPO interacts with the dynamic environment to obtain the state Based on the current state, the first proximal policy optimization algorithm PPO obtains the UAV flight action from its actor network To control the trajectory of the UAV; The second proximal policy optimization algorithm PPO interacts with the dynamic environment to obtain a state Based on the current state, the second proximal policy optimization algorithm PPO obtains the action of the RIS phase from the Actor network thereof To control the intelligent reflecting surface; finally, take the action Interact with the environment, calculate the reward r obtained from the environment t ; state state action action and reward r t stored in the experience pool; When the experiences in the experience pool accumulate to a certain amount, N d samples of this size are taken from it to update the parameters of the Critic and Actor networks. After each training round, the experience pool is emptied, and the above process is repeated until the model training converges, obtaining the UAV-mounted smart reflector assisted energy-saving emergency communication decision-making model. Step 7.3, decision-making stage The unmanned aerial vehicle-mounted intelligent reflecting surface-assisted energy-saving emergency communication decision-making model is used in a random dynamic unmanned aerial vehicle-mounted intelligent reflecting surface-assisted energy-saving emergency communication. In each time slot, the optimal reconfigurable intelligent surface (RIS) reflection coefficient matrix and unmanned aerial vehicle flight action are decided, so that the energy consumption of the communication system in the whole process is minimized, while ensuring the communication quality of all individual ground users.

8. An intelligent reflector assisted drone emergency communication system based on the method of any one of claims 1 to 7, characterized in that: It comprises: An unmanned aerial vehicle emergency communication scenario construction module for constructing an unmanned aerial vehicle-mounted intelligent reflecting surface-assisted energy-saving emergency communication scenario; A communication process modeling module for modeling the communication process of the unmanned aerial vehicle-mounted intelligent reflecting surface-assisted energy-saving emergency communication scenario in which the unmanned aerial vehicle carrying the intelligent reflecting surface system (UAV-RIS) assists multiple ground devices; A unmanned aerial vehicle propulsion energy consumption modeling module for modeling the propulsion energy consumption of the unmanned aerial vehicle; An optimization target and objective function modeling module for modeling the optimization target of the final energy consumption ratio by combining the communication process of the unmanned aerial vehicle-mounted intelligent reflecting surface-assisted energy-saving emergency communication scenario in which the unmanned aerial vehicle carrying the intelligent reflecting surface system (UAV-RIS) assists multiple ground devices and the propulsion energy consumption model of the unmanned aerial vehicle, and constructing the objective function when solving the model; A deep reinforcement learning algorithm model construction module for constructing a deep reinforcement learning algorithm model according to the proposed optimization target; A state space, action space and reward function setting module for setting the state space, action space and reward function of the deep reinforcement learning training model according to the deep reinforcement learning algorithm model, combining the unmanned aerial vehicle-mounted intelligent reflecting surface-assisted energy-saving emergency communication scenario and the objective function, and setting the state space, action space and reward function of the deep reinforcement learning training model; A deep reinforcement learning algorithm model training module for training the deep reinforcement learning training model to obtain the unmanned aerial vehicle-mounted intelligent reflecting surface-assisted energy-saving emergency communication decision-making model and the optimal solution of the optimization problem.

9. An intelligent reflecting surface assisted drone emergency communication device, characterized in that: It comprises: A memory for storing a computer program of the intelligent reflecting surface-assisted unmanned aerial vehicle emergency communication method according to any one of claims 1-7, which is a computer-readable device; A processor for implementing the intelligent reflecting surface-assisted unmanned aerial vehicle emergency communication method according to any one of claims 1-7 when the computer program is executed.

10. A computer-readable storage medium, characterized in that: A computer-readable storage medium stores a computer program, which can implement the intelligent reflecting surface-assisted unmanned aerial vehicle emergency communication method according to any one of claims 1-7 when executed by a processor.

Citation Information

Patent Citations

  • Energy-saving optimization method for nonorthogonal multiple access UAV air-to-ground communication network assisted by intelligent reflector

    CN113873575B

  • Intelligent reflecting surface assisted uplink NOMA unmanned aerial vehicle network energy efficiency optimization method

    CN117596616A