Energy efficiency optimization method for fixed-wing unmanned aerial vehicle communication coverage based on reflection intelligent surface assistance

By using reflective smart surfaces to assist fixed-wing UAV communication, and by using the MICH algorithm to plan smooth trajectories and combining the AS-VRE and NsHQPN algorithms to optimize speed and GN service scheduling, the energy efficiency problem of fixed-wing UAV communication systems in urban environments is solved, achieving more efficient GN coverage and energy efficiency optimization.

CN119521256BActive Publication Date: 2025-11-28SHENYANG AEROSPACE UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411446521.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2025-11-28
Estimated Expiration
2044-10-16

AI Technical Summary

Technical Problem

In urban environments, the energy efficiency of fixed-wing UAV-assisted communication systems is limited by obstructions. The defects of traditional trajectory planning and DRL algorithms lead to reduced communication quality and energy efficiency, especially when GN distribution is uneven, circular trajectories cause uneven GN coverage and energy waste.

Method used

A reflective smart surface is used to assist communication for fixed-wing UAVs. The MICH algorithm is used to plan smooth trajectories, and the AS-VRE mechanism and NsHQPN algorithm are combined to optimize UAV speed and GN service scheduling, thereby improving system energy efficiency.

Benefits of technology

It effectively improves the energy efficiency of fixed-wing UAVs in providing stable communication services to ground nodes in urban environments, extends the service time of communication tasks, and reduces UAV energy consumption and path loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119521256B_ABST
    Figure CN119521256B_ABST
Patent Text Reader

Abstract

The application discloses an energy efficiency optimization method for fixed-wing unmanned aerial vehicle communication coverage assisted by a reflective intelligent surface, comprising the following steps: S1: taking maximizing system energy efficiency as a target, constructing a problem P, and setting multiple constraint conditions; S2: dividing the problem P into two subproblems; S3: adopting an MICH algorithm to plan an unmanned aerial vehicle track to solve the first subproblem; and S4: firstly adopting an AS-VRE mechanism to improve the exploration ability of DRL, and secondly adopting an NsHQPN algorithm to improve the value precision of a mixed action space. The method can improve the system energy efficiency of GN communication assisted by the fixed-wing unmanned aerial vehicle as much as possible under the condition of limited energy, and can better and more durably provide communication services for GN in an emergency communication scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of emerging information technology, and particularly provides an energy efficiency optimization method for fixed-wing unmanned aerial vehicle communication coverage assisted by reflective intelligent surfaces. BACKGROUND

[0002] Unmanned aerial vehicles (UAVs) have attracted much attention due to their high mobility and flexible air deployment, and they are widely used in aerial patrol, emergency rescue, mobile edge computing, and other fields. In terms of wireless network communication, due to factors such as complex terrain and fixed communication distance, traditional communication methods are difficult to achieve stable and comprehensive signal coverage. In view of the limitations of traditional communication methods, UAVs can be used as wireless relay stations or mobile base stations to provide more effective communication services for ground nodes (GNs). However, since the UAVs fly at a certain height, the link between the UAVs and the ground nodes is often blocked by tall buildings, trees, and other obstacles. The probability of such blockage will significantly increase, especially in urban environments. Due to the blockage, the channel gain of the communication link will be greatly reduced, thereby affecting the communication quality. In recent years, reflective intelligent surfaces (RIS) based on 5G and 6G can effectively solve the influence of blockage on channel gain by adjusting the reflection angle and phase. In addition, this phase adjustment can also achieve signal reflection, thereby reducing energy consumption, improving communication efficiency and coverage, and maximizing signal-to-noise ratio. Many research works use RIS to assist UAVs to improve communication gain. Provide a reliable visual link.

[0003] The system energy efficiency is greatly influenced by the UAV flight trajectory. Therefore, it is necessary to focus on the optimization problem of trajectory planning. In UAV-assisted communication, UAVs can be divided into rotary-wing UAVs and fixed-wing UAVs according to the wing structure and power system. Compared with rotary-wing UAVs, fixed-wing UAVs have the advantage of long flight time and large coverage range. Since fixed-wing UAVs cannot hover, turning is completed by the centripetal force generated by the inclination of the fuselage, which will cause the turning angle to be insufficient, and the direction cannot be changed or turned around in time. Therefore, the flight trajectory should be continuous and smooth, rather than twisted. In addition, the trajectory needs to be designed according to the distribution of GNs, so that it is more suitable for the shape of the GN distribution, and more GNs can be covered while reducing the flight distance. In some works, the flight trajectory of the fixed-wing UAV is set to be circular. However, this circular trajectory has the following two shortcomings. First, when the GN distribution is relatively narrow relative to its length, the circular trajectory will have few or no GNs near some flight positions. This results in waste of setting some flight positions on the circular trajectory. Second, setting a circular trajectory can cause the flight radius of the UAV to be too large, making the UAV too far away from the center GN. Therefore, the communication quality of the center GN will be reduced. Therefore, finding a trajectory design method suitable for the GN distribution is a major challenge.

[0004] In addition, in addition to the flight trajectory of the UAV, the flight speed of the UAV and the service scheduling of the GN also have a profound influence on the system energy efficiency. In some works, DRL algorithms are often used for optimization. However, traditional DRL algorithms have two shortcomings. The first is that the exploration of effective experience is limited, and it is easy to learn a large amount of negative experience in the early stage of learning, thereby causing the training to fall into a local optimum. The second shortcoming is that when traditional DRL algorithms solve mixed action space (for example, the speed is a continuous action, and the scheduling of GN is a discrete action), the precision of the action value is insufficient. This will also cause the decision to fall into a local optimum, resulting in a decrease in system energy efficiency. Therefore, it is particularly important to solve the existing defects of DRL algorithms. SUMMARY

[0005] In view of this, the purpose of the present application is to provide an energy efficiency optimization method for fixed-wing UAV communication coverage assisted by reflective intelligent surfaces, so as to solve the problem of improving the system energy efficiency of fixed-wing UAV-assisted GN communication as much as possible under the condition of limited energy, prolonging the service time of the communication task, and providing better and more persistent communication services for GNs in emergency communication scenarios.

[0006] The technical scheme provided by the application is an energy efficiency optimization method for fixed-wing unmanned aerial vehicle communication coverage based on reflection intelligent surface assistance, which is realized by placing RIS on the outer wall of a ground building to reflect the link through communication coverage of GN by a fixed-wing unmanned aerial vehicle.

[0007] The energy efficiency optimization method comprises:

[0008] S1: constructing a problem P with the goal of maximizing system energy efficiency, and setting multiple constraint conditions;

[0009] S2: dividing the problem P into two sub-problems, sub-problem one: planning a smooth, ground node distribution-adaptive hovering trajectory; and sub-problem two: planning the flight speed of the unmanned aerial vehicle and the service scheduling of GN on the planned trajectory using a DRL algorithm to further improve energy efficiency;

[0010] S3: adopting a midpoint iterative convex hull (MICH) algorithm to plan the unmanned aerial vehicle trajectory to solve sub-problem one;

[0011] S4: solving sub-problem two: first, adopting an action screening virtual-real experience (AS-VRE) mechanism to improve the exploration ability of the DRL, and second, adopting a strategy network (NsHQPN) algorithm to improve the value accuracy of the mixed action space.

[0012] Specifically, the problem P in S1 is described as:

[0013]

[0014]

[0015]

[0016]

[0017]

[0018] (1) indicates that the goal is to maximize system energy efficiency, and there are constraint conditions (1a)-(1d) under the goal; wherein (1a) indicates that the service state of each GN in each time slot can only be a served state and an unserved state, which are recorded as 1 and 0, respectively; (1b) indicates that in one time slot, the unmanned aerial vehicle can only serve one GN, that is, the constraint of time division multiple access (TDMA); (1c) indicates that the cumulative data amount transmitted for each GN must be greater than the requested data amount (1d) indicates the upper and lower limits of the flight speed of the fixed-wing unmanned aerial vehicle.

[0019] Specifically, the MICH algorithm in S3 comprises:

[0020] Firstly, several key trajectory nodes are solved according to the distribution of GNs;

[0021] Secondly, trajectory fitting is performed on the key trajectory nodes to make the trajectory curve smoother to adapt to the special turning constraint of the fixed-wing UAV.

[0022] Specifically, the trajectory mode of the fixed-wing UAV is represented by the following equation:

[0023]

[0024] where a0∈[0,3000], b0∈[0,3000], x0∈[0,3000], y0∈[0,3000], and θ∈[-π,π] represent unknown parameters of the fitting curve;

[0025] The predicted data is calculated according to the defined fitting curve function, and the calculation process is as follows:

[0026]

[0027] Finally, gradient descent optimization is performed using the loss equation to obtain specific parameters;

[0028] A Markov decision process is constructed:

[0029] State space: The work of the fixed-wing UAV is to meet the communication needs of all GNs. Since the distance between the UAV and the selected GN directly affects the communication quality, a state space is defined, where the position of the UAV is represented by , and K reflects the service state of all GNs. The service state of a GN is a binary variable, indicating whether the service to the GN is completed in the current time slot. When , it indicates that the th GN has not completed the service, and the UAV continues to provide communication services for it. On the contrary, when , it indicates that the GN has completed the service and does not need to be served in this round; the service state of all GNs can be mapped by using a binary list to calculate K, i.e. , so as to determine whether the task for each GN is completed;

[0030] Action space: The action space A={v,Θ} contains the flight speed v of the UAV and the service scheduling Θ of the GN. In the k time slot, the specific action a(k) is defined as containing the flight speed v[k] in the current time slot and the service scheduling Θ, i.e. The flight speed v[k] of the UAV is limited in the range of [v max , v min ]; the selected GN in the k time slot is represented by ;

[0031] Reward function: the DRL algorithm is adopted, and the energy efficiency is taken as the reward function to evaluate the action a(k) performed by the UAV in the state s(k), so the reward r k can be expressed as

[0032]

[0033] The data transmission rate at k time slot is represented as ; the energy consumed by the fixed-wing UAV at k time slot is When the served GN is selected again, a penalty term is introduced

[0034] Specifically, the DRL is optimized by adopting the AS-VRE mechanism:

[0035] The AS-VRE mechanism contains two basic processes: the action screening process and the process of combining virtual and real experience;

[0036] First, the state observation value is input into the neural network, and the output of the neural network will evaluate the rewards of all possible actions;

[0037] The action screening step is performed to identify and exclude all illegal actions; when selecting an action, the action with the highest reward value is selected from the eligible actions as the action to be executed, i.e., the screening process;

[0038] A part of virtual experience, i.e., the experience of illegal actions that are not executed and have low reward values, is introduced into the experience pool, so that the neural network can fully learn the experience of illegal actions;

[0039] A virtual reward r v is added to the illegal action, so that the reward of the illegal action is lower than that of the legal action; the virtual reward is defined as follows:

[0040] r v = min(Q(s(k+1),a uc ))-1. (5)

[0041] The illegal action is marked as a uc ; since the illegal action is not executed when the illegal action is selected, r v does not constitute the true action value.

[0042] Specifically, the AS-VRE mechanism is that the neural network scores each action according to the input state; then, the neural network judges the illegal action, filters the qualified actions, and selects the action with the maximum reward from the qualified actions to execute; the experience after execution is added to the experience pool; at the same time, if the illegal action is encountered, the virtual experience that is not really executed needs to be added to the experience pool; finally, a batch of experiences are selected from the experience pool for neural network learning;

[0043] The NsHQPN algorithm is used to optimize the decision of service scheduling of the speed and GN:

[0044] The action set is composed of a discrete action set C and a continuous action set A c . Wherein, the action a(k) can be represented as a(k)=(c(k),a c (k)), wherein c(k) represents a discrete action, c∈C, a c (k) is a continuous action corresponding to the discrete action c(k), a c ∈A c The action a(k) taken in the state s(k) is evaluated by the state-action value function Q(s(k),c(k),a c (k)); the Bellman equation is represented as:

[0045]

[0046] Wherein γ∈(0,1] represents a discount factor;

[0047] The optimal continuous action corresponding to each discrete action is calculated, and the policy network is used to make optimal decisions;

[0048]

[0049] Wherein, pol(·) is a deterministic policy function, which is used to select the optimal action a c from the continuous action set A c ; the parameter set of the policy network is represented by Φ;

[0050] Then, the Q network makes decisions according to the continuous action and determines the discrete action;

[0051]

[0052] In each time slot, the agent starts from the current state s(k), accumulates the reward obtained by executing N steps of actions, and takes the Q value of the state s(k+N) reached after the N steps of actions as the target expected value y k , which is used to update the Q value of the state s(k) and the action a(k);

[0053]

[0054] where N represents the number of steps from the current state to the target state; the parameter set of the Q network is denoted by Ψ;

[0055] The loss function for updating the network parameters Φ and Ψ is defined as:

[0056]

[0057] The method provided by the application solves the energy efficiency optimization problem of the RIS-assisted fixed-wing unmanned aerial vehicle providing communication services for GN in an urban environment. The method solves the fixed-wing unmanned aerial vehicle flight trajectory planning problem based on the Graham scan MICH algorithm, which can consider the distribution shape of GN, thereby better covering GN. The application proposes placing the AS-VRE mechanism at the output end of the neural network to solve the problem that the DRL algorithm cannot add a taboo in the action space. The application also proposes an NsHQPN algorithm that can handle mixed action spaces to solve the problem of traditional DRL algorithms falling into local optimization due to insufficient action accuracy when processing mixed action spaces. Experiments show that the MICH algorithm, AS-VRE mechanism and NsHQPN algorithm proposed by the application can effectively improve the system energy efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0058] The application will be described in further detail below with reference to the accompanying drawings and embodiments:

[0059] Figure 1 The system scene diagram provided by the application;

[0060] Figure 2 The problem P solution process schematic diagram of the application;

[0061] Figure 3 The MICH algorithm solution schematic diagram of the application;

[0062] Figure 4 The AS-VRE mechanism schematic diagram of the application;

[0063] Figure 5 The NsHQPN algorithm principle diagram of the application;

[0064] Figure 6 The average reward of the reconstructed intelligent surface performance comparison in the embodiment of the application;

[0065] Figure 7 The energy efficiency CDF of the reconfigurable intelligent surface performance comparison in the embodiment of the application;

[0066] Figure 8 The comparison of different trajectories in the embodiment of the application;

[0067] Figure 9 Average reward for different trajectories in the embodiments of the present application;

[0068] Figure 10 Energy efficiency CDF for different trajectories in the embodiments of the present application;

[0069] Figure 11 Average reward for different mechanisms in the embodiments of the present application;

[0070] Figure 12 Energy efficiency CDF for different mechanisms in the embodiments of the present application;

[0071] Figure 13 Average reward for different DRL algorithms in the embodiments of the present application;

[0072] Figure 14 Energy efficiency CDF for different DRL algorithms in the embodiments of the present application. DETAILED DESCRIPTION

[0073] The present application will be further explained in conjunction with specific embodiments, but is not limited to the present application.

[0074] The present embodiment provides an energy efficiency optimization method for fixed-wing UAV communication coverage assisted by reflective intelligent surfaces, and a scene diagram is as shown in Figure 1 The fixed-wing UAV is used to communicate with GN, and RIS is placed on the outer wall of the ground building to reflect the link in order to avoid the influence of obstacles on the channel gain of the communication link. The purpose of the present application is to maximize the system energy efficiency. The problem description is as follows:

[0075]

[0076] (1) indicates that the target is to maximize the system energy efficiency. Under such a target, there are as many constraint conditions as (1a)-(1d). (1a) constrains the service state of each GN in each time slot to be only the served state and the unserved state, recorded as 1 and 0 respectively. (1b) indicates that in one time slot, the UAV can only serve one GN, that is, the constraint of time division multiple access technology (TDMA). (1c) indicates that the cumulative amount of data transmitted for each GN must be approximately equal to the amount of data requested by its request (1d) indicates the upper and lower limits of the flight speed of the fixed-wing UAV.

[0077] In problem P, the speed of the UAV and the service scheduling of GN have a profound impact on the system energy efficiency. However, due to the particularity of the flight trajectory of the fixed-wing UAV, that is, the fixed-wing UAV turns by relying on the centripetal force provided by the inclination of the UAV body to maintain turning, so its trajectory is smooth. The algorithm such as DRL according to time slot optimization will distort the trajectory. It cannot be truly implemented in actual situation. Therefore, problem P is divided into two sub-problems to solve. Sub-problem one: plan a smooth and adaptive hovering trajectory for the distribution of ground nodes to avoid excessive path loss and make the propulsion energy consumption larger, thereby improving the system energy efficiency. Sub-problem two: use the DRL algorithm to plan the flight speed of the UAV and the service scheduling of GN on the planned trajectory to further improve the energy efficiency. However, in sub-problem two, using the traditional DRL algorithm will cause the energy efficiency value to fall into local optimum due to the low precision of the value of the mixed action space and the low exploration ability of the effective experience.

[0078] Therefore, in order to make problem P obtain a decision with higher energy efficiency, the present application proposes an MICH algorithm for trajectory planning for sub-problem one. In view of the two defects of DRL in sub-problem two, firstly, the AS-VRE mechanism is proposed to improve the exploration ability of DRL, and secondly, the N-step hybrid deep Q and policy network (NsHQPN) algorithm is proposed to improve the value precision of the mixed action space. The problem P solution process diagram is as shown in Figure 2

[0079] MICH algorithm: considering the smoothness of the fixed-wing UAV trajectory, and in the energy efficiency optimization communication task, the trajectory needs to be as close to the distribution of GN as possible under the premise of reducing path loss. Therefore, the MICH algorithm is proposed. The solution diagram of MICH algorithm is as shown in Figure 3 . Specifically, it is divided into two steps, firstly, according to the distribution of GN, a plurality of key trajectory nodes are solved. Secondly, the trajectory fitting is performed on these key trajectory nodes, so that the trajectory curve is smoother and can adapt to the special turning constraint of the fixed-wing UAV. The specific steps are shown in algorithm 1.

[0080]

[0081]

[0082] The UAV trajectory is represented by the following function equation:

[0083]

[0084] ​where a0∈[0, 3000], b0∈[0, 3000], x0∈[0, 3000], y0∈[0, 3000], and θ∈[-π, π] represent unknown parameters of the fitting curve. The prediction data is calculated according to the defined fitting curve function, and the calculation process formula is as follows:

[0085]

[0086] Finally, gradient descent optimization is performed using the loss equation to obtain specific parameters.

[0087] Construction of Markov Decision Process:

[0088] State space: The work of the fixed-wing UAV is to meet the communication needs of all GNs. Since the distance between the UAV and the selected GN directly affects the communication quality, a state space is defined as The position of the UAV is represented by , and K reflects the service state of all GNs. The service state of the GN is a binary variable, indicating whether the service to the GN is completed in the current time slot. Specifically, when , it means that the ith GN has not completed the service, and the UAV can continue to provide communication services for it. On the contrary, when , it means that the GN has completed the service, and it does not need to be served in this round. The service state of all GNs can be mapped by using a binary list to calculate K, i.e. , so as to determine whether the task for each GN has been completed.

[0089] Action space: The action space A = {v, Θ} contains the flight speed v of the UAV and the service scheduling Θ of the GN. In the k time slot, the specific action a(k) is defined as containing the flight speed v[k] and the service scheduling Θ in the current time slot, i.e. The flight speed v[k] of the UAV is limited in the range of [v max ,v min ]. The GN selected in the k time slot is represented by .

[0090] Reward function: In the communication system, energy efficiency is an important indicator to evaluate the performance of the system, which is defined as the ratio between the amount of data transmitted and the energy consumed. In order to maximize the energy efficiency of the fixed-wing UAV-assisted wireless communication system, the DRL algorithm is adopted, and the energy efficiency is used as the reward function to evaluate the action a(k) performed by the UAV in the state s(k). Therefore, the reward r k can be represented as

[0091]

[0092] The data transmission rate at k time slot is denoted as The energy consumed by the fixed-wing UAV in k time slot is When the served GN is selected again, a penalty term

[0093] AS-VRE mechanism: The AS-VRE mechanism aims to improve the exploration efficiency of the agent to the environment, so that the deep neural network (DNN) can learn better. Each GN has different data service requirements, and the UAV needs to avoid repeatedly providing communication services to the GN that has been served. To solve this problem, action space planning is introduced in the DRL algorithm to optimize the scheduling of GNs.

[0094] In a round, each time slot will screen the required action in the action space, and taboo the screened illegal action in the round, so as to effectively avoid re-serving the GN that has met the required data volume. The illegal action is marked as a uc After the action screening (AS) mechanism, the experience available for learning in the experience pool is all compliant and positive, so the DNN will not learn negative experience. Simply performing the screening mechanism brings two problems: first, since the DNN will not learn negative experience, it will not give a low score on the illegal action. In fact, it is equally important for the DNN to give a low score to the illegal action as to give a high score to the compliant action. This effectively avoids negative factors, so that the agent can learn better. Second, simply screening illegal actions will not affect the DNN's scoring of illegal actions, because it is only an avoidance process. In addition, under the influence of the screening mechanism, the DNN has no difference in scoring whether the action is illegal. The only identifier of this difference is the state value K. The setting of the state value K is invalid due to its meaningless, and this confusion occurs in the DNN's scoring process of the action, resulting in poor training effect and falling into local optimum. Therefore, some virtual and bad experience that is not really executed needs to be added to the experience pool for the agent to learn.

[0095] Therefore, a virtual reward r v is added to the illegal action, so that the reward of the illegal action is lower than that of the compliant action. In this way, the learning effect of the neural network can be improved in the training process. The definition of the virtual reward is as follows:

[0096] r v = min(Q(s(k+1),a uc ))-1. (5)

[0097] Since the illegal action is not executed when the illegal action is selected, r v does not constitute the real value of the action.

[0098] The schematic diagram of the AS-VRE mechanism is as follows:Figure 4 As shown, its core contains two basic processes: action filtering process and virtual-real experience combination process. First, the state observation value is input to the neural network, and the output of the neural network will evaluate the rewards of all possible actions. Next, the action filtering step is performed to identify and exclude all illegal actions. When selecting an action, the action with the highest reward value will be selected from the eligible actions as the action to be executed, and this process is called the filtering process. However, since the action filtering process limits the experience pool to only include experiences of compliant actions, it may result in the neural network being unable to fully learn the experiences of illegal actions. To make up for this deficiency, a portion of virtual experiences (i.e., experiences of illegal actions that are not executed and have lower reward values) also need to be introduced into the experience pool. In this way, the neural network can be exposed to both compliant and non-compliant experiences during the learning process, allowing it to better understand the environment and make more accurate decisions.

[0099] NsHQPN algorithm: The NsHQPN algorithm is a DRL algorithm that can handle mixed action spaces. The NsHQPN algorithm aims to separate discrete actions and continuous actions and optimize them separately using Q-networks and policy networks. Subsequently, the results of these separate optimizations are integrated to obtain the overall strategy optimization result.

[0100] The action set is composed of a discrete action set C and a continuous action set A c . Wherein, the action a(k) can be represented as a(k) = (c(k), a c (k)), where c(k) represents a discrete action, c ∈ C, a c (k) is a continuous action corresponding to the discrete action c(k), a c ∈ A c The action a(k) taken in state s(k) is evaluated by the state-action value function Q(s(k), c(k), a c (k)). The Bellman equation is represented as:

[0101]

[0102] Where γ ∈ (0, 1] represents the discount factor.

[0103] In the NsHQPN algorithm, it is usually necessary to first calculate the optimal continuous action corresponding to each discrete action to implement the decision-making process of mixed actions. A policy network is used to make optimal decisions to achieve this process.

[0104]

[0105] Where pol(·) is a deterministic policy function used to select the optimal action a c from the continuous action set A cThe parameter set of the policy network is denoted by Φ. Then, the Q network makes decisions according to continuous actions and determines discrete actions.

[0106]

[0107] In each time slot, the agent starts from the current state s(k), accumulates the rewards obtained by performing N-step actions, and takes the Q value of the state s(k+N) reached after the N-step actions as the target expected value y k , which is used to update the Q value of the state s(k) and the action a(k).

[0108]

[0109] where N represents the number of steps from the current state to the target state. The parameter set of the Q network is denoted by Ψ.

[0110] Define the loss function for updating the network parameters Φ and Ψ:

[0111]

[0112] The principle diagram of the NsHQPN algorithm is shown in Figure 5 . The neural network scores each action according to the input state. Then, it needs to judge the violation action, filter out the qualified actions, and select the action with the maximum reward from the qualified actions to execute. The experience after execution must be added to the experience pool. At the same time, if a violation action is encountered, virtual experience that is not actually executed needs to be added to the experience pool. Finally, a batch of experiences are selected from the experience pool for neural network learning. The detailed algorithm process is as follows:

[0113]

[0114]

[0115] With the assistance of RIS, the MICH algorithm is used to plan the flight trajectory of the fixed-wing UAV. The AS-VER mechanism and the NsHQPN algorithm are used to plan the speed of the UAV and the scheduling of the GNs to improve the communication service of the GNs. Four groups of simulation experiments are conducted to verify the performance of the proposed scheme. First, the influence of RIS on the energy efficiency of the system is compared. Second, the influence of three different fixed-wing UAV flight trajectories on the energy efficiency of the system is compared. Then, the influence of the AS-VRE mechanism on the experimental results of the DRL algorithm is compared. Finally, the NsHQPN algorithm is compared with the traditional DRL algorithm.

[0116] Comparison of RIS performance: The specific influence of RIS on the reward and energy efficiency in the UAV-assisted communication system is discussed. As Figure 6 and Figure 7As shown, under the same flight trajectory and algorithm conditions, the UAV system equipped with RIS significantly improves communication performance. RIS reflects communication signals by adjusting phase shift, avoiding obstruction of the link between the UAV and GN by obstacles, thereby improving channel gain and system energy efficiency. Figure 6 and Figure 7 The reward and energy efficiency data in the data verified that the energy efficiency of the UAV system equipped with RIS is much higher than that of the system without RIS.

[0117] Comparison of different trajectories: The impact of different UAV trajectories on the experiment was compared. Trajectories planned by different algorithms are shown below. Figure 8 As shown, the circular trajectory scheme is less advantageous. Because GNs are distributed in a narrow pattern, the circular trajectory scheme results in a significant waste of trajectory location placement. A large portion of the trajectory locations are placed in areas with sparse or even no GN distribution. This is evident from… Figure 9 and Figure 10 The rewards and energy efficiency of this trajectory setup are also confirmed; it doesn't yield very high rewards or energy efficiency. The CH (Chain Rewards) scheme's trajectory has a characteristic: it can surround all GNs as much as possible. This surround is relative to the outermost GN, so there's still a certain distance from the center of the GN. Therefore, its rewards and energy efficiency are slightly higher than the circular trajectory. However, it still doesn't reach the optimal state. Compared to the CH algorithm's trajectory and the circular trajectory, the MICH algorithm's trajectory can further shrink towards the center based on the CH algorithm, reducing the average distance between the UAV and GN, lowering the path loss of the UAV, and improving channel gain. Therefore, this scheme is superior to the previous two schemes in terms of reward and energy efficiency.

[0118] Comparison of different mechanisms: This implementation compares the performance of AS-VRE, AS mechanism, and traditional penalty-based DRL, from... Figure 11 and Figure 12As can be seen, traditional punitive learning performs the worst, followed by DRL using the AS mechanism. DRL using the AS-VRE mechanism performs the best. This is because traditional punitive learning, due to its large solution space, struggles to discover effective actions during the exploration process. Therefore, after a period of training, it gets stuck in a local optimum. The AS mechanism, however, can significantly increase the probability of selecting effective actions during exploration by filtering out illegal actions, thus translating them into better rewards. The AS-VRE mechanism adds a virtual experience mechanism to the AS mechanism. These virtual experiences are low-reward, undesirable experiences that may not be executed. In other words, it not only tells the agent to seek effective decisions but also tells the agent which behaviors are bad and should be avoided, thus improving the efficiency of the DRL exploration process from both positive and negative perspectives. Therefore, DRL under the AS-VRE mechanism is also the best.

[0119] Comparison of Different DRL Algorithms: This implementation scheme evaluates the performance of four algorithms—NsHQPN, HQPN, Dueling DQN, and DDPG—in handling mixed action space problems. Compared to NsHQPN and HQPN, Dueling DQN and DDPG face the challenge of insufficient action acquisition accuracy when dealing with such problems. Because Dueling DQN discretizes continuous actions, the action space expands dramatically when the size of the discretized actions is small, affecting the algorithm's efficiency. Simultaneously, if the size of the discretized actions is too large, it reduces the accuracy of the actions. When handling discrete actions, DDPG may encounter some cases where the values ​​are meaningless. To make them meaningful, rounding is required, thus reducing accuracy. From... Figure 13 and Figure 14 As can be seen, during the exploration phase, the performance of the four algorithms did not differ significantly. However, after a period of training, the convergence performance of DuelingDQN and DDPG algorithms was inferior to that of NsHQPN and HQPN. Furthermore, Figure 13 and Figure 14 The reward and energy efficiency data validated that NsHQPN outperforms HQPN in both convergence speed and height. This is partly due to the limitations of the one-step update mechanism in handling delayed rewards, failing to fully utilize existing data and resulting in poor convergence. In contrast, NsHQPN's N-step update strategy more effectively utilizes information from multiple states, improving sample utilization and reducing bias during the update process. This mechanism allows the algorithm to better plan for the long term, leading to more stable training and higher convergence efficiency.

[0120] The specific embodiments of the present invention are written in a progressive manner, emphasizing the differences between the various implementation schemes, and the similar parts can be referred to each other.

[0121] The embodiments of the present application are described in detail above with reference to the accompanying drawings, but the present application is not limited to the above-described embodiments, and various changes can be made within the knowledge of those skilled in the art without departing from the spirit of the present application.

Claims

1. An energy efficiency optimization method for fixed-wing unmanned aerial vehicle communication coverage assisted by reflective intelligent surfaces, characterized in that, GN is communicated by a fixed-wing unmanned aerial vehicle, and a RIS is placed on the outer wall of a ground building to reflect the link; the energy efficiency optimization method comprises: S1: constructing a problem P with the goal of maximizing system energy efficiency and setting multiple constraint conditions; S2: dividing the problem P into two sub-problems, sub-problem one: planning a smooth, ground node distribution-adaptable hovering trajectory; sub-problem two: planning the flight speed of the unmanned aerial vehicle and the service scheduling of the GN on the planned trajectory to further improve the energy efficiency using a DRL algorithm; S3: using a MICH algorithm to plan the unmanned aerial vehicle trajectory to solve sub-problem one; S4: solving sub-problem two: first, using an AS-VRE mechanism to improve the exploration ability of DRL, and second, using an NsHQPN algorithm to improve the value accuracy of the mixed action space; the problem P in S1 is described as: (1) represents the goal is to maximize the system energy efficiency, under the constraint conditions (1a)-(1d); wherein (1a) represents the constraint of each GN in each time slot service state can only be served state and not served state, recorded with 1 and 0 respectively; (1b) represents that in a time slot, the UAV can only serve one GN, that is, the constraint of TDMA; (1c) represents that the cumulative transmission data volume of each GN must be greater than the requested data volume σ ι ;(1d) represents the upper and lower limits of the flight speed of the fixed-wing UAV.

2. The method for energy efficiency optimization of reflect- intelligent surface assisted fixed-wing drone communication coverage according to claim 1, wherein, the MICH algorithm in S3 comprises: first, solving a number of key trajectory nodes according to the distribution of GN; second, fitting the trajectory of the key trajectory nodes to make the trajectory curve smoother to adapt to the special turning constraint of the fixed-wing unmanned aerial vehicle.

3. The method for energy efficiency optimization of reflect- intelligent surface assisted fixed-wing drone communication coverage according to claim 1, wherein, The fixed-wing unmanned aerial vehicle trajectory mode is represented by the following equation: where a0∈[0,3000], b0∈[0,3000], x0∈[0,3000], y0∈[0,3000], and θ∈[-π,π] represent unknown parameters of the fitting curve; the predicted data is calculated according to the defined fitting curve function, and the calculation process is as follows: finally, gradient descent optimization is performed using a loss equation to obtain specific parameters; a Markov decision process is constructed: State space: The work of the fixed-wing UAV is to meet all GNs' communication demands, since the distance between the UAV and the selected GN directly affects the communication quality, a state space is defined The position of the UAV is denoted by K reflects the service state of all GNs, the service state of the GN K ι (k) is a binary variable, indicating whether the service to this GN is completed in the current time slot, when K ι (k) = 0, it means that the i-th GN has not completed the service, and the UAV continues to provide communication service for it, on the contrary, when K ι (k) = 1, it means that the GN has completed the service, and there is no need to provide service for it in this round; the service state of all GNs can be mapped by using a binary list to calculate K, that is Thus, it is determined whether the task for each GN is completed or not. Action space: Action space A = {v, Θ} contains the flight speed v of the UAV and the service scheduling Θ of the GN, and in the k time slot, the specific action a(k) is defined as containing the flight speed v[k] and the service scheduling Θ in the current time slot, that is The flight speed v[k] of the UAV is limited in the range of [v max ,v min ]; the GN selected in the k time slot is denoted as ; Reward function: The DRL algorithm is adopted, and the energy efficiency is taken as a reward function to evaluate the action a(k) performed by the UAV in the state s(k). Therefore, the reward r k may be expressed as The data transmission rate at k time slot is denoted by The energy consumed by the fixed-wing UAV in k time slot is When the served GN is selected again, a penalty term is introduced 4. The method for energy efficiency optimization of reflect- intelligent surface assisted fixed-wing drone communication coverage according to claim 1, wherein, the DRL is optimized using an AS-VRE mechanism: the AS-VRE mechanism includes two basic processes: an action screening process and a virtual-real experience combination process; first, the state observation value is input into the neural network, and the output of the neural network will evaluate the rewards of all possible actions; perform the action screening step to identify and exclude all illegal actions; when selecting an action, select the action with the highest reward value from the actions that meet the requirements as the action to be executed, i.e., the screening process; introduce a part of virtual experience, i.e., illegal action experience that has not been executed and has a low reward value, into the experience pool to enable the neural network to fully learn the experience of illegal actions; adding a virtual reward r to the violation action v making the reward of the violation action lower than the reward of the compliant action; the virtual reward is defined as follows: r v = min(Q(s(k + 1), a uc ))-1 (5) The violation action is labeled a uc Since the violation action is not performed when the violation action is selected, r v does not constitute a real action value.

5. The method for energy efficiency optimization of reflect- intelligent surface assisted fixed-wing drone communication coverage according to claim 1, characterized in that, the NsHQPN algorithm is: the neural network scores each action according to the input state; then, the neural network judges illegal actions, selects actions that meet the requirements, and selects the action with the highest reward from the actions that meet the requirements to execute; the experience after execution is added to the experience pool; at the same time, if an illegal action is encountered, the virtual experience that has not been truly executed needs to be added to the experience pool; finally, a batch of experience is selected from the experience pool for neural network learning; the NsHQPN algorithm is used to optimize the decision of the speed and GN service scheduling: The set of actions is composed of a discrete set of actions C and a continuous set of actions A c wherein an action a(k) can be represented as a(k) = (c(k), a c (k)) where c(k) represents a discrete action, c e C, a c (k) is a continuous action corresponding to the discrete action c(k), a c ∈ A c An action a(k) taken at state s(k) is evaluated by a state-action value function Q(s(k), c(k), a c (k)); the Bellman equation is represented as: where γ∈(0,1] represents a discount factor; the optimal continuous action corresponding to each discrete action is calculated, and a policy network is used to make the optimal decision; Here, pol(·) is a deterministic policy function used to select from the set of continuous actions A. c Select the optimal action a c The parameter set of the policy network is denoted by Φ. Then, the Q-network makes a decision according to the continuous action and determines the discrete action; In each time slot, the agent starts from the current state s(k), accumulates the rewards obtained by performing N-step actions, and takes the Q value of the state s(k+N) reached after the N-step actions as the target expected value y k , for updating the Q values of the state s(k) and the action a(k); Wherein, N represents the number of steps from the current state to the target state; the parameter set of the Q-network is denoted by Ψ; The loss function for updating the network parameters Φ and Ψ is defined as:

Citation Information

Patent Citations

  • Intelligent reflecting surface phase shifting and unmanned aerial vehicle path planning method

    CN114257298A

  • RIS-assisted multi-unmanned aerial vehicle high-energy-efficiency fair communication coverage method

    CN118764879A