Energy harvesting intelligent metasurface assisted uav edge computing energy efficiency optimization
Patent Information
- Application Number
- CN202610932611.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-22
AI Technical Summary
[0009]本发明的目的在于针对现有技术中无人机辅助移动边缘计算网络能耗高、资源耦合强、轨迹搜索复杂度高以及单层强化学习控制稳定性不足的问题,提出能量采集智能超表面辅助的无人机边缘计算能效优化方法
本发明针对多用户能量采集场景下的 STAR-IRS 辅助 UAV-MEC 网络,在综合考虑无人机飞行能耗、用户任务卸载需求、信道状态变化、STAR-IRS 反射/透射链路选择以及终端能量动态约束等实际因素的基础上,提出一种面向轨迹规划与通信计算资源联合优化的智能调度策略。该方法利用 STAR-IRS 同时透射与反射的全空间覆盖能力增强用户与基站之间的通信链路质量,并结合能量采集模型对用户终端的能量状态进行动态刻画,从而在保障能量因果约束和服务质量需求的前提下,提高系统资源利用效率。进一步地,本申请通过优先级感知的聚类轨迹规划方法生成无人机悬停点与参考路径,降低轨迹搜索复杂度;同时引入层级深度强化学习机制,由高层智能体学习能耗与服务性能之间的长期权衡关系,由低层智能体完成无人机运动控制、STAR-IRS 参数调节、发射功率分配和计算资源调度等连续决策,实现复杂网络环境下的协同优化与策略收敛。仿真结果表明,所提出方法在不同用户密度和动态信道条件下均具有较好的收敛速度、任务完成率和能效表现,相较于现有单一强化学习算法能够有效降低系统总能耗、提升服务覆盖能力和资源调度稳定性。该研究为绿色通信、空天地协同网络以及边缘智能计算中的资源管理问题提供了一种新的优化思路,为低能耗、高可靠 UAV-MEC 系统的高效部署提供了可行方案。
Smart Images

Figure CN122802972A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of mobile edge computing, green communication and wireless resource optimization, and in particular to an energy efficiency optimization method for UAV edge computing assisted by an intelligent metasurface for energy harvesting. Background Technology
[0002] With the rapid development of 6G mobile communication and Mobile Edge Computing (MEC), future wireless networks are evolving towards high reliability, low latency, and ubiquitous coverage. In complex terrain, emergency communication scenarios, and large-scale Internet of Things (IoT) applications, traditional fixed infrastructure often struggles to meet dynamic and high-density communication and computing demands. Unmanned aerial vehicles (UAVs), with their flexible deployment, high mobility, and natural line-of-sight (LoS) communication advantages, have been widely used to assist wireless communication and edge computing systems. However, UAVs are limited by their limited onboard battery capacity and high propulsion energy consumption; therefore, trajectory design and resource allocation become key factors directly affecting system energy efficiency and Quality of Service (QoS). Thus, minimizing total system energy consumption while ensuring mission completion and service fairness has become a fundamental research challenge.
[0003] In recent years, intelligent reflective surface (IRS) technology has become a key enabling technology for improving the performance of wireless systems. Compared with traditional IRS that only supports unilateral reflection, the Simultaneous Transmitting and Reflecting Intelligent Reflective Surface (STAR-IRS) can achieve dual-sided coverage and spatial resource reconfiguration by jointly controlling the amplitude and phase of the reflecting and transmitting units, thus offering significant advantages in improving link quality and expanding coverage. Integrating STAR-IRS into UAV-assisted MEC networks can effectively alleviate the obstruction problem between users and base stations and improve channel gain and overall energy efficiency through adaptive beam manipulation. However, the joint optimization of STAR-IRS configuration, UAV trajectory design, transmit power control, and computational resource allocation forms a dynamically changing, highly coupled, high-dimensional, non-convex problem. Traditional convex optimization or alternating optimization methods typically rely on precise model assumptions, making real-time adaptive decision-making difficult in such environments.
[0004] Deep Reinforcement Learning (DRL) offers a promising solution for dynamic wireless resource optimization because it can learn adaptive policies through continuous interaction with the environment without requiring iterative model-based optimization. Compared to traditional convex and alternating optimization methods, DRL is better suited for handling high-dimensional sequential decision problems involving UAV maneuver control, resource allocation, and energy-aware network management. However, the STAR-IRS-assisted UAV-MEC system considered in this invention exhibits a significant multi-timescale decision structure. On the one hand, the system needs to determine the long-term tradeoff between mission service performance and energy consumption at the macroscopic level; on the other hand, it needs to perform fine-grained continuous control over UAV motion, STAR-IRS parameters, transmit power, and computational resources at the microscopic level. A single-layer DRL framework may struggle to handle these coupled decisions simultaneously, leading to slow convergence, unstable training, or suboptimal energy-service performance.
[0005] Green MEC has been extensively studied to reduce energy consumption and latency in compute-intensive wireless applications by offloading tasks from resource-constrained devices to nearby edge servers. Existing research mainly focuses on task offloading, resource allocation, computation scheduling, and latency-energy tradeoff optimization. The paper "Research on MEC computing offload strategy for joint optimization of delay and energy consumption (Mathematical Biosciences and Engineering)" by Ni et al. studies the computation offloading problem in MEC systems and proposes an enhanced sine-cosine algorithm to jointly reduce latency and energy consumption. The paper "Optimal computation resource allocation in energy-efficient edge IoT systems with deep reinforcement learning (IEEE Transactions on Green Communications and Networking)" by Ansere et al. studies the computation resource allocation problem in energy-efficient edge IoT systems and proposes a method based on Lyapunov optimization and dual DRL to improve long-term energy efficiency under latency and stability constraints. Bishoyi and Misra's paper, "Enabling green mobile-edge computing for 5G-based healthcare applications (IEEE Transactions on Green Communications and Networking)," studies green MEC for 5G healthcare applications. They model the interaction between MEC servers and WBAN users as a Stackelberg game and develop the ADMM distributed algorithm to incentivize local computing and reduce the energy consumption cost of MEC servers. Taimoor et al.'s paper, "Resource optimization for minimizing latency and cost in UAV-assisted mobile edge computing networks (Computer Networks)," investigates resource optimization in UAV-assisted MEC networks and constructs a binary linear programming problem to jointly optimize computing, communication, and caching resources.While these heuristic and swarm intelligence methods can improve MEC performance, their convergence behavior often depends on parameter tuning and their online adaptive capabilities in dynamic wireless environments are limited.
[0006] UAV-assisted MEC further expands the service coverage and deployment flexibility of traditional MEC systems by leveraging UAV mobility and Loss-based air-to-ground links. Based on these advantages, recent research has focused on UAV trajectory design, user association, offloading decisions, and resource scheduling in UAV-enabled MEC networks. Xue et al.'s paper, "Joint optimization of trajectory and resource allocation in cellular-connected multi-uav MEC networks (Physical Communication)," studied multi-UAV MEC systems and minimized UAV energy consumption by decomposing the original non-convex problem into several subproblems and then solving them using continuous convex optimization and alternating updates. Wang et al.'s paper, "Multi-uav assisted air-ground collaborative MEC system: Drl-based joint task offloading and resource allocation and 3duav trajectory optimization (Drones)," proposed a UAV-assisted MEC framework in which offloading strategies and UAV trajectories are alternately optimized, while computational resource allocation is solved using convex optimization. Wang et al.'s paper, "Dynamic trajectory control and user association for UAV-assisted MEC: A deep reinforcement learning approach (Drones)," models UAV trajectory control and user association as Markov decision processes (MDPs) and develops PPO-based algorithms to optimize UAV trajectories and system performance. These works demonstrate the effectiveness of UAV trajectory optimization in MEC systems. However, most studies primarily consider direct communication links and separate or phase-process trajectory design and resource allocation, while occlusion effects and environmental perception link reconstruction have not been adequately considered.
[0007] IRS technology provides a programmable way to reconstruct the wireless propagation environment and improve energy efficiency by adjusting the phase shift and amplitude of passive reflective elements. In recent years, STAR-IRS has expanded upon traditional reflective IRS by adding simultaneous transmission and reflection capabilities, achieving full-space coverage and more flexible link enhancement. Zhang et al.'s paper, "Energyminimization for IRS-assisted Swipt-MEC system (Sensors)," studies an IRS-assisted Swipt-MEC system and jointly optimizes CPU frequency, transmit power, local computation bits, power split ratio, and IRS phase shift to minimize energy consumption. Compared to traditional IRS, STAR-IRS can simultaneously achieve reflection and transmission, providing full-space coverage and more flexible link enhancement. The paper by Chen et al., “Multi-irs assisted wireless-powered mobile edge computing for internet of things (IEEE Transactions on Green Communications and Networking),” studies a multi-IRS assisted wireless-powered MEC system for IoT, which jointly optimizes energy beamforming, multi-user detection, CPU frequency, transmit power, and IRS phase shift to maximize computing rate. The paper by Liu et al., “Star-ris-aided mobile edge computing: Computation rate maximization with binary amplitude coefficients (IEEE Transactions on Communications),” considers a STAR-IRS assisted MEC system and maximizes computing rate by jointly designing reflection / transmission coefficients, phase shift, receive beamforming, and resource allocation. The paper by Yan-Long C. et al., “Double reconfigurable intelligent surface-aided green IoT edge computing for computing capacity (Journal of Guangdong University of Technology),” studies a dual-IRS assisted system and optimizes the phase matrix, uplink transmit power, and local computing bits to improve throughput. Nevertheless, most IRS / STAR-IRS studies focus on static or quasi-static scenarios, where the IRS location is fixed and UAV mobility is not considered in conjunction.When STAR-IRS is integrated with UAV-enabled MEC, variables such as UAV trajectory, flight speed, STAR-IRS phase / amplitude coefficients, transmit power, and calculation frequency become highly coupled, making the problem more challenging.
[0008] DRL (Direct Resource Allocation) is increasingly used in wireless communication and MEC (Multi-access Edge Computing) optimization due to its ability to learn adaptive policies from dynamic environments. Compared to traditional optimization methods that typically rely on precise mathematical models and iterative solutions, DRL methods can directly learn resource allocation, offloading, and trajectory control policies through interaction with the environment. In particular, actor-commentator based DRL methods are suitable for continuous control problems and have been applied to UAV trajectory optimization, computational offloading, and green resource allocation in dynamic networks. The paper "Drl-based green resource provisioning for 5G and beyond networks (IEEE Transactions on Green Communications and Networking)" by M. Dieye et al. developed a DRL-based green resource provisioning scheme for 5G and beyond networks. The paper "Green drl-based embedding of 5G user plane network functions over heterogeneous platforms (IEEE Transactions on Green Communications and Networking)" by Shahsavand et al. proposed a GAT-enhanced DRL method for energy efficiency embedding of 5G user plane network functions over heterogeneous platforms. These studies demonstrate the potential of DRL in green resource optimization, but the multi-timescale joint control problem in STAR-IRS-assisted UAV-MEC networks remains unsolved. Summary of the Invention
[0009] The purpose of this invention is to address the problems of high energy consumption, strong resource coupling, high trajectory search complexity, and insufficient stability of single-layer reinforcement learning control in existing UAV-assisted mobile edge computing networks, and to propose an energy efficiency optimization method for UAV edge computing assisted by intelligent metasurfaces for energy harvesting.
[0010] The technical solution adopted by this invention to solve its technical problem is: an energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization method, which includes the following steps: Step 1: Establish a mobile edge computing network model for a UAV assisted by a smart metasurface that simultaneously transmits and reflects light. The network includes several user devices with energy harvesting capabilities, a UAV equipped with a smart metasurface that simultaneously transmits and reflects light, and a ground base station. Step 2: Establish the communication model of the network. Based on the geometric positional relationship between the user equipment, the UAV and the ground base station, determine whether the user equipment communicates with the ground base station through the reflection link or the transmission link of the smart metasurface that simultaneously transmits and reflects light. Calculate the channel gain and transmission rate of the corresponding link. Step 3: Establish the energy consumption model and battery energy model of the network, jointly consider the energy consumption of UAV flight, base station computing energy consumption, user equipment transmission energy consumption and solar energy collected by user equipment, and update the battery status of user equipment according to energy causal constraints. Step 4: Based on the communication model, energy consumption model and battery energy model, establish a joint optimization problem to minimize the total energy consumption of the system. By jointly optimizing the UAV trajectory, STAR-IRS reflection / transmission phase and amplitude, user equipment transmission power, base station computing resources and UAV flight speed, the system energy efficiency trajectory and resource joint optimization is achieved. Step 5: The joint optimization problem is solved using the hierarchical SAC-TD3 optimization algorithm HSTO-GC for green communication. The HSTO-GC includes a priority-aware UAV trajectory planning module and a hierarchical deep reinforcement learning energy optimization module. The priority-aware UAV trajectory planning module generates hovering points and reference paths, and the hierarchical deep reinforcement learning energy optimization module performs energy consumption-service trade-off decisions and continuous resource control through high-level SAC agents and low-level TD3 agents. Step 6: Output the optimized UAV trajectory, STAR-IRS reflection / transmission phase and amplitude, user equipment transmit power, base station computing resource allocation results, and system energy consumption optimization results.
[0011] Further, in step 1 of the present invention, let the set of user equipment be K={1,2,...,K}, and the coordinates of the ground base station, the UAV, and the k-th user equipment be represented as follows: and ,in, It fixes the flight altitude of the drone; when user equipment cannot communicate stably with the ground base station directly, the STAR-IRS carried by the drone helps to unload the computing tasks.
[0012] Furthermore, in step 2 of the present invention, STAR-IRS is composed of... Composed of several reflection / transmission units, in the first The reflection phase shift matrix and transmission phase shift matrix for each time slot are expressed as follows: (1) (2) in, Represents time slot The Middle The phase angle of each reflecting and transmitting unit. Represents time slot The Middle The amplitude of each reflecting and transmitting unit, and .
[0013] The phase update formula for STAR-IRS can be expressed as: (3) (4) in, Representing the Phase increment of each reflection and transmission unit.
[0014] The amplitude update formula for STAR-IRS can be expressed as: (5) (6) in, Representing the The amplitude increment of each reflection and transmission unit.
[0015] Furthermore, the method for determining the reflection link or transmission link in step 2 of the present invention includes: Calculate the user-to-drone vector and the vector from the base station to the drone The dot product of these two vectors can be expressed as: (7) Next, the included angle can be calculated using the dot product and the magnitude of the vector: (8) in, These represent the magnitudes of the vectors between the user and the drone, and between the base station and the drone, respectively.
[0016] Based on the calculated angle, the following conditions can be used to determine the outcome: (9) If the conditions are met, the user communicates with the base station through the reflection link; if the conditions are not met, the transmission link is used.
[0017] Furthermore, in step 2 of the present invention, the LoS component between the user and the STAR-IRS can be expressed as: (10) in, Represents the carrier wavelength. Represents the STAR-IRS cell spacing. Representing time slots Chinese users Angle of arrival at STAR-IRS.
[0018] Time slot The channel gain between the user and the STAR-IRS can be expressed as: (11) in, For the path loss at the reference distance, For users Distance to STAR-IRS For users Path loss index to STAR-IRS It is the Rice factor.
[0019] The LoS component between STAR-IRS and the base station can be expressed as: (12) in, Representing time slots The angle of incidence from the STAR-IRS to the base station.
[0020] Time slot The channel gain between STAR-IRS and the base station can be expressed as: (13) in, This refers to the distance from the STAR-IRS to the base station. This is the path loss index from STAR-IRS to the base station.
[0021] Therefore, time slot The channel gain from the user to the base station via the reflection link can be expressed as: (14) Similarly, time slots The channel gain from the user to the base station via the transmission link can be expressed as: (15) Therefore, time slot The transmission rate from the user to the base station can be expressed as: (16) (17) in, For channel bandwidth, For users Transmission power between the base station and the base station This represents noise power.
[0022] Furthermore, in step 3 of the present invention, the time slot The energy consumption for flight and transmission can be expressed as: (18) in, The drag coefficient, It is the front projected area. It is air density. It is flight speed. It refers to flight time.
[0023] Time slot The computing power consumption of a mid-base station can be expressed as: (19) in, It calculates the energy coefficient. The calculation frequency for the base station. To calculate the number of cycles.
[0024] Time slot Chinese users The transmission energy consumption can be expressed as: (20) in, This refers to the amount of data for the task.
[0025] User equipment In the time slot The amount of solar energy collected is expressed as: (twenty one) in, For solar energy conversion efficiency, For the collection area, Solar radiation intensity, The angle of incidence is denoted as .
[0026] User equipment The battery state is updated according to energy causality constraints: (twenty two) in, This is the maximum capacity of the battery.
[0027] Therefore, in time slots The total energy consumption is: (twenty three) Furthermore, the joint optimization problem in step 4 of the present invention is expressed as: (twenty four) Among them, C1 ensures that the user equipment battery power must always be higher than the safety threshold; C2 ensures that the cumulative energy consumption of all users does not exceed the sum of the accumulated collected energy and the initial power; C3 ensures that the horizontal flight position of the drone is limited to the preset service area; C4 and C5 limit the transmission power and the base station's calculation frequency to between the minimum and maximum allowable values, respectively; C5 and C6 are related to the reflection and transmission phase shift angles and coefficients of STAR-IRS, specifying that the phase shift angle is within the allowable range, and the sum of the reflection and transmission coefficients of each element is equal to 1.
[0028] Furthermore, in step 5 of the present invention, the priority-aware UAV trajectory planning module includes: A weighted distance or weighted feature is constructed based on user equipment location, remaining user energy, task priority, channel quality, UAV altitude, and service radius. User clusters are obtained using an adaptive weighted DBSCAN algorithm. A weighted centroid is calculated for each user cluster as the UAV hovering point, and the user cluster priority is calculated. A distance matrix is constructed based on the UAV's starting position, hovering point set, and target position. The distance matrix is adjusted according to the user cluster priority, and the access sequence is obtained through priority-aware nearest neighbor path search.
[0029] HSTO-GC performs energy-aware control through a two-layer DRL architecture. The higher-layer SAC agent focuses on learning long-term energy-service tradeoffs, while the lower-layer TD3 agent focuses on fine-grained continuous control.
[0030] Furthermore, in step 5 of the present invention, the high-level SAC agent in HSTO-GC is modeled as a Markov decision process: ,in and Let these represent the high-level state space, action space, reward function, and discount factor, respectively. The high-level state can be represented as: (25) in, The average remaining energy level for users. Indicates the urgency of the task. For average channel quality, For task completion rate, Indicates energy harvesting potential. Indicates system load. For the UAV horizontal position, For normalized time index, This represents the state disturbance term. This indicates the path progress characteristics.
[0031] High-level actions are defined as: (26) in, and These are the original continuous outputs related to energy-saving preferences and service-oriented preferences. Used to determine the service target preferences of high-level personnel.
[0032] High-level rewards are defined as: (27) in, This indicates the increase in the number of users served. Indicates the increment of task completion. This represents the current task completion rate. This represents the normalized energy penalty. This is a step-by-step punishment for high-ranking officials.
[0033] Furthermore, in step 5 of the present invention, the low-level TD3 agent in HSTO-GC is modeled as a Markov decision process: The lower-level state can be written as: (28) in, This indicates UAV motion information, such as position and heading; This indicates communication related to STAR-IRS; Indicates the state of resources and energy; Represents local user characteristics. Indicates an enhancement.
[0034] Low-level actions are represented as: (29) in, Control the STAR-IRS reflection / transmission phase and amplitude configuration; and Control the UAV's heading and speed; and Adjust the transmit power and calculate the frequency respectively; This represents the remaining auxiliary continuous control components.
[0035] Lower-level rewards are defined as: (30) in, Indicates path progress. Indicates the increment of service or task completion. To normalize energy consumption, This is a low-level step penalty.
[0036] Beneficial effects: This invention addresses STAR-IRS-assisted UAV-MEC networks in multi-user energy harvesting scenarios. Based on comprehensive consideration of practical factors such as UAV flight energy consumption, user task offloading requirements, channel state changes, STAR-IRS reflection / transmission link selection, and dynamic energy constraints of terminals, it proposes an intelligent scheduling strategy for joint optimization of trajectory planning and communication / computing resources. This method leverages the full-space coverage capability of STAR-IRS's simultaneous transmission and reflection to enhance the communication link quality between users and base stations. It also combines an energy harvesting model to dynamically characterize the energy state of user terminals, thereby improving system resource utilization efficiency while ensuring energy causal constraints and service quality requirements. Furthermore, this application generates UAV hovering points and reference paths through a priority-aware clustering trajectory planning method, reducing trajectory search complexity. Simultaneously, a hierarchical deep reinforcement learning mechanism is introduced, where higher-level agents learn the long-term trade-off between energy consumption and service performance, while lower-level agents complete continuous decisions such as UAV motion control, STAR-IRS parameter adjustment, transmit power allocation, and computing resource scheduling, achieving collaborative optimization and strategy convergence in complex network environments. Simulation results demonstrate that the proposed method exhibits superior convergence speed, task completion rate, and energy efficiency under varying user densities and dynamic channel conditions. Compared to existing single reinforcement learning algorithms, it effectively reduces total system energy consumption, enhances service coverage, and improves resource scheduling stability. This research provides a novel optimization approach for resource management in green communications, space-air-ground collaborative networks, and edge intelligent computing, offering a feasible solution for the efficient deployment of low-energy, high-reliability UAV-MEC systems. Attached Figure Description
[0037] Figure 1 STAR-IRS-assisted UAV-MEC network system architecture based on energy harvesting Figure 2 The proposed HSTO-GC algorithm's overall framework for joint optimization of energy efficiency trajectory and resources. Figure 3 Priority-aware UAV trajectory planning results based on adaptive user clustering and hovering point generation Figure 4 Comparison of training reward convergence performance of HSTO-GC, SAC, and TD3 algorithms Figure 5 Comparison of energy consumption and convergence performance of HSTO-GC, SAC, and TD3 algorithms Figure 6 Comparison of task completion rates of HSTO-GC, SAC, and TD3 algorithms Figure 7 Comparison of task completion rates with different numbers of users Figure 8 Comparison of energy consumption convergence performance under different numbers of users Figure 9 Comparison of training reward convergence performance with different numbers of users Detailed Implementation
[0038] The invention will now be further described in detail with reference to the accompanying drawings.
[0039] Example like Figure 1 As shown, the energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization method proposed in this invention includes the following steps: Step 1: Establish a mobile edge computing network model for a UAV assisted by a smart metasurface that simultaneously transmits and reflects light. The network includes several user devices with energy harvesting capabilities, a UAV equipped with a smart metasurface that simultaneously transmits and reflects light, and a ground base station.
[0040] This invention assumes the user equipment set is K={1,2,...,K}, and the coordinates of the ground base station, the UAV, and the k-th user equipment are respectively represented as: and ,in, It fixes the flight altitude of the drone; when user equipment cannot communicate stably with the ground base station directly, the STAR-IRS carried by the drone helps to unload the computing tasks.
[0041] Step 2: Establish the communication model of the network. Based on the geometric positional relationship between the user equipment, the UAV and the ground base station, determine whether the user equipment communicates with the ground base station through a reflective link or a transmission link of the smart metasurface that simultaneously transmits and reflects light. Calculate the channel gain and transmission rate of the corresponding link.
[0042] A STAR-IRS consists of M reflection / transmission units, which can establish a good channel between the user and the base station. The phase of each reflection / transmission unit in a STAR-IRS is controllable. Therefore, the phase shift matrix of reflection and transmission in the nth time slot of a STAR-IRS can be expressed as: (31) (32) in, Represents time slot The Middle The phase angle of each reflecting and transmitting unit. Represents time slot The Middle The amplitude of each reflecting and transmitting unit, and .
[0043] The phase update formula for STAR-IRS can be expressed as: (33) (34) in, Representing the Phase increment of each reflection and transmission unit.
[0044] The amplitude update formula for STAR-IRS can be expressed as: (35) (36) in, Representing the The amplitude increment of each reflection and transmission unit.
[0045] Calculate the user-to-drone vector and the vector from the base station to the drone The dot product of these two vectors can be expressed as: (37) Next, the included angle can be calculated using the dot product and the magnitude of the vector: (38) in, These represent the magnitudes of the vectors between the user and the drone, and between the base station and the drone, respectively.
[0046] Based on the calculated angle, the following conditions can be used to determine the outcome: (39) If the conditions are met, the user communicates with the base station through the reflection link; if the conditions are not met, the transmission link is used.
[0047] The LoS component between the user and STAR-IRS can be represented as: (40) in, Represents the carrier wavelength. Represents the STAR-IRS cell spacing. Representing time slots Chinese users Angle of arrival at STAR-IRS.
[0048] Time slot The channel gain between the user and the STAR-IRS can be expressed as: (41) in, For the path loss at the reference distance, For users Distance to STAR-IRS For users Path loss index to STAR-IRS It is the Rice factor.
[0049] The LoS component between STAR-IRS and the base station can be expressed as: (42) in, Representing time slots The angle of incidence from the STAR-IRS to the base station.
[0050] Time slot The channel gain between STAR-IRS and the base station can be expressed as: (43) in, This refers to the distance from the STAR-IRS to the base station. This is the path loss index from STAR-IRS to the base station.
[0051] Therefore, time slot The channel gain from the user to the base station via the reflection link can be expressed as: (44) Similarly, time slots The channel gain from the user to the base station via the transmission link can be expressed as: (45) Therefore, time slot The transmission rate from the user to the base station can be expressed as: (46) (47) in, For channel bandwidth, For users Transmission power between the base station and the base station This represents noise power.
[0052] Step 3: Establish the energy consumption model and battery energy model of the network, jointly consider the energy consumption of UAV flight, base station computing, user equipment transmission, and solar energy collected by user equipment, and update the battery status of user equipment according to energy causal constraints.
[0053] In this invention, the UAV acts as an aerial relay node of the system, and its energy consumption mainly consists of flight energy consumption. (Time slot) The energy consumption for flight and transmission can be expressed as: (48) in, The drag coefficient, It is the front projected area. It is air density. It is flight speed. It refers to flight time.
[0054] Time slot The computing power consumption of a mid-base station can be expressed as: (49) in, It calculates the energy coefficient. The calculation frequency for the base station. To calculate the number of cycles.
[0055] Assuming all tasks from ground users are uploaded to the drone for processing, the energy consumption of ground user equipment consists only of transmission energy. Therefore, time slots Chinese users The transmission energy consumption can be expressed as: (50) in, This refers to the amount of data for the task.
[0056] Time slot The energy collected by solar energy can be expressed as: (51) in, For solar energy conversion efficiency, For the collection area, Solar radiation intensity, The angle of incidence is denoted as .
[0057] User equipment The battery state is updated according to energy causality constraints: (52) in, This is the maximum capacity of the battery.
[0058] Therefore, in time slots The total energy consumption is: (53) Step 4: Based on the communication model, energy consumption model, and battery energy model, establish a joint optimization problem to minimize the total energy consumption of the system. By jointly optimizing the UAV trajectory, STAR-IRS reflection / transmission phase and amplitude, user equipment transmission power, base station computing resources, and UAV flight speed, the system energy efficiency trajectory and resource joint optimization is achieved.
[0059] This invention constructs a joint optimization problem that minimizes the total system energy consumption by coordinating the optimization of the UAV trajectory, STAR-IRS phase shift and amplitude, user transmit power, and UAV flight speed, thereby achieving a balance between energy efficiency and fairness. The optimization objective is as follows: (54) Among them, C1 ensures that the user equipment battery power must always be higher than the safety threshold; C2 ensures that the cumulative energy consumption of all users does not exceed the sum of the accumulated collected energy and the initial power; C3 ensures that the horizontal flight position of the drone is limited to the preset service area; C4 and C5 limit the transmission power and the base station's calculation frequency to between the minimum and maximum allowable values, respectively; C5 and C6 are related to the reflection and transmission phase shift angles and coefficients of STAR-IRS, specifying that the phase shift angle is within the allowable range, and the sum of the reflection and transmission coefficients of each element is equal to 1.
[0060] Step 5: The joint optimization problem is solved using the hierarchical SAC-TD3 optimization algorithm HSTO-GC for green communication. The HSTO-GC includes a priority-aware UAV trajectory planning module and a hierarchical deep reinforcement learning energy optimization module. The priority-aware UAV trajectory planning module generates hovering points and reference paths, while the hierarchical deep reinforcement learning energy optimization module performs energy consumption-service trade-off decisions and continuous resource control through high-level SAC agents and low-level TD3 agents.
[0061] Figure 2The overall framework of HSTO-GC is demonstrated, consisting of two coupled modules: priority-aware UAV trajectory planning and energy optimization based on hierarchical SAC-TD3. Specifically, the trajectory planning module first performs clustering, hover point generation, and path construction based on the location and service status of ground users. During clustering, user battery urgency, task priority, and channel quality are considered to identify key user groups and generate weighted centroids as hover points for each cluster. Then, based on the distance matrix between hover points and their corresponding priorities, a priority-aware access order is obtained, providing a structured reference trajectory for subsequent continuous control and reducing the exploration burden of deep reinforcement learning. Based on the generated reference path, the hierarchical deep reinforcement learning module performs closed-loop energy-aware control. At higher levels, SAC learns energy-service preference decisions at a slower timescale, adjusting the trade-off between energy saving and task service; while at lower levels, TD3 generates fine-grained continuous actions at a faster timescale, including UAV motion control, STAR-IRS parameter tuning, transmit power allocation, and computational resource control. In each training round, the environment is first reset, and the path planner generates an initial reference trajectory. At each time step, the higher-level SAC outputs weighted decisions based on the global system state and path progress, while the lower-level TD3 generates continuous control actions based on the local state, the next waypoint, and higher-level preferences. After the joint actions are executed, the mission completion progress, transmission performance, path progress, and energy consumption are evaluated, a hierarchical reward is constructed, and the corresponding transformations are stored and sampled for offline policy updates. Therefore, HSTO-GC forms an end-to-end optimization loop consisting of trajectory planning, hierarchical decision-making, continuous control, experience replay, and parameter updates.
[0062] A weighted distance or weighted feature is constructed based on user equipment location, remaining user energy, task priority, channel quality, UAV altitude, and service radius. An adaptive weighted DBSCAN algorithm is used to obtain user clusters. A weighted centroid is calculated for each user cluster as the UAV hovering point, and the user cluster priority is calculated. A distance matrix is constructed based on the UAV's starting position, hovering point set, and target position. The distance matrix is adjusted according to the user cluster priority, and the access sequence is obtained through priority-aware nearest neighbor path search. HSTO-GC performs energy-aware control through a two-layer DRL architecture. The higher-level SAC agent focuses on long-term energy-service tradeoff learning, while the lower-level TD3 agent focuses on fine-grained continuous control.
[0063] The high-level decision-making process is modeled as an MDP, denoted as ,in and Let these represent the high-level state space, action space, reward function, and discount factor, respectively. The high-level state can be represented as: (55) in, The average remaining energy level for users. Indicates the urgency of the task. For average channel quality, For task completion rate, Indicates energy harvesting potential. Indicates system load. For the UAV horizontal position, For normalized time index, This represents the state disturbance term. This indicates the path progress characteristics.
[0064] High-level actions are defined as: (56) in, and These are the original continuous outputs related to energy-saving preferences and service-oriented preferences. Used to determine high-level service target preferences. The first two components are decoded and normalized as follows: (57) here, and These represent energy-saving weights and service-oriented weights, respectively, and are transmitted to the lower-level controllers as guiding information.
[0065] High-level rewards are defined as: (58) in, This indicates the increase in the number of users served. Indicates the increment of task completion. This represents the current task completion rate. This represents the normalized energy penalty. This is a step-by-step punishment for high-ranking officials.
[0066] Low-level decision-making process modeling The lower-level state can be written as: (59) in, This indicates UAV motion information, such as position and heading; This indicates communication related to STAR-IRS; Indicates the state of resources and energy; This represents local user characteristics. The enhancement is defined as: (60) in, and These are respectively the energy-saving weight for high-rise buildings and the service-oriented weight; This represents the direction vector of the UAV to the next waypoint. This is the normalized distance to the next waypoint.
[0067] Low-level actions are represented as: (61) in, Control the STAR-IRS reflection / transmission phase and amplitude configuration; and Control the UAV's heading and speed; and Adjust the transmit power and calculate the frequency respectively; This represents the remaining auxiliary continuous control components.
[0068] Lower-level rewards are defined as: (62) in, Indicates path progress. Indicates the increment of service or task completion. To normalize energy consumption, This is a low-level step penalty.
[0069] Step 6: Output the optimized UAV trajectory, STAR-IRS reflection / transmission phase and amplitude, user equipment transmit power, base station computing resource allocation results, and system energy consumption optimization results.
[0070] Figure 3 The green dots represent the initial position of the UAV, the scattered dots represent user equipment to be served, the dashed circles represent the user cluster boundaries formed based on user spatial distribution, remaining energy, task priority, and channel quality, the square markers represent the center positions of each user cluster, the red crosses represent the UAV hovering points generated according to the weighted centroid mechanism, and the blue broken lines represent the flight trajectory of the UAV obtained by priority-aware path planning. Through this trajectory planning method, the UAV can preferentially approach user clusters with lower remaining energy, higher task priority, or poorer channel conditions, reducing invalid search distances across the entire service area and providing a structured reference path for subsequent hierarchical SAC-TD3 energy optimization, thereby reducing trajectory search complexity and improving task service efficiency and system energy efficiency.
[0071] The effects of the present invention will be further explained in detail below with reference to simulation experiments, specifically including: 1. Simulation hardware requirements The experimental platform of this invention is equipped with a 3.7 GHz Intel Core i5-12600KF CPU, an NVIDIA GeForce RTX 4060 Ti GPU, 32 GB of RAM, and a 1 TB solid-state drive. The SAC-TD3 hybrid model was implemented using PyCharm IDE in a PyTorch 2.6.0 and Python 3.9 environment to complete the deep reinforcement learning algorithm and was compared with two baseline methods. A 300 × 300 m² area was set in the simulation, with the UAV flying at a fixed altitude of 40 m. The base station was located at the center of the area, and the start and end points of the UAV were both set to [150, 150, 40]. The main simulation parameter settings are shown in Table 1.
[0072] 2. Simulation Content Figure 4 compares the training reward convergence of HSTO-GC, SAC, and TD3. Since the reward function jointly considers task completion, trajectory advancement, transmission performance, and energy consumption penalties, a reward closer to 0 indicates better overall performance. HSTO-GC converges to a higher reward level than SAC and TD3, indicating its stronger ability to learn long-term energy-service tradeoffs. The brief drop in reward around rounds 300–400 is due to an instantaneous policy rebalancing process; at this point, the higher-level SAC adjusts its energy-saving preferences, while the lower-level TD3 adapts to its continuous control policy. After this adjustment, HSTO-GC quickly recovers and maintains optimal convergence performance.
[0073] Figure 5 illustrates the energy consumption convergence of the three algorithms. All methods exhibit fluctuations in energy consumption during the early stages of training due to policy exploration. As training progresses, HSTO-GC achieves the lowest convergence energy consumption, demonstrating that the proposed hierarchical structure can learn more energy-efficient UAV trajectories and resource allocation strategies. The short-term energy consumption increase around rounds 300–400 corresponds to the reward decrease in Figure 4, which is attributed to the policy rebalancing between task completion and energy conservation. Compared to SAC and TD3, HSTO-GC better suppresses unnecessary motion and resource consumption through coordinated high-level preference learning and low-level continuous control.
[0074] Figure 6 compares the task completion rates of HSTO-GC, SAC, and TD3 under the same simulation settings. HSTO-GC achieves a task completion rate of 0.987, higher than SAC and TD3, whose completion rates are 0.829 and 0.784, respectively. This indicates that under limited flight time and energy constraints, HSTO-GC can complete more user tasks. The performance improvement mainly comes from the hierarchical design: the higher-level SAC captures global energy consumption-service preferences, while the lower-level TD3 performs fine-grained control over UAV motion, STAR-IRS configuration, transmit power, and computational resources. Therefore, HSTO-GC provides better service coverage and resource utilization compared to individual baseline methods.
[0075] Figure 7 shows the task completion rate of HSTO-GC under different user densities. When the number of users is 10, 15, and 20, the completion rate is close to 1, indicating that the proposed algorithm can effectively coordinate trajectory planning, STAR-IRS auxiliary link enhancement, and resource allocation in low to medium load scenarios. However, when the number of users increases to 25 and 30, the completion rate drops significantly. This is because the service area, task volume, and scheduling complexity increase simultaneously, while UAV mobility, bandwidth, service radius, and energy resources remain limited. Therefore, excessively high user density will cause the system to enter an overload state.
[0076] Figure 8 illustrates the energy efficiency convergence of HSTO-GC under different user numbers. When the number of users is 10 and 15, energy consumption decreases rapidly and stabilizes at a low level, indicating that the algorithm can quickly learn an energy-efficient trajectory under light load conditions. When the number of users is 20, energy consumption remains at a moderate level, but experiences a brief increase around rounds 300–400 due to policy rebalancing. When the number of users increases to 25 and 30, energy consumption is higher and fluctuates more significantly because the UAVs need to adjust their trajectories more frequently, and resource allocation becomes more complex. This result confirms that user density directly affects system energy efficiency.
[0077] Figure 9 shows the reward convergence of HSTO-GC under different user densities. When the number of users is 10, 15, and 20, the reward is higher and more stable, indicating that the system can achieve a good balance between task completion and energy consumption in medium-load scenarios. Conversely, when the number of users increases to 25 and 30, the reward decreases and the fluctuations become more pronounced. This is because high user density increases both the task incomplete penalty and the energy consumption penalty. Furthermore, a larger service set introduces stronger coupling between trajectory control, service ordering, and resource allocation, making policy learning more difficult. This result is consistent with the task completion rate and energy consumption results.
[0078] Based on the above analysis and discussion, this invention studies the energy efficiency trajectory and resource optimization problem in a STAR-IRS-assisted UAV-enabled MEC network for energy harvesting, jointly considering UAV trajectory, STAR-IRS reflection / transmission configuration, transmit power, computational resource allocation, propulsion energy consumption, and user energy harvesting constraints. To address the resulting high-dimensional non-convex optimization problem, this paper proposes the HSTO-GC algorithm. Specifically, firstly, a priority-aware trajectory planning module based on weighted adaptive clustering is developed to generate service-aware hovering points and reference paths according to user energy urgency, task priority, and channel quality. Subsequently, a hierarchical reinforcement learning framework is designed, in which the high-level SAC agent learns the long-term energy-service tradeoff, while the low-level TD3 agent performs fine-grained continuous control over UAV motion, STAR-IRS parameter adjustment, transmit power allocation, and computational resource optimization. Simulation results show that compared with SAC and TD3 alone, HSTO-GC achieves faster convergence, lower system energy consumption, higher long-term rewards, and better task completion rate. Results at different user densities further demonstrate that HSTO-GC maintains stable performance in low to medium load scenarios, while the performance degradation at excessively high user densities reveals the impact of UAV mobility and resource constraints. Future work will consider multi-UAV collaboration, imperfect channel state information, dynamic mission arrival, and more realistic energy harvesting environments.
Claims
1. Energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization, characterized in that, The method includes the following steps: Step 1: Establish a mobile edge computing network model for a UAV assisted by a smart metasurface that simultaneously transmits and reflects light. The network includes several user devices with energy harvesting capabilities, a UAV equipped with a smart metasurface that simultaneously transmits and reflects light, and a ground base station. Step 2: Establish the communication model of the network. Based on the geometric positional relationship between the user equipment, the UAV and the ground base station, determine whether the user equipment communicates with the ground base station through the reflection link or the transmission link of the smart metasurface that simultaneously transmits and reflects light. Calculate the channel gain and transmission rate of the corresponding link. Step 3: Establish the energy consumption model and battery energy model of the network, jointly consider the energy consumption of UAV flight, base station computing energy consumption, user equipment transmission energy consumption and solar energy collected by user equipment, and update the battery status of user equipment according to energy causal constraints. Step 4: Based on the communication model, energy consumption model and battery energy model, establish a joint optimization problem to minimize the total energy consumption of the system. By jointly optimizing the UAV trajectory, STAR-IRS reflection / transmission phase and amplitude, user equipment transmission power, base station computing resources and UAV flight speed, the system energy efficiency trajectory and resource joint optimization is achieved. Step 5: The joint optimization problem is solved using the hierarchical SAC-TD3 optimization algorithm HSTO-GC for green communication. The HSTO-GC includes a priority-aware UAV trajectory planning module and a hierarchical deep reinforcement learning energy optimization module. The priority-aware UAV trajectory planning module generates hovering points and reference paths, and the hierarchical deep reinforcement learning energy optimization module performs energy consumption-service trade-off decisions and continuous resource control through high-level SAC agents and low-level TD3 agents. Step 6: Output the optimized UAV trajectory, STAR-IRS reflection / transmission phase and amplitude, user equipment transmit power, base station computing resource allocation results, and system energy consumption optimization results.
2. The energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization according to claim 1, characterized in that, In step 1, let the set of user equipment be K={1,2,...,K}, and the coordinates of the ground base station, the UAV, and the k-th user equipment be represented as follows: and ,in, It fixes the flight altitude of the drone; when user equipment cannot communicate stably with the ground base station directly, the STAR-IRS carried by the drone helps to unload the computing tasks.
3. The energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization according to claim 1, characterized in that, In step 2, STAR-IRS is composed of Composed of several reflection / transmission units, in the first The reflection phase shift matrix and transmission phase shift matrix for each time slot are expressed as follows: (1); (2); in, Represents time slot The Middle The phase angle of each reflecting and transmitting unit. Represents time slot The Middle The amplitude of each reflecting and transmitting unit, and The phase update formula for STAR-IRS can be expressed as: (3); (4); in, Representing the Phase increment of each reflection and transmission unit, The amplitude update formula for STAR-IRS can be expressed as: (5); (6); in, Representing the The amplitude increment of each reflection and transmission unit.
4. The energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization according to claim 1, characterized in that, The method for determining the reflection link or transmission link in step 2 includes: Calculate the user-to-drone vector and the vector from the base station to the drone The dot product of these two vectors can be expressed as: (7); Next, the included angle can be calculated using the dot product and the magnitude of the vector. : (8); in, Let these represent the magnitudes of the vectors between the user and the drone, and between the base station and the drone, respectively. Based on the calculated angle The following conditions can be used to determine this: (9); If the conditions are met, the user communicates with the base station through the reflection link; if the conditions are not met, the transmission link is used.
5. The energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization according to claim 1, characterized in that, In step 2, the line-of-sight channel component from the user equipment to the STAR-IRS is represented as follows: (10) in, Represents the carrier wavelength. Represents the STAR-IRS cell spacing. Representing time slots Chinese users Angle of arrival at STAR-IRS Time slot The channel gain between the user and the STAR-IRS can be expressed as: (11); in, For the path loss at the reference distance, For users Distance to STAR-IRS For users Path loss index to STAR-IRS Rice factor, The LoS component between STAR-IRS and the base station can be expressed as: (12); in, Representing time slots The angle of incidence from the STAR-IRS to the base station. Time slot The channel gain between STAR-IRS and the base station can be expressed as: (13); in, This refers to the distance from the STAR-IRS to the base station. The path loss index from STAR-IRS to the base station. Therefore, time slot The channel gain from the user to the base station via the reflection link can be expressed as: (14); Similarly, time slots The channel gain from the user to the base station via the transmission link can be expressed as: (15); Therefore, time slot The transmission rate from the user to the base station can be expressed as: (16); (17); in, For channel bandwidth, For users Transmission power between the base station and the base station This represents noise power.
6. The energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization according to claim 1, characterized in that, In step 3, an energy consumption model and a battery energy model are established, including the UAV flight energy consumption, base station computing energy consumption, user equipment transmission energy consumption, user equipment solar energy collection, battery status update, and total system energy consumption. (18); in, The drag coefficient, It is the front projected area. It is air density. It is flight speed. It's flight time. Time slot The computing power consumption of a mid-base station can be expressed as: (19); in, It calculates the energy coefficient. The calculation frequency for the base station. To calculate the number of cycles, Time slot Chinese users The transmission energy consumption can be expressed as: (20); in, For the amount of task data, User equipment In the time slot The amount of solar energy collected is expressed as: (twenty one); in, For solar energy conversion efficiency, For the collection area, Solar radiation intensity, Angle of incidence User equipment The battery state is updated according to energy causality constraints: (22); in, This is the maximum battery capacity. Therefore, in time slots The total energy consumption is: (23)。 7. The energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization according to claim 1, characterized in that, The joint optimization problem in step 4 is expressed as: (24); Among them, C1 ensures that the user equipment battery power must always be higher than the safety threshold; C2 ensures that the cumulative energy consumption of all users does not exceed the sum of the accumulated collected energy and the initial power; C3 ensures that the horizontal flight position of the drone is limited to the preset service area; C4 and C5 limit the transmission power and the base station's calculation frequency to between the minimum and maximum allowable values, respectively; C5 and C6 are related to the reflection and transmission phase shift angles and coefficients of STAR-IRS, specifying that the phase shift angle is within the allowable range, and the sum of the reflection and transmission coefficients of each element is equal to 1.
8. The energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization according to claim 1, characterized in that, In step 5, the priority-aware UAV trajectory planning module includes: A weighted distance or weighted feature is constructed based on user equipment location, remaining user energy, task priority, channel quality, UAV altitude, and service radius. User clusters are obtained using an adaptive weighted DBSCAN algorithm. A weighted centroid is calculated for each user cluster as the UAV hovering point, and the user cluster priority is calculated. A distance matrix is constructed based on the UAV's starting position, hovering point set, and target position. The distance matrix is adjusted according to the user cluster priority, and the access sequence is obtained through priority-aware nearest neighbor path search.
9. The energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization according to claim 1, characterized in that, In step 5, the high-level SAC agent in HSTO-GC is modeled as a Markov decision process: ,in and Let these represent the high-level state space, action space, reward function, and discount factor, respectively. The high-level state can be represented as: (25); in, The average remaining energy level for users. Indicates the urgency of the task. For average channel quality, For task completion rate, Indicates energy harvesting potential. Indicates system load. For the UAV horizontal position, For normalized time index, This represents the state disturbance term. Indicates path progress characteristics. High-level actions are defined as: (26); in, and These are the original continuous outputs related to energy-saving preferences and service-oriented preferences. Used to determine the service target preferences of high-level personnel. High-level rewards are defined as: (27); in, This indicates the increase in the number of users served. Indicates the increment of task completion. This represents the current task completion rate. This represents the normalized energy penalty. This is a step-by-step punishment for high-ranking officials.
10. The energy harvesting intelligent metasurface-assisted UAV edge computing energy efficiency optimization according to claim 1, characterized in that, In step 5, the low-level TD3 agent in HSTO-GC is modeled as a Markov decision process: The lower-level state can be written as: (28); in, Indicates UAV motion information; This indicates communication related to STAR-IRS; Indicates the state of resources and energy; Represents local user characteristics; Indicates an enhancement term. Low-level actions are represented as: (29); in, Control the STAR-IRS reflection / transmission phase and amplitude configuration; and Control the UAV's heading and speed; and Adjust the transmit power and calculate the frequency respectively; Indicates the remaining auxiliary continuous control components. Lower-level rewards are defined as: (30); in, Indicates path progress. Indicates the increment of service or task completion. To normalize energy consumption, This is a low-level step penalty.