RIS-assisted UAV-MEC task unloading system and method based on heterogeneous agent
By using a heterogeneous intelligent agent RIS-assisted UAV-MEC system, combined with the HAPPO algorithm and solar charging, the trajectory optimization and energy constraints of UAVs and RIS in the Internet of Vehicles were solved, achieving efficient offloading of computing tasks and load balancing, and improving task success rate and solar energy utilization efficiency.
Patent Information
- Application Number
- CN202511035765.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-11
AI Technical Summary
In the vehicle-to-everything (V2X) environment, the combination of UAVs and RIS suffers from problems such as lack of trajectory optimization, energy constraints, and unbalanced load, resulting in low task offloading efficiency and inability to meet the response time and accuracy requirements of computationally intensive applications.
A RIS-assisted UAV-MEC task offloading system based on heterogeneous intelligent agents is adopted. The HAPPO algorithm is used for task offloading path planning, UAV trajectory optimization and RIS phase configuration. Combined with a solar charging system, energy consumption and latency are optimized.
It significantly improved the success rate of task completion and the efficiency of solar energy utilization, reduced system energy consumption and latency, and achieved load balancing and efficient offloading of computing tasks.
Smart Images

Figure CN120935658A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a RIS-assisted UAV-MEC task offloading system and method based on heterogeneous intelligent agents, belonging to the fields of mobile edge computing, vehicle networking and deep reinforcement learning. Background Technology
[0002] With the rapid development of 5G and 6G technologies, the Internet of Vehicles (IoV) is now able to support time-sensitive and computationally intensive applications, such as augmented reality navigation, real-time collaborative perception, and autonomous driving, which require millisecond-level response times. However, the inherent computing and storage capabilities of vehicles are limited, making it impossible to process large-scale data and complex tasks while meeting stringent time and accuracy requirements. Therefore, task offloading technology has emerged, offloading computational tasks from vehicles to external nodes with abundant resources (such as Roadside Units, RSUs), which has become a necessary approach to improve task processing efficiency and system performance.
[0003] In real-world IoV environments, the spatiotemporal heterogeneous distribution of vehicles, the bursty nature of tasks, and the suboptimal deployment of RSUs (Roadside Units) collectively lead to some RSUs being overloaded by computational demands, while others are underutilized. This unbalanced and bursty computational load severely degrades the efficiency and responsiveness of IoV, becoming a major challenge in applying task offloading in large-scale vehicle-to-everything (V2X) networks.
[0004] Recently, unmanned aerial vehicles (UAVs) have been introduced into vehicle-to-everything (VEC) edge computing networks to transfer tasks from busy nodes to remote, underloaded nodes. Furthermore, researchers have proposed a Smart Reflector (RIS) technology to collaborate with UAVs for IoV task offloading. Leveraging its reconfigurability, the RIS can dynamically adjust the network topology according to actual needs, addressing poor transmission quality and channel instability caused by signal fading and congestion in complex urban traffic networks.
[0005] While the invention of UAVs and RIS has brought about a paradigm shift in how computing tasks are offloaded and completed, significant challenges remain hindering their application and deployment in practical IoV. First, although integrating UAVs into task offloading architectures enables rapid task propagation, an effective mechanism for guiding UAV trajectory optimization is still lacking. Second, adopting RIS on UAVs requires complex design and control to fully leverage the combined advantages of both. Third, the long-term service capability of UAVs is limited due to their limited battery power and poor sustainability in harvesting energy from the environment. Summary of the Invention
[0006] The technical problem to be solved by this invention is to provide a RIS-assisted UAV-MEC task offloading system and method based on heterogeneous intelligent agents. The system adopts the Heterogeneous Intelligent Agent Proximal Policy Optimization (HAPPO) algorithm, which can effectively reduce system energy consumption and latency, significantly improve solar energy utilization efficiency and task completion success rate, and solve key technical problems such as load imbalance, channel instability and UAV energy limitation in the vehicle network environment.
[0007] The technical solution adopted in this invention is: a RIS-assisted UAV-MEC task offloading system and method based on heterogeneous intelligent agents, comprising:
[0008] The vehicle, as a mobile terminal that generates computationally intensive tasks, sends computational task requests and status information to the RIS-assisted UAVs in its grid via wireless channels;
[0009] RSU, an edge computing infrastructure deployed on the roadside, receives computing tasks from RIS-assisted UAV relays and performs processing.
[0010] The RIS-assisted UAV is a drone equipped with a smart reflector, used for mission relay and signal enhancement. It is responsible for collecting real-time status information of vehicles and RSUs within its grid and uploading it to the central controller. At the same time, it performs mission forwarding, trajectory planning and RIS phase configuration according to the instructions of the central controller.
[0011] The solar charging system, with solar panels mounted on the UAV, provides a sustainable energy supply for the RIS-assisted UAV;
[0012] The central controller, a global coordinator located in the cloud, maintains communication with all RIS-assisted UAVs. It is responsible for receiving global network status information, making decisions on task offloading paths and resource allocation strategies, and issuing control commands to each RIS-assisted UAV.
[0013] Specifically, the tasks generated by the vehicle are time-sensitive and computationally intensive, including real-time safety-critical tasks, augmented reality tasks, and batch analysis tasks.
[0014] Specifically, the RIS-assisted UAV is equipped with L reconfigurable reflective elements, each of which can dynamically adjust phase and amplitude parameters.
[0015] Specifically, the solar charging system includes three components: direct normal irradiance, atmospheric diffuse irradiance, and ground reflected irradiance.
[0016] Specifically, the communication between the RIS-assisted UAV and the vehicle and RSU adopts multiple-input multiple-output (MIMO) technology, which improves system capacity and link reliability through spatial diversity and multiplexing gain.
[0017] Specifically, the computing tasks generated by the vehicle are transmitted through RIS-assisted UAVs, including two modes: intra-network communication and inter-network communication. Intra-network communication refers to the computing tasks generated by the vehicle being directly transmitted to the associated RSU within the same route coverage area via a single-hop relay through the RIS-assisted UAV of its own grid, avoiding cross-grid transmission overhead. Inter-network communication refers to the vehicle transmitting tasks to RSUs within the same grid, which is relayed through the RIS-assisted UAV of that grid and then transmitted to the RIS-assisted UAVs of other grids through multi-hop RIS-assisted UAV cooperation. The RIS-assisted UAV of the target grid then forwards the task to the RSU within that grid for execution, thereby achieving cross-grid load balancing.
[0018] A RIS-assisted UAV-MEC task offloading method based on heterogeneous intelligent agents, the specific steps of which are as follows:
[0019] Step 1: System modeling and initialization, constructing a three-layer heterogeneous network architecture, including K RSUs and M vehicles in the terminal layer, N RIS-assisted UAVs in the edge layer, and a central controller in the cloud layer, and initializing the state parameters of each node;
[0020] Step 2: Communication channel modeling. The Rayleigh fading channel model is adopted to establish the channel coefficient matrices from the vehicle to the RIS-assisted UAV, between RIS-assisted UAVs, and from the RIS-assisted UAV to the RSU, taking into account line-of-sight and non-line-of-sight propagation components.
[0021] Step 3: Task and Energy Modeling, representing the computational tasks generated by the vehicle as triples. This includes task size, computational density, and maximum tolerable latency, while also establishing a UAV-RIS energy model, including propulsion energy consumption, RIS operation energy consumption, and solar energy harvesting.
[0022] Step 4: Construct a multi-objective optimization problem, establishing a joint optimization problem with the objectives of minimizing system energy consumption and delay and maximizing solar energy utilization efficiency;
[0023] Step 5: Model the multi-agent Markov decision process, defining the central controller as the global coordinator and the RIS-assisted UAV as the autonomous agent, and designing the state space, action space and reward function respectively.
[0024] Step 6: Layered strategy architecture design, defining the strategy mapping π between the central controller and the RIS-assisted UAV agent. c and and the corresponding value function approximation and
[0025] Step 7: Trust domain constraint policy optimization. A truncation replacement objective function is used to prevent excessive policy bias. Stable updates are achieved through importance sampling ratio and dominance function estimation.
[0026] Step 8: Update the strategy in layers. Update the strategy network parameters in the order of central controller → UAV agent, and use the respective learning rate for gradient descent optimization.
[0027] Step 9: Iterative optimization and convergence determination. Repeat Step 5 to Step 8 until the strategy converges, and output the optimal task unloading path, UAV trajectory and RIS phase configuration.
[0028] Specifically, in Step 2, the communication channel modeling uses a dual-slope path loss model to describe the vehicle-to-RIS-assisted UAV link and a free-space path loss model to describe the RIS-assisted UAV-to-RSU link.
[0029] Specifically, the formula for calculating the power collected by solar energy in Step 3 is as follows:
[0030]
[0031] Where η solar For energy conversion efficiency, A solar I represents the effective area of the solar panel. n (t) represents the total incident solar irradiance.
[0032] Specifically, Step 7 is as follows:
[0033] The core mechanism of the heterogeneous agent proximal policy optimization algorithm adopts a hierarchical decision-making mechanism. The central controller performs task offloading path planning based on global state information, while the RIS-assisted UAV agent performs joint optimization of trajectory, phase, and power based on local observations. The trust domain constraint is implemented through the truncation mechanism of the PPO algorithm to ensure the stability of policy updates and prevent performance degradation during training. Multi-objective optimization is achieved, while considering the trade-off between minimizing energy consumption, minimizing latency, and maximizing solar energy utilization. An adaptive learning rate adjustment strategy is adopted, with the central controller and UAV agent using different learning rates to ensure the convergence of hierarchical optimization.
[0034] The truncation and substitution objective function is:
[0035]
[0036] in This represents the state-action pair (s) of the central controller. c ,a c The mathematical expectation of ), where ∈ is the truncation parameter controlling the size of the trust region, θ cThese are the policy network parameters, r c (θ c () is the ratio of the probability of the new policy to that of the old policy on a given state-action pair. It is the dominance function.
[0037] The beneficial effects of this invention are as follows: It fully considers the dynamic and complex nature of the vehicle-to-everything (V2X) environment and constructs a collaborative optimization framework integrating task offloading path selection, UAV trajectory planning, RIS phase configuration, and energy management. Through a heterogeneous agent near-end policy optimization algorithm, it combines global coordination by the central controller with distributed execution by the UAV agent, effectively solving the problems of slow convergence speed and local optima in traditional methods when dealing with large-scale dynamic networks. The introduction of a solar power system significantly improves the sustainable service capability of UAVs and reduces reliance on traditional batteries. Simulation results show that compared with existing methods, this invention significantly improves task completion delay, system energy consumption, and success rate. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of the system model of the present invention;
[0039] Figure 2 This is a schematic diagram of the hierarchical multi-agent reinforcement learning framework of the HAPPO algorithm of this invention;
[0040] Figure 3 This is a detailed structural diagram of the RIS-assisted UAV of the present invention. Detailed Implementation
[0041] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0042] Example 1: As Figure 1-3 As shown, a RIS-assisted UAV-MEC task offloading system based on heterogeneous intelligent agents includes:
[0043] Vehicles, as mobile terminals that generate computationally intensive tasks, include real-time safety-critical tasks, augmented reality tasks, and batch processing analysis tasks.
[0044] RSU, an infrastructure with edge computing capabilities deployed on the roadside;
[0045] RIS-assisted UAVs are drones equipped with intelligent reflective surfaces for mission relay and signal enhancement. RIS-assisted UAVs are equipped with L reconfigurable reflective elements, each of which can dynamically adjust phase and amplitude parameters.
[0046] The solar charging system provides a sustainable energy supply for RIS-assisted UAVs. The solar charging system includes three components: direct normal irradiance, atmospheric diffuse irradiance, and ground reflected irradiance.
[0047] The central controller, a global coordinator located in the cloud, is responsible for task offloading path decisions and resource allocation.
[0048] The communication between the RIS-assisted UAV and the vehicle and RSU adopts MIMO technology, which improves system capacity and link reliability through spatial diversity and multiplexing gain.
[0049] The computing tasks generated by the vehicle are transmitted through RIS-assisted UAVs, including two modes: intra-network communication and inter-network communication. Intra-network communication refers to the computing tasks generated by the vehicle being directly transmitted to the associated RSU within the same route coverage area via a single-hop relay through the RIS-assisted UAV of its own grid, avoiding cross-grid transmission overhead. Inter-network communication refers to the vehicle transmitting tasks to RSUs within the same grid, which are then relayed through the RIS-assisted UAV of that grid and transmitted to the RIS-assisted UAVs of other grids through multi-hop RIS-assisted UAV cooperation. The RIS-assisted UAV of the target grid then forwards the task to the RSUs within that grid for execution, thereby achieving cross-grid load balancing.
[0050] A method for a RIS-assisted UAV-MEC task offloading system based on heterogeneous intelligent agents, the specific steps of which are as follows:
[0051] Step 1: System Modeling and Initialization. This invention constructs a three-layer heterogeneous network architecture to solve the task offloading problem in a RIS-assisted UAV-MEC network. This architecture includes: 1) Terminal layer, comprising K RSUs and M vehicles. RSUs are evenly deployed and equipped with MEC servers for computation offloading. Vehicles travel on roads and randomly generate computational tasks; 2) Edge layer with N UAVs equipped with RIS panels. Each RIS panel contains L reconfigurable reflective elements. Each RIS-assisted UAV is equipped with a solar charging system, collecting solar radiation including direct normal irradiance, atmospheric diffuse irradiance, and ground reflected irradiance through solar panels. Energy conversion enables sustainable power supply, ensuring the long-term service capability of the UAVs. The entire area is divided into multiple grids, each patrolled by a designated UAV; 3) Cloud layer, containing a central controller that aggregates real-time network status information for global task scheduling and resource allocation.
[0052] like Figure 1As shown, in each time slot, the UAV collects real-time status information of vehicles and RSUs within its assigned grid and sends it to the central controller along with its own real-time status information. Once a vehicle has a computational task, it first sends a request to the central controller via the UAV relay responsible for that area. Upon receiving this request, the central controller makes an optimal offloading decision based on global status information, specifying the offloading route and location. Each UAV then makes its own decisions based on the instructions from the central controller, including specific flight paths, RIS phase shifts, and transmission power.
[0053] The system employs MIMO technology to fully utilize spatial diversity and multiplexing gains to enhance system capacity and link reliability. Computational tasks generated by vehicles may be located within the same routing coverage area as their associated RSUs, enabling single-hop communication. Simultaneously, they can also be transmitted across meshes and offloaded to other RSUs to achieve load balancing across the entire network.
[0054] Step 2: Communication channel modeling. The communication model adopts the Rayleigh fading channel model. The channel coefficient matrices are established for vehicle to RIS-assisted UAV, between RIS-assisted UAVs, and from RIS-assisted UAV to RSU, taking into account line-of-sight and non-line-of-sight propagation components.
[0055] Vehicle-to-everything (V2X) communication environments exhibit both line-of-sight (LoS) paths and multipath scattering propagation characteristics. (Definition) Let m be the complex channel coefficient matrix between the m-th vehicle and the n-th RIS. This is the complex channel coefficient matrix between the nth RIS and the kth RSU.
[0056] The path losses from vehicle to RIS and from RIS to RSU are calculated using a double-slope path loss model and a free-space path loss model, respectively. The double-slope model introduces the breakpoint distance d. c To distinguish between near-field and far-field propagation characteristics, the path loss index approaches free-space propagation over short distances, while over long distances it reflects the effects of obstacles and multipath effects in urban environments.
[0057] Step 3: Task and Energy Modeling, representing the computational tasks generated by the vehicle as triples. This includes task size, computational density, and maximum tolerable latency, while also establishing a UAV-RIS energy model, including propulsion energy consumption, RIS operation energy consumption, and solar energy harvesting.
[0058] The formula for calculating the power collected by solar energy is: Where η solar For energy conversion efficiency, A solar I represents the effective area of the solar panel. n (t) represents the total incident solar irradiance.
[0059] The RIS-assisted UAV energy model consists of three components: propulsion energy consumption, RIS operation energy consumption, and solar energy harvesting. The propulsion energy consumption formula is: Where c1 and c2 are aerodynamic constants related to the physical properties of the UAV, v n (t) is the flight speed of the nth RIS-assisted UAV at time t, a n (t) represents the acceleration of n RIS-assisted UAVs at time t, where g represents the acceleration, and ΔK n (t) represents the change in kinetic energy of the UAV between consecutive time steps.
[0060] RIS operating power consumption includes phase shift adjustment power consumption and circuit maintenance power consumption. Phase shift adjustment power consumption quantifies the power consumption of the phase and amplitude parameters of each reflector during dynamic reconfiguration, while circuit maintenance power consumption refers to the baseline power consumption required to maintain the normal operation of the RIS system.
[0061] The formula for calculating total solar irradiance is:
[0062] Total solar irradiance includes direct normal irradiance G DNI (t), Atmospheric scattering irradiance G DIFF (t) and ground reflected irradiance G REFL (t) Three components.
[0063] Step 4: Construct a multi-objective optimization problem, establishing a joint optimization problem with the objectives of minimizing system energy consumption and delay and maximizing solar energy utilization efficiency;
[0064] Step 5: Model the multi-agent Markov decision process, defining the central controller as the global coordinator and the RIS-assisted UAV as the autonomous agent, and designing the state space, action space and reward function respectively.
[0065] Step 6: Layered strategy architecture design, defining the strategy mapping π between the central controller and the RIS-assisted UAV agent. c and and the corresponding value function approximation and
[0066] Step 7: Trust domain constraint policy optimization. A truncation replacement objective function is used to prevent excessive policy bias. Stable updates are achieved through importance sampling ratio and dominance function estimation.
[0067] Step 7.1: The core mechanism of the heterogeneous agent proximal policy optimization algorithm adopts a hierarchical decision-making mechanism. The central controller performs task offloading path planning based on global state information, and the RIS-assisted UAV agent performs joint optimization of trajectory, phase, and power based on local observations. The trust domain constraint is implemented through the truncation mechanism of the PPO algorithm to ensure the stability of policy updates and prevent performance degradation during training. Multi-objective optimization is achieved, while considering the trade-off optimization of minimizing energy consumption, minimizing latency, and maximizing solar energy utilization. An adaptive learning rate adjustment strategy is adopted, with the central controller and UAV agent using different learning rates to ensure the convergence of hierarchical optimization.
[0068] The heterogeneous agent proximal policy optimization algorithm transforms the original mixed-integer nonlinear programming problem into a hierarchical multi-agent reinforcement learning problem. The central controller, acting as a global coordinator, is responsible for task offloading decisions and dynamic RSU resource allocation strategies, while the UAV, as an autonomous agent, dynamically adjusts its flight trajectory, RIS phase configuration, and transmission power allocation.
[0069] The algorithm uses a truncation substitution for the objective function to prevent excessive policy bias.
[0070] in This represents the state-action pair (s) of the central controller. c ,a c The mathematical expectation of r) c (θ c () is the ratio of the probability of the new policy to that of the old policy on a given state-action pair. The truncation parameter θ controls the size of the trust region. c These are policy network parameters. It is the dominance function.
[0071] Step 8: Update the policy in layers, following the order of central controller → UAV agent: Where η c and η u These are the learning rates of the central controller and the UAV, θ. c These are the central controller policy network parameters. These are the network parameters of the nth RIS-assisted UAV policy, optimized using gradient descent with their respective learning rates.
[0072] Step 9: Iterative optimization and convergence determination. Repeat Step 5 to Step 8 until the strategy converges, and output the optimal task unloading path, UAV trajectory and RIS phase configuration.
[0073] The core technical features of heterogeneous agent proximal policy optimization algorithms:
[0074] Hierarchical decision-making mechanism: The central controller uses global state information to plan the task offloading path, and the RIS-assisted UAV agent performs joint optimization of trajectory, phase and power based on local observations;
[0075] Trust domain constraint: The truncation mechanism of the PPO algorithm ensures the stability of policy updates and prevents performance degradation during training.
[0076] Multi-objective optimization: Simultaneously considering the trade-offs between minimizing energy consumption, minimizing delay, and maximizing solar energy utilization;
[0077] Adaptive learning rate adjustment: The central controller and the UAV agent use different learning rates to achieve convergence guarantee for hierarchical optimization.
[0078] This invention, through the combination of the model and algorithm, can significantly reduce task completion delay and system energy consumption, improve task success rate and solar energy utilization efficiency, thereby achieving efficient processing and sustainable development of computationally intensive applications in the Internet of Vehicles.
[0079] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A RIS-assisted UAV-MEC task offloading system based on heterogeneous intelligent agents, characterized in that: include: The vehicle, as a mobile terminal that generates computationally intensive tasks, sends computational task requests and status information to the RIS-assisted UAVs in its grid via wireless channels; RSU, an edge computing infrastructure deployed on the roadside, receives computing tasks from RIS-assisted UAV relays and performs processing. The RIS-assisted UAV is a drone equipped with a smart reflector, used for mission relay and signal enhancement. It is responsible for collecting real-time status information of vehicles and RSUs within its grid and uploading it to the central controller. At the same time, it performs mission forwarding, trajectory planning and RIS phase configuration according to the instructions of the central controller. The solar charging system, with solar panels mounted on the UAV, provides a sustainable energy supply for the RIS-assisted UAV; The central controller, a global coordinator located in the cloud, maintains communication with all RIS-assisted UAVs. It is responsible for receiving global network status information, making decisions on task offloading paths and resource allocation strategies, and issuing control commands to each RIS-assisted UAV.
2. The RIS-assisted UAV-MEC task offloading system based on heterogeneous intelligent agents according to claim 1, characterized in that: The tasks generated by the vehicle are time-sensitive and computationally intensive, including real-time safety-critical tasks, augmented reality tasks, and batch analysis tasks.
3. The RIS-assisted UAV-MEC task offloading system based on heterogeneous intelligent agents according to claim 1, characterized in that: The RIS-assisted UAV is equipped with L reconfigurable reflective elements, each of which can dynamically adjust phase and amplitude parameters.
4. The RIS-assisted UAV-MEC task offloading system based on heterogeneous intelligent agents according to claim 1, characterized in that: The solar charging system comprises three components: direct normal irradiance, atmospheric diffuse irradiance, and ground reflected irradiance.
5. The RIS-assisted UAV-MEC task offloading system based on heterogeneous intelligent agents according to claim 1, characterized in that: The communication between the RIS-assisted UAV and the vehicle and RSU adopts MIMO technology, which improves system capacity and link reliability through spatial diversity and multiplexing gain.
6. The RIS-assisted UAV-MEC task offloading system based on heterogeneous intelligent agents according to claim 1, characterized in that: The computing tasks generated by the vehicle are transmitted through RIS-assisted UAVs, including two modes: intra-network communication and inter-network communication. Intra-network communication refers to the computing tasks generated by the vehicle being directly transmitted to the associated RSU within the same route coverage area via a single-hop relay through the RIS-assisted UAV of its own grid, avoiding cross-grid transmission overhead. Inter-network communication refers to the vehicle transmitting tasks to RSUs within the same grid, which are then relayed through the RIS-assisted UAV of that grid and transmitted to the RIS-assisted UAVs of other grids through multi-hop RIS-assisted UAV cooperation. The RIS-assisted UAV of the target grid then forwards the task to the RSUs within that grid for execution, thereby achieving cross-grid load balancing.
7. A RIS-assisted UAV-MEC task offloading method based on heterogeneous intelligent agents, characterized in that: The specific steps are as follows: Step 1: System modeling and initialization, constructing a three-layer heterogeneous network architecture, including K RSUs and M vehicles in the terminal layer, N RIS-assisted UAVs in the edge layer, and a central controller in the cloud layer, and initializing the state parameters of each node; Step 2: Communication channel modeling. The Rayleigh fading channel model is adopted to establish the channel coefficient matrices from the vehicle to the RIS-assisted UAV, between RIS-assisted UAVs, and from the RIS-assisted UAV to the RSU, taking into account line-of-sight and non-line-of-sight propagation components. Step 3: Task and Energy Modeling, representing the computational tasks generated by the vehicle as triples. This includes task size, computational density, and maximum tolerable latency, while also establishing a UAV-RIS energy model, including propulsion energy consumption, RIS operation energy consumption, and solar energy harvesting. Step 4: Construct a multi-objective optimization problem, establishing a joint optimization problem with the objectives of minimizing system energy consumption and delay and maximizing solar energy utilization efficiency; Step 5: Model the multi-agent Markov decision process, defining the central controller as the global coordinator and the RIS-assisted UAV as the autonomous agent, and designing the state space, action space and reward function respectively. Step 6: Layered strategy architecture design, defining the strategy mapping π between the central controller and the RIS-assisted UAV agent. c and and the corresponding value function approximation and Step 7: Trust domain constraint policy optimization. A truncation replacement objective function is used to prevent excessive policy bias. Stable updates are achieved through importance sampling ratio and dominance function estimation. Step 8: Update the strategy in layers, updating the strategy network parameters in the order of central controller → UAV agent, and using their respective learning rates for gradient descent optimization; Step 9: Iterative optimization and convergence determination. Repeat Step 5 to Step 8 until the strategy converges, and output the optimal task unloading path, UAV trajectory and RIS phase configuration.
8. The RIS-assisted UAV-MEC task offloading method based on heterogeneous intelligent agents according to claim 7, characterized in that: In Step 2, the communication channel modeling uses a dual-slope path loss model to describe the vehicle-to-RIS-assisted UAV link and a free-space path loss model to describe the RIS-assisted UAV-to-RSU link.
9. The RIS-assisted UAV-MEC task offloading method based on heterogeneous intelligent agents according to claim 7, characterized in that: The formula for calculating the power collected by solar energy in Step 3 is as follows: Where η solar For energy conversion efficiency, A solar I represents the effective area of the solar panel. n (t) represents the total incident solar irradiance.
10. The RIS-assisted UAV-MEC task offloading method based on heterogeneous intelligent agents according to claim 7, characterized in that: Step 7 specifically refers to: The core mechanism of the heterogeneous agent proximal policy optimization algorithm adopts a hierarchical decision-making mechanism. The central controller performs task offloading path planning based on global state information, while the RIS-assisted UAV agent performs joint optimization of trajectory, phase, and power based on local observations. The trust domain constraint is implemented through the truncation mechanism of the PPO algorithm to ensure the stability of policy updates and prevent performance degradation during training. Multi-objective optimization is achieved, while considering the trade-off between minimizing energy consumption, minimizing latency, and maximizing solar energy utilization. An adaptive learning rate adjustment strategy is adopted, with the central controller and UAV agent using different learning rates to ensure the convergence of hierarchical optimization. The truncation and substitution objective function is: in This represents the state-action pair (s) of the central controller. c ,a c The mathematical expectation of ), where ∈ is the truncation parameter controlling the size of the trust region, θ c These are the policy network parameters, r c (θ c () is the ratio of the probability of the new policy to that of the old policy on a given state-action pair. It is the dominance function.
Citation Information
Cited By
Satellite network resource self-adaptive allocation method based on layered intelligent agent
CN121462058A
Unmanned aerial vehicle group path planning and task unloading method and device for edge calculation
CN121680435A