Method for optimizing and scheduling computing power resources of Internet of Things based on RIS-UAV assistance
By employing a resource allocation method based on RIS-UAV-assisted THz communication and multi-agent reinforcement learning, the problems of limited THz communication coverage and dynamic environment adaptability in IoT systems are solved. This enables low-latency and low-energy resource collaborative configuration in IoT systems, improving the system's dynamic adaptability and computing performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF TECH
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-12
AI Technical Summary
In IoT systems, terahertz communication has limited coverage and is sensitive to occlusion. Fixed-deployment reconfigurable smart surfaces are difficult to adapt to dynamic environments, making it difficult for traditional optimization strategies to meet the dynamic and real-time requirements of THz communication systems. Furthermore, variables such as task offloading decisions, bandwidth allocation, edge computing resource allocation, and RIS-UAV trajectory control are strongly coupled, making it difficult to achieve stable and real-time joint decision-making.
A multi-agent reinforcement learning-based resource allocation method is adopted. By using RIS-UAV-assisted THz communication and combining a cloud-edge-device collaborative architecture, task offloading decision-making, bandwidth allocation, edge computing resource allocation, and joint optimization of RIS phase control and RIS-UAV motion control are performed. The RA-MASAC algorithm is used for deep iterative learning in the multi-agent Markov decision-making process to achieve collaborative resource allocation and low latency and low energy consumption in dynamic environments.
The system achieves global collaborative configuration of communication and computing resources in a dynamic environment, reducing system latency and energy consumption, improving the system's adaptability to large-scale access and complex load changes, and significantly improving the performance of low-latency and low-energy computing networks in IoT scenarios.
Smart Images

Figure CN122028078A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for optimizing and scheduling IoT computing resources based on RIS-UAV, specifically a cloud-edge-device joint resource allocation method assisted by a Reconfigurable Intelligent Surface Unmanned Aerial Vehicle (RIS-UAV) under a terahertz (THz) communication link, belonging to the field of communication network technology. Background Technology
[0002] In recent years, the Internet of Things (IoT) has shown a trend of large-scale expansion, with the number of terminals growing rapidly and application scenarios becoming increasingly widespread, ranging from smart transportation and telemedicine to smart agriculture. These new applications typically have more stringent Quality of Service (QoS) requirements, especially placing higher demands on communication capacity and computing power; while the traditional approach of "relying solely on local terminal computing" often struggles to meet the processing needs of low-latency tasks.
[0003] To reduce terminal energy consumption and improve task processing efficiency, computing power networks, as a new type of network architecture, connect heterogeneous computing resources through networking, enabling them to collaborate and be flexibly invoked, thereby supporting the offloading of tasks requiring strong computing power to edge or cloud computing nodes for execution.
[0004] However, existing CPN research often focuses on traditional communication methods such as millimeter waves. In scenarios involving high-density access and large-scale data transmission, the communication efficiency still struggles to meet the continuously growing demands. Terahertz communication, utilizing higher frequency resources, has the potential to significantly improve transmission performance and is considered one of the important candidate technologies to meet future high-throughput requirements. However, at the same time, THz base stations have relatively small coverage areas and are highly sensitive to obstruction, leading to limited link reliability and available coverage, which in turn affects the continuous service capability of THz-based CPN systems in dynamic scenarios.
[0005] Reconfigurable smart surfaces (RIS) consist of a large number of independently adjustable phase-shift reflective units. By adjusting the reflection phase, they can improve the propagation environment and increase the transmission rate, and are considered one of the important means to overcome THz obstruction and coverage problems. Existing work mostly focuses on fixed-deployment RIS; however, fixed RIS are difficult to adapt to time-varying system characteristics such as terminal mobility and dynamic changes in services, and therefore cannot meet the dynamic and real-time requirements of CPN systems based on THz communication.
[0006] With the development of unmanned aerial vehicle (UAV) technology, integrating a Rectifier Array (RIS) into a UAV to form a RIS-UAV can further enhance transmission performance by dynamically adjusting the spatial position of the RIS and combining it with phase control. This provides a new technical approach to solving the limitations of occlusion and coverage in THz scenarios. Based on this, integrating RIS-UAV, THz communication, and CPN architecture is expected to simultaneously improve network transmission and task computing performance, adapting to the ever-evolving real-time and complexity requirements of IoT scenarios.
[0007] However, in the aforementioned integrated system, variables such as task offloading decision, bandwidth allocation, edge computing resource allocation, RIS reflection phase control, and RIS-UAV trajectory control are strongly coupled with each other. Furthermore, the decision space is high-dimensional, non-convex, and subject to multiple constraints, making it difficult for traditional analytical optimization or step-by-step strategies to achieve stable and real-time joint decision-making in dynamic environments.
[0008] Against this backdrop, Multi-Agent Reinforcement Learning (MARL) is considered a feasible solution to this type of joint resource allocation problem because it has the ability to make distributed decisions by multiple agents and can learn effective policies in complex high-dimensional non-convex optimization spaces. Summary of the Invention
[0009] This invention targets dynamic and ever-changing IoT computing network systems, characterized by a large number of terminals, time-varying task arrival and terminal locations, heterogeneous and limited communication and computing resources, and the adoption of a cloud-edge-device three-layer collaborative architecture to support low-latency service processing. In this complex scenario, while terahertz communication offers the advantage of high bandwidth, it suffers from limited coverage and sensitivity to obstruction; simultaneously, fixedly deployed reconfigurable smart surfaces struggle to continuously adapt to dynamic environments. This method integrates the high bandwidth transmission capability of THz with the enhanced capabilities of RIS-UAV for occlusion-sensitive and coverage-limited links within the CPN architecture to achieve task offloading and coordinated allocation of computing and communication resources. Furthermore, it employs a resource-allocation multi-agent soft-actor-critic (RA-MASAC) strategy through multi-agent reinforcement learning to jointly optimize task offloading decisions, bandwidth allocation, edge computing resource allocation, reconfigurable intelligent surface (RIS) phase control, and RIS-UAV motion control, thereby reducing overall system latency and energy consumption.
[0010] To address this, this invention introduces a drone equipped with a RIS (Resource Identifier) as a mobile communication enhancement unit. Through dynamic position control and phase reconstruction, it improves the wireless transmission capability from the device to the edge layer. With minimizing overall system consumption as the joint optimization objective, it performs unified modeling and collaborative optimization of task offloading decisions, wireless bandwidth allocation, edge computing resource allocation, RIS phase control, and RIS-UAV trajectory control. Furthermore, this method deeply considers the task offloading needs of massive IoT devices, the synergistic gain of THz direct links and RIS cascaded links, the joint scheduling of heterogeneous computing resources across cloud, edge, and device, and the mobility overhead of RIS-UAV maneuvers. By introducing a multi-agent soft actor critic algorithm for resource allocation, it performs deep iterative learning and optimization within a multi-agent Markov decision process framework, thereby obtaining a joint resource allocation strategy usable for online inference. This process not only achieves global collaborative configuration of communication and computing resources but also continuously reduces the weighted average consumption of task execution latency and energy consumption in dynamic environments, significantly improving the system's adaptability to large-scale access and complex load changes. In summary, this invention effectively solves the joint optimization problem under conditions of limited resources, easily blocked links, and strong coupling of decision variables by adopting an integrated technical solution of "RIS-UAV-assisted THz communication enhancement + cloud-edge-device collaborative computing + multi-agent reinforcement learning joint control". It provides engineering-practical technical support and solutions for the deployment and intelligent operation of low-latency, low-energy computing networks for IoT scenarios.
[0011] The three-layer architecture of IoT cloud-edge-device based on computing power network and the RIS-UAV assisted communication scenario model adapted to this invention are shown in [link to invention]. Figure 1 .
[0012] The network structure diagram of the technical solution of this invention is shown below. Figure 2 .
[0013] The flowchart of the cloud-edge-device computing power network joint resource allocation method based on RIS-UAV-assisted terahertz communication described in this invention is shown below. Figure 3 .
[0014] A comparison chart of the total system consumption and the number of IoT devices in this invention is shown below. Figure 4 .
[0015] A comparison chart of the total system power consumption and the computing power of edge computing nodes in this invention is shown below. Figure 5 .
[0016] A comparison chart of the total system consumption and task computational complexity of this invention is shown below. Figure 6 .
[0017] like Figure 1As shown, this invention constructs an IoT system based on a computing network as a three-layer collaborative computing architecture of cloud, edge, and device. This architecture, from bottom to top, consists of a device layer, an edge layer, and a cloud server layer. At the device layer, a large number of IoT devices are deployed, widely distributed across different business scenarios and continuously generating computing tasks in various time slots. In this invention, the set of device nodes is defined as... in This indicates the total number of devices. Each device... In the time slot The generated task is denoted as The controller makes decisions on how to execute the system based on the system status.
[0018] At the edge layer, there are deployed Each edge computing node and its corresponding terahertz base station is used to handle offloaded tasks from the device layer. The set of edge computing nodes is... And introduce sets Using dynamic characterization of time slots The current remaining available computing resources of each edge computing node provide a constraint basis for subsequent computing resource allocation and offloading decisions.
[0019] Furthermore, to improve the transmission performance of terahertz communication in scenarios with obstruction and limited coverage, this invention introduces airborne... A drone equipped with a reconfigurable smart surface serves as a communication enhancement unit. The cloud server layer provides centralized computing power support, receiving and executing offloaded tasks when edge computing power is insufficient or task demands are high, thereby achieving cloud-edge-device collaborative computing and resource scheduling together with the edge layer.
[0020] When the system receives an IoT data processing task, the control agent coordinates the configuration of communication and computing resources based on the real-time environmental conditions to adapt to task characteristics such as data volume, computational complexity, and latency constraints. Specifically, the agent dynamically determines the task execution method and resource allocation strategy by combining information such as the task arrival status of each terminal device in the current time slot, the remaining local execution time, the available computing power of each edge computing node, the wireless link status, and the spatial location of the RIS-UAV. This includes: decisions on offloading tasks locally, at the edge, or in the cloud; bandwidth allocation from the terminal to the base station; computing resource allocation for edge nodes; RIS reflection phase parameter configuration; and motion control of the RIS-UAV, thereby achieving joint scheduling of THz link transmission capabilities and cloud-edge-device computing capabilities.
[0021] Subsequently, the agent constructs a multi-agent reinforcement learning framework. This framework includes: designing a state space to represent the system's operating state (such as task queues, link and resource reserves, RIS-UAV position information, etc.), defining an action space to cover the aforementioned joint control variables (unloading decisions, bandwidth, computing power, RIS phase, UAV trajectory), and establishing a reward function to quantify policy effectiveness. Typically, the core metric is the weighted combined consumption of task execution latency and energy consumption, and a UAV flight energy consumption term is introduced to constrain movement overhead, thereby driving the agent to find the optimal solution while satisfying bandwidth, computing power, and motion constraints.
[0022] Furthermore, based on the constructed system model and multi-agent Markov decision process, the agents are trained using a multi-agent soft actor critic algorithm for resource allocation: First, the policy network and value network parameters of each agent are initialized, and an experience replay mechanism is established; then, an iterative learning process is initiated, where each agent generates joint actions through interaction with the environment and receives environmental feedback, continuously updating network parameters in time slots, so that the policy gradually converges to a resource optimization policy that minimizes the overall system consumption. Finally, in the inference phase, the trained policy network is used to quickly map the real-time state and output joint decisions, achieving low-latency, low-energy cloud-edge-device collaborative resource allocation for dynamic scenarios. Specifically, this is implemented in the following steps:
[0023] Step (1), the system included in this invention includes IoT devices, a drone equipped with a reconfigurable smart surface (RIS), edge computing nodes, and cloud computing nodes. The system we model includes... The set of ... Each cell is configured with one base station and one edge computing node. Each base station communicates with devices within the cell using the terahertz frequency band. Each base station is equipped with a single antenna, where the antenna height of the base station in the nth cell is... Therefore, the location of the base station antenna can be represented as: Base stations in different cells establish communication through wired connections.
[0024] IoT device aggregation It indicates that, regarding the equipment In the time slot The position is represented as For each device, there may be one task to execute in each time slot. The task of device m in time slot t is... The task information is calculated using triples. It means that, among them This indicates the amount of data required to perform the task. Indicates the computing resources required for the task. This indicates the maximum acceptable execution time for the task.
[0025] Step (2): In this invention, a new generation of wireless communication technology is used to assist in information transmission. Terahertz technology, as one of the new generation of wireless communication technologies, has received great attention in recent years. In this invention, the RIS-UAV technology effectively solves the problem of THz's susceptibility to obstacles. Due to the presence of RIS, the communication link can be divided into two types: direct terahertz communication link and RIS-cascaded terahertz communication link.
[0026] Step (2.1) The direct terahertz communication link is the direct communication channel between the base station and the device. With base station The distance between antennas is denoted as Its expression is:
[0027]
[0028] in For the first The location of each base station antenna. For equipment exist The position at any given moment.
[0029] No. Base station antennas and equipment At any moment The terahertz direct channel gain is denoted as Its expression is:
[0030]
[0031] in, The absorption factor represents the medium absorption factor, used to quantify the attenuation effect per unit distance caused by molecular resonance absorption of terahertz electromagnetic waves in the propagation medium. Its value is related to environmental conditions such as the terahertz operating frequency, humidity, and temperature of the propagation medium. For example, under standard sea-level conditions, the absorption factor is approximately 70-140. .
[0032] Step (2.2) cascaded terahertz communication link is a channel formed after information is reflected by RIS-UAV, at time At that time, the coordinates of the drone equipped with RIS were defined as follows: .equipment The distance between the RIS-UAV and the RIS-UAV is denoted as:
[0033]
[0034] The distance between the RIS-UAV and the base station antenna is expressed as:
[0035]
[0036] The cascaded channel gain in terahertz transmission scenarios can be obtained by the following formula:
[0037]
[0038]
[0039] in, Indicates from device To RIS The RIS receiver array response matrix, Indicates from RIS to base station The RIS emission array response matrix.
[0040] Step (2.3) If the equipment The decision to upload data to a base station for execution of tasks on an Edge Computing Network (ECN) node is determined by the base station's received signal, which is influenced by direct channel conditions, cascaded channel conditions, transmit power, and noise. The uplink transmission rate is denoted as... This can be calculated using Shannon's formula. Therefore, the transmission rate can be expressed as:
[0041]
[0042] in, For channel bandwidth, For the receiver antenna gain, For the transmitting antenna gain, Let be the power spectral density of the noise.
[0043] Step (3) In order to unify the measurement of the system’s operating overhead under different execution paths (local execution, edge offloading execution, cloud offloading execution) and to serve as the basis for optimization objectives and reinforcement learning rewards, this invention models the task execution process and quantifies the latency and energy consumption generated by the task in the transmission and computation stages.
[0044] Step (3.1) Local calculation of consumption
[0045] Equipment Execute tasks locally The time taken is The energy consumption for performing this task locally is recorded as .in, Calculate the time consumption and time slots for the task. Startup equipment The sum of the remaining execution times of the tasks on the table is expressed as:
[0046]
[0047] In the formula, Indicates equipment Computational power; Indicates time slot Startup equipment The remaining execution time of the task.
[0048] Device m performs tasks locally The energy consumption of device m is related to its hardware characteristics, and its expression is:
[0049]
[0050] In the formula, This represents the energy consumption coefficient of device m, and this parameter is related to the equipment. This is related to the hardware characteristics. Therefore, the execution of the task... The total consumption is:
[0051]
[0052] In the formula, This is the time delay preference coefficient; The larger the value of , the higher the system's preference for minimizing latency.
[0053] Step (3.2) Edge computing consumption
[0054] In this invention, the device can offload its tasks to any ECN via the aforementioned communication network. Assume device m will offload its tasks... The unloading process, from the target edge computing node ECN k to the unloading point, can be summarized as follows:
[0055] 1. Equipment The task The data is transmitted to the base station associated with the device. ;
[0056] 2. If Edge computing nodes The above data is forwarded to the target edge computing node. ;
[0057] 3. Target edge computing node ( ) Perform tasks ;
[0058] 4. The edge computing node executing the task sends the execution result back to the device. .
[0059] If the device Upload execution task The required data, and the latency of the upload process. Energy consumption for:
[0060]
[0061]
[0062] In the formula, For equipment To base station Uplink transmission rate; for Time slot device m performs tasks The amount of data to be uploaded; For time slots equipment To edge computing nodes Time taken to upload task data For equipment Transmission power consumption, For base stations Receive power consumption.
[0063] If forwarding needs to be done to an edge node that is not part of the base station, the latency and energy consumption are as follows:
[0064]
[0065]
[0066] in, For base stations and The transmission rate of the wired link between them; For base stations Forwarding power consumption; For base stations Received power consumption
[0067] After being forwarded to the target edge node, the latency and energy consumption for the target edge node to perform computation on the task are as follows:
[0068]
[0069]
[0070] In the formula, For time slots Task Number of CPU cycles required; For time slots Edge computing nodes Assigned to device Computing resources; For edge nodes Energy consumption coefficient; For equipment Idle power while waiting for calculation results.
[0071] After completing the computation, the edge nodes send the results back to the device. In most scenarios, the downlink rate is high and the computation result is small; this part can be modeled as needed or ignored under simplifying assumptions. Therefore, the overall cost of the edge computing task is:
[0072]
[0073]
[0074]
[0075] In the formula, For equipment In the time slot Total latency for offloading to edge execution; This corresponds to the total energy consumption; The overall overhead of edge execution.
[0076] Step (3.3) Cloud computing consumption
[0077] When the device In the time slot When a task is offloaded to the cloud for execution, the task completion process includes: uplink transmission from the device to the base station, backlink transmission from the base station to the cloud server, and cloud computing execution.
[0078] During uplink transmission from the device to the base station, latency and energy consumption are denoted as follows: and :
[0079]
[0080]
[0081] In the formula, For equipment The transmission power; For base stations The received power.
[0082] After receiving the task data from the device, the base station forwards the data to the cloud server. Its latency and energy consumption are as follows:
[0083]
[0084]
[0085] In the formula, This refers to the transmission rate from the base station to the cloud server. This refers to the power consumption during the backhaul transmission process.
[0086] The cloud then performs calculations on the task and returns the results. The latency and energy consumption calculations for this stage are as follows:
[0087]
[0088]
[0089] In the formula, For time slots Cloud server assigns tasks Computing resources; This refers to the energy consumption coefficient of the cloud server.
[0090] In summary, the overall consumption of cloud computing tasks can be expressed as the weighted sum of total latency and total energy consumption:
[0091]
[0092]
[0093]
[0094] In the formula, For equipment In the time slot Total latency for unloading and executing to the cloud; This corresponds to the total energy consumption; This refers to the overall cost of execution in the cloud.
[0095] Step (3.4) Based on the aforementioned communication resource model and computing task consumption model, this invention formalizes the joint optimization problem of "task offloading decision, bandwidth allocation, edge computing resource allocation, RIS phase control, and RIS-UAV trajectory control" in the computing power network into the following constrained optimization problem. By minimizing the sum of the overall system consumption, low latency and low energy consumption collaborative optimization in the long term is achieved:
[0096]
[0097]
[0098]
[0099]
[0100]
[0101] In the formula, O(t) represents the time slot. The overall system consumption is used to uniformly measure task execution latency and energy consumption; For the unloading decision vector, where This indicates that the task will be executed locally. Indicates unloading to edge computing nodes , This indicates that the software will be uninstalled to the cloud. For the bandwidth allocation matrix, its elements Indicates time slot base station Assigned to device Uplink bandwidth, constraints This means that the total bandwidth allocated to each base station's serving equipment must not exceed the total available bandwidth. ; To compute the resource allocation matrix, its elements Indicates time slot Edge computing nodes Assigned to device The computational power of the task, constraints This means that the total computing power allocated to all tasks by an edge node must not exceed its remaining available computing power. ; For the phase control set of all RIS-UAVs, where Indicates the first The phase shift matrix of the RIS-UAV; Let be the set of motion vectors of all RIS-UAVs, where Indicates the first RIS-UAV in time slot Motion control variables, constraints This indicates that its range of motion must not exceed the maximum motion constraint. , For the number of devices, This represents the number of edge nodes (i.e., cells). This refers to the number of RIS-UAVs.
[0102] By solving the above optimization problem, this invention achieves joint decision-making on task offloading, bandwidth allocation, computing power allocation, RIS phase, and UAV trajectory under the constraints of communication bandwidth, edge computing power, and UAV motion, thereby minimizing the overall system consumption in the long-term time domain.
[0103] Step (4) Based on steps (1)-(3), and in conjunction with the environment and optimization objectives, set the state space, action space and reward function of the algorithm. Due to the Markov characteristics of the task unloading computation process, we model this process as a Markov process.
[0104] The state space of the Markov process in step (4.1) is:
[0105]
[0106] in, This represents the set of task information for all IoT devices in time slot t; where It is the first Each device in the time slot Task description; This indicates that all RIS-UAVs are in the time slot. The initial set of positions is used to provide a baseline for trajectory adjustment in this time slot; Indicates the time slot of each device The set of remaining execution times for local tasks that were not completed at the start. This amount affects local computation latency, and thus affects the unloading decision. This represents the set of remaining available computing power for each edge computing node after the previous round of resource allocation, used to ensure the feasibility and constraint satisfaction of computing power allocation in this time slot.
[0107] Step (4.2) This algorithm sets up five cooperative agents, and the division of labor among the agents is as follows:
[0108] Trajectory Agent: Adjusting the motion vector matrix of the UAV based on the observed system state This enables the optimization of drone positioning;
[0109] Unloading Agent: Generating an unloading decision matrix based on observed system states Determine the task processing mode and select the optimal solution from local computing, edge computing, and cloud computing based on task requirements;
[0110] Bandwidth agent: Allocating bandwidth resource matrix based on observed system state Under the premise of not exceeding the bandwidth limit of the base station, bandwidth resources should be allocated reasonably;
[0111] Phase agent: Obtaining the reflection phase matrix of a tuned smart metasurface based on the observed system state. Increase the gain of the cascaded channel;
[0112] Computing power agent: Allocates computing resources to each task offloaded to the edge node based on the observed system state. To maximize efficiency and achieve the lowest possible task timeout rate when there are many tasks.
[0113] Step (4.3) To obtain better long-term returns, it is necessary to flexibly adjust the action strategy according to changes in the environment. Therefore, the action space of each agent is set as follows:
[0114]
[0115]
[0116]
[0117]
[0118]
[0119] Indicates the trajectory agent in the time slot The set of action spaces; where Indicates the first RIS-UAV in time slot The motion vector is used to adjust the position of the RIS-UAV to improve coverage and link quality; R represents the number of RIS-UAVs.
[0120] This indicates that the agent is unloaded in the time slot. The set of action spaces; where Indicates the first Each device task in a time slot The unloading decision (generally, the value means: Indicates local execution. This indicates that the data will be unloaded to the corresponding edge computing node. (Indicates uninstallation to the cloud). Indicates the number of devices. This indicates the number of edge nodes.
[0121] Indicates the bandwidth of the agent in the time slot The set of action spaces; where Indicates in time slot For the first The uplink communication bandwidth allocated to each device is used to determine the upload rate and is constrained by the total bandwidth of the base station.
[0122] This indicates that the phase agent is in the time slot. The set of action spaces; where Indicates the first Phase control variables for RIS-UAV.
[0123] Indicates the computing power of the intelligent agent in the time slot The set of action spaces; where Indicates in time slot For the first Each task is allocated edge computing resources and is constrained by the total computing power of the edge nodes.
[0124] Step (4.4) In order to optimize the system's energy consumption and task processing latency, the reward function is set to the negative of the total system consumption:
[0125]
[0126] Step (5): Based on the state space, action space, and reward function constructed in step (4), this invention uses the Multi-Agent Soft Actor-Critic (MASAC) algorithm to train the joint resource allocation strategy and completes the structure and training parameter settings of the policy network (actor) and value network (critic). MASAC belongs to the Actor-Critic method based on maximum entropy reinforcement learning. Its core idea is to explicitly maximize policy entropy while maximizing expected reward, so as to improve exploration ability and enhance training stability.
[0127] Step (5.1) Multi-agent policy and action sampling
[0128] In this invention, there are a total of five intelligent agents, the first being... The policy network of an agent is denoted as . ,in The sub-actions output by the agent ( In step (4) ), This represents the global state. During inference and interaction sampling, each agent samples actions from its own policy in parallel:
[0129]
[0130] This results in a combined action:
[0131]
[0132] To handle continuous actions and achieve differentiable sampling, this invention employs a reparameterization technique:
[0133]
[0134] in and These are the mean and standard deviation of the policy network output, respectively.
[0135] Step (5.2) Maximum Entropy Optimization
[0136] The optimization objective of MASAC is to maximize the expected return including the entropy regularization term. Its policy objective is:
[0137]
[0138] in As a discount factor, For temperature coefficient, The entropy regularization term is defined as follows:
[0139]
[0140] For state In this regard, its state value function and Q function are as follows:
[0141]
[0142]
[0143] in, In the state Execute joint operations The instant reward obtained afterward For the next state, This is the discount factor.
[0144] Step (5.3) Network parameter update
[0145] This invention introduces an experience replay pool during the training process. Store the interaction sample quadruple:
[0146]
[0147] from Small batch sampling Construct a Q-network time-difference objective:
[0148]
[0149] in, This represents the training mini-batch set sampled from the replay pool; This represents the target value used to update the network of the j-th critic.
[0150] The algorithm uses two networks: a critic network and an actor network. The critic network consists of two Q-networks, employs minimum mean squared error loss and gradient descent, and is updated at the end of each training step. Its loss function and update formula are as follows:
[0151]
[0152]
[0153] in, Indicates the first The loss function of an agent-actor network. For actors' online learning rate; Indicates the parameter The gradient operator. This update makes the policy tend to choose actions that yield larger Q values while maintaining a certain degree of randomness, thereby reducing overall cost.
[0154] The actor updates by maximizing the Q-value. For the ... For each agent, the policy loss and parameter updates are as follows:
[0155]
[0156]
[0157] in, Indicates the first The loss function of the agent network; For actors' online learning rate; Indicates the parameter The gradient operator is used. This update makes the policy tend to choose actions that yield larger soft Q values while maintaining a certain degree of randomness, thereby reducing overall cost.
[0158] Step (6) Based on the multi-agent policy network trained in step (5), in each time slot The system state S(t) is obtained and input into the actor network of each agent, generating sub-actions such as trajectory vector, unloading decision, bandwidth allocation, RIS phase, and edge computing power allocation, which in turn form a joint action. Treat this joint action as a time slot. The optimal executable action is then assigned and sent to the cloud-edge-device computing network controller and computing entities for execution. After execution, the system updates to the next state based on environmental feedback, repeating the closed-loop process of "state acquisition - strategy reasoning - action execution" to continuously execute the optimal action in each state until the task is completed or the preset time domain termination condition is met.
[0159] The advantages of this invention lie in its focus on large-scale, highly dynamic IoT computing network scenarios with high real-time requirements. It constructs a cloud-edge-device collaborative computing network architecture, achieving flexible collaboration and task offloading of heterogeneous computing resources through networking. Simultaneously, it introduces terahertz communication to support high-capacity data transmission and employs RIS-UAV as a mobile intelligent reflection enhancement unit to alleviate the problems of limited THz coverage and susceptibility to obstruction, thereby simultaneously improving system availability and performance limits on both the communication and computing sides. Furthermore, to address the joint optimization challenges in a continuous / hybrid action space with strong coupling of multiple variables such as RIS-UAV trajectory / phase, task offloading, and bandwidth and computing power allocation, this invention adopts the RA-MASAC multi-agent reinforcement learning method. It sets up a trajectory agent, an offloading agent, a bandwidth agent, a phase agent, and a computing power agent to make parallel collaborative decisions. Under constraints such as bandwidth and computing power, the optimization objective is to minimize the weighted comprehensive consumption of system latency and energy consumption, thereby achieving the control capability to quickly infer and output joint resource allocation strategies in dynamic environments. Simulation results show that the proposed method exhibits lower total system consumption under different conditions such as the number of devices, different edge computing power, and different task complexity, and is significantly better than the baseline method. Attached Figure Description
[0160] Figure 1 A schematic diagram of a three-layer architecture for the Internet of Things (IoT) cloud-edge-device based on computing power networks, including the IoT device layer, the edge layer (terahertz base stations and edge computing nodes), the cloud server layer, and the RIS-UAV auxiliary communication unit and controller.
[0161] Figure 2 The network structure diagram of the present invention is shown below.
[0162] Figure 3 The flowchart of the cloud-edge-device computing power network joint resource allocation method based on RIS-UAV assisted terahertz communication described in this invention.
[0163] Figure 4 The graph shows the relationship between total system consumption and the number of IoT devices. It is used to compare the total system consumption of the method described in this invention and other comparative methods under different device scales.
[0164] Figure 5 The graph shows the relationship between total system consumption and task computational complexity, used to compare the total system consumption performance of the method described in this invention and other comparative methods under different task complexity conditions.
[0165] Figure 6 A comparison chart of the relationship between total system consumption and the computing power of edge computing nodes is shown in the figure. The chart is used to compare the total system consumption performance of the method described in this invention and various comparative methods under different edge computing power configurations. Detailed Implementation
[0166] The technical solution of the RIS-UAV-assisted IoT computing resource optimization and scheduling method will be further explained below with reference to the accompanying drawings and examples.
[0167] The flowchart of the method described in this invention is as follows: Figure 3 As shown, it includes the following steps:
[0168] Step 1: Set system information such as the number of IoT device nodes, the number of edge server nodes, cloud server configuration, and the number of RIS-UAVs, and model infrastructure parameters such as device computing power, edge computing power, cloud computing power, and base station available bandwidth;
[0169] Step 2: Terahertz communication and RIS-UAV reconfigurable intelligent reflection enhancement technology are introduced to model the wireless communication process from IoT devices to edge servers, characterize the equivalent transmission capabilities of direct links and RIS-assisted cascaded links, and provide communication performance characterization for bandwidth scheduling and phase configuration.
[0170] Step 3: Model the system task execution process, quantify the task latency and energy consumption under three modes: local execution, edge offloading execution, and cloud offloading execution, and incorporate the mobility overhead introduced by RIS-UAV in communication enhancement into the overall system consumption to form a unified task consumption evaluation index and establish a joint optimization problem.
[0171] Step 4: Model a multi-agent Markov decision process, set up a state space to reflect system states such as task arrival and resource availability, define an action space to cover joint decision variables such as unloading, bandwidth, computing power, phase and trajectory, and construct a reward function with the overall system consumption as the core.
[0172] Step 5: Use a multi-agent soft actor critic algorithm to complete the execution strategy and deep neural network training parameter settings;
[0173] Step 6: Based on the policy network trained in Step 5, generate and select the optimal joint action in each state, and issue the action as the optimal execution instruction in the current state to the communication and computing entities for continuous execution. The action is then continuously updated to the next state based on environmental feedback until the preset termination condition is met or the task execution ends.
[0174] Figure 4This is a comparison chart showing the relationship between total system consumption and the number of IoT devices. It can be noted that as the number of devices increases, the number of tasks the system needs to process also increases, resulting in an overall upward trend in the total system consumption for all solutions. Meanwhile, comparing different solutions longitudinally, the solution adopted in this invention maintains the lowest total consumption under different device counts, demonstrating its advantages in resource coordination and joint optimization in large-scale device access scenarios.
[0175] Figure 5 This is a comparison chart showing the relationship between total system consumption and device task computational complexity. It can be observed that increased task computational complexity generally leads to a rise in total system consumption for all methods; however, the solution adopted in this invention consistently maintains the lowest total consumption across different complexity ranges, indicating that it has a stronger adaptive capability to task load fluctuations and complexity changes, and can continuously output better offloading and resource allocation strategies under dynamic task requirements.
[0176] Figure 6 This is a comparison chart showing the relationship between total system consumption and the computing power of edge computing nodes. As can be seen from the chart, with the improvement of edge computing node computing power, tasks offloaded to the edge can be improved in terms of computational latency and energy consumption. Therefore, except for schemes such as "all local computing" that do not rely on edge computing power, the total consumption of other schemes usually decreases as the computing power of edge computing nodes increases. Based on this, the scheme adopted in this invention outperforms other baseline methods under various edge computing node computing power settings, indicating that it can more effectively utilize edge computing resources to achieve joint decision-making.
Claims
1. A method for optimizing and scheduling IoT computing resources based on RIS-UAV, the method comprising the following steps: Step 1: Set the system information for the number of IoT device nodes, edge server nodes, cloud server configuration, and RIS-UAV quantity, and model the device computing power, edge computing power, cloud computing power, and base station available bandwidth infrastructure parameters. Step 2: Terahertz communication and RIS-UAV reconfigurable intelligent reflection enhancement technology are introduced to model the wireless communication process from IoT devices to edge servers, characterize the equivalent transmission capabilities of direct links and RIS-assisted cascaded links, and provide communication performance characterization for bandwidth scheduling and phase configuration. Step 3: Model the system task execution process, quantify the task latency and energy consumption under three modes: local execution, edge offloading execution, and cloud offloading execution, and incorporate the mobility overhead introduced by RIS-UAV in communication enhancement into the overall system consumption to form a unified task consumption evaluation index and establish a joint optimization problem. Step 4: Model a multi-agent Markov decision process, set up a state space to reflect the system state of task arrival and resource reserve, define an action space to cover joint decision variables such as unloading, bandwidth, computing power, phase and trajectory, and construct a reward function with the overall system consumption as the core. Step 5: Use the multi-agent soft actor critic algorithm to complete the execution strategy and deep neural network training parameter settings; Step 6: Based on the policy network trained in Step 5, generate and select the optimal joint action in each state, and issue the action as the optimal execution instruction in the current state to the communication and computing entities for continuous execution. The action is then continuously updated to the next state based on environmental feedback until the preset termination condition is met or the task execution ends.
2. The method for optimizing and scheduling computing network resources for the Internet of Things according to claim 1, characterized in that: The system information, including the number of IoT device nodes, edge server nodes, cloud server configuration, and the number of RIS-UAVs, is set, and the computing power of the devices, edge computing power, cloud computing power, and base station bandwidth infrastructure parameters are modeled. The specific steps are as follows: The system includes IoT devices, drones equipped with reconfigurable smart surface RIS, edge computing nodes, and cloud computing nodes; the modeling system includes... The set of _ ... Each cell is equipped with one base station and one edge computing node; each base station uses the terahertz frequency band to communicate with devices within the cell; each base station is equipped with a single antenna, where the antenna height of the base station in the nth cell is... The location of the base station antenna is indicated as follows: Base stations in different cells establish communication through wired connections; IoT device aggregation It indicates that, regarding the equipment In the time slot The position is represented as For each device, there may be one task to execute in each time slot. The task of device m in time slot t is... The task information is calculated using triples. It means that, among them This indicates the amount of data required to perform the task. Indicates the computing resources required for the task. This indicates the maximum acceptable execution time for the task; The system has a total of The RIS-UAV is mounted, and the position of the UAV and the phase of the reflective unit mounted on the RIS are dynamically adjusted.
3. The method for optimizing and scheduling computing network resources for the Internet of Things according to claim 2, characterized in that: In step two, terahertz communication and RIS-UAV reconfigurable intelligent reflection enhancement technology are introduced to model the wireless communication process from IoT devices to edge servers, characterize the equivalent transmission capabilities of direct links and RIS-assisted cascaded links, and provide communication performance representations for bandwidth scheduling and phase configuration; specifically as follows: The direct terahertz communication link is a direct communication channel between the base station and the equipment. With base station The distance between antennas is denoted as ; No. Base station antennas and equipment At any moment The terahertz direct channel gain is denoted as ; Cascaded terahertz communication links are channels formed after information is reflected by RIS-UAV, at specific times. At that time, the coordinates of the drone equipped with RIS were defined as follows: ; If the device The decision to upload its data to the base station for execution of its tasks on the edge computing network node (ECN) is determined by the direct channel conditions, concatenated channel conditions, transmit power, and noise; the uplink transmission rate is denoted as... The calculation is performed using Shannon's formula.
4. The method for optimizing and scheduling computing network resources for the Internet of Things according to claim 3, characterized in that: In step three, the system task execution process is modeled, and the task latency and energy consumption are quantified under three modes: local execution, edge offloading execution, and cloud offloading execution. The mobility overhead introduced by RIS-UAV in communication enhancement is included in the overall system consumption to form a unified task consumption evaluation index and establish a joint optimization problem. In order to uniformly measure the system's operating overhead under different execution paths and use it as the basis for optimization objectives and reinforcement learning rewards, the task execution process is modeled, and the latency and energy consumption generated by the task in the transmission and computing stages are quantified. Equipment Execute tasks locally The time taken is The energy consumption for performing this task locally is recorded as ;in, Calculate the time consumption and time slots for the task. Startup equipment The sum of the remaining execution times of the tasks on the table is expressed as: ; In the formula, Indicates device Computational power; Indicates time slot Startup equipment Remaining execution time for the task; Device m performs tasks locally The energy consumption of device m is related to its hardware characteristics, and its expression is: ; In the formula, This represents the energy consumption coefficient of device m, and this parameter is related to the equipment. Related to hardware characteristics; executing tasks The total consumption is: ; In the formula, This is the time delay preference coefficient; The device offloads its tasks to any ECN via the aforementioned communication network; device m will then... Unload to the target edge computing node ECNk; If forwarding needs to be done to an edge node that is not part of the base station, the latency and energy consumption are as follows: ; ; in, For base stations and The transmission rate of the wired link between them; For base stations Forwarding power consumption; For base stations Received power consumption After being forwarded to the target edge node, the latency and energy consumption for the target edge node to perform computation on the task are as follows: ; In the formula, For time slots Task Required number of CPU cycles; For time slots Edge computing nodes Assigned to device Computing resources; For edge nodes Energy consumption coefficient; For equipment Idle power while waiting for calculation results; The overall cost of edge computing tasks is: In the formula, For equipment In the time slot Total latency for offloading to the edge for execution; This corresponds to the total energy consumption; The overall overhead of edge execution; When the device In the time slot When a task is offloaded to the cloud for execution, the task completion process includes: uplink transmission from the device to the base station, backlink transmission from the base station to the cloud server, and cloud computing execution. During uplink transmission from the device to the base station, latency and energy consumption are denoted as follows: and : After receiving the task data from the device, the base station forwards the data to the cloud server to calculate latency and energy consumption. The cloud then performs calculations on the task and returns the results, including calculations of latency and energy consumption; The total consumption of cloud computing tasks is expressed as the weighted sum of total latency and total energy consumption.
5. The method for optimizing and scheduling computing network resources for the Internet of Things according to claim 4, characterized in that: In step four, a multi-agent Markov decision process is modeled. A state space is set up to reflect the system state of task arrival and resource reserves. An action space is defined to cover joint decision variables of unloading, bandwidth, computing power, phase, and trajectory. A reward function with the overall system consumption as its core is constructed. Specifically, as follows: Based on the environment and optimization objectives, the state space, action space and reward function of the algorithm are set. Due to the Markov property of the task unloading computation process, this process is modeled as a Markov process. Five collaborative agents were set up, and the roles of each agent were as follows: Trajectory Agent: Adjusting the motion vector matrix of the UAV based on the observed system state This enables the optimization of drone positioning; Unloading Agent: Generating an unloading decision matrix based on observed system states Determine the task processing mode and select the optimal solution from local computing, edge computing, and cloud computing based on task requirements; Bandwidth agent: Allocating bandwidth resource matrix based on observed system state Under the premise of not exceeding the bandwidth limit of the base station, bandwidth resources should be allocated reasonably; Phase agent: Obtaining the reflection phase matrix of a tuned smart metasurface based on the observed system state. Increase the gain of the cascaded channel; Computing power agent: Allocates computing resources to each task offloaded to the edge node based on the observed system state. To maximize efficiency and achieve the lowest possible task timeout rate when there are many tasks.
6. The method for optimizing and scheduling computing network resources for the Internet of Things according to claim 5, characterized in that: In step five, the multi-agent soft actor critic algorithm is used to complete the execution strategy and deep neural network training parameter settings, as follows: Based on the state space, action space, and reward function constructed in step four, the joint resource allocation strategy is trained using the multi-agent soft actor critic MASAC algorithm, and the structure and training parameter settings of the policy network actor and the value network critic are completed. There are two types of networks: critic networks and actor networks; the critic network consists of two Q networks, uses minimum mean squared error loss and gradient descent, and is updated at the end of each training step; The actor network is updated by maximizing the Q-value.
7. The method for optimizing and scheduling computing network resources for the Internet of Things according to claim 6, characterized in that: In step six, based on the policy network trained in step five, the optimal joint action is generated and selected in each state. This action is then sent to the communication and computing entities as the optimal execution instruction in the current state for continuous execution. The action is then updated to the next state based on environmental feedback until the preset termination condition is met or the task execution ends.