A complex dependent task offloading method based on multi-agent reinforcement learning

By constructing an edge-cloud MEC system model and adopting a multi-agent reinforcement learning approach, the timing constraints and resource conflicts of task offloading strategies in a multi-user environment are solved. The optimal task offloading decision is achieved under multi-agent collaborative learning, reducing system latency and energy consumption.

CN122132129APending Publication Date: 2026-06-02GUILIN UNIVERSITY OF TECHNOLOGY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUILIN UNIVERSITY OF TECHNOLOGY
Filing Date
2026-01-23
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In dynamic mobile edge computing environments with multiple users and multiple servers, existing technologies struggle to effectively handle the mobility of terminal devices, the time-varying nature of wireless channels, and the complex dependencies between tasks. This results in task offloading strategies failing to effectively address the timing constraints and resource conflicts caused by dependent tasks while ensuring overall system performance.

Method used

A multi-agent reinforcement learning-based approach is adopted to construct an edge-cloud MEC system model, establishing a mobile model, a communication model, and a task computing model for intelligent terminals. The task priority scheduling mechanism is integrated through the MAPPO algorithm to optimize the task offloading strategy and minimize the total system cost.

Benefits of technology

In a multi-user environment, the system can dynamically select the optimal edge server based on real-time status, efficiently utilize heterogeneous computing resources, meet task dependencies while reducing latency and energy consumption, and optimize the long-term average computing cost of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132129A_ABST
    Figure CN122132129A_ABST
Patent Text Reader

Abstract

This invention provides a method for offloading complex dependency tasks based on multi-agent reinforcement learning, belonging to the field of mobile edge computing technology. It designs a three-tiered mobile edge computing system model with an edge-cloud architecture and proposes a multi-agent near-end policy optimization algorithm that integrates task priority scheduling mechanisms to optimize the task offloading strategy for multi-users with dependency relationships in an MEC environment. This method supports collaborative learning among multiple agents within a framework of centralized training and distributed execution, thereby improving policy coordination and convergence efficiency, and effectively reducing the overall average user cost of the system. The system combines local devices, edge servers, and cloud servers to provide task execution services, enhancing the flexibility and efficiency of task processing. In an MEC system containing multiple edge servers, this method allows users to dynamically select the optimal edge server based on real-time status, achieving efficient utilization of heterogeneous computing resources while satisfying task dependencies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mobile edge computing technology, and more particularly to a method for offloading complex dependency tasks based on multi-agent reinforcement learning. It is particularly suitable for optimizing the offloading of tasks with dependencies in systems comprising cloud servers, multiple edge servers, and multiple agents. Background Technology

[0002] In recent years, with the rapid development of 5G communication and mobile internet technologies, a large number of computationally intensive and latency-sensitive applications have emerged on various smart terminals, such as mobile AR / VR and real-time video analytics. These applications typically consist of multiple dependent tasks with a strict logical order, which can be modeled as a directed acyclic graph. Due to the limited computing resources of mobile terminal devices, it is difficult to independently and efficiently process all tasks, so tasks are often offloaded to edge servers or cloud centers for execution. However, in dynamic mobile edge computing environments with multiple users and multiple servers, the mobility of terminal devices, the time-varying nature of wireless channels, and the complex dependencies between tasks pose significant challenges to designing efficient task offloading strategies. Existing single-agent reinforcement learning methods struggle to achieve collaborative decision-making in multi-user competitive environments, while some multi-agent algorithms suffer from deficiencies in policy training stability and joint action space exploration, failing to effectively address the temporal constraints and resource conflicts caused by dependent tasks while ensuring overall system performance. Summary of the Invention

[0003] The purpose of this invention is to provide a complex dependency task offloading method based on multi-agent reinforcement learning, solving the technical problem that existing offloading methods cannot effectively handle the timing constraints and resource conflicts caused by dependent tasks while ensuring the overall system performance. It aims to find the optimal offloading scheme under the constraints of the mobility of intelligent terminals, the time-varying nature of channels, and the finiteness of complex task dependencies, thereby minimizing the total system cost composed of latency and energy consumption to adapt to the dynamic MEC environment with multiple time slots and multiple applications.

[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0005] A method for unloading complex dependency tasks based on multi-agent reinforcement learning, the method comprising the following steps:

[0006] Step 1: Build an edge-cloud MEC system model, which includes a local computing layer, an MEC layer, and a cloud computing layer;

[0007] Step 2: Establish the intelligent terminal mobile model, communication model, and task calculation model that comprehensively considers latency and energy consumption in the MEC system model;

[0008] Step 3: Calculate the total computational cost of each smart terminal executing tasks under different offloading decisions, and construct an optimization function for the task offloading strategy with the goal of minimizing the average computational cost of all smart terminals in the MEC system;

[0009] Step 4: Model the final optimization objective as a Markov decision process;

[0010] Step 5: Construct several agent-based proximal policy optimization algorithms that integrate task priority scheduling mechanism. The proximal policy optimization algorithm first designs task priority scheduling based on directed acyclic graph, generates task queue, and uses several agent-based proximal policy optimization algorithms for collaborative optimization training to solve the optimal task unloading decision.

[0011] Furthermore, in step 1, the set of intelligent agent terminals for all tasks awaiting execution at the local layer is: Smart terminal devices The generated application in the time slot It can be modeled as a DAG , where the set of nodes This represents all dependent subtasks in the application, and the edges... This means that a task can only begin computation after its predecessor tasks have been completed. In other words, a task can only begin computation after all its predecessor tasks have been completed. Corresponding to a triple ,in This indicates the number of CPU cycles required to process each bit of task. Indicates the data size of the subtask. This indicates the tolerance latency for the subtask.

[0012] Furthermore, in step 1, the MEC layer consists of several edge base stations, each connected to an MEC server via optical fiber, and the set of MEC servers is as follows: The cloud computing layer consists of several servers, which have computing power and storage resources for processing data.

[0013] Furthermore, the specific process of step 2 is as follows:

[0014] Step 2.1: Construct the mobile model and communication model of the smart terminal, and obtain the signal-to-interference-plus-noise ratio (SIR) of the smart terminal within the time slot. ;

[0015] Step 2.2: Construct a computational model, using... Indicates task The uninstallation strategy, when Time indicates smart terminal Execute the task locally when This indicates that the task is executed on the MEC server. The time factor indicates that the task is executed in the cloud. The latency factors involved at each execution end include waiting latency and execution latency. The energy consumption factors consider the user's local energy consumption and the energy consumption during task transmission. The computational model dependent on the task is represented as follows:

[0016] When the task's offloading decision is to execute locally, the task's latency and energy consumption are as follows:

[0017]

[0018] in This indicates the latency of the task execution locally. The waiting time for local execution to begin. Switched capacitors for smart terminal devices. For local execution energy consumption, For local computing power;

[0019] When the task offloading decision is to execute at the edge layer, the task completion latency and energy consumption are as follows:

[0020]

[0021] in This indicates the delay before the edge server begins execution. This indicates the execution latency of the task on the MEC server. Indicates task The transmission latency for uploading to the MEC server is used express arrive The transmission latency is negligible due to the high downlink transmission rate;

[0022]

[0023] in and These represent the energy consumption for uploading tasks to the edge server and, respectively. arrive Transmission energy consumption;

[0024] When the task offloading decision is to be executed in the cloud, the task completion latency and energy consumption are as follows:

[0025]

[0026] in, For cloud execution latency, This indicates the delay before the task begins execution in the cloud. For the transmission latency of tasks from the edge server to the cloud, Total energy consumption when the task is offloaded to the cloud.

[0027] Furthermore, in step 3, the optimization objective of the optimization function is to optimize the computation offloading decision of the requested tasks in all smart terminals, and to ensure that the system cost consisting of task completion latency and energy consumption is minimized while satisfying the task tolerance latency constraint.

[0028] Furthermore, in step 4, the dependency task unloading problem is optimized into a Markov decision process that requires determining its global state space. Local observation space Action space Reward Space .

[0029] Furthermore, in step 5, when executing several agent proximal policy optimization algorithms, the Actor network... Critic network is used to output action policies. The strategies learned by several agents are used to evaluate the current state. To indicate, among which and These represent the parameters of the Actor network and the Critic network, respectively. Each user seeks the optimal task offloading decision by training a proximal optimization algorithm based on several agents according to the current environment state. Considering that different task scheduling queues will lead to different completion delays of DAG tasks, the proximal policy optimization algorithm of several agents incorporates a task priority scheduling mechanism.

[0030] Furthermore, the specific process of several agent proximal policy optimization algorithms is as follows:

[0031] Step 5.1: Randomly initialize each agent Actor network parameters and Critic network parameters Initialize the old strategy, target Critic network, and initialize the experience replay buffer. ;

[0032] Step 5.2: Initialize the total number of iterations (episode);

[0033] Step 5.3: Obtain the initial system state ;

[0034] Step 5.4: Set the number of time steps for each iteration ;

[0035] Step 5.5: For DAG tasks generated by smart terminals, design a task priority scheduling method based on DAG topology. Priority information is based on DAG topology information and task tolerance latency, while offloading decisions are made in ascending order of priority metrics. The specific steps are as follows:

[0036] The average latency of the subtask is calculated as follows:

[0037]

[0038] Based on the task's tolerance delay as well as Calculate the task priority for each task. It is represented as follows:

[0039]

[0040] Step 5.6: Each agent generates an action based on the currently observed state. ;

[0041] Step 5.7: The agent executes a joint action. Each agent receives a reward. and new status ;

[0042] Step 5.8: Collect trajectory data for each agent. And store it in the experience replay pool. ;

[0043] Step 5.9: Repeat steps 5.5 to 5.8 until completion. Up to the next time step;

[0044] Step 5.10: Calculate the advantage value for each agent. The formula is as follows:

[0045]

[0046] in, To balance the discount factor between the importance of current and future rewards, The parameter is used to balance the bias and variance. It's a cumulative discount reward, therefore the Critic network is applied to the state. The value estimation function is expressed as ;

[0047] Step 5.11: Begin Each training cycle shuffles the data in the buffer. Following the order of the data, and randomly selecting small batches, update the network parameters as follows;

[0048] Step 5.12: Update the Critic network for each agent. The loss function is calculated as follows:

[0049]

[0050] Update the Actor network for each agent. The loss function is calculated as follows:

[0051]

[0052] in, This is a shearing factor used to limit the magnitude of policy updates and prevent training instability. This represents the old strategy before the update. This is the updated strategy.

[0053] Step 5.13: Complete After one training cycle, use the updated policy parameters of all agents. Update the corresponding old strategy parameters. And the updated Critical parameters after using all agents. Update its corresponding target Critic parameters ;

[0054] Step 5.14: End the total number of iterations.

[0055] In a MEC system containing multiple edge servers, this method allows users to dynamically select the optimal edge server based on real-time status, achieving efficient utilization of heterogeneous computing resources while satisfying task dependencies.

[0056] The present invention, by adopting the above-described technical solution, has the following beneficial effects:

[0057] This invention constructs an MEC system model comprising a cloud server, multiple edge servers, and multiple intelligent agents, and establishes a mobile model for intelligent terminals, a communication model, and a task computation model that comprehensively considers latency and energy consumption. Based on this, a task offloading optimization function is constructed with the objective of minimizing the system's long-term average computational cost. Furthermore, the optimization problem is modeled as a Markov decision process. Considering that different task scheduling queues lead to different completion latencies for DAG tasks, a MAPPO algorithm integrating a task priority scheduling mechanism is proposed to optimize the task computation offloading strategy for multi-users with dependencies in an MEC environment. Through a framework of centralized training and distributed execution, multiple agents can collaboratively learn the optimal task offloading decision based on environmental conditions. Attached Figure Description

[0058] Figure 1 This is a flowchart of the task unloading method of the present invention. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and preferred embodiments. However, it should be noted that many details listed in the specification are merely to provide the reader with a thorough understanding of one or more aspects of the present invention, and these aspects of the invention can be implemented even without these specific details.

[0060] like Figure 1 As shown, a method for unloading complex dependency tasks based on multi-agent reinforcement learning includes the following steps:

[0061] S1: Build a “device-edge-cloud” MEC system model, which includes a local computing layer, a MEC layer, and a cloud computing layer;

[0062] S2: Establish the intelligent terminal mobile model, communication model, and task calculation model that comprehensively considers latency and energy consumption in the MEC system model;

[0063] S3: Calculate the total computational cost of each user executing tasks under different offloading decisions, and construct an optimization function for the task offloading strategy with the goal of minimizing the long-term average computational cost of all users in the MEC system.

[0064] S4: Model the final optimization objective as a Markov decision process;

[0065] S5: A MAPPO algorithm integrating task priority scheduling mechanism is proposed. The algorithm first designs task priority scheduling based on DAG to generate task queues; on this basis, the MAPPO algorithm is used for collaborative optimization training to solve the optimal task offloading decision.

[0066] S1. Construct an "edge-cloud" MEC system model, which includes a local computing layer, an MEC layer, and a cloud computing layer. Within the local layer, the set of intelligent agent terminals awaiting execution is... Smart terminal devices The generated application in the time slot It can be modeled as a DAG , where the set of nodes This represents all dependent subtasks in the application, and the edges... This indicates that a task cannot begin computation until all its predecessor tasks have been completed. Each task... Corresponding to a triple ,in This indicates the number of CPU cycles required to process each bit of task. Indicates the data size of the subtask. This represents the tolerable latency of the subtask. The MEC layer consists of multiple edge base stations, each connected to an MEC server via fiber optic cable. The set of MEC servers is denoted as [missing information - likely a typo]. The cloud computing layer comprises multiple servers with powerful computing capabilities and massive storage resources, which can be used to process large-scale data.

[0067] S2, during the execution of step 2, establish the intelligent terminal mobile model, communication model, and task calculation model that comprehensively considers latency and energy consumption in the MEC system. The specific steps are as follows:

[0068] Step 2.1: Construct a mobile model for smart terminals

[0069] Smart terminal The velocities are independent and identically distributed, typically following a Gaussian-Markov random process. Let... and The speeds of the smart terminal in the current time slot and the next time slot are represented by the following expressions:

[0070] ;

[0071] in, This represents an uncorrelated Gaussian random process with a mean of zero and a variance of 1. Indicates smart terminal The memory depth parameter of velocity during motion; and They represent smart terminals respectively. asymptotic mean and asymptotic standard deviation;

[0072] Step 2.2: Constructing the communication model

[0073] Signal-to-interference-plus-noise ratio (SINR) from the smart terminal to the MEC server:

[0074] ;

[0075] in, and They represent smart terminals respectively. and smart terminals The transmission power; This indicates the channel gain between the smart terminal and the MEC server; This represents the power of additive white Gaussian noise;

[0076] Step 2.3: Constructing the computational model

[0077] use Indicates task The uninstallation strategy, when Time indicates smart terminal Execute the task locally; when This indicates that the task is being executed on the MEC server; The time factor indicates that the task is executed in the cloud. The latency factors involved at each execution endpoint include waiting latency and execution latency. Energy consumption primarily considers the local energy consumption of the smart terminal and the energy consumed during task transmission. Therefore, the task-dependent computational model is represented as follows:

[0078] When the task's offloading decision is to execute locally, the task's latency and energy consumption are as follows:

[0079]

[0080] in This indicates the latency of the task execution locally. This represents the local computing power, where Indicates task The set of precursor tasks. , ,and Representing tasks Execute locally, or on the server. The execution time and the completion time when executed in the cloud computing center. This represents the minimum completion latency of a task executed by a smart terminal at a certain moment, denoted as the idle latency of the local device. The waiting time for local execution to begin. Switched capacitors for smart terminal devices. Energy consumption for local execution.

[0081] When the task offloading decision is to execute at the edge layer, the task latency is represented as follows:

[0082]

[0083] in Indicates task Transmission latency to the MEC server, This represents the allocated communication bandwidth, assuming the task... If the target server is not a MEC server covered by the current location, then it is necessary to calculate the distance from the current server to the target server specified in the unloading decision. express arrive Transmission delay, This indicates the transmit power of the smart terminal. express arrive transmission rate express arrive The number of jumps, This indicates the delay before the edge server begins execution. This indicates the minimum time required to complete the task on the target server at a given moment. It means the task will complete when it arrives at the target server and there are available resources. , This indicates the execution latency of the task on the MEC server. This indicates the execution latency of the task on the edge server.

[0084] Due to the high downlink transmission rate, this paper ignores downlink transmission time and energy consumption. Therefore, the energy consumption when the task is offloaded to the edge server is expressed as follows:

[0085]

[0086] in and These represent the energy consumption for uploading tasks to the edge server and, respectively. arrive Transmission energy consumption.

[0087] When the task offloading decision is to be executed in the cloud, the task's latency and energy consumption are as follows:

[0088]

[0089] in For the transmission latency of tasks from the edge server to the cloud, This indicates the transmission rate between the edge server and the cloud. For cloud computing power, For cloud execution latency, This indicates the delay before the task begins execution in the cloud. Total energy consumption for offloading tasks to the cloud.

[0090] S3. Establish the system cost optimization objective function. Based on step 2, combined with the optimization model and corresponding constraints, establish the system cost optimization objective function for dependent task unloading. The total system cost in this paper is expressed as follows:

[0091]

[0092] in and It is a control coefficient used to balance the total system delay and energy consumption, and satisfies Therefore, the final optimization objective of this paper can be expressed as follows:

[0093]

[0094] This indicates that the total latency of the computation task is less than the maximum tolerable latency. This represents the constraint on the computing power of each smart terminal. This means that the total computing resources allocated to each task cannot exceed the computing power of the MEC server.

[0095] S4 abstracts each smart terminal device as an intelligent agent. Since the intelligent agent itself does not have the ability to acquire all environmental information, the optimization problem needs to be transformed into POMDP. The POMDP interaction process uses... It means that among them Represents the global state space. Represents the local observation set. The set representing the action space. This indicates a reward. In each time slot... At the beginning, intelligent agents From the environment Obtain environmental observation information Thus, actions can be selected based on existing strategies. The joint action of all intelligent agents To transition to the next state The intelligent system receives a reward value from the environment. .

[0096] The environmental conditions are used to make offloading decisions, including the total data size of all tasks in the entire MEC system, the location of smart terminals, the signal-to-noise ratio, the local waiting start time of tasks, the waiting start time of edge servers, and the waiting start time of cloud computing. Therefore, The environmental state at any given time is represented as follows:

[0097]

[0098] Local observation includes intelligent agents In the time slot The unloading decision is made based on the task data size, smart terminal location, signal-to-noise ratio, local task start time, edge server start time, and cloud computing start time.

[0099]

[0100] An action refers to the action taken by an intelligent agent after obtaining an observation. In the edge-cloud computing system described in this chapter, the intelligent agent can choose to keep the task running locally, offload it to a heterogeneous edge server, or offload it to the cloud center for execution. Therefore, the intelligent agent... The body's movement is:

[0101]

[0102] The optimization objective is to minimize the latency and energy consumption for all users while satisfying the maximum tolerable latency. This is the function of the intelligent agent. The reward can be set as follows:

[0103]

[0104] S5, MAPPO is a multi-agent deep reinforcement learning algorithm based on the AC framework. In the MAPPO algorithm, the Actor network... Critic network is used to output action policies. The strategy used to evaluate the current state in multi-agent learning is... To indicate, among which and These represent the parameters of the Actor network and the Critic network, respectively. Each user seeks the optimal task offloading decision by training a multi-agent proximal optimization algorithm based on the current environment state. Considering that different task scheduling queues will lead to different completion delays for DAG tasks, the MAPPO algorithm incorporates a task priority scheduling mechanism. The algorithm specifically includes the following steps:

[0105] Step 5.1: Randomly initialize each agent Actor network parameters and Critic network parameters Initialize the old strategy, target Critic network, and initialize the experience replay buffer. ;

[0106] Step 5.2: Initialize the total number of iterations (episode);

[0107] Step 5.3: Obtain the initial system state

[0108] Step 5.4: Set the number of time steps for each iteration ;

[0109] Step 5.5: For DAG tasks generated by smart terminals, design a task priority scheduling method based on DAG topology. Priority information is based on DAG topology information and task tolerance latency, while offloading decisions are made in ascending order of priority metrics. The specific steps are as follows:

[0110] The average latency of the subtask is calculated as follows:

[0111]

[0112] Based on the task's tolerance delay as well as Calculate the task priority for each task. It is represented as follows:

[0113]

[0114] Step 5.6: Each agent generates an action based on the currently observed state. ;

[0115] Step 5.7: The agent executes a joint action. Each agent receives a reward. and new status ;

[0116] Step 5.8: Collect trajectory data for each agent. And store it in the experience replay pool. ;

[0117] Step 5.9: Repeat steps 5.5 to 5.8 until completion. Up to the next time step;

[0118] Step 5.10: Calculate the advantage value for each agent. The formula is as follows:

[0119]

[0120] in, To balance the discount factor between the importance of current and future rewards, The parameter is used to balance the bias and variance. It's a cumulative discount reward, therefore the Critic network is applied to the state. The value estimation function is expressed as ;

[0121] Step 5.11: Begin Each training cycle shuffles the data in the buffer. Following the order of the data, and randomly selecting small batches, update the network parameters as follows;

[0122] Step 5.12: Update the Critic network for each agent. The loss function is calculated as follows:

[0123]

[0124] Update the Actor network for each agent. The loss function is calculated as follows:

[0125]

[0126] in, This is a shearing factor used to limit the magnitude of policy updates and prevent training instability. This represents the old strategy before the update. This is the updated strategy.

[0127] Step 5.13: Complete After one training cycle, use the updated policy parameters of all agents. Update the corresponding old strategy parameters. And the updated Critical parameters after using all agents. Update its corresponding target Critic parameters ;

[0128] Step 5.14: End the total number of iterations.

[0129] Matters not covered in this invention are common knowledge.

[0130] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for unloading complex dependency tasks based on multi-agent reinforcement learning, characterized in that: The method includes the following steps: Step 1: Build an edge-cloud MEC system model, which includes a local computing layer, an MEC layer, and a cloud computing layer; Step 2: Establish the intelligent terminal mobile model, communication model, and task calculation model that comprehensively considers latency and energy consumption in the MEC system model; Step 3: Calculate the total computational cost of each smart terminal executing tasks under different offloading decisions, and construct an optimization function for the task offloading strategy with the goal of minimizing the average computational cost of all smart terminals in the MEC system; Step 4: Model the final optimization objective as a Markov decision process; Step 5: Construct several agent-based proximal policy optimization algorithms that integrate task priority scheduling mechanism. The proximal policy optimization algorithm first designs task priority scheduling based on directed acyclic graph, generates task queue, and uses several agent-based proximal policy optimization algorithms for collaborative optimization training to solve the optimal task unloading decision.

2. The method for unloading complex dependency tasks based on multi-agent reinforcement learning according to claim 1, characterized in that: In step 1, the set of intelligent agent terminals for all tasks waiting to be executed in the local layer is: Smart terminal devices The generated application in the time slot It can be modeled as a DAG , where the set of nodes This represents all dependent subtasks in the application, and the edges... This means that a task can only begin computation after its predecessor tasks have been completed. In other words, a task can only begin computation after all its predecessor tasks have been completed. Corresponding to a triple ,in This indicates the number of CPU cycles required to process each bit of task. Indicates the data size of the subtask. This indicates the tolerance latency for the subtask.

3. The method for unloading complex dependency tasks based on multi-agent reinforcement learning according to claim 1, characterized in that: In step 1, the MEC layer consists of several edge base stations, each connected to an MEC server via optical fiber. The set of MEC servers is as follows: The cloud computing layer consists of several servers, which have computing power and storage resources for processing data.

4. The method for unloading complex dependency tasks based on multi-agent reinforcement learning according to claim 1, characterized in that, The specific process of step 2 is as follows: Step 2.1: Construct the mobile model and communication model of the smart terminal, and obtain the signal-to-interference-plus-noise ratio (SIR) of the smart terminal within the time slot. ; Step 2.2: Construct a computational model, using... Indicates task The uninstallation strategy, when Time indicates smart terminal Execute the task locally when This indicates that the task is executed on the MEC server. The time factor indicates that the task is executed in the cloud. The latency factors involved at each execution end include waiting latency and execution latency. The energy consumption factors consider the user's local energy consumption and the energy consumption during task transmission. The computational model dependent on the task is represented as follows: When the task's offloading decision is to execute locally, the task's latency and energy consumption are as follows: in This indicates the latency of the task execution locally. The waiting time for local execution to begin. Switched capacitors for smart terminal devices. For local execution energy consumption, For local computing power; When the task offloading decision is to execute at the edge layer, the task completion latency and energy consumption are as follows: in This indicates the delay before the edge server begins execution. This indicates the execution latency of the task on the MEC server. Indicates task The transmission latency for uploading to the MEC server is used express arrive The transmission latency is negligible due to the high downlink transmission rate; in and These represent the energy consumption for uploading tasks to the edge server and, respectively. arrive Transmission energy consumption; When the task offloading decision is to be executed in the cloud, the task completion latency and energy consumption are as follows: in, For cloud execution latency, This indicates the delay before the task begins execution in the cloud. For the transmission latency of tasks from the edge server to the cloud, Total energy consumption when the task is offloaded to the cloud.

5. A method for unloading complex dependency tasks based on multi-agent reinforcement learning according to claim 1, characterized in that: In step 3, the optimization objective of the optimization function is to optimize the computation offloading decision of the requested tasks in all smart terminals, and to minimize the system cost consisting of task completion latency and energy consumption while satisfying the task tolerance latency constraint.

6. A method for unloading complex dependency tasks based on multi-agent reinforcement learning according to claim 1, characterized in that: In step 4, the dependency unloading problem is optimized into a Markov decision process that requires determining its global state space. Local observation space Action space Reward Space .

7. A method for unloading complex dependency tasks based on multi-agent reinforcement learning according to claim 1, characterized in that: In step 5, when executing several agent proximal policy optimization algorithms, the Actor network... Critic network is used to output action policies. The strategies learned by several agents are used to evaluate the current state. To indicate, among which and These represent the parameters of the Actor network and the Critic network, respectively. Each user seeks the optimal task offloading decision by training a proximal optimization algorithm based on several agents according to the current environment state. Considering that different task scheduling queues will lead to different completion delays of DAG tasks, the proximal policy optimization algorithm of several agents incorporates a task priority scheduling mechanism.

8. A method for unloading complex dependency tasks based on multi-agent reinforcement learning according to claim 7, characterized in that, The specific process of several agent proximal policy optimization algorithms is as follows: Step 5.1: Randomly initialize each agent Actor network parameters and Critic network parameters Initialize the old strategy, target Critic network, and initialize the experience replay buffer. ; Step 5.2: Initialize the total number of iterations (episode); Step 5.3: Obtain the initial system state ; Step 5.4: Set the number of time steps for each iteration ; Step 5.5: For DAG tasks generated by smart terminals, design a task priority scheduling method based on DAG topology. Priority information is based on DAG topology information and task tolerance latency, while offloading decisions are made in ascending order of priority metrics. The specific steps are as follows: The average latency of the subtask is calculated as follows: Based on the task's tolerance delay as well as Calculate the task priority for each task. It is represented as follows: Step 5.6: Each agent generates an action based on the currently observed state. ; Step 5.7: The agent executes a joint action. Each agent receives a reward. and new status ; Step 5.8: Collect trajectory data for each agent. And store it in the experience replay pool. ; Step 5.9: Repeat steps 5.5 to 5.8 until completion. Up to the next time step; Step 5.10: Calculate the advantage value for each agent. The formula is as follows: in, To balance the discount factor between the importance of current and future rewards, The parameter is used to balance the bias and variance. It's a cumulative discount reward, therefore the Critic network is applied to the state. The value estimation function is expressed as ; Step 5.11: Begin Each training cycle shuffles the data in the buffer. Following the order of the data, and randomly selecting small batches, update the network parameters as follows; Step 5.12: Update the Critic network for each agent. The loss function is calculated as follows: Update the Actor network for each agent. The loss function is calculated as follows: in, This is a shearing factor used to limit the magnitude of policy updates and prevent training instability. This represents the old strategy before the update. This is the updated strategy. Step 5.13: Complete After one training cycle, use the updated policy parameters of all agents. Update the corresponding old strategy parameters. And the updated Critical parameters after using all agents. Update its corresponding target Critic parameters ; Step 5.14: End the total number of iterations.