Dependent task-oriented intelligent edge cooperative computing unloading method

By building a multi-agent deep reinforcement learning framework and collaborative computing model, the computational offloading complexity of dependent tasks in mobile edge computing is solved, and the system service delay optimization and resource utilization are achieved.

CN120302439APending Publication Date: 2025-07-11CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510438309.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When handling dependent tasks of complex structures, existing mobile edge computing methods have high computational offload complexity, low resource utilization, and existing algorithms are difficult to adapt to the dynamic changes of computing requests and complex execution dependencies.

Method used

Build a multi-agent deep reinforcement learning framework, combine end-to-end collaboration and edge-to-edge collaborative computing models, optimize task scheduling and unloading strategies, and minimize the system's long-term time average service delay through the MADRL-Dep algorithm.

Benefits of technology

In different MEC scenarios, efficient collaborative computing between terminal devices and edge nodes is realized, system service delay is optimized, and resource utilization is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302439A_ABST
    Figure CN120302439A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of mobile communication, and particularly relates to a dependent task-oriented intelligent edge cooperative computing unloading method, which comprises the following steps of: constructing a system model comprising a plurality of users and a plurality of base stations; constructing an application model, a calculation unloading model, a task scheduling model, a communication model, a local equipment calculation model, a local base station calculation model, an end-to-end cooperative calculation model and an edge-to-edge cooperative calculation model based on the system model; constructing an optimization problem model by taking the minimization of the long-term average service time delay of the system as a target; constructing a deep reinforcement learning framework, and obtaining an environment state, local observation, an action and a reward function; an MADRL-Dep algorithm is adopted to solve the optimization problem to obtain an unloading strategy; according to the method, efficient cooperative computing between terminal equipment and edge nodes under the condition of diversified task execution dependency and cache dependency is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of mobile communications, and particularly relates to an intelligent edge collaborative computing offloading method for dependent tasks. Background Art

[0002] With the rapid development of technologies such as the Internet of Things, artificial intelligence, and big data, emerging applications with the main characteristics of computing-intensive and latency-sensitive, such as augmented / virtual reality, driverless, and intelligent healthcare, have put forward more stringent requirements for real-time performance, computing reliability, and network robustness. Mobile Edge Computing (MEC) has become the core solution to meet these extreme computing requirements, and as one of the key technologies in the field of MEC, computing offloading has become a hot research area for a large number of domestic and foreign scholars. However, the booming development of technologies such as artificial intelligence and the Internet of Things has also spawned a large number of emerging applications mainly characterized by complex structures, such as face recognition, autonomous driving, and intelligent healthcare systems. These applications usually consist of multiple tasks, and there are complex execution dependencies between tasks. For example, a face recognition application usually consists of multiple tasks such as image acquisition, face detection, preprocessing, feature extraction, and classification. Image acquisition provides raw data, face detection locates the face based on the acquired image, preprocessing optimizes the detected image, feature extraction analyzes the preprocessed image to obtain key features, and classification uses these features for identity recognition. These tasks are closely linked, and the output of the previous task is the input of the next task, with a high degree of dependency. At the same time, the calculation of the task also requires the processor to cache data related to the task type. For example, the feature extraction and classification tasks in face recognition can only be executed on a processor that has cached the relevant machine learning model. However, since edge nodes usually only cache some popular services, it further increases the complexity of computing offloading. Some of the main achievements are as follows:

[0003] Researchers at home and abroad have conducted research on this problem. Some scholars have proposed a task offloading method for dependency and service caching in mobile edge computing. This method aims at minimizing service latency for computing tasks with execution and caching dependencies, establishes an optimization problem model for collaborative computing offloading of multiple edge nodes, and proposes a one-shot decision-making computing offloading method based on convex optimization to solve this problem. However, in the actual MEC network environment, users' computing requests usually change dynamically. Therefore, the above one-shot decision-making computing offloading method based on convex optimization is not applicable to scenarios where computing requests change dynamically. Some other scholars have proposed a dynamic offloading method for dependent tasks based on deep reinforcement learning in mobile edge computing. This method considers the computing offloading problem in a multi-user scenario, aims at minimizing the weighted sum of service latency and energy consumption, and proposes a deep reinforcement learning method based on the actor-critic network. However, in this method, computing tasks can only be computed on local devices or edge nodes, that is, the resources of nearby idle devices or neighbor edge nodes are not utilized for collaborative computing, which will lead to low resource utilization in the edge system and thus poor computing performance and other problems. Summary of the Invention

[0004] To solve the above problems, the present invention provides an intelligent edge collaborative computing offloading method for dependent tasks, including the following steps:

[0005] S1. Construct a system model including multiple users and multiple base stations;

[0006] S2. Based on the system model, construct an application model, a computing offloading model, a task scheduling model, a communication model, a local device computing model, a local base station computing model, an end-to-end collaborative computing model, and an edge-to-edge collaborative computing model;

[0007] S3. With the goal of minimizing the long-term time-average service latency of the system, construct an optimization problem model;

[0008] S4. Based on steps S1 - S3, construct a deep reinforcement learning framework to obtain the environmental state, local observation, action, and reward function;

[0009] S5. Use the MADRL-Dep algorithm to solve the optimization problem to obtain the offloading strategy.

[0010] Advantages of the present invention:

[0011] The present invention comprehensively considers factors such as the dynamics of diverse computing tasks, execution and cache dependencies, etc., designs an end-to-end collaborative computing model among users and an edge-to-edge collaborative computing model among edge nodes. On this basis, the service delay of users is analyzed based on the execution dependency relationship between tasks; then, with the goal of minimizing the system service delay, an optimization problem model of collaborative computing offloading for dependent tasks is established, and the optimization goal is to minimize the long-term average system delay under the constraints of task execution dependency relationships and edge node cache capacities; a computing offloading optimization method driven by multi-agent deep reinforcement learning is proposed, enabling the MEC system to efficiently coordinate the computing and storage resources of edge nodes and users and realizing an intelligent collaborative computing offloading strategy for dependent tasks.

[0012] The simulation results show that, compared with existing algorithms, the proposed algorithm has the optimal system delay in different MEC scenarios, and realizes efficient collaborative computing between terminal devices and between edge nodes under diverse task execution dependencies and cache dependencies. Description of the Drawings

[0013] Figure 1 It is a system model diagram of the present invention;

[0014] Figure 2 It is a task dependency model diagram of the present invention;

[0015] Figure 3 It is an example diagram of the task scheduling process of the DAG application of the present invention;

[0016] Figure 4 It is a framework diagram of the MADRL-Dep algorithm of the present invention;

[0017] Figure 5 It is a curve diagram of the convergence process of different algorithms of the present invention;

[0018] Figure 6 It is a curve diagram of the influence of the number of users of the present invention on the performance of different algorithms;

[0019] Figure 7 It is a curve diagram of the influence of the number of neighboring base stations of the present invention on the performance of different algorithms;

[0020] Figure 8 It is a curve diagram of the influence of the base station storage capacity of the present invention on the performance of different algorithms;

[0021] Figure 9 It is a curve diagram of the influence of the number of tasks of the present invention on the performance of different algorithms;

[0022] Figure 10 It is a curve diagram of the influence of the user computing resources of the present invention on the performance of different algorithms;

[0023] Figure 11It is the influence curve of the application request probability of the present invention on the performance of different algorithms. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0025] The present invention provides an intelligent edge collaborative computing offloading method for dependent tasks, including the following steps:

[0026] S1. Construct a system model including multiple users and multiple base stations.

[0027] Specifically, as Figure 1 shown, a MEC server is equipped at each base station, and each base station can act as both a local base station and a neighbor base station. In the embodiments of the present invention, it is considered that M users (UEs) are randomly distributed within the coverage range of a single base station, and the user set is denoted as The neighbor base stations are represented by NBS, and the local base stations are represented by LBS. The base station set in the system model is denoted as

[0028] {LBS}, where LBS represents a local base station, represents the neighbor base station set of LBS, and NBS n represents the nth (n = 1, 2,..., N) neighbor base station of LBS, and N represents the number of neighbor base stations; in the system model, time is discretized into a set of time slots with equal lengths ξ m (t) ∈ {0, 1} represents the application request status of user at time slot . If user m does not generate an application at time slot t, it means it is in an idle state, then ξ m (t) = 0. If user m generates an application at time slot t, then ξ m (t) = 1.

[0029] S2. Based on the system model, construct an application model, a computing offloading model, a task scheduling model, a communication model, a local device computing model, a local base station computing model, an end-to-end collaborative computing model, and an edge-to-edge collaborative computing model.

[0030] Specifically, the application model is used to model applications that generally have task dependencies. The present invention uses DAG to describe applications with task dependencies:

[0031] Define the set of task types in the system model Denote the number of task types; the application G requested by user m at time slot t m (t) = {V m (t), E m (t)}, Denote the set of application tasks, and Denote the virtual start task and the virtual end task; |V m | + 1 denotes the total number of tasks of user m at time slot t, because the index of the last task is |V m |, so the total number is |V m | + 1; E m (t) represents the execution dependency between tasks, The i-th task requested by user m at time slot t, where Denote the task Data volume, Denote the task CPU cycles required per bit of data, Denote the task Type; for the virtual start task And the virtual end task The data volume is 0; specifically, use Denote the task And the task The execution dependency between them, that is, the task Is the predecessor task of the task ; The present invention indicates the dependency between two tasks by the edge weight size, that is, use δ i,j (t) represents the calculation result size from the task To the task , Let Denote the task Is the predecessor task of the task , The edge weight size between the two tasks is δ i,j (t), if δ i,j (t) = 0, it indicates that there is no dependency between the two tasks, that is, there is no edge.

[0032] Specifically, the calculation offloading model is used to model the offloading decision variables of users in each time slot, including:

[0033] Generally, the computing tasks of users can be computed on local devices or offloaded to the local base station to which the user belongs for computing. The present invention additionally considers several types of horizontal collaborative computing offloading models such as the end-to-end collaborative computing between the remaining idle users and the local device, and the edge-to-edge collaborative computing between the local base station and the neighboring base stations.

[0034] For the end-to-end collaborative computing mode, assume that users in the idle state can provide computing services for users who have requested applications in the system. Let denote the task and the set of users who can provide computing services at time slot t. Then

[0035]

[0036] where indicates that user m has not requested a computing application.

[0037] For the edge-to-edge collaborative computing mode, to improve the utilization rate of storage resources, the MEC server deployed on the base station side usually only caches the relevant data (such as program code and databases) required for some popular task types. For simplicity of description, the following content uses "cached task type" to refer to "cached data related to this task type". Let denote the task and the set of offloadable edge nodes at time slot t. Φ x (t) represents the cached task type of base station x at time slot t. Then

[0038]

[0039] where indicates that user m has not requested a computing application.

[0040] Let denote the computing offloading decision variable of task at time slot t. Then

[0041]

[0042] where, if task is computed on the local device, then if task is offloaded to the local base station for computing, then if user m has not requested the application, then if task is offloaded to the neighbor base station NBS n for computing, then if the task is offloaded to another user m' for computing, then the virtual start task and the virtual end task can only be computed on the local device. Therefore

[0043] In summary, the offloading decision variable of user m at time slot t

[0044] Specifically, the task scheduling model is used to model the task scheduling order and process of DAG tasks:

[0045] Due to the complex execution dependencies among tasks, define to represent the task scheduling order of G m (t), to represent the scheduling priority of task ; The smaller the value, the higher the scheduling priority of task . The task scheduling order Q m (t) satisfies the execution dependencies among tasks. Under this condition, an application may have multiple feasible task scheduling orders, and different task scheduling orders will result in different application computing delays. Due to the constraints of the execution dependencies among tasks, define the start time ST to represent the earliest time when a task can start execution, and define the finish time FT to represent the earliest time when a task is computed to completion. In addition, considering the limited communication resources of users, it is set that a user can only send one task at the same time.

[0046] As Figure 2 shown, the scheduling process of each task in the DAG application is as Figure 3 shown, where

[0047] Q(t) = (0, 1, 2, 3, 4, 5, 6, 7, 8)

[0048] α(t) = (1, NBS1, 1, 1, LBS, NBS2, LBS, 2, 1).

[0049] Specifically, in the MEC system of the present invention, the OFDMA technology is adopted, and the communication model is used to model the transmission rate of the system:

[0050] The uplink transmission rate from user m to the local base station in time slot t is expressed as

[0051]

[0052] where B up represents the uplink bandwidth, P m represents the transmission power of user m, G m,LBS (t) represents the channel gain between user m and the local base station, and σ 2 represents the background noise power;

[0053] The downlink transmission rate from the local base station to user m in time slot t is expressed as

[0054]

[0055] where B down represents the downlink bandwidth, P LBSDenote the transmission power of the local base station as G LBS,m (t) represents the channel gain between the local base station and user m;

[0056] In addition, in the end-to-end collaborative computing mode considered in the present invention, the device-to-device (D2D) communication model is adopted between users. Then, the transmission rate from user m to user m' at time slot t is expressed as

[0057]

[0058] where B D2D denotes the D2D bandwidth, and G m,m’ represents the channel gain between user m and user m'.

[0059] Specifically, the local device computing model is used to model the computing process of tasks on local devices:

[0060] According to the foregoing definition, if the task of user m is computed on its local device, then user m will provide CPU computing resources to execute this task. The present invention defines the end time when the task is computed on the local device corresponding to user m as

[0061]

[0062] denotes the start time when the task is executed on the local device. According to the task scheduling model described in the present invention, it is also the maximum value of the transmission completion time of the computed results of the predecessor task of the task ;

[0063] denotes the computing delay when the task is executed on the local device. The specific calculation formula is

[0064]

[0065]

[0066] where denotes the end time when the predecessor task of the task is executed on the local device, that is, the earliest time when the computation is completed; denotes the end time when the predecessor task of the task is executed at ; Denote the task The set of users who can provide computing services in time slot t Denote the task The set of offloadable edge nodes in time slot t; Denote the task When computing on the local device, the previous task The data transmission delay for sending the computation result to the local device; Denote the task The set of previous tasks of ; 1 {·} Is an indicator function, if the event {·} is true, then 1 {·} = 1, otherwise 1 {·} = 0; Denote the task The data volume of the task Denote the task The number of CPU cycles required per bit of data, f m (t) represents the computing resources allocated by user m to the task in time slot t In the present invention, it is considered that any user evenly allocates computing resources to the arriving tasks; F m Represents the total computing resources of user m; Represents the available communication time of the device of cooperative user m' when sending the computation result in time slot t Represents the transmission rate from user m to user m' in time slot t; Represents the downlink transmission rate from the local base station to user m in time slot t; Represents the average transmission rate between neighbor base station n and the local base station; Represents the set of computation results sent before the computation result Represents the size of the computation result of task to task on user m. If an idle user receives a computation request for a task, the user will allocate computing resources to execute it. When the task computation is completed, the computation result enters the communication scheduling queue of the user in a FIFO manner. When there are multiple computation results for a computation task (i.e., the task has multiple subsequent tasks), they enter the communication scheduling queue in sequence based on the task scheduling order of the application.

[0067] Specifically, the local base station computing model is used to model the computing process of tasks at the local base station:

[0068] According to the foregoing definitions, if the task of user m is computed at the local base station associated with it, then ​​User m uploads the task to the local base station via the wireless network, and the MEC server configured by the local base station executes the task. According to the communication model described in the present invention, the task The uplink transmission delay during local base station calculation is After the task is uploaded to the local base station and can be executed, the MEC server configured in the local base station provides computing resources to execute the task.

[0069] Define the task The end time during the calculation at the local base station corresponding to user m Denoted as

[0070]

[0071] Denote the task The start time of execution at the local base station. According to the task scheduling model described in the present invention, Is the maximum value of the completion time of the calculation result transmission of the predecessor task and the completion time of the data transmission; Denote the task The computing delay during execution on the local device. The specific calculation formula is

[0072]

[0073] Among them, Denote the task Of the predecessor task The end time during execution at the local base station, Denote the task During the calculation at the local base station, the predecessor task The data transmission delay for sending the calculation result of to the local base station; Denote the communication resource available time when user m sends the task at time slot t ; Denote the task set scheduled before the task obtained based on the task scheduling order Q m (t) ; Denote the task The data volume of, Denote the uplink transmission rate from user m to the local base station at time slot t, Denote the transmission rate from user m to user m' at time slot t; Denote the task Of the predecessor task At The end time of execution at; f LBS Denote the CPU clock frequency of the local base station. Denote the collaborative user The device sends the calculation result at time slot t The available time of communication resources when sending the calculation result at time slot t

[0074] Specifically, the end-to-end collaborative computing model is used to model the computing process of task selection for end-to-end collaboration;

[0075] When the task selects end-to-end collaborative computing, that is At this time, user m first sends the task to collaborative user m' based on D2D communication. After obtaining the calculation result of the predecessor task, collaborative user m' will provide computing resources to execute the task.

[0076] Define the end time when the task is calculated on the device corresponding to collaborative user m Denoted as

[0077]

[0078] Denote the start time when the task is executed on the device of collaborative user m, Denote the start time when the task is executed on the device of collaborative user m, the calculation delay, and the specific calculation formula is

[0079]

[0080] Wherein, Denote the end time when the predecessor task of the task is executed on the device of collaborative user m; Denote the data transmission delay when the calculation result of the predecessor task is sent to the device of collaborative user m when the task is calculated on the device of collaborative user m ; Denote the end time when the predecessor task of the task is executed at ; Denote the transmission delay when the task is executed on the device of collaborative user m, Denote the available time of communication resources when user m sends the task at time slot t; Denote the transmission rate from user m to user m' at time slot t; Denote the available time of communication resources when sending the task at time slot t to the task ; Denote at time slot t from The transmission rate from the location to user m'. Indicates the downlink transmission rate from the local base station to the collaborative user m' in time slot t. Indicates the average transmission rate between the neighbor base station n and the local base station; f m’ Indicates the computing resources allocated to the task by the collaborative user m' in time slot t for the task.

[0081] Specifically, the edge-edge collaborative computing model is used to model the computing process of task selection for edge-edge collaboration:

[0082] If the local base station does not cache the corresponding task type or the computing resources do not meet the computing requirements, the task can be migrated to a nearby base station for edge-edge collaborative computing. When the task selects edge-edge collaborative computing, that is user m first sends the task to the local base station, and then the local base station LBS forwards it to the neighbor base station n (NBS n ). After the task can start execution, the MEC server configured by NBS n allocates computing resources to execute the task.

[0083] Define the end time when the task is being computed at the neighbor base station denoted as

[0084]

[0085] Indicates the start time when the task is being executed at the neighbor base station n, Indicates the computing delay when the task is being executed at the neighbor base station n, and the specific calculation formula is

[0086]

[0087] where Indicates the end time when the predecessor task of the task is being executed at the neighbor base station n, Indicates the data transmission delay when the computing result of the predecessor task of the task is sent to the neighbor base station n during the computing at the neighbor base station n; Indicates the end time when the predecessor task of the task is being executed at the location, Indicates the uplink transmission delay when the task is being computed at the neighbor base station n, Indicates that user m sends a task in time slot t for the available time of communication resources; Indicates the computing resources allocated by neighbor base station n to the task in time slot t.

[0088] In the formula, T m (t) = 0 indicates that user m does not request an application in time slot t, and the service delay is defined as 0 at this time.

[0089] S3. With the goal of minimizing the long-term time-average service delay of the system, an optimization problem model is constructed.

[0090] Specifically, define T m (t) to represent the service delay of user m in time slot t,

[0091]

[0092] Through optimizing the task scheduling order and task offloading strategy, the present invention minimizes the long-term time-average service delay of the system. The optimization problem model is as follows:

[0093]

[0094] In the formula, represents the decision of the task scheduling order; represents the task offloading decision, represents the total number of time slots. The constraint condition means that the task can only be executed on an idle UE or an edge node that has cached the task type. This optimization problem is a sequential decision problem, and the service delay of the MEC system can be minimized by finding the optimal Q and α for each time slot.

[0095] S4. Based on steps S1 - S3, a deep reinforcement learning framework is constructed to obtain the environmental state, local observation, action, and reward function.

[0096] When using the traditional centralized optimization algorithm to formulate the computation offloading strategy, the centralized control node needs to obtain global information through frequent information interaction. However, due to the characteristics of distributed deployment of edge resources and limited communication resources, etc., it will face challenges such as low reliability and difficulty in obtaining a large amount of real-time system state information. Therefore, it is difficult to effectively solve this problem under the system constraint conditions.

[0097] To address the above challenges, the present invention proposes a Multi-Agent Deep Reinforcement Learning-based Computation Offloading for Dependent Tasks (MADRL-Dep) algorithm to solve the optimization problem.

[0098] First, transform the optimization problem into a Dec-POMDP. In the sequential decision-making optimization problem, the states of the users' computing resource status, idle status, application information status, and the task type cache status of the base stations at time slot t + 1 are only related to the states at time slot t, and are independent of the states before time slot t. That is, the computing offloading process has the Markov property. Therefore, the present invention transforms the optimization problem into a Dec-POMDP problem, where each user represents an agent. A Dec-POMDP can be described as a tuple where S represents the environmental state of all agents; O m represents the local observation space of agent m; A m represents the action space of agent m; R represents the reward function; γ ∈ [0, 1) represents the discount factor. Next, define the environmental state, local observation, action, and reward function of the Dec-POMDP of the present invention.

[0099] (1) Environmental state: According to the system model in the present invention, the environmental state includes the computing resource status, idle status, application information of all users in the MEC system, and the task type cache status of all base stations. Therefore, the environmental state s(t) ∈ S at time slot t is defined as:

[0100]

[0101] In the formula, the request information of application G m (t) includes the DAG dependency relationship between tasks, the data volume size of the tasks the number of CPU cycles required per bit of data and the task type

[0102] (2) Local observation: Agent m can observe the system device idle status, task type cache status, and its own requested application information and computing resource status at each time slot. Therefore, the local observation o m (t) ∈ O m of agent m at time slot t is defined as:

[0103]

[0104] (3) Action: According to the optimization problem and the Dec-POMDP process defined in the present invention, define the action a m (t) ∈ A m of agent m at time slot t as:

[0105]

[0106] (4) Reward function: According to the optimization problem, the common goal of each agent is to minimize the service delay of the MEC system. Therefore, the reward function of agent m at time slot t is defined as:

[0107]

[0108] S5. The MADRL-Dep algorithm is used to solve the optimization problem to obtain the offloading strategy.

[0109] The algorithm framework of MADRL-Dep is as Figure 4 shown. Each UE / agent is configured with a DDPG network and a task scheduling module, which are used to optimize the computing offloading strategy and obtain the dependent task scheduling order respectively. In the algorithm framework of MADRL-Dep, agent m receives a local observation o m (t) ∈ O m , selects an action a m (t) ∈ A m , and at the same time inputs am(t) into a task scheduling module. The task scheduling module returns the scheduling order Q m (t) of agent m based on the selected am(t). Then, agent m executes the action am(t) based on Q m (t), returns the reward rm(t) according to the reward function R, and transfers the local observation to the next state o m (t + 1). The ultimate goal of each agent is to explore and improve its policy μ m , so as to maximize the cumulative return .

[0110] Specifically, in the centralized training stage (the process of the black solid line), given the local observation o of agent m at the current time slot m , the actor network will take an action a m (o m ) according to the current policy μ m . Then, agent m inputs the action a m into its task scheduling module, and the task scheduling module outputs the DAG application scheduling order Q m of the agent. In the task scheduling module, for any application G m , first calculate the scheduling priority of each task in the application. For task its scheduling priority is obtained according to the task processing time cost and the internal execution dependency relationship of the DAG. Among them, the task processing time cost of is defined as:

[0111]

[0112] Furthermore, the task The scheduling priority of

[0113]

[0114] In the formula, represents the subsequent task set of task . If task has no subsequent tasks, then Then, sort the scheduling priorities of each task. The larger is, the higher the scheduling priority of task , that is, m The optimal task scheduling order of G is Next, the agent m will use the scheduling order Q m and the offloading action a m to interact with the MEC environment and obtain the reward r according to the reward function m , and at the same time transfer the local observation to the next state o′ m . The critic network evaluates the actions taken by the actor network according to a centralized action-value function, so as to guide the actor network to learn the optimal policy. Specifically, define and to represent the sets of actor networks and critic networks of all agents respectively, and the corresponding parameters are and For the actor network, any agent m updates the policy in the way of policy gradient, and the policy gradient is as follows:

[0115]

[0116] In the formula, X represents the number of mini-batch samples randomly sampled from the experience replay pool; x represents the xth sample; represents the centralized action-value function described by the parameter , that is, the critic network. The experience replay pool contains the sequence (o, o′, a, r), where o represents the observation information of all agents; o’ represents the local observations of all agents in the next time slot; a represents the actions of all agents; r represents the rewards of all agents.

[0117] The critic network evaluates the actions selected by the actor network to guide the actor network to learn the optimal offloading strategy. Agent m updates the corresponding parameter by minimizing the loss function. The loss function is expressed as follows:

[0118]

[0119] In the formula, μ′m represents the target actor network described by the parameter ; \(Q'\) m represents the target critic network described by the parameter . Finally, each agent updates the parameters of the target actor network and the target critic network in a soft update manner, which are respectively expressed as follows:

[0120]

[0121]

[0122] where represents the update rate.

[0123] During the centralized training process, the MADRL-Dep algorithm can adaptively coordinate the computing and storage resources of neighboring base stations and nearby idle users according to the current system state information and other user behaviors, so as to explore the optimal cooperative computing offloading strategy. When the training of the model parameters is completed, as Figure 4 shown, in the distributed execution stage (the red dotted line process), each UE / agent only needs to obtain the application information it requests, the computing resource status, as well as the MEC system service cache status and the UE idle status, and input them into the corresponding actor network to obtain its optimal offloading action.

[0124] The training process of MADRL-Dep is as shown in Table 1 of the following algorithm.

[0125] Table 1

[0126]

[0127] The time complexity of the MADRL-Dep algorithm in the training stage mainly depends on the number of agents \(M\), the number of training rounds \(T\) e , the number of time steps per round \(T\) s , the number of mini-batch samples \(X\), the number of edges \(E\) of the DAG application, and the structures of the actor network and the critic network of each agent. Assuming that the actor network and the critic both adopt fully connected networks with the same structure, the time complexity calculation during the MADRL-Dep training is as follows:

[0128]

[0129] where \(L\) a and \(L\) c respectively represent the number of layers of the actor network and the critic network; \(l\) a and \(l\) c respectively represent the \(l\) a -th layer of the actor network and the \(l\) c -th layer of the critic network; and respectively represent the number of neurons in the l-th a layer actor network and the number of neurons in the lc-th layer critic network; O(MT e T s E) represents the time complexity of the task scheduling module in the training phase. In the distributed execution phase, each agent only executes actions based on the actor network. Therefore, its time complexity is

[0130] In one embodiment, the present invention conducts simulation experiments to verify the performance of the proposed MADRL-Dep algorithm. The comparison algorithms include:

[0131] (1) no-coop. The no-coop algorithm does not consider the collaborative computing between UEs and between BSs, that is, tasks can only be computed on local devices or LBSs.

[0132] (2) no-E2E. The no-E2E algorithm does not consider the edge-edge collaborative computing between BSs, that is, tasks can only be computed on local devices, idle devices or LBSs.

[0133] (3) no-End2End. The no-End2End algorithm does not consider the end-to-end collaborative computing between UEs, that is, tasks can only be computed on local devices, LBSs or NBSs.

[0134] (4) Single-Agent Deep Reinforcement Learning-based Computation Offloading for Dependent Tasks (SADRL-Dep). Each agent in the SADRL-Dep algorithm is configured with a DDPG network and a task scheduling module. SADRL-Dep adopts the DTDE framework, that is, agents are only trained and executed independently according to their local observations.

[0135] It should be noted that the no-coop, no-E2E and no-End2End algorithms are only different from the MADRL-Dep algorithm in the computation offloading mode, and are exactly the same in other aspects.

[0136] This simulation uses the Pytorch 1.13 deep learning framework and the Adam optimizer to implement MADRL-Dep and all DRL-based comparison algorithms. All experiments are completed based on the Windows 10 operating system (CPU: Intel(R) Core(TM) i9-10920X CPU @ 3.5GHz, GPU: NVIDIA GeForce RTX 3090).

[0137] The MEC simulation scenario of the present invention includes one LBS and two NBSs. There are a total of five UEs evenly distributed within the coverage range of the LBS, and the coverage radius of the base station is 200m. Set the uplink radio channel bandwidth B up to be 10MHz; the downlink channel bandwidth B down is 20MHz; the D2D channel bandwidth B D2D is 10MHz; the channel gain is set to d -3 , where d is the distance between the transmitter and the receiver; the transmission power of the UE is set to 0.5W; the background noise power is set to -50dBm; the average transmission rate between BSs is set to 50Mbps; the total computing resources of the UE are set to 3GHz; the computing resources provided by the BS to any arriving task in each time slot are set to 2.5GHz; set the number K of task types in the MEC system to be 20, and each BS randomly caches task types, and the caching capacity is set to 40%.

[0138] Each UE generates a DAG application in each time slot according to a Bernoulli distribution with a parameter of 0.5. The present invention uses a DAG application synthesizer to generate diverse DAG applications. The parameters of the DAG synthesizer include n, fat, and density, where n controls the number of tasks of the DAG application; fat controls the width and height; density controls the number of edges; the number of tasks in the simulation is default set to n = 8; fat is randomly selected from {0.4, 0.5, 0.6, 0.7, 0.8}; density is randomly selected from {0.4, 0.5, 0.6, 0.7, 0.8}; the data volume size of each task is evenly distributed between [0.1, 0.5] Mbit; the data volume size of the calculation result is evenly distributed between [0.01, 0.05] Mbit; the type of each task is randomly determined, and the number of CPU cycles required per bit for each task type is evenly distributed between [1000, 1200] cycles / bit.

[0139] For all deep reinforcement learning algorithms, set the discount factor γ = 0.99; the learning rate of the actor network is 0.0001; the learning rate of the critic network is set to 0.001; to reduce the time complexity of the algorithm, both the actor network and the critic network use two-layer fully connected neural networks, and the dimension of the hidden layer is set to 128; the size of the experience replay pool is 5×10 5 ; the number of mini-batch samples is X = 256; the update rate is The number of training rounds is 4000, and the number of time steps in each round is 100; the exploration probability decreases from 1 to 0.05 within 1×10 4 time steps.

[0140] The simulation experiment first compared the convergence of different algorithms; then, it presented the impacts of network parameters such as the number of users, the number of neighboring base stations, the storage capacity of base stations, the number of tasks, the computing resources of users, and the application request probability on the performance of different algorithms.

[0141] Figure 5 It presented the convergence processes of different DRL-based algorithms as the number of training rounds increased. It can be seen from the figure that all algorithms finally converged to a stable value as the number of training rounds increased. Compared with other benchmark algorithms, the proposed MADRL-Dep algorithm in the present invention has the optimal system service delay. The no-coop algorithm does not adopt a collaborative computing mechanism, and tasks can only be computed on local devices or LBSs. However, since LBSs usually only cache a small number of popular task types in the system, most tasks are computed on local devices with a high time delay cost, so the system service delay is the highest. The no-E2E algorithm does not consider edge-edge collaborative computing, but utilizes the computing resources of idle UEs to greatly reduce the service delay. Although the no-End2End algorithm does not utilize the computing resources of idle UEs, it can offload tasks to neighboring edge nodes for computing, solving the problem of insufficient cache capacity of a single edge node, and reducing the system service delay to a greater extent compared with the no-E2E algorithm. Although the SADRL-Dep algorithm simultaneously adopts the end-to-end and edge-edge collaborative computing mechanisms, this algorithm is based on the DTDE framework, and each agent / UE conducts distributed independent training, resulting in the agent being unable to effectively identify whether the reward change at each time step is caused by its own actions or the actions of other agents in the MEC system, making it difficult for the agent to learn the globally optimal offloading strategy. In contrast, the MADRL-Dep algorithm adopts the CTDE learning framework, and the agent considers the global state of the MEC system and the offloading actions of other agents during the training process, solving the problems existing in the training process of the SADRL-Dep algorithm. Therefore, the MADRL-Dep algorithm can more effectively coordinate idle devices and edge nodes for multi-node collaborative computing.

[0142] Figure 6 It demonstrated the impacts of the number of users in the MEC system on the performance of different algorithms. In the experiment, the number of users was set to {1, 2, 3, 4, 5, 6, 7, 8}, and other parameters were all default values. From Figure 6It can be seen that as the number of users increases, the system service delay of all algorithms increases accordingly. Moreover, the proposed MADRL-Dep algorithm in the present invention always has the optimal system service delay under different numbers of users. Compared with the no-coop and no-E2E algorithms that do not consider edge-edge cooperation, the MADRL-Dep algorithm always has significant performance advantages. Compared with the no-End2End algorithm that does not consider end-to-end cooperation, as the number of users increases, the MADRL-Dep algorithm gradually demonstrates the advantages of end-to-end collaborative computing. Compared with the SADRL-Dep algorithm, the performance improvement of MADRL-Dep is not obvious when the number of users is small, but as the number of users increases, its performance improvement also increases. This is because the MADRL-Dep algorithm adopts a CTDE framework and considers the local observations and actions of all users during training, so it can effectively cope with the increase in the number of users. The SADRL-Dep algorithm adopts a DTDE framework and is trained only based on local observations. The increase in the number of users / agents will lead to a more unstable training environment, making it impossible to handle collaborative computing offloading when the number of users is large. This shows that the MADRL-Dep algorithm can better coordinate the complementary computing resources among users in a multi-user MEC scenario.

[0143] Figure 7 shows the impact of the number of neighboring base stations in the MEC system on the performance of different algorithms. In the experiment, the number of NBS is set to {0, 1, 2, 3, 4, 5, 6}, and other parameters are default values. Observe Figure 7 It can be found that since the no-E2E and no-coop algorithms do not consider the collaborative computing between edge nodes, the system service delay is not affected by the change in the number of NBS. For the no-End2End, SADRL-Dep, and MADRL-Dep algorithms, as the number of NBS increases, the system service delay decreases accordingly. This is because the increase in the number of NBS makes the types of tasks cached in the MEC system increase, and the probability that the task types requested by UEs are executed at edge nodes with lower computing delays increases, thus reducing the system service delay.

[0144] In addition, as Figure 7As shown, the performance gap between the algorithms without edge-edge collaborative computing (no-E2E and no-coop algorithms) and the algorithms with edge-edge collaborative computing (no-End2End, SADRL-Dep, and MADRL-Dep algorithms) increases with the increase in the number of NBSs. Specifically, the more NBSs that can collaborate, the more significant the advantage of the algorithm system service delay based on edge-edge collaborative computing. This is because multi-edge node collaborative computing can effectively make up for the problem of incomplete cached task types in a single base station, which further verifies the necessity and effectiveness of edge-edge collaborative computing. Compared with the no-End2End algorithm, the MADRL-Dep algorithm utilizes the computing resources of nearby idle UEs. Compared with the SADRL-Dep algorithm, the MADRL-Dep algorithm efficiently coordinates the multi-node collaborative computing among idle UEs and base stations. Therefore, it has the optimal system service delay under different numbers of NBSs.

[0145] Figure 8 shows the influence of the base station storage capacity on the performance of different algorithms. In the experiment, each base station randomly caches task types with the same probability, and the storage capacity is set to {0%, 20%, 40%, 60%, 80%, 100%}, and other parameters are default values. From Figure 8 it can be observed that the system service delay of all algorithms decreases with the increase in the storage capacity of each base station. This is because the increase in the base station cache capacity enables each base station to cache more task types, so the probability that the task types requested by UEs are executed at edge nodes with lower computing delay increases, thereby reducing the system service delay. When the base station cache capacity is 100%, all task types can be calculated at the local base station with lower computing delay, and at this time, the system service delay differences of each algorithm are smaller. From Figure 8 it can be seen that the MADRL-Dep algorithm always shows the optimal system service delay under different base station storage capacities. Compared with other benchmark algorithms, especially when the base station storage capacity is small, it has obvious performance advantages. This indicates that the MADRL-Dep algorithm can not only efficiently collaborate the computing resources of idle UEs in the scenario of limited MEC storage resources, but also utilize the task types cached by neighbor edge nodes to make up for problems such as incomplete cached task types in a single base station.

[0146] Figure 9 gives the influence of the number of tasks on the performance of different algorithms. In the experiment, the number of tasks for each application is set to {4, 5, 6, 7, 8, 9, 10}, and other parameters are default values. From Figure 9It can be observed that as the number of tasks increases, the system service delay of all algorithms shows an increasing trend. This is because the completion time of an application depends on the completion time of the exiting task. When the number of tasks that make up the application increases, it means that the number of predecessor tasks of the exiting task increases. Due to the execution dependence between tasks, the exiting task can only be executed after the completion of its predecessor tasks. Therefore, the increase in the number of predecessor tasks leads to an increase in the computing delay, which in turn leads to an increase in the completion time of the exiting task, that is, an increase in the application computing delay. In addition, as Figure 9 shown, compared with all benchmark algorithms, the MADRL-Dep algorithm proposed by the present invention exhibits the optimal system service delay in scenarios with different numbers of tasks, further verifying that the algorithm can efficiently coordinate multi-node collaborative computing between terminal devices and edge nodes when dealing with dependent tasks.

[0147] Figure 10 shows the impact of user computing resources on the performance of different algorithms. In the experiment, the total computing resources of each user are set to {2, 3, 4, 5, 6, 7, 8} GHz, and other parameters are default values. From Figure 10 it can be seen that the system service delay of all algorithms decreases as the user computing resources increase. This is because the present invention considers that any UE evenly allocates computing resources to the arriving tasks. When the total computing resources of the user increase, the computing resources allocated to each task increase accordingly, thus reducing the computing delay of the tasks on the UE. From Figure 10 it can also be observed that when the UE computing resources are relatively large (7 - 8 GHz), the system service delay of the no-E2E algorithm begins to be better than that of the no-End2End algorithm. This is because the no-E2E algorithm adopts end-to-end collaborative computing. When the computing resources of the terminal device are sufficient, the collaborative computing delay with the terminal device is lower than the collaborative computing delay with the edge node. While the no-End2End algorithm does not consider end-to-end collaborative computing, so its performance is inferior to that of the no-E2E algorithm. This further highlights the advantage of adopting the end-to-end collaborative computing mechanism. The MADRL-Dep algorithm proposed by the present invention utilizes both end-to-end collaborative and edge-to-edge collaborative computing mechanisms, so it has the optimal system service delay.

[0148] Figure 11 shows the impact of application request probability on the performance of different algorithms. In the experiment, each UE generates a DAG application in each time slot according to a Bernoulli distribution with parameters from 0.1 to 1, and other parameters are default values. From Figure 11 it can be observed that the system service delay of all algorithms increases as the application request probability increases. The MADRL-Dep algorithm proposed by the present invention has the optimal system service delay under different application request probabilities. In addition, from Figure 11It can be found that when the application request probability is relatively small (0.1, 0.2, 0.3) and relatively large (0.8, 0.9, 1), the performance gap between the MADRL-Dep and SADRL-Dep algorithms is relatively small. This is because when the application request probability is small, most of the UEs in the system are in the idle state and the total number of tasks is small. At this time, both the MADRL-Dep and SADRL-Dep algorithms can effectively coordinate the idle UEs for collaborative computing. As the application request probability increases, the number of idle UEs in the system decreases and the total number of tasks increases, making the system environment more complex. At this time, the SADRL-Dep algorithm based on the DTDE framework cannot handle the more complex MEC environment, so its performance is inferior to that of the MADRL-Dep algorithm. As the application request probability continues to increase, the number of idle UEs available for collaborative computing decreases, and there may even be no idle UEs. At this time, most of the tasks in the MADRL-Dep and SADRL-Dep algorithms can only be executed on local devices or edge nodes and cannot utilize and coordinate the computing resources of idle UEs for collaborative computing, so the performance gap gradually decreases.

[0149] In the present invention, unless otherwise clearly defined and limited, terms such as "installation", "setting", "connection", "fixation", "rotation" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components or the interaction relationship between two components. Unless otherwise clearly defined, for those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0150] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent edge collaborative computing offloading method for dependency-based tasks, characterized in that, It includes the following steps: S1. Construct a system model including multiple users and multiple base stations, where: There is an MEC server equipped at each base station; denote the set of base stations in the system model as where LBS represents a local base station, denotes the set of neighbor base stations of LBS, NBS n denotes the n = 1, 2, …, Nth neighbor base station of LBS, and N represents the number of neighbor base stations; M users are randomly distributed within the coverage area of LBS, and let the set of users be In the system model, time is discretized into a set of time slots with equal lengths ξ m (t) ∈ {0, 1} represents the application request status of user at time slot . If user m does not generate an application at time slot t, then ξ m (t) = 0, and if user m generates an application at time slot t, then ξ m (t) = 1; S2. Based on the system model, construct an application model, a computing offloading model, a task scheduling model, a communication model, a local device computing model, a local base station computing model, an end-to-end collaborative computing model, and an edge-to-edge collaborative computing model; S3. With the goal of minimizing the long-term time-average service delay of the system, construct an optimization problem model; S4. Based on steps S1 - S3, construct a deep reinforcement learning framework to obtain the environmental state, local observation, action, and reward function; S5. Use the MADRL-Dep algorithm to solve the optimization problem to obtain the offloading strategy.

2. The intelligent edge collaborative computing offloading method for dependency-based tasks according to claim 1, wherein The application model uses a DAG to describe an application with task dependencies: Define the set of task types in the system model Denote the number of task types; application G requested by user m at time slot t m (t) = {V m (t), E m (t)}, Denote the set of application tasks, and Denote the virtual start task and the virtual end task; |V m | + 1 denotes the total number of tasks, E m (t) represents the execution dependency between tasks, Denote the i-th task requested by user m at time slot t, Denote the task data volume, Denote the task Number of CPU cycles required per bit of data, Denote the task type; for the virtual start task and the virtual end task the data volume is 0; adopt Denote the task and the task the execution dependency between them, where Denote the i-th task is the predecessor task of the j-th task , δ i,j (t) represents the size of the calculation result from task to task ; The computing offloading model is used to model the offloading decision variables of users in each time slot: Among them, represents the set of users who can provide computing services for task in time slot t, represents the set of offloadable edge nodes for task in time slot t, Φ x (t) represents the type of cached task of base station x in time slot t; represents the computing offloading decision variable of task in time slot t. If task is offloaded to neighbor base station NBS n for computing, then If the task is offloaded to other user m' for computing, then If task is computed on the local device, then If task is offloaded to the local base station for computing, then If user m does not request an application, then The virtual start task and the virtual end task can only be computed on the local device. Therefore In summary, the offloading decision variable of user m in time slot t The task scheduling model is used to model the task scheduling order and process of DAG tasks: Definition Denote G m (t)'s task scheduling order, Denote the task Scheduling priority of.

3. An intelligent edge collaborative computing offloading method for dependency-based tasks according to claim 1, characterized in that, The communication model is used to model the transmission rate of the system: The uplink transmission rate from user m to the local base station at time slot t is denoted as Among them, B up represents the uplink bandwidth, P m represents the transmit power of user m, G m,LBS (t) represents the channel gain between user m and the local base station, σ 2 represents the background noise power; Downlink transmission rate from the local base station to user m in time slot t Denoted as Among them, B down represents the downlink bandwidth, P LBS represents the transmission power of the local base station, G LBS,m (t) represents the channel gain between the local base station and user m; The transmission rate from user m to user m' at time slot t is denoted as Among them, B D2D represents the D2D bandwidth, and G m,m’ represents the channel gain between user m and user m'.

4. An intelligent edge collaborative computing offloading method for dependency-based tasks according to claim 1, characterized in that The local device computing model is used to model the computing process of tasks on local devices: Define the task The end time when calculating on the local device corresponding to user m Denoted as Indicates the task The start time of execution on the local device, Indicates the task The computing delay of execution on the local device, and the specific calculation formula is Among them, represents the end time when the predecessor task of the task is executed on the local device; represents the end time when the predecessor task of the task is executed at ; represents the set of users who can provide computing services for the task in time slot t, represents the set of offloadable edge nodes for the task in time slot t; represents the data transmission delay for sending the computing result of the predecessor task to the local device when computing on the local device ; represents the set of predecessor tasks of the task ; 1 {·} is an indicator function. If the event {·} is true, then 1 {·} = 1, otherwise 1 {·} = 0; represents the data volume of the task , represents the number of CPU cycles required per bit of data for the task , f m (t) represents the computing resource allocated by user m to the task in time slot t; F m represents the total computing resource of user m; represents the available communication resource time when the device of collaborative user m' sends the computing result in time slot t, represents the transmission rate from user m to user m' in time slot t; represents the downlink transmission rate from the local base station to user m in time slot t; represents the average transmission rate between neighbor base station n and the local base station; represents the set of computing results sent before the computing result ; represents the size of the computing result of the task to the task on user m; The local base station computing model is used to model the computing process of tasks on local base stations: Define task End time during local base station calculation corresponding to user m Denoted as Indicates the task The start time executed at the local base station, Indicates the task The computing delay executed on the local device, and the specific calculation formula is Among them, represents the predecessor task of the end time when executed at the local base station, represents the data transmission delay when the predecessor task is calculated at the local base station and sent to the local base station; represents the upward transmission delay when the task is calculated at the local base station, represents the available communication resource time for user m to send the task in time slot t; represents the task set scheduled before the task m (t) obtained based on the task scheduling order Q ; represents the data volume of the task, represents the uplink transmission rate of user m to the local base station in time slot t, represents the transmission rate from user m to user m' in time slot t; represents the predecessor task of the task at the end time when executed at ; f LBS represents the CPU clock frequency of the local base station.

5. The intelligent edge collaborative computing offloading method for dependent tasks according to claim 1, wherein The end-to-end collaborative computing model is used to model the computing process of tasks selecting end-to-end collaboration; Define task End time during calculation on the device corresponding to collaborative user m Denoted as Represents a task At the start time of execution on the device of collaborative user m Represents a task The computing delay of execution on the device of collaborative user m, and the specific calculation formula is Among them, represents the predecessor task of the end time when executed on the device of collaborative user m; represents the data transmission delay when the predecessor task is calculated on the device of collaborative user m and sent to the device of collaborative user m; represents the predecessor task of the end time when executed at ; represents the transmission delay when executed on the device of collaborative user m, represents the available communication resource time when user m sends task in time slot t; represents the transmission rate from user m to user m' in time slot t; represents the available communication resource time when sending task to task in time slot t; represents the transmission rate from to user m' in time slot t; represents the downlink transmission rate from the local base station to collaborative user m' in time slot t; represents the average transmission rate between neighbor base station n and the local base station; f m’ represents the computing resource allocated to task by collaborative user m' in time slot t; The edge-to-edge collaborative computing model is used to model the computing process of tasks selecting edge-to-edge collaboration: Define task End time during neighbor base station calculation Expressed as Indicates the task The start time of execution at neighbor base station n Indicates the task The computing delay of execution at neighbor base station n, and the specific calculation formula is Among them, represents the end time when the predecessor task of the task is executed at the neighbor base station n; represents the data transmission delay for sending the calculation result of the predecessor task to the neighbor base station n when the task is calculated at the neighbor base station n; represents the end time when the predecessor task of the task is executed at ; represents the uplink transmission delay when the task is calculated at the neighbor base station n; represents the available communication resource time for user m to send the task in time slot t; represents the computing resource allocated by the neighbor base station n to the task in time slot t.

6. The intelligent edge collaborative computing offloading method for dependency-based tasks according to claim 1, wherein The optimization problem model is as follows: In the formula, represents the task scheduling order decision; represents the task offloading decision; represents the total number of time slots, T m (t) represents the service delay of user m at time slot t, represents the task the set of users who can provide computing services for the task at time slot t, represents the task the set of edge nodes where the task can be offloaded at time slot t, represents the task scheduling priority; |V m (t)| + 1 represents the total number of tasks of user m at time slot t; represents the task computing offloading decision variable at time slot t.

7. The intelligent edge collaborative computing offloading method for dependent tasks according to claim 1, characterized in that Define the environmental state of time slot t Define the local observation of time slot t Define the action of agent m at time slot t Define the reward function of agent m at time slot t Among them, G m (t) represents the application requested by user m at time slot t, and f m (t) represents the computing resources allocated to task by user m at time slot t, and Φ x (t) represents the cache task type of base station x at time slot t; represents the computing offloading decision variable of task at time slot t; T m (t) represents the service delay of user m at time slot t.

8. An intelligent edge collaborative computing offloading method for dependency-based tasks according to claim 1, characterized in that In the MADRL-Dep algorithm framework, each agent is equipped with a DDPG network and a task scheduling module; The training process of the MADRL-Dep algorithm includes: S51. Initialize the parameters θ of the main actor network μ , the parameters θ of the target actor network μ′ , the parameters θ of the main critic network Q , the parameters θ of the target critic network Q′ and the experience replay pool S52. In each episode, after obtaining the local observation o, perform time step iteration, including: S521. Each agent selects an action according to the current policy, and then inputs the action into the task scheduling module corresponding to the agent to obtain the application scheduling order; S522. The environment returns a reward based on the action and the task scheduling order, and obtains the next local observation; S523. Put the local observation, the next local observation, the action, and the reward into the experience replay pool, and then perform the following operations for each agent respectively: Sample a small batch of sample data of size X from the experience replay pool; Update the parameters of the main critic network; Update the parameters of the main actor network; Update the parameters of the target actor network and the target critic network in a soft update manner.