Task unloading method based on graph neural network and deep reinforcement learning

By adopting the task unloading method of graph neural network and deep reinforcement learning in the cloud edge collaborative environment, the challenges of environmental dynamics and hybrid action space in the task unloading problem are solved, and efficient and accurate task unloading effect is achieved.

CN120128985APending Publication Date: 2025-06-10FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510267631.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the cloud-edge collaborative environment, the task unloading problem faces environmental dynamic challenges. Traditional methods are difficult to effectively deal with the relationship between hybrid action space and network topology, resulting in increased computing complexity and increased exploration difficulty.

Method used

Using a task offload method based on graph neural network and deep reinforcement learning, a cloud-edge collaborative system of multi-terminal devices, multi-edge servers and single-cloud servers is modeled, a three-layer task offload model is built, and optimization problems are transformed into Markov decision-making process. Graph neural network is used to capture complex environment information to assist in deep reinforcement learning.

Benefits of technology

It effectively solves the problem of task offloading under multi-terminal devices, multi-edge servers and single-cloud servers, improves the efficiency and accuracy of task offloading, and reduces computing complexity and energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128985A_ABST
    Figure CN120128985A_ABST
Patent Text Reader

Abstract

The invention relates to a task unloading method based on a graph neural network and deep reinforcement learning, and belongs to the technical field of Internet of Things and artificial intelligence. The method comprises the following steps: firstly, modeling a cloud side-end cooperative system of multiple terminal devices, multiple edge servers and a single cloud server, and constructing a three-layer task unloading model of the cloud server, the edge servers and the terminal devices; then, defining an optimization problem which aims at minimizing task time delay and equipment energy consumption weighted sum, and further converting the optimization problem into a Markov decision process; and finally, a task unloading algorithm based on the graph neural network and deep reinforcement learning is provided, and the algorithm captures complex environment information by using the graph neural network, so that more accurate input is provided for deep reinforcement learning. Experimental results show that the task unloading problem under multiple terminal devices, multiple edge servers and a single cloud server can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of Internet of Things and artificial intelligence, and particularly relates to a task offloading method based on graph neural network and deep reinforcement learning. Background Art

[0002] With the rapid development of wireless communication technology and semiconductor technology, terminal devices such as tablet computers, smart phones, and Internet of Things devices have been widely popularized. At the same time, more and more new applications designed for terminal devices have emerged, such as face recognition, virtual reality (VR), natural language processing, etc. These applications usually have the characteristics of being computationally intensive and latency-sensitive, while the computing power and battery capacity of terminal devices are often limited and difficult to meet the requirements of such tasks.

[0003] In recent years, the rise of cloud computing technology has provided a new solution for terminal devices. By offloading computationally intensive tasks to the cloud for execution, the problem of limited device resources has been effectively alleviated. However, the processing mode centered on cloud computing has limitations such as high network latency and weak privacy protection ability, and it is difficult to meet the requirements of a large number of terminal devices, large-scale data, and diverse services. To solve these problems, edge computing (EC) has emerged. Edge computing significantly reduces network latency by deploying edge servers near users and data sources. However, the computing power of edge servers is limited and it is difficult to serve a large number of devices simultaneously. For this reason, the cloud-edge-end collaborative computing mode has been proposed. This mode combines the powerful computing power of cloud servers and the advantage of edge servers being close to users, and can provide more efficient services for users.

[0004] However, in the cloud-edge-end collaborative environment, the task offloading problem faces many challenges. One of the main challenges is the dynamicity of the environment, such as changes in channel conditions, dynamic generation of device tasks, and fluctuations in the number of devices under the same base station. To cope with these dynamic changes, many scholars have proposed task offloading algorithms based on reinforcement learning. Reinforcement learning can adaptively adjust strategies according to changes in the environment, so as to achieve efficient task offloading. In addition, the reinforcement learning algorithm can make decisions quickly after training, meeting the real-time requirements of the system. However, the variables in the task offloading problem often include both discrete values (such as offloading decisions) and continuous values (such as resource allocation). Traditional solutions usually convert all variables into the same type, but this will increase the computational complexity and exploration difficulty. Therefore, how to effectively handle the problem of mixed action space has become a key challenge in current research.

[0005] In a cloud-edge-terminal collaborative network, devices within the coverage area of the same base station need to compete for the computing and communication resources of the edge server. Changes in the number of devices will directly affect the intensity of resource competition. In addition, the cloud-edge-terminal collaborative network can be modeled as a topological structure. However, most of the existing task offloading methods based on deep reinforcement learning use multi-layer fully connected neural networks to extract the features of terminal devices. Since fully connected neural networks cannot effectively learn network topological relationships, it is difficult to capture the impact of changes in the number of devices on the intensity of resource competition. Therefore, how to accurately model the dynamic relationships among terminal devices, edge servers, cloud servers, and other terminal devices has become a key problem to be solved urgently. Summary of the Invention

[0006] The purpose of the present invention is to provide a task offloading method based on graph neural networks and deep reinforcement learning, which can effectively solve the task offloading problem under multiple terminal devices, multiple edge servers, and a single cloud server.

[0007] To achieve the above object, the technical solution of the present invention is: A task offloading method based on graph neural networks and deep reinforcement learning, including:

[0008] Step S1, model the cloud-edge-terminal collaborative system of multiple terminal devices, multiple edge servers, and a single cloud server;

[0009] Step S2, construct a three-layer task offloading model for the cloud server, edge server, and terminal device;

[0010] Step S3, define an optimization problem with the goal of minimizing the weighted sum of task latency and device energy consumption;

[0011] Step S4, transform the optimization problem defined in Step S3 into a Markov decision process;

[0012] Step S5, construct a task offloading method based on graph neural networks and deep reinforcement learning to solve the task offloading problem under multiple terminal devices, multiple edge servers, and a single cloud server.

[0013] In an embodiment of the present invention, in Step S1, the process of constructing the cloud-edge-terminal collaborative system of multiple terminal devices, multiple edge servers, and a single cloud server is as follows:

[0014] S11. System time slot

[0015] Assume that the system time is discretized into T time slots of equal duration, and use the set to represent;

[0016] S12. Terminal device

[0017] Assume that there are D terminal devices, and use the set to represent, and each terminal device The maximum computing power is denoted as f d , and the idle power is denoted as In each time slot t, the task of the terminal device d is described as a triple, namely denotes the task data size; denotes the CPU cycles required per bit of task data; denotes the maximum tolerable delay of the task, denotes the offloading indicator. When , it means the task is processed locally; when , it means the task is offloaded to the edge server for processing; when , it means the task is offloaded to the cloud server for processing;

[0018] S13. Edge server

[0019] Assume there are M base stations, and an edge server is deployed near each base station; the base stations are connected to the edge servers through point-to-point optical fibers; since the relationship between the base stations and the edge servers is one-to-one, the set is used to uniformly represent the base station-edge server pairs; the edge server has a computing power defined as f m ;

[0020] S14. Cloud server

[0021] Assume there is a central cloud server, and the computing power of the cloud server is defined as f c .

[0022] In an embodiment of the present invention, step S2 includes the following steps:

[0023] S21. Establish a local computing model:

[0024] Assume that each terminal device has different computing powers, denoted by ; denotes the task size generated by the terminal device d in the t time slot; denotes the CPU cycles required per bit of task data; thus, the time for local computing is defined as:

[0025]

[0026] The energy consumption for local computing is defined as:

[0027]

[0028] k is an energy consumption calculation coefficient related to the chip architecture;

[0029] S22. Establish an edge computing model

[0030] Assume that the total computing power on the edge server is f m ; Considering that the edge server needs to provide services to multiple terminal devices simultaneously, the computing power allocated by server m to terminal d at time t is expressed as The time of edge computing is defined as:

[0031]

[0032] Since the edge server has sufficient power supply, the energy consumption of the edge server when executing tasks is not considered; when the task is executed on the edge server, the idle energy consumption of terminal device d is defined as:

[0033]

[0034] Where is the idle power of device d;

[0035] When the task is offloaded to the edge server, it needs to be transmitted to the base station through the wireless channel, and its transmission rate is defined as:

[0036]

[0037] Where represents the bandwidth between terminal device d and base station m at time slot t, represents the transmission power of terminal device d at time slot t, N 0 represents the power spectral density of noise; represents the channel gain between terminal device d and base station m at time slot t;

[0038] The time for terminal device d to transmit to base station m is defined as:

[0039]

[0040] The energy consumption for terminal device d to transmit to the base station is defined as:

[0041]

[0042] S23. Establish a cloud computing model

[0043] Similar to the edge server, the time for the task to be executed on the cloud server is defined as:

[0044]

[0045] When the task is executed on the cloud server, terminal device d is in an idle state and still needs to consume energy. Therefore, the idle energy consumption of terminal device d is defined as:

[0046]

[0047] The cloud server is connected to the base station through a wide area network. Therefore, when the terminal device communicates with the cloud server, data needs to be first transmitted from the terminal device to the base station and then forwarded by the base station to the cloud server. The time for the terminal device to transmit to the cloud server is defined as:

[0048]

[0049] where is the time for the terminal device to transmit to the base station, is the rate for the base station m to allocate the transmission of the terminal device d to the cloud server at time slot t;

[0050] The energy consumption of the terminal device d to transmit to the cloud server is defined as:

[0051]

[0052] In an embodiment of the present invention, step S3 includes the following steps:

[0053] S31. Define the total processing delay of the task

[0054] The task The total processing delay is defined as:

[0055]

[0056] S32. Define the energy consumption of the terminal device

[0057] According to the above, the energy consumption of the terminal device is defined as:

[0058]

[0059] S33. Define the optimization problem

[0060] The optimization problem is defined as:

[0061]

[0062] where η is a weighting parameter used to trade off between time and energy.

[0063] In an embodiment of the present invention, step S4 includes the following steps:

[0064] S41. Define the system state, including the cloud server state, edge server state, terminal device state, and network state;

[0065] S42. Define the actions of the terminal device, including the offloading location, transmission power, and local computing ability;

[0066] S43. Define a reward function that includes latency, energy consumption, and task completion rate.

[0067] In an embodiment of the present invention, step S4 is specifically implemented as follows:

[0068] S41. The system state consists of the terminal device state, the edge server state, the cloud server state, and the network topology of the cloud-edge-end system; the terminal device state is defined as: The edge server state is defined as: The state of the cloud server is defined as: Assume that in a training, the topology structure of the cloud-edge-end collaborative network remains unchanged; the cloud-edge-end collaborative network is modeled as an undirected graph structure, denoted as G=(V, E), where V is the set of nodes, including three types of nodes: terminal devices, edge servers, and cloud servers; E is the set of edges, representing the connection relationships between these nodes; therefore, the entire system state is defined as:

[0069]

[0070] S42. The actions of terminal device d consist of the following parts: the offloading location of the task The transmission power of the device The computing resources allocated by the device to the task For simplicity of representation, the continuous actions of device d are defined as Therefore, at time slot t, the action of device d is defined as Furthermore, the set of actions of all devices at time slot t is defined as

[0071] S43. The reward function consists of three parts, aiming to comprehensively consider the task processing latency, the energy consumption of the terminal device, and the satisfaction of the task latency constraint; the first part is the task processing latency; the second part is the energy consumption of the terminal device; the third part is the penalty for violating the maximum tolerable latency; the penalty term is defined as:

[0072]

[0073] In summary, at time slot t, the reward function of device d is defined as:

[0074]

[0075] At time slot t, the rewards of all devices are defined as

[0076]

[0077] In an embodiment of the present invention, step S5 includes the following steps:

[0078] S51. Design the structure of the graph neural network and process the input and output of the network;

[0079] S52. Construct the actor network and explain the generation process of continuous actions;

[0080] S53. Construct the critic network and explain the generation process of state-action value pairs and discrete actions;

[0081] S55. Design the loss functions of the actor network and the critic network.

[0082] In an embodiment of the present invention, the structure of the graph neural network adopts a two-layer graph attention network.

[0083] The present invention also provides a task offloading system based on a graph neural network and deep reinforcement learning, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the method steps as described above can be implemented.

[0084] The present invention also provides a computer-readable storage medium, on which computer program instructions capable of being run by the processor are stored. When the processor runs the computer program instructions, the method steps as described in any of the above can be implemented.

[0085] Compared with the prior art, the present invention has the following beneficial effects: In the method of the present invention, first, a cloud-edge-end collaborative system of multiple terminal devices, multiple edge servers, and a single cloud server is modeled, and a three-layer task offloading model of the cloud server, edge server, and terminal device is constructed; then, an optimization problem with the goal of minimizing the weighted sum of task latency and device energy consumption is defined, and the optimization problem is further transformed into a Markov decision process; finally, a task offloading algorithm based on a graph neural network and deep reinforcement learning is proposed. The algorithm uses the graph neural network to capture complex environmental information, thereby providing more accurate input for deep reinforcement learning. Experimental results show that the present invention can effectively solve the task offloading problem under multiple terminal devices, multiple edge servers, and a single cloud server. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 is a flowchart of the present invention.

[0087] Figure 2 is an architecture diagram of the cloud-edge-end collaborative system of the present invention.

[0088] Figure 3 is a system structure diagram of the present invention based on reinforcement learning. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0089] The technical solutions of the present invention will be specifically described below with reference to the accompanying drawings.

[0090] The present invention provides a task offloading method based on graph neural network and deep reinforcement learning, including:

[0091] Step S1, model the cloud-edge-end collaborative system of multiple terminal devices, multiple edge servers and a single cloud server;

[0092] Step S2, construct a three-layer task offloading model for the cloud server, edge server and terminal device;

[0093] Step S3, define an optimization problem aiming to minimize the weighted sum of task latency and device energy consumption;

[0094] Step S4, transform the optimization problem defined in Step S3 into a Markov decision process;

[0095] Step S5, construct a task offloading method based on graph neural network and deep reinforcement learning to solve the task offloading problem under multiple terminal devices, multiple edge servers and a single cloud server.

[0096] The following is the specific implementation process of the present invention.

[0097] As Figures 1-3 shown, this embodiment provides a task offloading method based on graph neural network and deep reinforcement learning, and the detailed steps are as follows:

[0098] Step 1: Model the cloud-edge-end collaborative system of multiple terminal devices, multiple edge servers and a single cloud server

[0099] Step 11: Assume that the system time is discretized into T time slots with equal durations, and use the set to represent.

[0100] Step 12: Assume that there are D terminal devices, and use the set to represent. The maximum computing power of each terminal device is represented as f d , and the idle power is represented as At each time slot t, the task of terminal device d can be described as a triple, that is, represents the task data size (unit: bit); represents the CPU cycles required for each bit of task data; represents the maximum tolerable latency of the task. The offloading location set is defined as represents the offloading indicator, and when , it means that the task is processed locally; when , it means that the task is offloaded to the edge server for processing; when It indicates that the task is offloaded to the cloud server for processing.

[0101] Step 13: Assume that there are M base stations, and an edge server is deployed near each base station. Since the relationship between the base station and the edge server is one-to-one, we use the set to represent the base station-edge server pair uniformly.

[0102] Edge server 's computing power is represented by f m . Each terminal device can only communicate with one base station, and the communication between the two is through a wireless channel, so it will bring a relatively high transmission delay. In contrast, the communication between the base station and the edge server is through optical fiber, and its transmission delay can be ignored compared with the wireless transmission delay.

[0103] Step 14: Assume that there is a central cloud server, which is connected to the base station through a wide area network. Assume that the network transmission rate between each base station and the cloud server is the same, and use R m,c to represent. The terminal device communicates with the cloud server through the base station, so the terminal devices under the same base station share the network bandwidth with the cloud server. The computing power of the cloud server is represented by f c .

[0104] Step 2: In this embodiment, a three-layer task offloading model of the cloud server, edge server and terminal device is constructed, which specifically includes the following steps:

[0105] Step 21: Establish a local computing model. The time and terminal energy consumption of local computing are respectively defined as:

[0106]

[0107] where k is the energy consumption calculation coefficient related to the chip architecture.

[0108] Step 22: Establish an edge computing model. Considering that the edge server needs to provide services for multiple terminal devices at the same time, we represent the computing power allocated by server m to terminal d at time t as The time of edge computing is defined as:

[0109]

[0110] Considering and the larger they are, the more time is required to process the task, and the smaller it is, the more sensitive the task is to delay. Therefore the calculation formula of

[0111]

[0112] where is an indicator function whose value depends on When the value of the indicator function is 1, and when the value of the indicator function is 0. link(m) refers to the set of terminal devices connected to base station m.

[0113] Since the edge server has sufficient power supply, the energy consumption of the edge server for executing tasks is not considered in this paper. When the task is executed on the edge server, the idle energy consumption of terminal device d is defined as:

[0114]

[0115] Offloading the task to the edge server requires transmission to the base station through the wireless channel, and its transmission rate is defined as:

[0116]

[0117] where represents the bandwidth between terminal device d and base station m at time slot t, represents the transmission power of terminal device d at time slot t, N 0 represents the power spectral density of the noise. represents the channel gain between terminal device d and base station m at time slot t, and its calculation formula is defined as:

[0118]

[0119] The time for terminal device d to transmit to base station m is defined as:

[0120]

[0121] The energy consumption for terminal device d to transmit to the base station is defined as:

[0122]

[0123] Step 23: Establish a cloud computing model

[0124] Similar to the edge server, the time for the task to be executed on the cloud server is defined as:

[0125]

[0126] When the task is executed on the cloud server, terminal device d is in an idle state and still needs to consume energy. Therefore, the idle energy consumption of terminal device d is defined as:

[0127]

[0128] The cloud server is connected to the base station via a wide area network. Therefore, when the terminal device communicates with the cloud server, data needs to be first transmitted from the terminal device to the base station and then forwarded by the base station to the cloud server. The time for the terminal device to transmit to the cloud server is defined as:

[0129]

[0130] where is the time for the terminal device to transmit to the base station, is the rate at which the base station m allocates the transmission of the terminal device d to the cloud server in time slot t.

[0131] The energy consumption of the terminal device d transmitting to the cloud server is defined as:

[0132]

[0133] Step 3: An optimization problem is defined with the goal of minimizing the weighted sum of task delay and device energy consumption, which specifically includes the following steps:

[0134] Step 31: Define the total task processing delay. There are three cases for the total task processing delay. When the task is processed locally, it only includes the task computing delay. When the task is executed on the edge server, it includes the task computing delay and the task transmission delay. When the task is executed on the cloud server, it includes the task computing delay and the task transmission delay. Therefore, the task processing delay is defined as:

[0135]

[0136] Step 32: Terminal device energy consumption. There are three cases for the terminal device energy consumption. When the task is processed locally, it includes the computing energy consumption. When the task is executed on the edge server, it includes the transmission energy consumption and the idle device energy consumption. When the task is executed on the cloud server, it includes the transmission energy consumption and the idle device energy consumption. Therefore, the device energy consumption is defined as:

[0137]

[0138] Step 33: According to Step 31 and Step 32, the optimization problem is defined as:

[0139]

[0140] where η is the weighting parameter, which is used to trade off between time and energy.

[0141] Step 4 approximates the optimization problem as a Markov decision process (MDP), and the MDP includes states, actions, and rewards. The specific steps are as follows:

[0142] Step 41: The system state consists of the terminal device state, the edge server state, the cloud server state, and the network topology of the cloud-edge-end system. The terminal device state is defined as: The edge server state is defined as: The state of the cloud server is defined as: Assume that the topology of the cloud-edge collaborative network remains unchanged during a training. The cloud-edge collaborative network can be modeled as an undirected graph structure, denoted as G=(V, E), where V is the set of nodes, including three types of nodes: terminal devices, edge servers, and cloud servers. E is the set of edges, representing the connection relationships between these nodes. Therefore, the entire system state is defined as:

[0143]

[0144] Step 42: The actions of terminal device d consist of the following parts: the offloading location of the task The transmission power of the device The computing resources allocated by the device to the task To simplify the representation, the continuous action of device d is defined as Therefore, at time slot t, the action of device d is defined as Furthermore, we define the set of actions of all devices at time slot t as

[0145] Step 43: The reward function consists of three parts, aiming to comprehensively consider the task processing delay, the energy consumption of the terminal device, and the satisfaction of the task delay constraint. The first part is the task processing delay; the second part is the energy consumption of the terminal device; the third part is the penalty for violating the maximum tolerable delay. The penalty term is defined as:

[0146]

[0147] In summary, at time slot t, the reward function of device d is defined as:

[0148]

[0149] At time slot t, the rewards of all devices are defined as

[0150] Step 5: A task offloading method based on graph neural network and deep reinforcement learning, the specific steps are as follows:

[0151] Step 51: The graph neural network adopts a two-layer graph attention network. Since the state dimensions of the terminal device, the edge server, and the cloud server are different, directly inputting the features into the graph neural network will cause the problem of dimensional mismatch. To solve this problem, in this paper, zero-padding operations are performed on the features of the edge server and the cloud server to make their feature dimensions the same as those of the terminal device. After the zero-padding operation, the features of the terminal device, the edge server, and the cloud server are represented by The process of the graph attention network is as follows:

[0152] The first step is to calculate the attention coefficients between nodes. In the l-th layer of the graph attention network, the calculation formula for the attention coefficient between node i and its adjacent node j is defined as:

[0153]

[0154] where N(i) represents the set composed of all nodes adjacent to node i, and "||" represents the concatenation operation. a(·) maps the concatenated vector to a real number. represents the feature of node i in the (l - 1)-th layer. w l represents the weight matrix multiplied by the feature and is a trainable parameter.

[0155] The second step is neighbor aggregation, through which a node can learn the information of its neighbor nodes. The feature after node aggregation is calculated as follows:

[0156]

[0157] where σ(·) represents the activation function, is the attention coefficient between node i and its adjacent node j. is the feature of the (l - 1)-th layer.

[0158] Step 52: The continuous actions are generated using the actor network which consists of a GAT neural network and a Gaussian policy network. represents all the trainable parameters of the actor network. The continuous actions are generated using the Gaussian policy network. First, a multi-layer perceptron is used to approximate the mean and standard deviation of the distribution of the continuous actions, which are defined as:

[0159]

[0160] Through and a continuous action distribution for the terminal device d is constructed, and a continuous action

[0161] Step 53: Similar to the actor network, the critic network also needs to focus on the relationship between the terminal device and the remaining devices and servers. However, since the key focuses of the two networks may be different, independent graph attention networks are used respectively to learn the relationships between devices. In the critic network, the relationship between devices after learning by the graph neural network is defined as

[0162] The state-action value function of the terminal device d is defined as:

[0163]

[0164] where is the state value function, and is the advantage function.

[0165] The discrete action is defined as:

[0166]

[0167] Step 54: Definition of the loss function. The higher the continuous action value output by the actor network, the better its performance. Therefore, the loss function of the actor network is defined as:

[0168]

[0169] To reduce the problem of value overestimation, this paper introduces a target critic network θ'. We use the target critic network and the reward at time slot t to construct the target value, which is defined as:

[0170]

[0171] where γ is the discount factor.

[0172] Using the target value and the value evaluated by the critic network to construct the mean squared error loss function, which is defined as:

[0173]

[0174] The present invention also provides a task offloading system based on a graph neural network and deep reinforcement learning, including a memory, a processor, and computer program instructions stored on the memory and capable of being run by the processor. When the processor runs the computer program instructions, the method steps as described above can be implemented.

[0175] The present invention also provides a computer-readable storage medium, on which computer program instructions capable of being run by the processor are stored. When the processor runs the computer program instructions, the method steps as described in any of the above can be implemented.

[0176] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, as long as the functions and effects produced do not exceed the scope of the technical solution of the present invention, fall within the protection scope of the present invention.

Claims

1. A task offloading method based on graph neural network and deep reinforcement learning, characterized in that: include: Step S1, modeling a cloud-edge-end collaborative system of multiple terminal devices, multiple edge servers and a single cloud server; Step S2: construct a three-layer task offloading model of cloud server, edge server and terminal device; Step S3, defining an optimization problem with the goal of minimizing the weighted sum of task delay and device energy consumption; Step S4, converting the optimization problem defined in step S3 into a Markov decision process; Step S5: construct a task offloading method based on graph neural network and deep reinforcement learning to solve the task offloading problem under multiple terminal devices, multiple edge servers and a single cloud server.

2. According to claim 1, a task offloading method based on graph neural network and deep reinforcement learning is characterized in that: In step S1, the cloud-edge-end collaborative system of multiple terminal devices, multiple edge servers and a single cloud server is constructed as follows: S11, system time slot Assume that the system time is discretized into T time slots of equal length, and use the set To express; S12. Terminal equipment Assume that there are D terminal devices, using the set To indicate that each terminal device The maximum computing power is denoted as f d , the idle power is expressed as In each time slot t, the task of terminal device d is described as a triple, namely Indicates the task data size; Indicates the CPU cycles required for each bit of task data; represents the maximum tolerable delay of the task, Indicates the uninstall indicator. When , it means that the task is processed locally; when When , it means that the task is offloaded to the edge server for processing; when , it indicates that the task is offloaded to the cloud server for processing; S13, Edge Server Assume that there are M base stations, and an edge server is deployed near each base station. The base station is connected to the edge server through a point-to-point optical fiber. Since there is a one-to-one relationship between the base station and the edge server, the collection to uniformly represent the base station-edge server pair; edge server The computational power is defined as f m ; S14, Cloud Server Assume that there is a central cloud server, and the computing power of the cloud server is defined as f c .

3. The task offloading method based on graph neural network and deep reinforcement learning according to claim 2 is characterized in that: Step S2 includes the following steps: S21. Establish a local computing model: Assuming that each terminal device has different computing power, use express; It represents the task size generated by terminal device d in time slot t; Represents the CPU cycles required for each bit of task data; therefore, the local computation time is defined as: The energy consumption of local computation is defined as: k is the energy consumption calculation coefficient related to the chip architecture; S22. Establish edge computing model Assume that the total computing power on the edge server is f m Considering that the edge server needs to provide services to multiple terminal devices at the same time, the computing power allocated by server m to terminal d at time t is expressed as The time of edge computing is defined as: Since the edge server has sufficient power supply, the energy consumption of the edge server executing tasks is not considered; when the task is executed on the edge server, the idle energy consumption of the terminal device d is defined as: in is the idle power of device d; The task offloading to the edge server needs to be transmitted to the base station through the wireless channel, and its transmission rate is defined as: in represents the bandwidth between terminal device d and base station m in time slot t, represents the transmission power of terminal device d at time slot t, and N0 represents the power spectral density of noise; represents the channel gain between terminal device d and base station m at time slot t; The transmission time from terminal device d to base station m is defined as: The energy consumption of terminal device d transmitted to the base station is defined as: S23. Establish cloud computing model Similar to the edge server, the time for executing a task on the cloud server is defined as: When the cloud server is executing tasks, the terminal device d is in an idle state and still needs to consume energy. Therefore, the idle energy consumption of the terminal device d is defined as: The cloud server and the base station are connected via a wide area network. Therefore, when the terminal device communicates with the cloud server, the data needs to be transmitted from the terminal device to the base station first, and then forwarded to the cloud server by the base station. The time for the terminal device to transmit to the cloud server is defined as: in is the time it takes for the terminal device to transmit to the base station, is the rate at which base station m allocates the transmission of terminal device d to the cloud server in time slot t; The energy consumption of terminal device d transmitted to the cloud server is defined as:

4. The task offloading method based on graph neural network and deep reinforcement learning according to claim 3 is characterized in that: Step S3 includes the following steps: S31. Define the total processing delay of the task Task The total processing delay is defined as: S32. Define the energy consumption of terminal equipment According to the above, the energy consumption of the terminal device is defined as: S33. Define the optimization problem The optimization problem is defined as: where η is a weighting parameter used to make a trade-off between time and energy.

5. The task offloading method based on graph neural network and deep reinforcement learning according to claim 1 is characterized in that: Step S4 includes the following steps: S41. Define system status, including cloud server status, edge server status, terminal device status, and network status; S42, defining the action of the terminal device, including the offloading location, transmission power and local computing capability; S43. Define a reward function, including delay, energy consumption and task completion rate.

6. The task offloading method based on graph neural network and deep reinforcement learning according to claim 4 is characterized in that: Step S4: The specific implementation steps are as follows: S41. The system status consists of the terminal device status, edge server status, cloud server status, and the network topology of the cloud-edge system. The terminal device status is defined as: The edge server status is defined as: The status of a cloud server is defined as: Assume that in one training, the topology of the cloud-edge-end collaborative network remains unchanged; the cloud-edge-end collaborative network is modeled as an undirected graph structure, represented by G = (V, E), where V is a node set, including three types of nodes: terminal devices, edge servers, and cloud servers; E is an edge set, representing the connection relationship between these nodes; therefore, the entire system state is defined as: S42, the action of terminal device d consists of the following parts Composition: The unloading location of the task Transmit power of the device The computing resources allocated by the device to the task To simplify the representation, the continuous action of device d is defined as Therefore, at time slot t, the action of device d is defined as Furthermore, the set of actions for all devices in time slot t is defined as S43, the reward function consists of three parts, which aims to comprehensively consider the task processing delay, the energy consumption of the terminal device and the satisfaction of the task delay constraint; the first part is the task processing delay; the second part is the energy consumption of the terminal device; the third part is the penalty for violating the maximum tolerable delay; the penalty term is defined as: In summary, the reward function of device d at time slot t is defined as: At time slot t, the reward of all devices is defined as 7. The task offloading method based on graph neural network and deep reinforcement learning according to claim 1 is characterized in that: Step S5 includes the following steps: S51. Design the structure of the graph neural network and process the input and output of the network; S52, construct an actor network and explain the generation process of continuous actions; S53. Construct a critic network and explain the generation process of state-action value pairs and discrete actions; S55. Design loss functions for actor and critic networks.

8. The task offloading method based on graph neural network and deep reinforcement learning according to claim 7 is characterized in that: The structure of the graph neural network adopts a two-layer graph attention network.

9. A task offloading system based on graph neural network and deep reinforcement learning, characterized in that: The method comprises a memory, a processor and computer program instructions stored in the memory and executable by the processor. When the processor executes the computer program instructions, the method steps as claimed in any one of claims 1 to 8 can be implemented.

10. A computer-readable storage medium storing computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 8 can be implemented.