A Task Replication and Offloading Method Based on Game Theory in Mobile Edge Computing
By adopting game theory and multi-agent reinforcement learning methods in mobile edge computing, the task replication and unloading strategies are optimized, and the problem of insufficient system reliability and availability in the existing technology is solved, and efficient task execution and performance optimization in dynamic environments are achieved.
Patent Information
- Application Number
- CN202411308188.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-09-19
AI Technical Summary
The existing mobile edge computing task offloading strategies are difficult to ensure high availability and system reliability when facing device failures, network instability, etc., and traditional methods have additional costs and insufficient theoretical analysis when acquiring global information.
The task replication and unloading method based on game theory is adopted, combined with the multi-agent reinforcement learning algorithm, and the mobile edge system architecture is built. The Markov decision-making process and the multi-agent dual-latency depth deterministic strategy gradient algorithm can be observed through the multi-agent part to optimize the task replication and unloading strategies to improve system performance and reliability.
In the face of network failures and insufficient computing resources, it provides strong fault tolerance to ensure successful execution of computing tasks and improves the overall service performance and reliability of the system.
Smart Images

Figure CN119212106B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of mobile edge computing, and particularly relates to a task replication and offloading method based on game theory in mobile edge computing. Background Art
[0002] The popularization of the fifth-generation mobile communication technology (5G) has enabled the Internet of Things (IoT) to enter a rapid development era. In this era, a large number of computationally intensive and latency-sensitive applications have become ubiquitous. Applications such as real-time video analysis, Augmented Reality (AR), and autonomous driving require a large amount of computing resources and ultra-low latency to operate effectively. This surge in demand has prompted the emergence of a new paradigm, mobile edge computing (MEC), which brings computing power to the network edge to make computing resources closer to the data source. MEC offers many advantages, including reduced latency, improved bandwidth utilization, and enhanced user experience by offloading tasks from a centralized cloud computing center to an edge server (MEC server) closer to the end device.
[0003] In the context of MEC, task offloading strategies have received much attention because they are used to optimize the allocation of computing tasks between end devices and MEC servers. These strategies are crucial for maximizing the efficiency and performance of the edge computing environment. By intelligently deciding which tasks need to be offloaded and where to offload them, the global task execution time can be significantly reduced, the energy consumption of end devices can be lowered, and the overall system throughput can be increased.
[0004] However, the successful implementation of task offloading technology not only requires efficient decision algorithms and communication protocols but also fault tolerance to ensure the reliability of task execution in the event of network failures, communication delays, or computational node failures. In a mobile edge computing environment, the fault tolerance of task offloading is particularly important because end devices and edge nodes often encounter unstable network connections, device failures, limited computing resources, and unpredictable working environments. In addition, many application scenarios, such as autonomous driving, remote healthcare, and intelligent manufacturing, require the system to have high availability and continuity. Therefore, designing task offloading technology with fault tolerance is crucial for ensuring the high availability of edge computing systems.
[0005] Task offloading technology in MEC systems has attracted great attention from academia and industry in the past few years. From the perspective of the fault tolerance of tasks, it can be divided into two categories: task offloading strategies without fault tolerance and those with fault tolerance.
[0006] (I) Task offloading strategies without fault tolerance
[0007] The task offloading strategy without fault tolerance, as the name implies, does not consider the fault tolerance of the system as an indicator for evaluating the performance of the offloading strategy when considering the task offloading strategy in the MEC environment. Instead, it takes the overall execution delay or average execution delay of all tasks as the main optimization goal. This type of method can be divided into the following two subcategories:
[0008] (1.1) Decentralized offloading strategy
[0009] This type is also known as general edge computing, which refers to edge computing that relies only on edge devices with sensing, storage, and communication capabilities to achieve task offloading without centralized management. A typical technology is to convert the problem into a stochastic game theory model based on a complete observation of the system state, and derive the Nash equilibrium among edge devices. On this basis, methods such as general adversarial imitation learning or deep reinforcement learning are used to solve the Nash equilibrium point, so as to obtain the optimal task offloading strategy.
[0010] (1.2) Centralized offloading strategy
[0011] This type is also known as edge computing based on a centralized controller. Its basic idea is to rely on a central controller that masters the global information of the network, and combines information such as network state information and edge device computing resources to decide how to offload tasks. Typical technologies include a learning algorithm using Deep Deterministic Policy Gradient (DDPG) to implement the strategy of multi-priority task scheduling for the characteristics of mutual dependence among multiple tasks; another technology mainly aims at the computing offloading and service caching in a three-layer mobile cloud edge computing architecture (edge devices, MEC servers, and cloud computing centers). In this structure, mobile users subscribe to the cloud service center to obtain computing offloading services and pay relevant fees monthly or annually. In addition, the cloud service center can purchase some computing and communication resources from MEC servers with limited cache capacity and computing resources to help mobile users offload computing.
[0012] However, these studies are usually limited to the assumption of a reliable edge network environment, and potential failure situations are ignored when making task offloading decisions, such as sudden failures of devices executing tasks, power outages, etc.
[0013] (2) Task offloading strategies with fault tolerance
[0014] At present, in the academic and industrial circles, there have emerged some task offloading strategies that consider fault tolerance, which are summarized into two categories: system-level task fault tolerance solutions and task-level task fault tolerance solutions.
[0015] (2.1) System-level task fault tolerance offloading solutions
[0016] The characteristics of the system-level task fault-tolerant offloading scheme are that it requires global information to support offloading decisions. There are several typical technologies as follows: One is based on the auction mechanism, which models the interaction between MEC participants and the probability of successful user offloading to ensure fault-tolerant edge services in unreliable situations; another is based on modeling the dynamic mobile edge environment, user costs, fault penalties, and diverse Quality of Service (QoS) requirements, transforming the task offloading problem into an online decision-making problem in a stochastic process and implementing a fault recovery strategy during the decision-making process to handle different types of faults; the third is to use the Dueling Deep Q network algorithm to determine user offloading behavior and an adaptive checkpoint mechanism to improve task reliability, thereby implementing a semi-online fault-tolerant offloading method to optimize service offloading efficiency and system reliability.
[0017] However, the above-mentioned technologies are all fault-tolerant task offloading schemes implemented at the system global level, all of which require obtaining corresponding global information and will incur additional costs because it is difficult to easily obtain global network, device, or task information when actually implemented.
[0018] (2.2) Task-level task fault-tolerant offloading scheme
[0019] The second type of offloading scheme with fault-tolerant capabilities is the task-level offloading scheme, that is, from the perspective of the edge device computing task itself, to ensure the fault-tolerant capabilities of the MEC system without the perspective of global information. Typical technologies include: One is the fault-tolerant method based on checkpoints and replication, whose basic idea is to use intelligent checkpoints for IoT application tasks executed in the edge network and improve system reliability by replicating checkpoint files on nearby alternative MEC servers, thereby improving the overall availability of the system. Another is based on task offloading and service replication technologies, which usually model the offloading strategy problem as an integer linear programming problem, aiming to minimize the response time of all users and modeling with the constraints of simultaneously meeting user time and time difference, and then using linear relaxation programming based on Lagrangian analysis to solve the model to obtain the optimal offloading strategy to improve system availability; the third is to model the problem as a non-linear integer programming problem based on the probability of a single replicated task interruption, and jointly optimize the replication decision of the task under the constraints of the hardware resources of the user terminal and the MEC server to minimize the probability of system interruption to the greatest extent. However, these technologies mainly have the following two deficiencies: One is achieved through heuristic solutions; the other is the lack of comprehensive theoretical analysis support. These deficiencies make these schemes difficult to adapt to the dynamic characteristics of the edge computing scenario.
[0020] As can be seen from the above analysis, the existing edge computing task replication and offloading strategies have the following deficiencies:
[0021] 1) The task offloading strategy without fault tolerance does not consider unexpected situations such as sudden failures of devices executing tasks or task transmission delays due to network reasons when making task offloading decisions, resulting in task offloading or execution failures. Such task offloading strategies are not applicable to business scenarios that require high availability to users;
[0022] 2) The existing system-level task fault-tolerant offloading solutions assume the existence of a centralized controller when performing task fault-tolerant offloading, which can obtain global information such as network status information, device status information, and task status information. However, this assumption is too ideal, and it is difficult to obtain global information in actual implementation. From a technical implementation perspective, obtaining global information will introduce additional costs, such as communication delays, etc.;
[0023] 3) The commonality of the technical principles of the existing task-level task fault-tolerant offloading solutions is achieved through heuristic solutions. At the same time, these solutions lack rigorous theoretical analysis support and are difficult to adapt to changes when the edge network environment changes. Therefore, it is difficult to achieve good results. In the MEC environment, changes in the edge network status are relatively common phenomena. Summary of the Invention
[0024] In view of the defects existing in the prior art, the present invention provides a task replication and offloading method based on game theory in mobile edge computing, which can effectively solve the above problems.
[0025] The technical solution adopted by the present invention is as follows:
[0026] The present invention provides a task replication and offloading method based on game theory in mobile edge computing, including the following steps:
[0027] Step S1, constructing a mobile edge system architecture; the mobile edge system architecture includes N IoT devices D1, D2,..., D N forming a set D = {D1, D2,..., D N}, M MEC servers E1, E2,..., E M forming a set ε = {E1, E2,..., E M}, and a cloud computing center r; where, and respectively represent the sets {1, 2,..., N} and {1, 2,..., M};
[0028] Step S2. During the period T, taking the task replication and offloading strategy of each IoT device as the one to be solved, where the task replication and offloading strategy of each IoT device includes the original task replication strategy x, the original task offloading strategy y, and the replicated task offloading strategy z; taking the maximization of the overall service performance of the mobile edge system in the period T as the objective function, and combining the constraint conditions, an optimization model for task replication and offloading is constructed;
[0029] Step S3. Transform the optimization model of task replication and offloading into a game theory model;
[0030] Step S4. Adopt a Nash equilibrium solving method based on multi-agent reinforcement learning to solve the game theory model, and obtain the optimal task replication and offloading strategy of each IoT device.
[0031] Preferably, step S2 is specifically as follows:
[0032] Step S2.1. In the mobile edge environment, define the parameters and meanings related to task replication and offloading:
[0033] Step S2.1.1. Assume that IoT device D n generates an original task at time t forming a set where Q(n, t) represents the number of all original tasks generated by IoT device D n at time t;
[0034] Step S2.1.2. For the original task n generated by IoT device D at time t q ∈ {1, 2,..., Q(n, t)}, its input data volume and computational complexity are respectively and The number of copies it is replicated at time t is expressed as Therefore, copied tasks are generated; where G represents the maximum copy number limit, and the number of task copies Therefore, if the original task is not replicated, then
[0035] Step S2.1.3. In the mobile edge environment, each original task or each replicated task generated by IoT device D n at time t can only be executed on one node, that is: either on the local IoT device D n or be offloaded to a certain MEC server E in the MEC server set ε m or be offloaded to the cloud computing center r for execution;
[0036] Therefore, define the original task for the decision result variable and the decision result variable for the replicated task x i ∈ {r ∪ n ∪ ε}, where i represents the node executing the task, and the meaning is as follows:
[0037] If the original task is executed on the local IoT device D n , then at this time i = n. Therefore,
[0038] If the replicated task x is executed on the local IoT device D n , then at this time i = n. Therefore,
[0039] If the original task is executed in the cloud computing center r, then at this time i = r. Therefore,
[0040] If the replicated task x is executed in the cloud computing center r, then at this time i = r. Therefore,
[0041] If the original task is offloaded to the MEC server E m for execution, then at this time i = m. Therefore,
[0042]
[0043] If the replicated task x is offloaded to the MEC server E m for execution, then at this time i = m. Therefore,
[0044]
[0045] Since each original task or each replicated task x can only be executed on one node i, therefore, and
[0046] In step S2.2, evaluate to obtain the transmission delay of the IoT device D n at time t
[0047] In step S2.3, evaluate to obtain the execution delay of the IoT device D n at time t
[0048] In step S2.4, evaluate to obtain the reliability Υ n of the IoT device D at time t n (t);
[0049] Step S2.5, using formula (1), for IoT device D n The transmission delay Execution delay and reliability Υ n (t) are weighted and summed to obtain the service cost u n of IoT device D at time t n (t):
[0050]
[0051] where: α1 and α2 are the weight parameters for weighted summation;
[0052] Step S2.6, using formula (2), to obtain the overall service cost U g (t) of the mobile edge system at time t:
[0053]
[0054] Step S2.7, with the goal of minimizing the overall service cost of the mobile edge system over the period T, that is, maximizing the overall service performance, and combined with the constraint conditions, an optimization model for task replication and offloading is constructed.
[0055] Preferably, step S2.2 is specifically as follows:
[0056] Step S2.2.1, the original task generated by IoT device D n at time t The transmission delay from IoT device D n to the MEC server E m is determined using formula (3):
[0057]
[0058] where: r n,m (t) represents the wireless link transmission rate from IoT device D n to the MEC server E m and is determined using formula (4):
[0059]
[0060] where: W represents the link bandwidth; N0 is the noise power spectral density; h n,m (t) and p n,m (t) respectively represent the wireless channel gain and transmit power from IoT device D n to the MEC server E m ;
[0061] Step S2.2.2, the IoT device D n The original task generated at time t The transmission delay from the IoT device D n To the cloud computing center r Is determined using Equation (5):
[0062]
[0063] Where: Represents the transmission rate between the IoT device D n And the cloud computing center r;
[0064] Step S2.2.3, using Equation (6), obtain the total transmission delay at time t of all the original tasks of the IoT device D n Of all the original tasks of the IoT device D
[0065]
[0066] The IoT device D n The total transmission delay at time t of all the original tasks Is the transmission delay of the IoT device D n At time t
[0067] Preferably, Step S2.3 is specifically as follows:
[0068] Step S2.3.1, if the original task generated by the IoT device D n At time t Is executed on the local IoT device D n Then its execution delay when executed locally Is obtained using Equation (7):
[0069]
[0070] Where: f n (t) represents the computing power of the IoT device D n At time t;
[0071] Step S2.3.2, if the IoT device D n Offloads the original task at time t To the MEC server E m For execution, then its execution delay when executed on the MEC server E m Is obtained using Equation (8):
[0072]
[0073] Where: f n,q,m (t) represents the computing power allocated by the MEC server E m to the original task ;
[0074] Step S2.3.3, if the IoT device D n offloads the original task to the cloud computing center r for execution at time t, its execution delay is calculated as 0;
[0075] Step S2.3.4, using formula (9), obtain the total execution delay of all the original tasks of the IoT device D n at time t
[0076]
[0077] The IoT device D n at time t The total execution delay of all the original tasks n is the execution delay of the IoT device D
[0078] Preferably, step S2.4 is specifically as follows:
[0079] Step S2.4.1, the failure rate of the MEC server E m is represented by δ m (t), and the failure rate follows a Poisson distribution. Therefore, at time t, the probability that the MEC server E m has r m failures The probability that the MEC server E m has no failure, that is, r m = 0 is
[0080] Therefore, the probability that the MEC server E m has a failure at time t
[0081] Therefore, for the IoT device D n at time t the cumulative failure probability of the original task
[0082] Step S2.4.2, the reliability of the original task of the IoT device D n at time t is determined by formula (10):
[0083]
[0084] Step S2.4.3, IoT device D n The reliability γ n (t) of all the original tasks at time t is determined by formula (11):
[0085]
[0086] IoT device D n The reliability γ n (t) of all the original tasks at time t is the reliability γ n (t) of IoT device D at time t. n (t).
[0087] Preferably, in step S2.7, the optimization model for task replication and offloading is as follows:
[0088] Objective function P1:
[0089] The constraint conditions C1 - C7 are respectively:
[0090] C1:
[0091] C2:
[0092] C3:
[0093] C4:
[0094] C5: 0 ≤ p n,m (t) ≤ P n (t)
[0095] C6: q ∈ {1, 2, …, Q(n, t)}
[0096] C7: i ∈ {r ∪ n ∪ ε}
[0097] Where:
[0098] C1 represents the value range of .
[0099] C2 represents that and are both binary variables;
[0100] C3 represents that for any one original task it can only be executed on one node, that is: the local IoT device D n executes, the cloud computing center r executes, or the MEC server Em Execute, that is: the original task It cannot be split and executed on multiple nodes;
[0101] In C4, F m F(t) represents the overall computing power of the MEC server E m ; f n,q,m f(t) represents the computing power allocated to the original task in the MEC server E m ; Therefore, the meaning of C4 is: in the MEC server E the allocated computing resources cannot exceed its total computing resources; m
[0102] In C5, p n,m p(t) represents the transmission power of the IoT device D n to the MEC server E m ; P n P(t) represents the maximum power of the IoT device D n ; The meaning of C5 is: the transmission power of the IoT device D n to the MEC server E m cannot exceed the maximum power of the IoT device D n ;
[0103] C6 represents the value range of the number of IoT devices, the number of MEC servers, and the number of original tasks;
[0104] C7 represents the value range of the number of task replicas and the value range of the node types for task execution.
[0105] Preferably, step S3 is specifically:
[0106] Regarding the N IoT devices D1, D2,..., D N as players in game theory; for player D n , that is, the IoT device D n , its task replication and offloading strategy is represented as π n = {x n , y n , z n}; where x n represents the original task replication strategy of the IoT device D n , y n represents the offloading strategy of the original task, and z n represents the offloading strategy of the replicated task;
[0107] When the IoT device D n adopts the task replication and offloading strategy π n , its service cost u n (π(t), its service cost un ) As the payoff in game theory, a game theory model is thus established.
[0108] Preferably, step S4 is specifically as follows:
[0109] Step S4.1, convert the game theory model into a partially observable Markov decision process (POMDP) for multi-agent systems.
[0110] Specifically, each player D in the game theory model n is regarded as an agent D n ; The POMDP consists of five parts, denoted as
[0111] represents the state space, which represents the global state of the entire system;
[0112] represents the observation space, which represents the system state observable by the agent and is part of the global state; At time t, the observation space of the agent where the observation state of agent D n L L n (t) represents the position information of agent D n f n (t) represents the computing and processing power of agent D n ; represents the number of original tasks generated by agent D n and the size of the original task input data, Σ ε represents the global information of the MEC server, Σ r represents the global information of the cloud computing center;
[0113] represents the action space, which contains all possible actions of all agents, i.e., a n (t) = {x n (t), y n (t), z n (t)}; where a n (t) represents the action taken by agent D n at time t; x n (t), y n (t), z n (t) respectively represent the actions taken by agent D n at time t, including the original task replication strategy, the offloading strategy of the original task, and the offloading strategy of the replicated task;
[0114] p represents the observation state transition probability function, indicating the likelihood of receiving a specific observation given the current state and the action taken;
[0115] Represents the reward space, denoted as R n (t) = u n (t); where R n (t) represents the reward obtained by agent D n at time t; u n (t) represents the service cost of agent D n at time t;
[0116] Based on the five parts included in the POMDP The problem of multi-task replication and offloading is expressed as:
[0117] Objective:
[0118]
[0119] where: t represents time t, T represents the time interval; γ(t) represents the decay coefficient, 0 ≤ γ(t) ≤ 1, representing the influence of the reward value at a future time on the current reward value; R n (t) represents the reward received by agent D n at time t; In this objective function Objective, the goal of the task replication and offloading strategy is to find appropriate x, y, and z to maximize the cumulative discounted reward value of all agents within the time interval T; Represents the optimal task replication and offloading strategy found through optimization, which are respectively the optimal original task replication strategy, the offloading strategy of the original task, and the offloading strategy of the replicated task for IoT device D n ;
[0120] Step S4.2, use the task replication and offloading Nash equilibrium solving method based on MATD3 to solve the objective function Objective:
[0121] Specifically, each agent D n has a MATD3 controller, and each agent D n interacts with the environment by executing its own action a n (t) to obtain the corresponding reward R n (t) and observe the next state; The algorithm includes an experience replay pool D, a critic network, and an actor network;
[0122] Experience replay pool Used to store existing experience samples, denoted as Among them, Denotes the agent D n The current observation state, a n (t) represents the action taken currently, Denotes the reward obtained currently, Denotes the state observed next; The experience samples stored in the experience replay pool D are used to train the critic network and the actor network;
[0123] The actor network includes an Evaluation Network and a Target Network; The Evaluation Network is used to receive the current observation state of the agent D n The current observation state Through the policy network Generate the action a n (t); This action is continuously improved through the gradient update optimizer; The Target Network adopts a policy Used to add noise to the next observation state To generate the target action And synchronize the parameters from the Evaluation Network through soft update to ensure the stability and consistency of the policy network ; Through the experience samples in the experience replay pool The actor network continuously learns and optimizes the policy to achieve efficient decision-making for edge computing tasks;
[0124] The critic network includes two independent Evaluation Networks and two independent Target Networks. The parameters of the two Evaluation Networks are respectively And Optimized through the gradient descent algorithm, with the goal of minimizing the loss function This loss function measures the difference between the Q value output by each Evaluation Network and the target Q value; Each Evaluation Network receives the current observation state of the agent D n The current observation state And the action a n (t) taken, and calculates the target Q value respectively, that is: And Update its parameters through the gradient descent optimizer And The two Target Networks generate the target Q values respectively: And Among them, and respectively represent the parameters of two target networks, and the parameters of the evaluation network are synchronized through soft update.
[0125] A task replication and offloading method based on game theory in mobile edge computing provided by the present invention has the following advantages:
[0126] The present invention proposes a task replication and offloading method based on game theory in mobile edge computing, which effectively combines a robust game theory model with a multi-agent reinforcement learning algorithm, and can effectively improve system service performance and system reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0127] Figure 1 is a schematic flow chart of a task replication and offloading method based on game theory in mobile edge computing provided by the present invention;
[0128] Figure 2 is an architecture diagram of the mobile edge system provided by the present invention;
[0129] Figure 3 is a schematic diagram of the principle of the task replication and offloading Nash equilibrium solving method based on the MATD3 algorithm provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0130] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0131] The present invention proposes a task replication and offloading method based on game theory in mobile edge computing, which effectively combines a robust game theory model with a multi-agent reinforcement learning (Multi-Agent Reinforcement Learning, MARL) algorithm, and can effectively improve system service performance and system reliability.
[0132] The present invention proposes a game-theoretic task replication and offloading method (Game Theoretical Task Replication and Offloading, GTRO) in MEC. Specifically, GTRO first models the proposed task replication and offloading problem using a multi-party game theory model, analyzes the Nash equilibrium (NE) points in the game model, and studies the existence and uniqueness conditions of NE to ensure that the game model has a unique solution. Secondly, GTRO adopts the Multi-agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm to overcome the problem of various resource dynamic changes in the edge network, and uses multiple agents to learn and adjust strategies through interaction in a simulated environment.
[0133] The method proposed by the present invention improves the robustness of the MEC system through flexible, efficient, and optimized task replication and offloading strategies. In the face of intermittent network failures, hardware failures, and lack of computing resources, it can ensure the fault tolerance of computing tasks, thereby improving the overall availability of the system.
[0134] Refer to Figure 1 , the present invention provides a game-theoretic task replication and offloading method in mobile edge computing, including the following steps:
[0135] Step S1, construct a mobile edge system architecture; the mobile edge system architecture includes a set formed by N IoT devices D1, D2,..., D N formed set a set ε formed by M MEC servers E1, E2,..., E M formed set ε = {E1, E2,..., E M}, and a cloud computing center r; where, and respectively represent the sets {1, 2,..., N} and {1, 2,..., M};
[0136] Refer to Figure 2 , which is a specific example diagram of a mobile edge system architecture. In Figure 2 , the IoT devices are represented by ED1, ED2, ED3, etc. The MEC servers are represented by "MEC"; the cloud computing center r is represented by "cloud". The communication between the IoT devices and the MEC servers is based on the local area network (LAN), and the communication between the IoT device / MEC server and the cloud computing center r is based on the wide area network.
[0137] Step S2. Within the period T, the task replication and offloading strategy of each IoT device is to be solved. Among them, the task replication and offloading strategy of each IoT device includes the original task replication strategy x, the original task offloading strategy y, and the replicated task offloading strategy z. With the overall service performance maximization of the mobile edge system within the period T as the objective function, combined with the constraint conditions, an optimization model for task replication and offloading is constructed.
[0138] Step S3. Transform the optimization model of task replication and offloading into a game theory model.
[0139] Step S4. Use the Nash equilibrium solution method based on multi-agent reinforcement learning to solve the game theory model, and obtain the optimal task replication and offloading strategy of each IoT device.
[0140] The following details Steps S2 to S4:
[0141] Step S2. Within the period T, the task replication and offloading strategy of each IoT device is to be solved. Among them, the task replication and offloading strategy of each IoT device includes the original task replication strategy x, the original task offloading strategy y, and the replicated task offloading strategy z. With the overall service performance maximization of the mobile edge system within the period T as the objective function, combined with the constraint conditions, an optimization model for task replication and offloading is constructed.
[0142] Step S2 is specifically implemented through Steps S2.1 to S2.7:
[0143] Step S2.1. In the mobile edge environment, define the parameters and meanings related to task replication and offloading:
[0144] Step S2.1.1. Assume that IoT device D n generates original tasks at time t forming a set where Q(n,t) represents the number of all original tasks generated by IoT device D n at time t.
[0145] Step S2.1.2. For the original task n generated by IoT device D at time t q ∈ {1, 2, …, Q(n,t)}, its input data volume and computational complexity are respectively and The number of copies it is replicated at time t is denoted as Therefore, copied tasks are generated; where G represents the maximum copy number limit, and the task copy number Therefore, if the original task is not replicated, then
[0146] Step S2.1.3, in the mobile edge environment, IoT device D n For each original task or each replicated task generated at time t, it can only be executed on one node, that is: either on the local IoT device D n for execution, or offloaded to a certain MEC server E in the set ε of MEC servers m for execution, or offloaded to the cloud computing center r for execution;
[0147] Therefore, define the decision result variable of the original task and the decision result variable of the replicated task x, where i ∈ {r ∪ n ∪ ε}, and i represents the node executing the task, meaning:
[0148] If the original task is executed on the local IoT device D n then at this time i = n, so
[0149] If the replicated task x is executed on the local IoT device D n then at this time i = n, so
[0150] If the original task is executed in the cloud computing center r, then at this time i = r, so
[0151] If the replicated task x is executed in the cloud computing center r, then at this time i = r, so
[0152] If the original task is offloaded to the MEC server E m for execution, then at this time i = m, so
[0153] If the replicated task x is offloaded to the MEC server E m for execution, then at this time i = m, so
[0154] Since each original task or each replicated task x can only be executed on one node i, so and
[0155] Step S2.2, evaluate to obtain the transmission delay of IoT device D n at time t
[0156] In Figure 2 the system architecture, for IoT device D n , there are two communication links, namely the communication link between IoT device D n and MEC server E m , and the communication link between IoT device D n and cloud computing center r. Therefore, calculate the transmission delays of these two communication links separately and then sum them up to obtain the transmission delay of IoT device D n at time t
[0157] The specific steps are as follows:
[0158] Step S2.2.1, the transmission delay caused by the communication link between IoT device D n and MEC server E m :
[0159] For the original task generated by IoT device D n at time t its transmission delay from IoT device D n to MEC server E m is determined using formula (3):
[0160]
[0161] where: r n,m (t) represents the wireless link transmission rate between IoT device D n and MEC server E m , which is determined using formula (4):
[0162]
[0163] where: W represents the link bandwidth; N0 is the noise power spectral density; h n,m (t) and p n,m (t) respectively represent the wireless channel gain and transmit power between IoT device D n and MEC server E m ;
[0164] Step S2.2.2, the transmission delay caused by the communication link between IoT device D n and cloud computing center r:
[0165] Specifically, when MEC server E m cannot execute IoT device D nWhen generating a task, IoT device D n needs to offload the task to cloud computing center r. Since cloud computing center r has sufficient computing power to ensure the successful execution of the task and does not consider task replication, there is no need to consider the transmission delay of the replicated task.
[0166] Therefore, IoT device D n The original task generated at time t from IoT device D n The transmission delay to cloud computing center r is determined using formula (5):
[0167]
[0168] where: represents the transmission rate between IoT device D n and cloud computing center r;
[0169] Step S2.2.3, using formula (6), obtain the total transmission delay of all original tasks of IoT device D n at time t
[0170]
[0171] The total transmission delay of all original tasks of IoT device D n at time t is the transmission delay of IoT device D n at time t
[0172] Step S2.3, evaluate and obtain the execution delay of IoT device D n at time t
[0173] In Figure 2 the system architecture, for the task generated by IoT device D n at time t, there are three task execution modes, namely: executed locally by IoT device D n executed, offloaded to MEC server E m executed, or offloaded to cloud computing center r for execution. Therefore, the following three task execution delays are generated; then, sum the three task execution delays to obtain the execution delay of IoT device D n at time t
[0174] The specific steps are as follows:
[0175] Step S2.3.1, the execution delay of local IoT device D n :
[0176] If the original task generated by IoT device D n at time t is executed locally on IoT device D n then the execution delay of its local execution is obtained using formula (7):
[0177]
[0178] where: f n (t) represents the computing power of IoT device D n at time t;
[0179] Step S2.3.2, the execution delay of MEC server E m Execution delay:
[0180] If IoT device D n unloads the original task to MEC server E m at time t and executes it, then the execution delay of its execution on MEC server E m is obtained using formula (8):
[0181]
[0182] where: f n,q,m (t) represents the computing power allocated by MEC server E m to the original task ;
[0183] Step S2.3.3, the execution delay of cloud computing center r:
[0184] If IoT device D n unloads the original task to cloud computing center r for execution at time t, since cloud computing center r has sufficient powerful and rich computing resources, the delay brought by task execution in cloud computing center r can be ignored. Therefore, its execution delay is calculated as 0;
[0185] Step S2.3.4, using formula (9), the total execution delay of all original tasks of IoT device D n at time t
[0186]
[0187] The total execution delay of all original tasks of IoT device D n at time t is the total execution delay of all original tasks of IoT device D nExecution delay at time t
[0188] Step S2.4, evaluate to obtain IoT device D n Reliability γ n (t) at time t;
[0189] Step S2.4 is specifically as follows:
[0190] Step S2.4.1, the failure rate of MEC server E m is represented by δ m (t). The failure rate follows a Poisson distribution. Therefore, at time t, the probability that MEC server E m has r m failures is:
[0191] The probability that MEC server E m does not fail, i.e., r m = 0, is
[0192] Therefore, the probability that MEC server E m fails at time t
[0193] Therefore, for IoT device D n the original task at time t The cumulative failure probability at time t is expressed as
[0194] Step S2.4.2, the reliability of the original task of IoT device D n at time t is determined by formula (10):
[0195]
[0196] Step S2.4.3, the reliability Υof all the original tasks of IoT device D n at time t n is determined by formula (11):
[0197]
[0198] The reliability Υ n of all the original tasks of IoT device D n at time t is the reliability γ n of IoT device D n (t).
[0199] Step S2.5, using formula (1), for IoT device D n The transmission delay at time t Execution delay and reliability Υ n (t) are weighted and summed to obtain the service cost u of IoT device D n at time t: n (t):
[0200]
[0201] where: α1 and α2 are the weight parameters for weighted summation;
[0202] Step S2.6, using formula (2), to obtain the overall service cost U of the mobile edge system at time t g (t):
[0203]
[0204] Step S2.7, aiming at minimizing the overall service cost of the mobile edge system in period T, that is, maximizing the overall service performance, and combining the constraint conditions, an optimization model for task replication and offloading is constructed.
[0205] Specifically, the present invention proposes a task replication and offloading method based on game theory to minimize the service cost of each IoT device, that is: maximizing the service performance, so as to maximize the service performance of the overall system. Specifically, the present invention defines the service cost u n of IoT device D n (t) as the weighted sum of reliability, execution delay and transmission delay, and defines the service cost U g (t) of the overall system as the sum of the service costs of all IoT devices in the edge computing system, that is: where: α1 and α2 are the weight parameters for weighted summation.
[0206] The present invention models the task replication and offloading problem as an optimization problem of finding the maximum service performance within time T for the original task replication strategy x, the original task offloading strategy y, and the replicated task offloading strategy z (i.e., the service cost minimum problem of the following P1 model); the specific objective function and constraint conditions are as follows:
[0207] Objective function P1:
[0208] The constraint conditions C1 - C7 are respectively:
[0209] C1:
[0210] C2:
[0211] C3:
[0212] C4:
[0213] C5: 0 ≤ p n,m (t) ≤ P n (t)
[0214] C6: q ∈ {1, 2, …, Q(n, t)}
[0215] C7: i ∈ {r ∪ n ∪ ε}
[0216] Where:
[0217] C1 represents the value range of
[0218] C2 represents and are both binary variables;
[0219] C3 represents for any one of the original tasks it can only be executed on one node, that is: the local IoT device D n executes, the cloud computing center r executes, or the MEC server E m executes, that is: the original task cannot be split and executed on multiple nodes;
[0220] In C4, F m (t) represents the overall computing power of the MEC server E m ; f n,q,m (t) represents the computing power allocated to the original task m in the MEC server E ; Therefore, the meaning of C4 is: in the MEC server E m the allocated computing resources cannot exceed its total computing resources;
[0221] In C5, p n,m (t) represents the transmit power of the IoT device D n to the MEC server E m ; P n (t) represents the maximum power of the IoT device D n ; The meaning of C5 is: the transmit power of the IoT device D n to the MEC server E m cannot exceed the maximum power of the IoT device D n ;
[0222] C6 represents the value range of the number of IoT devices, the number of MEC servers, and the number of original tasks.
[0223] C7 represents the value range of the number of task replications and the value range of the node types for task execution, i.e., execution by local IoT devices, execution by MEC servers, or execution by the cloud computing center r.
[0224] Step S3, transform the optimization model of task replication and offloading into a game theory model.
[0225] Regarding the solution problem of the above P1 model, its essence is to find appropriate x, y, and z such that P1 can obtain the minimum value. The present invention proposes a solution based on multi-party game theory. In this solution, N IoT devices D1, D2,..., D N are used as players in the game theory; for player D n , that is, IoT device D n , its task replication and offloading strategy is the strategy in the game theory, denoted as π n ={x n , y n , z n}; where x n represents the original task replication strategy of IoT device D n , y n represents the offloading strategy of the original task, and z n represents the offloading strategy of the replicated task; the player (i.e., D n ) can select its own strategy according to their respective goals; the present invention uses the maximization of the service performance of IoT devices as an index to represent the payoffs in the game theory.
[0226] Therefore, when IoT device D n adopts the task replication and offloading strategy π n , its service cost u n (π n ) is used as the payoff in the game theory, and thus a game theory model is established For this game theory model, finding the Nash Equilibrium of the model represents finding π n ={x n , y n , z n}.
[0227] Step S4, adopt a Nash equilibrium solving method based on multi-agent reinforcement learning to solve the game theory model, and obtain the optimal task replication and offloading strategy for each IoT device.
[0228] In the mobile edge computing environment, factors such as the dynamic variability of the computing resources of IoT devices and the dynamic variability of the edge network environment lead to obvious deficiencies in the solution methods of traditional game theory models in capturing these dynamics. The present invention proposes a method based on multi-agent reinforcement learning (MARL) to solve the Nash equilibrium of the game theory model.
[0229] Specifically, MARL first transforms the game theory model into a multi-agent partially observable Markov decision process (POMDP). In the POMDP, the game theory players (i.e., the aforementioned IoT devices) are regarded as agents. The agents make decisions on task replication and offloading based on their own observations of the environment. Here, the environment refers to the network delay state, the task execution state, and the device computing resources. The state is determined by the actions of this agent and other agents because the actions of the agents will change the current state of the system and cause the system to enter the next state. Secondly, the present invention proposes a multi-agent twin delayed deep deterministic policy gradient algorithm (MATD3) to interact with the environment to obtain the optimal offloading and replication strategies. The specific method is described as follows:
[0230] Step S4 is specifically implemented through steps S4.1 to S4.2:
[0231] Step S4.1, transform the game theory model into a multi-agent partially observable Markov decision process POMDP;
[0232] Specifically, each player D in the game theory model n is regarded as agent D n ; The POMDP includes five parts, denoted as
[0233] represents the state space, representing the global state of the entire system; it should be noted that in the scenario targeted by the present invention, the agent cannot obtain the global state of the system.
[0234] represents the observation space, representing the system state that the agent can observe, which is a part of the global state; at time t, the observation space of the agent where agent Dn Observation state L n (t) represents the position information of agent D r and f n (t) represents the computing and processing ability of agent D n ; represents agent D n to generate the number of original tasks and the size of the original task input data, Σ ε represents the global information of the MEC server, Σ r represents the global information of the cloud computing center;
[0235] represents the action space, which contains all possible actions of all agents, that is where a n (t) represents the action taken by agent D n at time t; x n (t), y n (t), z n (t) respectively represent the actions taken by agent D n at time t, including the original task replication strategy, the offloading strategy of the original task, and the offloading strategy of the replicated task;
[0236] represents the observation state transition probability function, which represents the likelihood of receiving a specific observation given the current state and the action taken;
[0237] represents the reward space, expressed as R n (t) = u n (t); where R n (t) represents the reward obtained by agent D n at time t; u n (t) represents the service cost of agent D n at time t;
[0238] Based on the five parts included in POMDP The problem of multi-task replication and offloading is expressed as:
[0239] Objective:
[0240]
[0241] where: t represents the moment t, and T represents the time interval; γ(t) represents the attenuation coefficient, where 0 ≤ γ(t) ≤ 1, representing the impact of the reward value at a future moment on the current reward value; R n (t) represents the reward received by the agent D n at moment t; in this objective function Objective, the goal of the task replication and offloading strategy is to find appropriate x, y, and z such that the cumulative discounted reward value of all agents within the time interval T is maximized; represents the optimal task replication and offloading strategy found through optimization, which are respectively the optimal original task replication strategy, the offloading strategy for the original task, and the offloading strategy for the replicated task of IoT device D n ;
[0242] Step S4.2, use the MATD3-based task replication and offloading Nash equilibrium solution method to solve the objective function Objective:
[0243] The present invention proposes a MATD3-based task replication and offloading Nash equilibrium solution method, and the overall architecture diagram of the method is as shown in Figure 3 shown.
[0244] Specifically, each IoT device D n is an agent D n , and the agent D n has a MATD3 controller. Each agent D n interacts with the environment by executing its own action a n (t) to obtain the corresponding reward R n (t) and observe the next state. Figure 3 shows the algorithm design details of one of the IoT devices. The algorithm includes an experience replay pool a critic network, and an actor network;
[0245] The experience replay pool is used to store existing experience samples, represented as where, represents the current observed state of the agent D n , a n (t) represents the currently taken action, represents the currently obtained reward, represents the next observed state; the experience replay pool Stored experience samples are used to train the critic network and the actor network, improving the learning efficiency and performance of the algorithm.
[0246] The actor network includes an Evaluation Network and a Target Network; the Evaluation Network is used to receive the agent D n The current observation state Through the policy network Generate action a n (t); this action is continuously improved through the gradient update optimizer to enhance the effect of the policy; the Target Network adopts the policy For the next observation state Add noise to generate the target action And synchronize the parameters from the Evaluation Network through soft update to ensure the stability and consistency of the policy network Through the experience samples in the experience replay pool D, the actor network continuously learns and optimizes the policy to achieve efficient decision-making for edge computing tasks;
[0247] The critic network includes two independent Evaluation Networks and two independent Target Networks for more stable and efficient policy learning. The parameters of the two Evaluation Networks are respectively And Optimized through the gradient descent algorithm with the goal of minimizing the loss function This loss function measures the difference between the Q value output by each Evaluation Network and the target Q value; each Evaluation Network receives the agent D n The current observation state And the action a taken n (t), and calculates the target Q value respectively, that is: And Update its parameters through the gradient descent optimizer And The two Target Networks respectively generate the target Q values: And Among them, And Respectively represent the parameters of the two Target Networks, and synchronize the parameters of the Evaluation Network through soft update to ensure stability.
[0248] The task replication and offloading algorithm based on MATD3 uses MARL to optimize the task allocation among multiple devices and obtain the Nash equilibrium in game theory. The algorithm first initializes the actor network and the critic network, and uses an experience replay pool to store experience tuples. During training, each player will choose actions that increase exploration noise to encourage diverse experiences. The environment will execute these actions and return the next observation, reward, and completion signal. The player updates its critic network by minimizing the mean squared error between the predicted Q-value and the target Q-value, using the minimum value of the target critic network during calculation to reduce overestimation bias. Using deterministic policy gradients, the update frequency of the actor network is reduced to ensure stable learning. The target network is updated less frequently to slowly track the learned network, thereby improving the stability of training. This method can efficiently make task replication and offloading decisions in a dynamic edge computing environment, thus improving the overall performance of the system.
[0249] The present invention proposes a method for task replication and offloading in a mobile edge computing environment, and the key points are as follows:
[0250] 1) Aiming at the task replication and offloading problem in the mobile edge computing scenario and with the goal of improving the overall service performance of the system, the present invention comprehensively considers factors such as transmission delay, computing delay, and system reliability, and proposes a new problem modeling method, which transforms this problem into a model for optimization solution under certain constraints (i.e., the above P1 model).
[0251] 2) The present invention proposes a new method for solving the model based on game theory. The above P1 model is transformed into a game theory model. The players in game theory are IoT devices, the strategies are task replication and offloading strategies, and the payoff is the overall service cost of the edge computing system. Therefore, obtaining the Nash equilibrium of the game theory model represents obtaining the optimal task replication and offloading strategy.
[0252] 3) The present invention proposes a new method for solving the Nash equilibrium of game theory. Aiming at the dynamic uncertainty problem existing in game theory, this method solves the Nash equilibrium based on multi-agent reinforcement learning technology. Specifically, this method transforms the game theory model into a partially observable Markov decision process, and designs a method based on the multi-agent double delayed deep deterministic policy gradient algorithm to obtain the Nash equilibrium.
[0253] The present invention proposes a method for task replication and offloading in a mobile edge computing environment, and it has the following effects:
[0254] 1) This method can support the edge computing system to provide strong fault tolerance when facing network failures, device failures, insufficient computing resources or other unpredictable failures, enabling edge computing tasks to be successfully executed and completed.
[0255] 2) The task replication and offloading method proposed in the present invention can comprehensively consider three factors, namely, the computing delay, the transmission delay and the system fault tolerance of the task, to form the overall service performance of the system, and based on the proposed modeling method and model solving method, the overall service performance of the edge computing system can reach the optimal.
[0256] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A task replication and offloading method based on game theory in mobile edge computing, characterized in that Including the following steps: Step S1, construct a mobile edge system architecture; the mobile edge system architecture includes N IoT devices D1, D2, …, D N formed set M MEC servers E1, E2, …, E M formed set ε = {E1, E2, …, E M}, and a cloud computing center r; where and respectively represent the set {1, 2, …, N} and the set {1, 2, …, M}; Step S2, within the period T, taking the task replication and offloading strategy of each IoT device as the problem to be solved. Among them, the task replication and offloading strategy of each IoT device includes the original task replication strategy x, the original task offloading strategy y, and the replicated task offloading strategy z; taking the maximization of the overall service performance of the mobile edge system in the period T as the objective function, combined with the constraint conditions, an optimization model for task replication and offloading is constructed; Step S2 is specifically as follows: Step S2.1, in the mobile edge environment, define the parameters and meanings related to task replication and offloading: Step S2.1.1, assume IoT device D n generates an original task at time t formed set where Q(n, t) represents IoT device D n the number of all original tasks generated at time t; Step S2.1.2, for IoT device D n The original task generated at time t whose input data volume and computational complexity are respectively and The number of copies it is replicated at time t is denoted as Therefore, generate copied tasks; where G represents the maximum number of copy limits, and the number of task copies Therefore, if the original task is not replicated, then Step S2.1.3, in a mobile edge environment, IoT device D n For each original task or each replicated task generated by the IoT device D at time t, it can only be executed on one node, that is: either on the local IoT device D n for execution, or offloaded to a certain MEC server E in the set ε of MEC servers m for execution, or offloaded to the cloud computing center r for execution; Therefore, define the original task for the decision result variable and the decision result variable for the replicated task x where i represents the node that executes the task, meaning: If the original task is executed on the local IoT device D n then at this time i = n, so If the replication task x is executed on the local IoT device D n at this time, i = n. Therefore If the original task is executed at the cloud computing center r, then at this time i = r. Therefore, If the replication task x is executed at the cloud computing center r, then at this time i = r. Therefore, If the original task is offloaded to MEC server E m is executed, then at this time i = m, so If the copy task x is offloaded to the MEC server E m is executed, then at this time i = m, so Since each original task or each replicated task x can only be executed on one node i, therefore and Step S2.2, evaluate to obtain the IoT device D n transmission delay at time t Step S2.3, evaluate to obtain the IoT device D n Execution delay at time t Step S2.4, evaluate and obtain the reliability γ n of the IoT device D n (t) at time t; Step S2.5, using formula (1), for IoT device D n The transmission delay at time t Execution delay And reliability γ n (t) are weighted and summed to obtain the service cost u n Of IoT device D at time t n (t): Among them: α1 and α2 are the weight parameters for weighted summation; Step S2.6, using formula (2), obtain the overall service cost U of the mobile edge system at time t g (t): Step S2.7, taking the minimization of the overall service cost of the mobile edge system in the period T, that is, maximizing the overall service performance as the objective, and combined with the constraint conditions, an optimization model for task replication and offloading is constructed; Step S3, transform the optimization model of task replication and offloading into a game theory model; Step S4, adopt a Nash equilibrium solving method based on multi-agent reinforcement learning to solve the game theory model, and obtain the optimal task replication and offloading strategy of each IoT device.
2. The task replication and offloading method based on game theory in mobile edge computing according to claim 1, wherein Step S2.2 is specifically as follows: Step S2.2.1, IoT device D n The original task generated at time t which is from IoT device D n to MEC server E m The transmission delay between them Is determined by formula (3): where: r n,m (t) represents the wireless link transmission rate between the IoT device D n and the MEC server E m which is determined by Equation (4): Where: W represents the link bandwidth; N0 is the noise power spectral density; h n,m (t) and p n,m (t) respectively represent the wireless channel gain and transmit power between the IoT device D n and the MEC server E m ; Step S2.2.2, IoT device D n The original task generated at time t Its transmission delay from the IoT device D n to the cloud computing center r is determined using Equation (5): Wherein: represents the transmission rate between the IoT device D n and the cloud computing center r; Step S2.2.3, using formula (6), obtain the total transmission delay at time t of all the original tasks of IoT device D n IoT device D n The total transmission delay of all original tasks of at time t is the transmission delay of IoT device D n at time t 3. A task replication and offloading method based on game theory in mobile edge computing according to claim 1, characterized in that Step S2.3 is specifically as follows: Step S2.3.1, if the IoT device D n generates an original task at the local IoT device D n and executes it, then the execution delay of its local execution is obtained using Equation (7): where: f n (t) represents the computing power of IoT device D n at time t; Step S2.3.2, if IoT device D n unloads the original task to MEC server E m for execution at time t, then the execution delay m of its execution on MEC server E is obtained using Equation (8): where: f n,q,m (t) represents the computing power m allocated by the MEC server E to the original task; Step S2.3.3, if IoT device D n offloads the original task to cloud computing center r for execution at time t, and its execution delay is calculated as 0; Step S2.3.4, using formula (9), obtain the IoT device D n The total execution delay of all original tasks at time t IoT device D n Total execution delay of all original tasks at time t That is, IoT device D n Execution delay at time t 4. A task replication and offloading method based on game theory in mobile edge computing according to claim 1, characterized in that Step S2.4 is specifically as follows: Step S2.4.1, MEC server E m has a failure rate represented by δ m (t). The failure rate follows a Poisson distribution. Therefore, at time t, the probability that MEC server E m has r m failures The probability that MEC server E m does not fail, i.e., r m = 0, is Therefore, MEC server E m The probability of failure at time t Therefore, for IoT device D n the original task at time t the cumulative failure probability at time t is expressed as Step S2.4.2, IoT device D n The original task at time t Reliability Is determined by formula (10): Step S2.4.3, IoT device D n The reliability Υ n (t) of all the original tasks at time t is determined by Equation (11): IoT device D n Reliability γ of all original tasks at time t n (t), which is the reliability γ n of IoT device D at time t n (t).
5. A task replication and offloading method based on game theory in mobile edge computing according to claim 1, characterized in that, In Step S2.7, the optimization model of task replication and offloading is: Objective function P1: The constraint conditions C1 - C7 are respectively: C1: C2: C3: C4: C5: 0 ≤ p n,m (t) ≤ P n (t) C6: C7: Among them: C1 represents the range of values; C2 represents and are both binary variables; C3 represents any one of the original tasks which can only be executed on one node, i.e., the local IoT device D n execute, the cloud computing center r executes, or the MEC server E m execute, i.e., the original task cannot be split and executed on multiple nodes; In C4, F m (t) represents the overall computing power of MEC server E m ; f n,q,m (t) represents the computing power allocated to the original task in MEC server E m ; therefore, the meaning of C4 is: in MEC server E , the allocated computing resources cannot exceed its total computing resources; m In C5, p n,m (t) represents the transmission power of IoT device D n to MEC server E m ; P n (t) represents the maximum power of IoT device D n ; The meaning of C5 is: the transmission power of IoT device D n to MEC server E m shall not exceed the maximum power of IoT device D n ; C6 represents the value range of the number of IoT devices, the number of MEC servers, and the number of original tasks; C7 represents the value range of the number of task replications and the value range of the node types for task execution.
6. The task replication and offloading method based on game theory in mobile edge computing according to claim 1, wherein Step S3 is specifically as follows: Take N IoT devices D1, D2, …, D N as players in game theory; for player D n , that is, IoT device D n , its task replication and offloading strategy is represented as π n = {x n , y n , z n}; where x n represents the original task replication strategy of IoT device D n , y n represents the offloading strategy of the original task, and z n represents the offloading strategy of the replicated task; When IoT device D n adopts the task replication and offloading strategy π n its service cost u n (π n ) as the payoff in game theory, and thus a game theory model is established 7. A task replication and offloading method based on game theory in mobile edge computing according to claim 6, characterized in that, Step S4 is specifically as follows: Step S4.1, transform the game theory model into a multi-agent partially observable Markov decision process POMDP; Specifically, each player D in the game theory model n is regarded as an agent D n ; The POMDP consists of five parts, denoted as Denote the state space, representing the global state of the entire system; Represents the observation space, which represents the system state observable by the agent and is part of the global state; at time t, the observation space of the agent Among them, agent D n The observed state L n (t) represents the position information of agent D n The position information, f n (t) represents the computing and processing ability of agent D n The computing and processing ability, Represents the number of original tasks generated by agent D n And the size of the original task input data, Σ ε Represents the global information of the MEC server, Σ r Represents the global information of the cloud computing center; Denotes the action space, which contains all possible actions of all agents, that is a n (t) = {x n (t), y n (t), z n (t)}; where a n (t) represents the action taken by agent D n at time t; x n (t), y n (t), z n (t) respectively represent the actions taken by agent D n at time t, including the original task replication strategy, the offloading strategy of the original task, and the offloading strategy of the replicated task; Represents the observation state transition probability function, which represents the likelihood of receiving a particular observation given the current state and the action taken; Represents the reward space, denoted as R n (t) = u n (t); where R n (t) represents the reward obtained by agent D n at time t; u n (t) represents the service cost of agent D n at time t; Based on the five parts included in the POMDP The problem of multi-task replication and offloading is expressed as: Objective: where: t represents the moment t, and T represents the time interval; γ(t) represents the attenuation coefficient, where 0 ≤ γ(t) ≤ 1, representing the influence of the reward value at a future moment on the current reward value; R n (t) represents the reward received by the agent D n at the moment t; in this objective function Objective, the goal of the task replication and offloading strategy is to find appropriate x, y, and z such that the cumulative discounted reward value of all agents within the time interval T is maximized; represents the optimal task replication and offloading strategy found through optimization, which are respectively the optimal original task replication strategy, the offloading strategy of the original task, and the offloading strategy of the replicated task for the IoT device D n ; Step S4.2, adopt a task replication and offloading Nash equilibrium solving method based on MATD3 to solve the objective function Objective: Specifically, each agent D n has a MATD3 controller. Each agent D n interacts with the environment by executing its own action a n (t) to obtain the corresponding reward R n (t) and observes the next state; the algorithm includes an experience replay pool a critic network and an actor network; Experience replay pool For storing existing experience samples, represented as ( a n (t), ), where, represents the current observation state of the agent D n a n (t) represents the action taken currently, represents the reward obtained currently, represents the state observed next; The experience replay pool stores experience samples for training the critic network and the actor network; The actor network includes an Evaluation Network and a Target Network; the Evaluation Network is used to receive the current observation state of the agent D n Current observation state Through the policy network Generate action a n (t); this action is continuously improved through a gradient update optimizer; the Target Network adopts a policy For the next observation state Add noise to generate a target action And synchronize parameters from the Evaluation Network through soft update to ensure the stability and consistency of the policy network Through the experience samples in the experience replay pool The actor network continuously learns and optimizes the policy to achieve efficient decision-making for edge computing tasks; The critic network, including two independent evaluation networks and two independent target networks, with the parameters of the two evaluation networks being and are optimized by the gradient descent algorithm with the goal of minimizing the loss function This loss function measures the difference between the Q-values output by each evaluation network and the target Q-values; each evaluation network receives the current observation state n of the agent D and the action a n (t) taken, and calculates the target Q-values respectively, that is: and and updates its parameters through the gradient descent optimizer and The two target networks respectively generate the target Q-values: and where and represent the parameters of the two target networks respectively, and synchronize the parameters of the evaluation network through soft update.