Random task-oriented intelligent edge collaborative computing unloading method
By constructing the MEC system model and combining Lyapunov optimization theory and MADRL algorithm, the LyMADRL algorithm framework is designed, which solves the problems of mutual influence of computing requests and resources in the MEC system, and the unloading strategy and service cache strategy are highly coupled, and the computational performance stability and effective satisfaction of resource constraints in the MEC system are achieved.
Patent Information
- Application Number
- CN202510339790.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-10
AI Technical Summary
In the large-scale MEC scenario where diversified computing requests burst randomly, the MEC system faces the problem of mutual influence between computing requests and required resources, as well as the high coupling of unloading policies and service cache policies in time and space dimensions, resulting in unstable computing performance.
By building a MEC system model, including users, MEC servers, base stations and remote central clouds, combined with Lyapunov optimization theory and distributed multi-agent reinforcement learning (MADRL) algorithm, it is transformed into a single time slot deterministic optimization problem, and the LyMADRL algorithm framework is designed, and the CTDE framework is used to adaptively explore the optimal computing offload strategy in the centralized training stage, and the distributed decisions of computing offloading are realized in the distributed execution stage.
It realizes that while ensuring the stability of the computing performance of the MEC system, meets the computing and storage resource constraints, formulates a computing offload and service cache scheme for distributed intelligent decision-making, and verifies the performance of the LyMADRL algorithm through simulation experiments, showing the optimal system overhead and task queue stability.
Smart Images

Figure CN120128989A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of mobile communication technology, and in particular relates to an intelligent edge collaborative computing offloading method for random tasks. Background Art
[0002] Mobile Edge Computing (MEC) is a technology that pushes computing and storage resources from traditional cloud computing centers to the edge of the network. With the development of artificial intelligence and machine learning technologies, IoT devices are showing an explosive growth trend. Edge nodes with limited computing and storage resources will face massive computing requests with high random dynamics, burstiness, diversity and unpredictability, which will lead to problems such as unstable computing performance of MEC systems. In addition, since the calculation of diversified applications usually requires caching data related to the application type (such as machine learning models, databases, etc.), in order to effectively utilize the limited storage resources of edge nodes, caching popular services in edge servers in advance is an effective means. However, in large-scale MEC scenarios with random bursts of diversified computing requests, MEC systems will face challenges such as the mutual influence of computing requests and required resources, and the high coupling of offloading strategies and service caching strategies in time and space dimensions.
[0003] The distributed deployment characteristics of edge nodes, the high coupling of offloading strategies and service caching strategies, and the random arrival of diversified computing requests have significantly increased the difficulty of formulating computing offloading strategies and edge node service caching. Therefore, how to formulate computing offloading and service caching solutions for distributed intelligent decision-making under the constraints of ensuring the computing performance stability of the MEC system and satisfying computing and storage resources is still a hot research topic that needs to be solved in the MEC field. Domestic and foreign researchers have conducted research on this issue and proposed two types of methods. The first type is a computing offloading strategy based on deep reinforcement learning. Although this type of method can adapt to dynamic environments, it usually requires a large amount of training data and computing resources, and may have insufficient generalization capabilities in practical applications. The second type is an offloading method based on graph neural networks. This type of method can handle network topology structures, but it has problems such as high computational complexity and slow convergence when dealing with large-scale networks. Therefore, how to formulate computing offloading and service caching solutions for distributed intelligent decision-making under the constraints of ensuring the computing performance stability of the MEC system and satisfying computing and storage resources is still a hot research topic that needs to be solved in the MEC field. Summary of the invention
[0004] To solve the above problems, the present invention provides an intelligent edge collaborative computing offloading method for random tasks, comprising the following steps:
[0005] S1. Build a MEC system model, which includes users, MEC servers, base stations, and remote center clouds;
[0006] S2. Based on the MEC system model, construct the random task arrival model, service cache model, local task queue model, collaborative task queue model and overhead model;
[0007] S3. Construct a problem model with the optimization goal of minimizing the long-term average time overhead of the system under the long-term stability constraint of the task queue;
[0008] S4. The problem model is transformed into a single-slot deterministic optimization problem through Lyapunov optimization theory;
[0009] S5. Convert the single-slot deterministic optimization problem into a Dec-POMDP problem, build the LyMADRL algorithm framework, and obtain the environment state, local observation, action, and reward function;
[0010] S6. Based on the LyMADRL algorithm framework, the LyMADRL algorithm based on the CTDE framework is used to solve the optimal strategy.
[0011] Beneficial effects of the present invention:
[0012] The present invention comprehensively considers the randomness and diversity of user computing requests in a multi-user multi-edge node MEC system, as well as the dynamic coupling relationship between offloading strategy and service cache decision, designs a general multi-task queue model, and proposes a random optimization problem of cloud-edge collaborative computing offloading and service cache; then, the original random optimization problem is transformed into a single time slot deterministic optimization problem based on Lyapunov optimization, and a MADRL algorithm based on Lyapunov optimization is proposed; the LyMADRL algorithm adopts the CTDE framework. In the centralized training phase, each edge node adaptively explores and learns the optimal computing offloading strategy with the goal of ensuring the stability of system computing performance and minimizing system overhead. In the distributed execution phase, the edge node only needs to implement distributed decision-making for computing offloading based on its local state information. Finally, the performance of the LyMADRL algorithm is verified by simulation experiments. The simulation results show that compared with the existing algorithms, the LyMADRL algorithm has the optimal system overhead and can ensure the stability of the system task queue. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a system model diagram of the present invention;
[0014] Figure 2 It is the local task queue diagram of the present invention;
[0015] Figure 3 It is the collaborative task queue diagram of the present invention;
[0016] Figure 4 It is the LyMADRL algorithm framework of the present invention;
[0017] Figure 5 It is the convergence process curve of different algorithms of the present invention;
[0018] Figure 6 This is the influence curve of the Lyapunov control parameters of the present invention on the performance of different algorithms. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] The present invention provides an intelligent edge collaborative computing unloading method for random tasks, comprising the following steps:
[0021] S1. Build a MEC system model, which includes users, MEC servers, base stations, and remote center clouds.
[0022] Specifically, Figure 1 As shown in Figure 2, the MEC system model includes N base stations and 1 remote center cloud, where the base station set is defined Base stations are connected through wired links, and each base station is connected to the remote central cloud through the core network; at the same time, each base station is equipped with a MEC server with storage and computing capabilities to provide computing services for users associated with it; definition Indicates base station The associated user collection, where Indicates the number of users associated with base station n; in the MEC system model, time is discretized into a set of time slots of equal length The length of each time slot is τ, and it is assumed that each user generates one random computation request in each time slot.
[0023] S2. Based on the MEC system model, we build a random task arrival model, a service cache model, a local task queue model, a collaborative task queue model, and an overhead model.
[0024] Specifically, the random task arrival model is used to model the random task arrival situation; assuming that there are K types of computing services in the MEC system model, let the service type set Define that each user randomly selects one of K computing service types in each time slot to generate a random computing request. Assuming that the process of the user's computing task (random computing request) arriving at its associated local base station follows independent and identical distribution, then:
[0025]
[0026] in, represents the total number of tasks belonging to computing service type k arriving at base station n in time slot t, Indicates time slot t user m n The amount of random computing requests arriving at base station n, represents time slot t, user m n The minimum and maximum amount of tasks for random computing requests to reach base station n; 1 {·} represents the indicator function. If the event {·} is true, then 1 {·} =1, otherwise 1 {·} =0; Indicates base station The associated user collection, where represents the number of users associated with base station n; Indicates time slot t user m n The computing service type for random computing requests.
[0027] In the present invention, computing tasks of all computing service types in the MEC system model adopt a partial offloading mode, that is, the computing tasks can be divided into subtasks of any size, and each subtask can be collaboratively and parallelly calculated by the local base station, neighboring base stations and the remote central cloud.
[0028] Specifically, the service cache model is used to model cache decision variables. Processing a computing task with a specific computing service type requires the processor to configure data related to the specific computing service type, such as program code and database. Since the storage resources of a single base station are very limited, only some computing service types can be cached at the same time. Define binary variables represents the cache decision of base station n for computing service type k in time slot t. If computing service type k is cached in base station n, then otherwise The present invention assumes that the computing service type cached during time slot t can be used to process the corresponding computing task at the start time of time slot t+1. Considering that the computing service type cached by each base station cannot exceed its storage capacity, the following cache capacity constraints are imposed:
[0029]
[0030] Among them, z k represents the storage space required for cache computing service type k, Zn Represents the storage capacity of base station n.
[0031] Specifically, the local task queue model is used to model the local task queue variables, where the local task queue is used to receive random computing requests from users associated with the local base station. Based on the amount of arriving tasks and the amount of leaving tasks, the dynamic update equation of the local task queue of base station n belonging to computing service type k in each time slot is expressed as follows:
[0032]
[0033] in, Indicates the local task queue of base station n belonging to computing service type k in time slot t+1, that is, the total number of tasks in the local task queue of this time slot; Indicates that time slot t, base station n belongs to the local task queue of computing service type k; represents the number of tasks leaving the local task queue of base station n belonging to computing service type k in time slot t, represents the number of tasks arriving at the local task queue of base station n belonging to computing service type k at time slot t. According to the random task arrival model, the number of tasks arriving at the local task queue of base station n belonging to computing service type k at time slot t is In particular, the present invention Both refer to the number of tasks at the beginning of time slot t+1. No tasks will arrive or leave at the beginning.
[0034] The base station needs to decide how to handle the tasks in the local task queue and how much tasks to handle. Figure 2 In the diagram on the left, when base station n caches computing service type k in time slot t, the computing task can be executed at the local base station n, or the local base station n can forward the computing task to other base stations or remote center cloud for collaborative computing; Figure 2 As shown in the diagram on the right, when base station n does not cache computing service type k in time slot t, the computing task will be forwarded by base station n to other base stations or remote center cloud for collaborative computing. Therefore, the local task queue is defined Uninstall decision express The computing tasks in the process are outsourced to the remote central cloud for execution; express The computation tasks in are executed at the local base station n. express The computing tasks in are forwarded to other base stations n' for execution.
[0035] Departure tasks is the amount of tasks processed by the local task queue in time slot t. In order to avoid the computing task not being processed for a long time after entering the local task queue, and considering the The amount of tasks cannot exceed the length of the local task queue and the maximum amount of tasks to be processed, so The following conditions should be met:
[0036]
[0037] and They respectively represent the minimum and maximum task amounts of the local task queue of computing service type k leaving base station n in each time slot.
[0038] Specifically, the collaborative task queue model is used to model collaborative task queue variables, where the collaborative task queue will receive tasks forwarded from other base stations; the dynamic update equation of the collaborative task queue of base station n belonging to computing service type k in each time slot is expressed as follows:
[0039]
[0040] in, Indicates that in time slot t+1, base station n belongs to the collaborative task queue of computing service type k, that is, the total number of tasks in the collaborative task queue of this time slot; Indicates that time slot t, base station n belongs to the collaborative task queue of computing service type k; represents the number of tasks leaving the collaborative task queue of base station n belonging to computing service type k in time slot t, represents the amount of arriving tasks of the collaborative task queue of base station n belonging to computing service type k in time slot t; in particular, the present invention proposes Both refer to the number of tasks at the beginning of time slot t+1. No tasks will arrive or leave at the beginning.
[0041] According to the above definition, we can get the time slot t arrival The amount of tasks is:
[0042]
[0043] Similarly, the base station needs to decide how to handle the workload in the collaborative task queue and how much workload to handle. Figure 3 In the diagram on the left, when base station n caches computing service type k in time slot t, the task can be executed in the local base station or through collaborative computing in the remote center cloud, such as Figure 3 In the diagram on the right, when base station n does not cache computing service type k in time slot t, the task will be forwarded by base station n to the remote center cloud for collaborative computing. To this end, a collaborative task queue is defined Task offloading decision in express Outsourcing tasks to remote cloud centers for execution; express The tasks in are executed at the local base station. In addition, the number of tasks leaving is the amount of tasks processed by the collaborative task queue in time slot t, which satisfies the following conditions:
[0044]
[0045] In the formula, and They respectively represent the minimum and maximum task amounts of the collaborative task queue of computing service type k leaving base station n in each time slot.
[0046] In order to ensure the reliability of the computing performance of the MEC system, it is necessary to consider the long-term stability of the task queue. That is, the task queue should not increase indefinitely, otherwise the computing tasks will wait too long in the task queue. Therefore, the local task queue of the MEC system and collaborative task queues The following constraints should be satisfied in terms of time average:
[0047]
[0048] In the formula, and They respectively represent the expected lengths of the local task queue and the collaborative task queue of base station n belonging to computing service type k.
[0049] Specifically, in a MEC network with distributed resource deployment, the communication, computing and storage resources of edge nodes are usually very limited, and their business carrying capacity, computing power and storage capacity are far inferior to those of the resource-rich remote center cloud. At the same time, edge nodes are faced with a large number of random computing requests for diverse computing-intensive and delay-sensitive tasks, which leads to problems such as high computing costs and unstable computing performance of the MEC system. Therefore, the present invention establishes a system overhead model for base stations to guide the formulation of computing offloading strategies in scenarios where diverse random tasks arrive. The system overhead mainly includes task computing overhead and cloud-edge collaboration overhead, which are defined as follows:
[0050] (1) Task computing overhead. When the tasks in the queue are at the base station, the base station needs to spend computing resources to process the tasks executed at the local base station. To this end:
[0051] Base station n processes the local task queue in time slot t The computational overhead of the task for
[0052]
[0053] Base station n processes the collaborative task queue in time slot t The computational overhead of the task for
[0054]
[0055] Indicates the unit task calculation overhead weight coefficient,
[0056] (2) Cloud-edge collaboration overhead. For any base station, if the corresponding computing service type is not cached in the current time slot or the computing resources are insufficient, the base station will spend a certain cost to migrate the tasks in the queue to nearby base stations or remote central clouds for execution. The amount of tasks that base station n migrates to other base stations in time slot t is The amount of tasks outsourced to the remote center cloud is Accordingly:
[0057] Base station n processes the local task queue in time slot t Cloud-edge collaboration overhead for tasks in for
[0058]
[0059] Base station n processes the collaborative task queue in time slot t Cloud-edge collaboration overhead for tasks in for
[0060]
[0061] represents the weight coefficient of the edge-to-edge coordination overhead of a unit task, Represents the weight coefficient of the cloud-edge collaboration overhead per task.
[0062] Define the sum of task computing overhead and cloud-edge collaboration overhead as the overhead model of the MEC system model, then the total system overhead C(t) at time slot t is expressed as:
[0063]
[0064] S3. Construct a problem model with the optimization goal of minimizing the long-term average time overhead of the system under the long-term stability constraint of the task queue.
[0065] Specifically, the optimization goal of the present invention is to minimize the long-term average time overhead of the system by optimizing the calculation offloading and service caching strategy under the constraint of ensuring the long-term stability of the MEC system task queue. The problem model is modeled as follows:
[0066]
[0067] in, Indicates the base station service cache decision; Indicates the task offloading strategy; and They represent the processing task volume strategies of the local task queue and the collaborative task queue respectively. Constraint C6 indicates that the computing service type cached by any base station n cannot exceed its maximum storage capacity; Constraint C7 represents the long-term stability constraint of the task queue of the MEC system.
[0068] S4. The problem model is transformed through Lyapunov optimization theory to obtain a single-slot deterministic optimization problem.
[0069] Specifically, the problem model constructed by the present invention is a sequential decision-making problem. The base station needs to formulate the optimal service caching decision and task offloading decision for each time slot to minimize the system overhead. However, the strong coupling characteristics between the service caching decision and the task offloading decision increase the difficulty of policy formulation. In addition, solving the problem model requires complete task arrival information, but the arrival of diversified tasks is highly random and unpredictable, and it is difficult and impractical to predict the task arrival information in advance. At the same time, in the MEC system where resources are highly dispersed and limited, the use of traditional centralized optimization algorithms will face frequent information interaction between base stations, which will make it difficult to obtain real-time status information of the system, and it is difficult to effectively solve the problem model while meeting the system constraints.
[0070] Lyapunov optimization theory is widely used in the task queue control problem of MEC system for random task arrival. It aims to optimize the predefined target problem while ensuring the stability constraint of the task queue. The core idea is to transform the original random optimization problem into a deterministic optimization problem of a single time slot. Based on this, in order to solve the above challenges, the present invention adopts Lyapunov optimization theory and distributed multi-agent reinforcement learning to solve the problem model. The present invention proposes a MADRL algorithm based on Lyapunov optimization (Lyapunov Optimizationbased Multi-Agent DeepReinforcement Learning, LyMADRL) to solve the problem model. First, the problem model is transformed into a single time slot deterministic optimization problem based on Lyapunov optimization; the algorithm framework and detailed principle of LyMADRL are given later.
[0071] The process of converting the problem model into a single time-slot deterministic optimization problem based on Lyapunov optimization is as follows:
[0072] Based on the problem model constructed by the present invention, define represents the length of all task queues of the MEC system at time slot t, based on which the Lyapunov function L(Θ(t)) is defined as follows:
[0073]
[0074] Where, the Lyapunov function L(Θ(t)) represents the measure of the length of all task queues. Next, in order to describe the degree of change of task queues in adjacent time slots, the conditional Lyapunov drift Δ(Θ(t)) is defined as follows:
[0075]
[0076] Express expectations.
[0077] It can be seen that minimizing Δ(Θ(t)) can ensure the stability of the MEC system task queue. In order to minimize the system overhead while ensuring the stability of the task queue, the present invention combines Δ(Θ(t)) with the MEC system overhead and defines the Lyapunov drift plus penalty function Δ V (Θ(t)) is as follows:
[0078]
[0079] In the formula, Π k,Q (t) and Π k,Y (t) represent the Lyapunov dynamic control parameters that weigh the stability of the local task queue and the collaborative task queue against the system overhead, respectively, and are defined as:
[0080]
[0081] Where V k represents the Lyapunov static control parameter. The dynamic control parameter Π k,Q (t) and Π k,Y The purpose of (t) is to control the task queues of the same service type to maintain the same queue backlog upper bound during the evolution process. It should be noted that the present invention sets corresponding Lyapunov control parameters for diversified service types to meet the queue stability requirements of different service types.
[0082] It can be seen that by minimizing Δ V (Θ(t)) can minimize the MEC system overhead while ensuring the stability of the task queue. According to the Lyapunov optimization theory, Δ V (Θ(t)) and solve the problem by minimizing this upper bound. Specifically, Δ V The upper bound of (Θ(t)) can be obtained by the following Theorem 1:
[0083] Theorem 1:
[0084] Drift plus penalty function Δ V The upper bound of (Θ(t)) is:
[0085]
[0086] In the formula, B represents a constant, and the specific calculation is as follows:
[0087]
[0088] In the formula, Indicates user m n The maximum average task arrival rate. It can be seen that minimizing Δ V The upper bound of (Θ(t)) can achieve task queue stability while minimizing system overhead. According to Theorem 1 and combined with the Opportunistic Expectation Minimization (OEM) technique, the problem model can be transformed into a deterministic optimization problem for a single time slot:
[0089]
[0090] S5. Convert the single-slot deterministic optimization problem into a Dec-POMDP problem, build the LyMADRL algorithm framework, and obtain the environment state, local observation, action, and reward function.
[0091] Specifically, solving the above-mentioned single-time-slot deterministic optimization problem does not require knowing the future information of the arrival of random tasks, and obtaining the optimal task offloading and service caching strategy only requires knowing the length of the current time slot task queue. Despite this, the use of traditional centralized optimization algorithms to solve single-time-slot deterministic optimization problems still faces severe challenges. Traditional centralized optimization algorithms, such as convex optimization, cannot effectively guarantee the convergence of algorithms in highly dynamic and distributed densely deployed MEC network environments. To solve this problem, the present invention proposes a distributed intelligent random task offloading algorithm.
[0092] In the single-slot deterministic optimization problem of the present invention, the service cache state, local task queue state and collaborative task queue state of the base station in time slot t+1 are only related to the state of time slot t, but not to the state before time slot t, that is, the computation offloading process has a Markov property. Therefore, the present invention transforms the single-slot deterministic optimization problem into a Dec-POMDP problem, where each base station represents an agent, and there are N agents in total. Dec-POMDP can be described as a tuple Where S represents the environmental state of all agents; O n represents the local observation space; A n represents the action space of the nth agent; R represents the reward function; γ∈[0,1) represents the discount factor. Specifically, the environment state, local observation, action and reward function of Dec-POMDP are defined as follows:
[0093] (1) Environmental state. According to the above system model, the environmental state includes the current service cache status of all base stations in the MEC system, as well as the status of all local task queues and collaborative task queues. Therefore, the environmental state s(t)∈S at time slot t is defined as
[0094]
[0095] (2) Local observation. In a distributed partially observable MEC network, each base station / agent only needs to observe its own status, including the local service cache status, and the local task queue and collaborative task queue status of the local BS. Therefore, agent n locally observes o in time slot t. n (t)∈O n Defined as
[0096]
[0097] (3) Action. According to the optimization problem and the Dec-POMDP process defined in the present invention, the action of each agent includes a service cache action and a task offloading action, that is, the action a of agent n in time slot t n (t)∈A n Defined as
[0098]
[0099] (4) Reward function. The optimization goal of the present invention is to ensure the stability of the task queue of the MEC system while minimizing the system overhead. In the LyMADRL algorithm, the task queue stability module ensures the stability of the task queue of the MEC system, and the system overhead is minimized by the agent through exploration and learning of task offloading strategies and service caching strategies. According to the single-slot deterministic optimization problem, the agents are in a completely cooperative relationship, that is, the common goal of each agent is to minimize the system overhead, so the reward function is defined as
[0100] r n (t)=-C(t)
[0101] Agent n receives a local observation o at time slot t n (t)∈O n , select action a n (t)∈A n , and at the same time n (t) is input to a task queue stabilization module based on Lyapunov optimization, which is based on the selected a n (t) Returns the optimal processing task volume of the local task queue and the collaborative task queue and The agent performs action an (t), and the post - environment returns a reward r according to R n (t), and transfers the local observation to the next state o n (t + 1). The ultimate goal of each agent is to explore and improve its policy μ n , such that the cumulative return is maximized.
[0102] The framework of the LyMADRL algorithm is as Figure 4 shown, where each base station / agent is configured with a DDPG network and a task queue stability module. The DDPG network is a deep reinforcement learning algorithm based on AC, which combines the advantages of policy gradient and DQN. The task queue stability module of each agent uses Lyapunov optimization theory to control the stability of its task queue. LyMADRL adopts the CTDE framework, and its advantage is that when any base station executes its respective task offloading and caching policies, it does not need to obtain the task queue information and service cache information of other base stations in the MEC system, and can obtain the globally optimal action only based on its local observation, thus solving challenges such as frequent information interaction between base stations and difficulty in obtaining global information during the execution stage.
[0103] S6. Based on the LyMADRL algorithm framework, the LyMADRL algorithm based on the CTDE framework is used to solve and obtain the optimal policy.
[0104] Specifically, the principle of the LyMADRL algorithm framework as Figure 4 shown is introduced in detail. In the centralized training stage (the process of the black solid line), at the current time slot, given the local observation o n of agent n, the actor network will take an action a n (o n ) according to the current policy μ n . Then, the agent inputs the action a n into its task queue stability module, and the task queue stability module outputs the optimal task processing amounts and
[0105] Specifically, the and output by the task queue stability module are obtained by solving the following deterministic optimization problem at the current time slot:
[0106]
[0107] In the above optimization problem, the agents are independent of each other, and the task queues of each agent are also independent of each other. At the same time, the objective function is about and The affine functions, and the constraint conditions are respectively convex constraints on and Therefore, the above optimization problem is a convex optimization problem. The present invention uses the Lagrange multiplier method to solve this optimization problem, and obtains the optimal processing task volume of the local task queue and the optimal processing task volume of the collaborative task queue, which are respectively as follows:
[0108]
[0109] Based on this, the optimal processing task volume and can be obtained. The task queue stability module obtains the optimal processing task volume based on the Lyapunov optimization theory, thus ensuring the stability of the task queue.
[0110] Next, the agent will and a n interact with the MEC environment to obtain the reward r n and transfer the local observation to the next state o n ′. All agents store the samples in the experience replay pool to form a shared sample. The experience replay pool contains the sequence (o, o′, a, r), where o = (o 1 , o 2 , …, o N ) represents the observation information of all agents; o’ represents the local observation of all agents in the next time slot; a = (a 1 , a 2 , …, a N ) represents the actions of all agents; r = (r 1 , r 2 , …, r N ) represents the rewards of all agents. The critic network samples from the experience replay pool and evaluates the selected actions through a centralized action-value function, which contains a and o.
[0111] Define and to represent the sets of the actor networks and critic networks of all agents respectively, and the corresponding parameters are and For the actor network, any agent n updates the policy by means of policy gradient, and the policy gradient is as follows:
[0112]
[0113] Where X represents the number of mini - batch samples randomly sampled from the experience replay pool; x represents the x - th sample; represents the centralized action - value function described by the parameter i.e., the critic network.
[0114] For the critic network, agent n updates the corresponding parameters by minimizing the loss function The loss function is expressed as follows:
[0115]
[0116] Where μ n ′ represents the target actor network described by the parameter and Q n ′ represents the target critic network described by the parameter respectively.
[0117] Finally, each agent updates the parameters of the target actor network and the target critic network in a soft - update manner, which are expressed as follows:
[0118]
[0119] Where ζ represents the update rate.
[0120] During the centralized training process, the LyMADRL algorithm can adaptively explore and learn the computing offloading strategy according to the current MEC system task queue status, service cache status, and other base station behaviors, so as to ensure the stability of the system task queue and minimize the system overhead. As Figure 4 shown, after the model training is completed, in the distributed execution stage (the red dotted - line process), each agent only needs to know its local observation and input it into the corresponding actor network to obtain the learned action.
[0121] The training process of the LyMADRL algorithm is shown in Table 1 below.
[0122] Table 1 Training process of the random task computing offloading algorithm based on LyMADRL
[0123]
[0124]
[0125] The time complexity of the LyMADRL algorithm in the training stage mainly depends on the number of agents N, the number of training rounds T e , the number of time steps per round T s, the number of small - batch samples \(X\), the number of computing service type \(K\), and the actor network and critic network structures of each agent. Assuming that both the actor network and the critic adopt fully - connected networks with the same structure, the time complexity during the LyMADRL training is calculated as follows:
[0126]
[0127] Among them, \(L\) a and \(L\) c respectively represent the number of layers of the actor network and the critic network; \(l\) a and \(l\) c respectively represent the \(l\) a -th layer of the actor network and the \(l\) c -th layer of the critic network; and respectively represent the number of neurons in the \(l\) a -th layer of the actor network and the number of neurons in the \(l_c\) - th layer of the critic network; \(O(NT\) e \(T\) s \(K)\) represents the time complexity of the task queue stability module in the training phase. In the distributed execution phase, each agent only executes actions based on the actor network. Therefore, its time complexity is
[0128] In one embodiment, the present invention conducts a simulation experiment to verify the performance of the proposed LyMADRL algorithm. The comparison algorithms include:
[0129] Greedy algorithm: The Greedy algorithm does not consider the stability of the MEC system task queue and always processes the tasks in the queue backlog based on the maximum processing task volume. Specifically, for the local task queue at the current time slot, the optimal processing task volume is For the cooperative task queue, the optimal processing task volume is
[0130] Edge - cloud cooperation only (ECC - only) algorithm that only considers cloud - edge cooperation for computing offloading and service caching: The ECC - only algorithm does not consider cooperative computing between BSs. For the computing tasks requested by users, if the current base station does not cache the corresponding service type or has insufficient computing power, the task is directly sent to the remote central cloud for execution through the local base station.
[0131] This simulation implements LyMADRL and all comparison algorithms using the Pytorch 1.13 deep learning framework and the Adam optimizer. All experiments were completed based on the Windows 10 operating system (CPU: Intel(R) Core(TM) i9-10920X CPU @ 3.5GHz, GPU: NVIDIA GeForce RTX 3090).
[0132] The MEC simulation scenario of the present invention includes 5 base stations. The total number of types of computing services in the network is K = 10, and the storage capacity size z required for each type of computing service k is randomly distributed between [10, 20] GB. Unless otherwise specified, the storage capacity of each base station is set to 100 GB. The length of each time slot is set to 5 minutes. It is assumed that the generation of computing tasks is based on user preferences, and a user generates only one type of computing service task within one time slot. Define φ n,k to represent the popularity of computing service type k within the coverage area of base station n. This popularity follows a Zipf distribution, which is specifically expressed as follows:
[0133]
[0134] In the formula, represents the popularity ranking of computing service type k in base station n, and l n,k The smaller the value of, the higher the popularity of service type k; δ n ≥0 represents the skewness coefficient. When δ n = 0, it means that users have equal preferences for requests for all types of computing services. The larger the skewness coefficient, the greater the preference of users for requests for computing service types with higher popularity.
[0135] To simulate the differentiated computing requests of users between different base stations, different popularity rankings are randomly generated for each base station in the simulation. For all base stations, the skewness coefficient δ n is set to 0.8 for all;
[0136] The minimum and maximum task processing amounts of each base station for each task queue in each time slot are set to 30 MB and 150 MB respectively; the weight coefficient of the unit task computing overhead is set to 0.001 $ / MB; the weight coefficient of the unit task edge-edge collaboration overhead is set to 0.01 $ / MB; the weight coefficient of the unit task cloud-edge collaboration overhead is set to 0.1 $ / MB.
[0137] In addition, unless otherwise specified, the number of associated users of each base station is set to 25; the average task arrival rate at which the computing tasks requested by each user arrive at their local base station Subject to a uniform distribution within [0.02, 0.08] MB / s, the Lyapunov control parameter V k is default set to 600.
[0138] For all deep reinforcement learning algorithms, the discount factor γ is set to 0.99; to reduce the time complexity of the algorithm, both the actor network and the critic network adopt two-layer fully connected neural networks, and the hidden layer dimension is set to 64; the learning rate of the actor network is 0.0001; the learning rate of the critic network is set to 0.001; the size of the experience replay pool is 5×10 5 ; the number of mini-batch samples is X = 256; the update rate is The number of training episodes is 1000; the number of time steps per episode is 500; the initial lengths of the local task queue and the cooperative task queue are 0; the exploration probability decreases from 1 to 0.05 within 5×10 5 time steps. The task queue backlog is defined as the average task queue length per time slot.
[0139] Figure 5 shows the convergence process of different algorithms as the number of training episodes increases. From Figure 5 (a), it can be seen that the system overhead of each algorithm decreases as the number of training episodes increases and gradually converges to a stable value. Compared with the other two groups of benchmark algorithms, the proposed LyMADRL algorithm in the present invention has the smallest system overhead. The Greedy algorithm does not consider the stability of the MEC task queue and greedily processes the tasks in the queue backlog based on the maximum task processing capacity in each time slot, so the system overhead is higher than that of the LyMADRL algorithm. In the ECC-only algorithm, if the local base station does not cache the corresponding computing service type, the ECC-only algorithm outsources the tasks to the remote central cloud computing with a large overhead, so the system overhead is the largest.
[0140] Figure 5 (b) and Figure 5 (c) respectively show the convergence processes of the local task queue backlog and the cooperative task queue backlog. From Figure 5 (b), it can be seen that the local task queue backlog of all algorithms does not change with the number of episodes. This is because the LyMADRL algorithm and the ECC-only algorithm adopt the task queue stability module, enabling the edge nodes to obtain the optimal task processing amount based on the Lyapunov control parameter and the offloading decision in each time slot of each episode, so it does not change with the number of training episodes. Although the Greedy algorithm does not adopt the task queue stability module, the task arrivals are always within the processing capacity of the MEC server, and the MEC server executes or forwards tasks with the same maximum processing capacity in each episode, so the local queue backlog does not change with the number of episodes. In addition, from Figure 5As can be seen from (c), as the number of rounds increases, the backlog of the collaborative task queue of the LyMADRL algorithm and the Greedy algorithm finally converges to a stable value. Since the Greedy algorithm processes the arriving tasks based on the maximum processing capacity, the backlog of the collaborative task queue is always smaller than that of the LyMADRL algorithm.
[0141] Figure 6 shows the influence of different control parameters V k on the performance of different algorithms. In the experiment, the control parameter V k is set to {0, 300, 600, 1000, 2000, 5000, 10000, 20000, 50000}, and other parameters are the aforementioned default values. As can be observed from Figure 6 (a), the system overhead of the LyMADRL algorithm and the ECC-only algorithm decreases as the control parameter increases. This is because the LyMADRL algorithm and the ECC-only algorithm adopt a task queue stabilization module based on Lyapunov optimization. According to Lyapunov optimization theory, the larger the control parameter, the larger the backlog of the task queue (as shown in Figure 6 (b)) and the smaller the system overhead, which indicates a trade-off relationship between the backlog of the task queue and the system overhead. At the same time, as can be observed from Figure 6 (b), since both the LyMADRL algorithm and the ECC-only algorithm adopt a task queue stabilization module, they have the same local task queue backlog under the same control parameter, enabling each edge node to obtain the optimal task processing volume of the local task queue for each time slot based on the control parameter and the offloading decision. Therefore, they have the same local task queue backlog under the same control parameter. The Greedy algorithm greedily processes the arriving tasks based on the principle of maximum task processing volume and does not adopt a task queue stabilization module. Therefore, the local task queue backlog is the smallest and is not affected by the control parameter. Figure 6 (c) shows the backlog of the collaborative task queue under different control parameters. As can be seen from the figure, the backlog of the collaborative task queue of the LyMADRL algorithm and the Greedy algorithm remains basically unchanged under different control parameters. This is because the task arrival volume of the collaborative task queue is always within the minimum processing range of the MEC server. Therefore, the queue backlog will not continue to increase and is not affected by the control parameter. As shown in Figure 6 (a), compared with the benchmark algorithm, the proposed LyMADRL algorithm of the present invention has the optimal system overhead under different control parameters.
[0142] In the present invention, unless otherwise clearly specified or limited, the terms "install", "set", "connect", "fix", "rotate" and the like shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements. Unless otherwise clearly limited, for those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0143] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent edge collaborative computing offloading method for random tasks, characterized in that: The following steps are involved: S1. Build a MEC system model, which includes users, MEC servers, base stations, and remote center clouds; S2. Based on the MEC system model, construct the random task arrival model, service cache model, local task queue model, collaborative task queue model and overhead model; S3. Construct a problem model with the optimization goal of minimizing the long-term average time overhead of the system under the long-term stability constraint of the task queue; S4. The problem model is transformed into a single-slot deterministic optimization problem through Lyapunov optimization theory; S5. Convert the single-slot deterministic optimization problem into a Dec-POMDP problem, build the LyMADRL algorithm framework, and obtain the environment state, local observation, action, and reward function; S6. Based on the LyMADRL algorithm framework, the LyMADRL algorithm based on the CTDE framework is used to solve the optimal strategy.
2. According to the method of claim 1, the method is characterized in that: The MEC system model includes N base stations and 1 remote center cloud, where the base station set is defined Base stations are connected through wired links, and each base station is connected to the remote central cloud through the core network; at the same time, each base station is equipped with a MEC server with storage and computing capabilities to provide computing services for users associated with it; definition Indicates base station The associated user set, where Indicates the number of users associated with base station n; in the MEC system model, time is discretized into a set of time slots of equal length The length of each time slot is τ, and it is assumed that each user generates one random computation request in each time slot.
3. The method for intelligent edge collaborative computing offloading for random tasks according to claim 1 is characterized in that: The random task arrival model is used to model the random task arrival situation; assuming that there are K types of computing services in the MEC system model, let the service type set Define that each user randomly selects one of K computing service types in each time slot to generate one random computing request, then: in, represents the total number of tasks belonging to computing service type k arriving at base station n in time slot t, Indicates time slot t user m n The amount of random computing requests arriving at base station n, represents time slot t, user m n The minimum and maximum amount of tasks for random computing requests to reach base station n; 1 {·} represents the indicator function. If the event {·} is true, then 1 {·} =1, otherwise 1 {·} =0; Indicates base station The associated user collection, where represents the number of users associated with base station n; Indicates time slot t user m n The computing service type for random computing requests.
4. The method for intelligent edge collaborative computing offloading for random tasks according to claim 1 is characterized in that: The service cache model is used to model cache decision variables and define binary variables represents the cache decision of base station n for computing service type k in time slot t. If computing service type k is cached in base station n, then otherwise 5. The method for intelligent edge collaborative computing offloading for random tasks according to claim 1 is characterized in that: The local task queue model is used to model the local task queue variables. The dynamic update equation of the local task queue of base station n belonging to computing service type k in each time slot is expressed as follows: in, Indicates that in time slot t+1, base station n belongs to the local task queue of computing service type k; Indicates that time slot t, base station n belongs to the local task queue of computing service type k; represents the number of tasks leaving the local task queue of base station n belonging to computing service type k in time slot t, represents the amount of arriving tasks in the local task queue of base station n belonging to computing service type k in time slot t; definition Uninstall decision express The computing tasks in the process are outsourced to the remote central cloud for execution; express The computation tasks in are executed at the local base station n. express The computing tasks in are forwarded to other base stations n' for execution.
6. The method for intelligent edge collaborative computing offloading for random tasks according to claim 1, characterized in that: The collaborative task queue model is used to model collaborative task queue variables; the dynamic update equation of the collaborative task queue of base station n belonging to computing service type k in each time slot is expressed as follows: in, Indicates that in time slot t+1, base station n belongs to the collaborative task queue of computing service type k; Indicates that time slot t, base station n belongs to the collaborative task queue of computing service type k; represents the number of tasks leaving the collaborative task queue of base station n belonging to computing service type k in time slot t, represents the amount of arriving tasks in the collaborative task queue of base station n belonging to computing service type k in time slot t; definition Task offloading decision in express Outsourcing tasks to remote cloud centers for execution; express The tasks in are executed at the local base station.
7. The method for intelligent edge collaborative computing offloading for random tasks according to claim 1 is characterized in that: The cost model is used to model the task computing cost and cloud-edge collaboration cost of the system, where: Base station n processes the local task queue in time slot t The computational overhead of the task for Base station n processes collaborative task queue Y in time slot t n k The computational cost of the task in (t) for Indicates the unit task calculation overhead weight coefficient, Indicates the offloading decision of the local task queue, represents the unloading decision of the collaborative task queue, represents the number of tasks leaving the local task queue of base station n belonging to computing service type k in time slot t, represents the number of tasks leaving the collaborative task queue of base station n belonging to computing service type k in time slot t, a binary variable represents the cache decision of base station n for computing service type k in time slot t, 1 {·} represents the indicator function; Base station n processes the local task queue in time slot t Cloud-edge collaboration overhead for tasks in for Base station n processes the collaborative task queue in time slot t Cloud-edge collaboration overhead for tasks in for represents the weight coefficient of the edge-to-edge coordination overhead of a unit task, Represents the weight coefficient of the cloud-edge collaboration overhead per task; The sum of the task computing overhead and the cloud-edge collaboration overhead is defined as the total system overhead C(t), which can be expressed as:
8. The method for intelligent edge collaborative computing offloading for random tasks according to claim 1, characterized in that: The problem model is expressed as in, represents a set of time slots, represents the number of time slots, C(t) represents the total system overhead of time slot t, represents the cache decision of base station n for computing service type k in time slot t, Represents the local task queue The uninstall decision Represents a collaborative task queue Task offloading decision; represents the number of tasks leaving the local task queue of base station n belonging to computing service type k in time slot t, and They represent the minimum and maximum task amounts of the local task queue of computing service type k that leaves base station n in each time slot respectively; represents the number of tasks leaving the collaborative task queue of base station n belonging to computing service type k in time slot t, and They represent the minimum and maximum task amounts of the local task queue of computing service type k that leaves base station n in each time slot respectively; Represents a set of service types, z k represents the storage space required for cache computing service type k, Z n represents the storage capacity of base station n, represents the base station set, and They represent the expected lengths of the local task queue and the collaborative task queue of base station n belonging to computing service type k respectively; Indicates the base station service cache decision; Indicates the task offloading strategy; and They represent the task volume strategies of the local task queue and the collaborative task queue respectively; The problem model is transformed through Lyapunov optimization theory to obtain a single-slot deterministic optimization problem, which can be expressed as Among them, k,Q (t) and Π k,Y (t) represents the Lyapunov dynamic control parameter.
9. The method for intelligent edge collaborative computing offloading for random tasks according to claim 1, characterized in that: Define the state of the environment at time slot t Define the local observation of agent n in time slot t Define the actions of agent n in time slot t Define the reward function r of agent n in time slot t n (t)=-C(t) in, represents the cache decision of base station n for computing service type k in time slot t-1, Indicates that time slot t, base station n belongs to the local task queue of computing service type k; Indicates that time slot t, base station n belongs to the collaborative task queue of computing service type k; express The uninstall decision express The task offloading decision is represented by C(t), and the total system overhead is represented by C(t).