Online Learning-based Scheduling Method for Edge Computing Based on Container Layer Dependencies
Through the method based on factorization and policy gradient reinforcement learning, the problem that the container layer dependency relationship in traditional scheduling algorithms is not considered, efficient resource utilization and task scheduling optimization are achieved, and it is suitable for edge computing systems.
Patent Information
- Application Number
- CN202110808603.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-07-16
AI Technical Summary
Traditional container image scheduling algorithms cannot effectively consider the dependencies of the container layer, resulting in duplicate downloads and resource waste, and cannot adapt to the heterogeneity and resource dynamics of edge nodes. Deep learning technology cannot extract hidden dependencies between layers.
Factor-based algorithms are used to extract high-dimensional and low-dimensional sparse dependencies, combine reinforcement learning algorithms with strategy gradients for task scheduling, and design an online learning scheduling method based on container layer dependencies, taking into account the heterogeneity and resource dynamics of edge nodes.
It effectively reduces the overall task overhead and container image file download overhead in edge computing systems, optimizes the task scheduling results, and is suitable for heterogeneous edge computing environments.
Smart Images

Figure CN113641447B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of distributed systems, edge computing, and machine learning in computer networks, and relates to resource scheduling in edge computing and distributed systems, and a method of deep reinforcement learning for machine learning. Background Art
[0002] Traditional algorithms based on container image scheduling cannot consider the existence relationship of the container layer at a finer granularity level, which may lead to a large number of repeated downloads, wasting limited bandwidth and storage resources in edge nodes. Secondly, existing layer-based scheduling algorithms only consider the size of the upper layer of edge nodes, and cannot well consider the heterogeneity of edge nodes, the dynamics of available resources, and the hidden dependency relationships between front and back tasks during scheduling. In addition, for traditional deep learning techniques, the hidden dependency relationships between layers cannot be well extracted, so they cannot be directly applied to layer-based scheduling problems.
[0003] To overcome the above problems, we propose a learning-based scheduling algorithm based on container layer dependencies, fully consider the distribution relationship of the container layer in edge nodes, and consider the heterogeneity and dynamics and real-time nature of resources in edge nodes. Use methods such as factorization to extract the hidden relationships therein, and use reinforcement learning-based techniques to implement task scheduling to obtain long-term benefits. Summary of the Invention
[0004] In view of the above problems, the present invention proposes to better plan the resources in edge computing, reduce the overall overhead of user tasks in the edge computing system, and the overhead required to download container image files during container runtime in edge computing. To achieve the above-mentioned purposes and solve the above-mentioned problems, first, model edge computing at the container layer level, considering the user task completion time in edge computing, including the download time of containers required by user tasks and the running time of user tasks. On this basis, an algorithm based on factorization is proposed to extract the dependency relationships of the container layer in edge computing and extract the high-dimensional and low-dimensional sparse dependency features therein. Finally, based on the extracted dependency relationships and task and node resource characteristics, a learning-based task scheduling algorithm based on policy gradient is designed, and the entire process is verified with real data.
[0005] Specifically, it includes the following steps:
[0006] 1) Build a new model for the edge computing system, and the new model includes: remote cloud, user tasks, containers, image files, layer files;
[0007] 2) Model the main overhead of task scheduling in edge computing,
[0008] 3) Model the task scheduling problem based on layer dependencies in edge computing;
[0009] 4) Use a model-free policy gradient reinforcement learning algorithm to solve the problem. The content of the reinforcement learning includes an advantage function of a policy optimization algorithm:
[0010] A π (s, a) = Q π (s, a) - V π (s)
[0011] And a method for maximizing long-term rewards:
[0012] 5) Design a policy network for extracting container layer dependencies,
[0013] 6) Perform action constraints and selections,
[0014] 7) Train the policy network of the reinforcement learning.
[0015] Furthermore, for the online learning-based scheduling method in edge computing based on container layer dependencies, the relationships among the remote cloud, user tasks, containers, image files, and layer files in 1) are as follows: In edge computing, users generate different tasks and request different containers. The operation of each container requires an image file. Each image file requires several layer files, and the layer files can be shared by different image files. If the requested image file does not exist locally, it is downloaded from the repository of the remote cloud.
[0016] Furthermore, for the online learning-based scheduling method in edge computing based on container layer dependencies, when modeling the task scheduling problem based on layer dependencies in 3), it also includes: the download time of the container image, and the execution time of the user task and some constraint conditions for task scheduling,
[0017] The constraint conditions for the task scheduling include: the limit on the number of containers running simultaneously on each node; the storage space occupied by each layer file cannot exceed the total storage size limit; and each task can only be scheduled to one node and cannot be scheduled to multiple nodes simultaneously.
[0018] Furthermore, for the online learning-based scheduling method in edge computing based on container layer dependencies, when using a model-free policy gradient reinforcement learning algorithm in 4), the reinforcement learning also includes: the state space of the task, the action space of the task, and the reward function;
[0019] The state space of the task includes the states of the task nodes and task resources; expressed as:
[0020]
[0021] The action space described above is the set of all edge nodes and the cloud, denoted as: a t ∈ N ∪ {n |N|+1}
[0022] The reward function is denoted as: r t = -T k
[0023] Furthermore, for the online learning-based scheduling method in edge computing based on container layer dependencies, in step 4), a model-free policy gradient reinforcement learning algorithm is adopted, and the advantage function
[0024] A π (s, a) = Q π (s, a) - V π (s)
[0025] where V π (s) is the value function, defined as:
[0026] V π (s) = E τ~π [R(τ)|s0 = s]
[0027] In addition, Q π (s, a) is the state-action function, defined as:
[0028] Q π (s, a) = E τ~π [R(τ)|s0 = s, a0 = a]
[0029] Finally, the loss function L(θ) is defined as:
[0030]
[0031] where is the estimated advantage function, calculated by the following method:
[0032]
[0033] Furthermore, for the online learning-based scheduling method in edge computing based on container layer dependencies, in step 5), the policy network is: effectively combining the factorization algorithm into the neural network to form a policy network structure based on factorization and vector embedding layers;
[0034] Furthermore, for the online learning-based scheduling method in edge computing based on container layer dependencies, in step 6), the action limit and selection include the following steps:
[0035] First, limit the node load so that it does not exceed its computing power;
[0036] Second, limit the node storage resources so that they do not exceed its storage space.
[0037] Finally, overall, both of the above restrictions need to be satisfied simultaneously. If both of the above restrictions are satisfied, then this action is reasonable; otherwise, the action will be reselected.
[0038] Furthermore, for the online learning-based scheduling method in edge computing based on container layer dependencies, the step 7) of training the reinforcement learning policy network includes the following steps:
[0039] First, initialize a memory space to store the historical data required for training.
[0040] Second, obtain the initial state from the environment and perform scheduling using the existing policy.
[0041] Finally, store the result obtained from the scheduling in the memory space. When the historical data exceeds 1000, start training the network and update the network parameters. After the first round of training is completed, enter the next loop.
[0042] Furthermore, for the online learning-based scheduling method in edge computing based on container layer dependencies, the step of obtaining the initial state from the environment and performing scheduling using the existing policy is specifically as follows:
[0043] First, after extracting the state information from the environment, it is input into a policy network. After passing through the vector embedding layer, factorization layer, linear layer, and output layer, a random policy is obtained.
[0044] Second, select an action according to this policy. After obtaining the action, obtain the reward function of the action from the environment and input it into a value function.
[0045] Finally, the value function calculates a value and calculates the reward function to obtain a loss function, and updates the policy network according to the loss function.
[0046] Furthermore, for the online learning-based scheduling method in edge computing based on container layer dependencies, the task resources include the required CPU resources and the estimated information of the requested containers and tasks. The estimated information is defined as:
[0047]
[0048]
[0049] Description of the Drawings
[0050] Figure 1 shows the relationship among user tasks, containers, and layers in the edge computing of the present invention
[0051] Figure 2 is the flowchart of the learning-based scheduling algorithm based on layer dependency relationship of the present invention
[0052] Figure 3 is the cumulative distribution function graph of different algorithms of the present invention
[0053] Figure 4 shows the convergence of the learning-based algorithm based on layer dependency relationship of the present invention
[0054] Figure 5 shows the performance of the algorithm in heterogeneous edge computing under different numbers of tasks (tasks are generated based on uniform distribution)
[0055] Figure 6 shows the performance of the algorithm in heterogeneous edge computing under different numbers of nodes (tasks are generated based on uniform distribution)
[0056] Figure 7 shows the performance of the algorithm in homogeneous edge computing under different numbers of tasks (tasks are generated based on uniform distribution)
[0057] Figure 8 shows the performance of the algorithm in homogeneous edge computing under different numbers of nodes (tasks are generated based on uniform distribution)
[0058] Figure 9 shows the performance of the algorithm in homogeneous edge computing under different bandwidths (tasks are generated based on uniform distribution)
[0059] Figure 10 shows the performance of the algorithm in homogeneous edge computing under different numbers of containers, CPU frequencies, and storage resources (tasks are generated based on uniform distribution)
[0060] Figure 11 shows the performance of the algorithm in heterogeneous edge computing under different numbers of tasks, nodes, and random number seeds (tasks are generated based on Zipf distribution)
[0061] Figure 12 shows the performance of the algorithm in homogeneous edge computing under different numbers of tasks (tasks are generated based on Zipf distribution)
[0062] Figure 13 shows the performance of the algorithm in homogeneous edge computing under different numbers of nodes (tasks are generated based on Zipf distribution) Detailed Implementation Manner
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0064] 1. Model the edge computing system, fully considering the user task completion time. Let N = {n1, n2, …, n |N|} represent a set of edge nodes, where |N| is the number of edge nodes. In addition, there is a remote cloud, denoted by n |N|+1 . A series of user tasks K = {k1, k2, …, k |K|} are requesting different containers and need to be scheduled to an edge node or the cloud. To process user tasks, different containers C = {c1, c2, …, c |C|} must be created on the edge nodes. Each container requires an image file to run, and each image file must contain a series of specific layer files. We use L = {l1, l2, …, l L} to represent the set of layers. For each node n in the set N, it has a set of running containers a set of local layers In addition, each node also has a CPU frequency f n , a bandwidth b n , and a storage resource d n . The number of containers that each node can run is also limited, and at most C n containers can run simultaneously.
[0065] In addition, the set of layers required by the container c ∈ C is If then the container c contains the layer l. Conversely, if then the container c does not contain the layer l. The size of each layer l is d l . For the user task k ∈ K generated at time t, the CPU resources it requests are p k , and the requested container is c k. After scheduling, the node to which the task is assigned is represented as where indicates that task k is scheduled to edge node n, otherwise
[0066] 2. Then, model the main overhead of task scheduling in edge computing. The user task completion time mainly includes the initialization time of the container requested by the user, mainly the download time of the container image, and the execution time of the user task. To calculate the download time, a variable is introduced. If layer l exists on node n at time t, then Otherwise, In addition, the variable is used to represent the download completion time of each layer l. If layer l already exists on node n or has not started downloading, then
[0067] After the above definitions, the download time and running time can be expressed as follows:
[0068] 1. Download time:
[0069] 2. Computation time:
[0070] 3. Total time:
[0071] In addition, there are some constraints during task scheduling. First, the computing resources of each node are limited, so the number of containers running simultaneously on each node is restricted:
[0072]
[0073] Second, the storage resources of each node are also limited, so the storage space occupied by the layer files cannot exceed the total storage size limit:
[0074]
[0075] Finally, each task can only be scheduled to one node and cannot be scheduled to multiple nodes simultaneously. This constraint can be expressed as:
[0076] 3. With these overheads and constraints, we can finally model the task scheduling problem based on layer dependencies in edge computing. The problem can be modeled as follows:
[0077]
[0078]
[0079] 4. After modeling the above problems, we use a model-free policy gradient reinforcement learning algorithm to solve the problems. Reinforcement learning mainly includes several aspects such as state space, action space, reward function, policy, etc. First, for the state space, since task scheduling is highly relevant to the resources of nodes and the requests of the tasks themselves, both of these aspects are considered in the state space. For each node, it first includes the existence of layers on each node. And there are real-time situations of some resources on each node, including bandwidth, CPU frequency, and the total remaining required download time in real time.
[0080] Then for each node, the state can be represented as:
[0081] And the states of all nodes can be represented as:
[0082] In addition, for the tasks to be processed at each moment, the resource requests of the tasks themselves are very important information, including the required CPU resources and the requested containers. In addition, some estimated information of the tasks is also very important, including:
[0083]
[0084]
[0085]
[0086] Finally, the state space of the tasks can be represented as:
[0087]
[0088] And the state of the entire edge computing system is defined as:
[0089] After that, for the definition of the action space, we define the action space as the set of all edge nodes and the cloud, which can be represented as: a t ∈N∪{n |N|+1}
[0090] Finally, for the reward function, it can be defined as r t =-T k
[0091] The goal of reinforcement learning is to maximize the long-term reward, and the long-term reward can be defined as:
[0092] To train reinforcement learning to obtain an optimal policy, we need a policy optimization algorithm. For policy gradient algorithms, an advantage function is very important and is defined as:
[0093] A π (s, a) = Q π (s, a) - V π (s)
[0094] where V π (s) is the value function and is defined as:
[0095] V π (s) = E τ~π [R(τ)|s0 = s]
[0096] In addition, Q π (s, a) is the state-action function and is defined as:
[0097] Q π (s, a) = E τ~π [R(τ)|s0 = s, a0 = a]
[0098] Finally, the loss function is defined as:
[0099]
[0100] where is the estimated advantage function, which is calculated by the following method:
[0101]
[0102] 6. After these definitions, we also need to design a policy network that can well extract the container layer dependencies. Common structures based on convolutional neural networks or recurrent neural networks cannot well extract sparse features, while using a vector embedding layer cannot well extract features from different dimensions. To solve these problems, we effectively combine the factorization algorithm into the neural network and design a policy network structure based on factorization and vector embedding layer.
[0103] Finally, when selecting actions, reinforcement learning may select some relatively poor or obviously unreasonable actions, such as scheduling tasks to nodes with high loads or nodes with obviously insufficient storage resources. To avoid these situations as much as possible, we also need to impose some restrictions on the actions. First, some restrictions on node loads:
[0104]
[0105] Secondly, some restrictions on node storage resources:
[0106]
[0107] Finally, the overall limit:
[0108] If then this action is reasonable, otherwise, the action will be re - selected.
[0109] 7. Start training the policy network of reinforcement learning. First, initialize a memory space to store the historical data required for training. Then, obtain the initial state from the environment and schedule using the existing policy. The result of the scheduling is stored in the memory space. If there is enough historical data, then start training the network and update the network parameters. If the training is completed, enter the next loop. The whole process is as follows:
[0110] Input: s0
[0111] Output: π *
[0112] For each loop:
[0113] Initialize D as an empty set, D n = 0,
[0114] Reset the environment and obtain the initial state s0.
[0115] For each time step:
[0116] Run the policy in the environment
[0117] Store the obtained information (s t , a t , r t , r t (θ), s t+1 ) into D,
[0118] D n ← D n + 1
[0119] If D n mod |D| = 0
[0120] Calculate and L(θ)
[0121] Train the network and update the parameter θo ld ← θ
[0122] If the training is completed: break.
[0123] Return π *
[0124] And the action selection process is as follows:
[0125] Input: a t
[0126] Output: a t
[0127] Initialize u n = 0
[0128] Calculate
[0129] Loop while
[0130] If u n ≥ u N :
[0131] a t = N n+1 , interrupt
[0132] Resample a t
[0133] u n = u n + 1
[0134] Calculate new
[0135] Return a t Specific embodiments
[0137] We used Python for programming and simulated classes such as edge nodes, containers, images, layers, user tasks, and schedulers in the edge computing system. Based on these classes, a simulated edge computing environment was implemented. On this basis, functions required for learning algorithms such as state acquisition, action selection, environment update, policy network training, policy gradient calculation, and value function update were implemented. The policy network mainly includes a vector embedding layer, a factorization layer, and a linear layer. The experimental data came from data crawled from a real container repository. After data cleaning and preprocessing, a total of 70 image files and 337 layer files were obtained.
[0138] In addition, the main parameter settings for some experiments are as follows. The storage space range for each node is 5GB to 15GB and is set randomly. The bandwidth for each node is randomly set to 70Mbps to 90Mbps. The CPU frequency for each node is randomly set between 7GHz and 9GHz. For the cloud, the bandwidth is 100Mbps and the CPU frequency is 10GHz. In addition, the default number of user tasks is 2000.
[0139] To better compare the algorithm performance, we selected the following baseline algorithms for comparison and reference in the environments where tasks are generated based on the uniform distribution and the Zipf distribution. The details are as follows:
[0140] The first one is the currently state-of-the-art layer-based scheduling algorithm Dep. This algorithm calculates a score based on the size of the upper layer of each node and then performs scheduling according to the score.
[0141] The second algorithm is Dep-Soft, which is modified based on Dep. This algorithm sets a threshold and randomly selects one node from all nodes with scores exceeding the threshold for scheduling.
[0142] The third algorithm is Kube, which is one of the current default container scheduling algorithms. It schedules based on the image file and does not consider the relationship between layer dependencies.
[0143] The fourth algorithm is Monkey, which is a random algorithm.
[0144] The fifth algorithm is Dep-Down, which is based on the Dep algorithm and additionally considers the download time.
[0145] The sixth algorithm is Dep-Wait, which modifies the size of the container layer considered by the Dep algorithm to the waiting time.
[0146] The seventh algorithm is Dep-Comp, which modifies the size of the layer considered in the Dep algorithm to the computing time. Finally, the experimental results are as follows:
[0147] First, the comparison in the case where tasks are generated based on the uniform distribution is as follows:
[0148] 1. The performance of the algorithms in heterogeneous edge computing under different numbers of tasks, as shown in the attached instructions Figure 5 as shown
[0149] 2. The performance of the algorithms in heterogeneous edge computing under different numbers of nodes, as shown in the attached instructions Figure 6 as shown
[0150] 3. The performance of the algorithms in homogeneous edge computing under different numbers of tasks, as shown in the attached instructions Figure 7 as shown
[0151] 4. The performance of the algorithms in homogeneous edge computing under different numbers of nodes, as shown in the attached instructions Figure 8 as shown
[0152] 5. The performance of the algorithms in homogeneous edge computing under different bandwidths, as shown in the attached instructions Figure 9 as shown
[0153] 6. The performance of the algorithm in homogeneous edge computing under different numbers of containers, CPU frequencies, and storage resources is shown in the attached drawings of the specification, Figure 10 as shown;
[0154] Then, there is a comparison of tasks generated based on the Zipf distribution, specifically as follows:
[0155] 1. The performance of the algorithm in heterogeneous edge computing under different numbers of tasks, nodes, and random number seeds is shown in the attached drawings of the specification, Figure 11 as shown,
[0156] 2. The performance of the algorithm in homogeneous edge computing under different numbers of tasks is shown in the attached drawings of the specification, Figure 12 as shown,
[0157] 3. The performance of the algorithm in homogeneous edge computing under different numbers of nodes is shown in the attached drawings of the specification, Figure 13 as shown.
[0158] From the above, it can be found that the method described in this invention has the following beneficial effects:
[0159] 1) This method models from the perspective of container layer dependencies and schedules at the granularity of layers, which can effectively solve the problem of excessive task latency in edge computing and effectively reduce the overall completion time of tasks.
[0160] 2) This method proposes a set of methods for mining and extracting container layer dependencies and a policy gradient learning-based scheduling algorithm based on layer dependencies, which can schedule users' tasks to the most suitable nodes according to the layer distribution of nodes.
[0161] 3) The training method based on deep reinforcement learning proposed by this method can effectively consider layer dependencies and distribution, as well as the hidden relationship between the layers of front and back task requests, and optimize the scheduling results as a whole.
[0162] 4) By analyzing the specific situation of the edge computing system, this method sets a reward function that helps to select actions, which can effectively reduce the completion time of user tasks in the entire edge computing system and improve the user experience.
[0163] 5) This method is tested based on real datasets and has strong generalization ability, and can be applied to various heterogeneous edge computing environments.
[0164] The above are only embodiments of the present invention, and do not limit the protection scope of the present invention. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related system fields, shall be included in the protection scope of the present invention by the same token.
Claims
1. An online learning-based scheduling method for container layer dependencies in edge computing, characterized in that: Specifically, it includes the following steps: 1) Build a new model for the edge computing system, and the new model includes: a remote cloud, user tasks, containers, image files, and layer files; 2) Model the main overhead of task scheduling in edge computing; 3) Model the task scheduling problem based on layer dependencies in edge computing; 4) Use a model-free policy gradient reinforcement learning algorithm to solve the problem, and the content of the reinforcement learning includes an advantage function of a policy optimization algorithm: a π (s,a) = Q π (s,a) - V π (s) where Vπ(s) is the value function and Qπ(s,a) is the state-action function; And a method for maximizing long-term benefits: 5) Design a policy network for extracting container layer dependencies, 6) Perform action restriction and selection, 7) Train the policy network of the reinforcement learning; The step 2) of modeling the main overhead of task scheduling in edge computing also includes: the download time of the container image, as well as the execution time of the user task and the task scheduling constraint conditions After these definitions, the download time and the running time can be expressed as: a. Download time: b. Calculation time: c. Total time: In addition, there are some constraint conditions during task scheduling. First, the computing resources of each node are limited, so the number of containers running on each node at the same time is limited: Second, the storage resources of each node are also limited, so the storage space occupied by the layer files cannot exceed the total storage size limit: Finally, each task can only be scheduled to one node and cannot be scheduled to multiple nodes simultaneously. This restriction can be expressed as: Finally, model the task scheduling problem based on layer dependencies in edge computing, and model the problem as follows:
2. The online learning-based scheduling method based on container layer dependencies in edge computing according to claim 1, wherein: The relationship among the remote cloud, user tasks, containers, image files, and layer files in the step 1) is: in edge computing, users generate different tasks and request different containers, and the running of each container requires an image file. Each image file contains several layer files, and the layer files can be shared by different image files. If the requested image file does not exist locally, it will be downloaded from the repository of the remote cloud.
3. The online learning-based scheduling method based on container layer dependencies in edge computing according to claim 1, wherein: The task scheduling constraint conditions include: the limit on the number of containers running on each node at the same time; the storage space occupied by each layer file cannot exceed the total storage size limit; and each task can only be scheduled to one node and cannot be scheduled to multiple nodes at the same time.
4. The online learning-based scheduling method based on container layer dependencies in edge computing according to claim 1, characterized in that: The step 4) uses a model-free policy gradient reinforcement learning algorithm, and the reinforcement learning also includes: the state space of the task, the action space of the task, and the reward function; The state space of the task includes the states of the task nodes and task resources; it is expressed as: For each node, the state can be represented as: And the states of all nodes can be represented as: The action space is the set of all edge nodes and the cloud, and is expressed as: a t ∈N∪{n |N|+1} 5. The online learning-based scheduling method based on container layer dependencies in edge computing according to claim 1, characterized in that: The step 4) uses a model-free policy gradient reinforcement learning algorithm, and the advantage function: A π (s,a) = Q π (s,a) - V π (s) where, V π (s) is the value function, defined as: V π (s) = E τ~π [R(τ)|s0 = s] In addition, Q π (s,a) is the state-action function, defined as: Q π (s,a) = E τ~π [R(τ)|s0 = s,a0 = a] Finally, the loss function L(θ) is defined as: wherein, is the estimated advantage function, which is calculated by the following method:
6. The online learning-based scheduling method based on container layer dependencies in edge computing according to claim 1, characterized in that: The policy network in the step 5) is: combine the factorization algorithm into the neural network to form a policy network structure based on factorization and vector embedding layers.
7. The online learning-based scheduling method based on container layer dependencies in edge computing according to claim 1, characterized in that: The step 6) of action restriction and selection includes the following steps: First, restrict the node load and do not exceed its computing power; Second, restrict the node storage resources and do not exceed its storage space; Finally, overall, both of the above restrictions need to be satisfied at the same time. If both of the above restrictions are satisfied at the same time, then this action is reasonable. Otherwise, the action will be reselected.
8. The online learning-based scheduling method based on container layer dependencies in edge computing according to claim 1, wherein: The steps for training the policy network of reinforcement learning are as follows: First, initialize a memory space to store the historical data required for training; Second, obtain the initial state from the environment and schedule it using the existing policy; Finally, store the result of the scheduling in the memory space. When the historical data exceeds 1000, start training the network and update the network parameters. After the first round of training is completed, enter the next loop.
9. The online learning-based scheduling method based on container layer dependencies in edge computing according to claim 8, characterized in that: The specific steps for obtaining the initial state from the environment and scheduling it using the existing policy are as follows: First, after extracting the state information from the environment, it is input into a policy network. After passing through the vector embedding layer, factorization layer, linear layer, and output layer, a stochastic policy is obtained; Second, select an action according to this policy. After obtaining the action, obtain the reward function of the action from the environment and input it into a value function; Finally, the value function calculates a value and calculates the reward function to obtain a loss function, and updates the policy network according to the loss function.
10. The online learning-based scheduling method based on container layer dependencies in edge computing according to claim 4, characterized in that: The task resources include the required CPU resources and the estimated information of the requested containers and tasks. The estimated information is defined as:
Citation Information
Patent Citations
Learning-based low-delay task scheduling method in edge computing network
CN109976909A
Multi-copy-based task scheduling method and system for edge computing environment
CN111381950A