Scheduling and distributing method and device for multi-layer heterogeneous computing power resources of industrial internet
By adopting meta-learning and DDPG algorithms in the industrial Internet, a multi-layer heterogeneous computing resource model is built, which solves the optimization problem under resource constraints, and realizes efficient resource scheduling and allocation across tasks and devices, improving the generalization ability and efficiency of the system.
Patent Information
- Application Number
- CN202510292247.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology has the problem of optimization under resource limitation in the industrial Internet. Traditional reinforcement learning algorithms lack the generalization capabilities of multiple heterogeneous tasks and different devices, and cannot effectively solve the resource scheduling and allocation across tasks and devices.
The meta-learning method is used to combine the deep deterministic strategy gradient algorithm (DDPG) to build an industrial system model. Through internal and external loop optimization, task-specific strategies and meta-learning strategies are generated to realize the scheduling and allocation of multi-layer heterogeneous computing power resources, including the collaborative work of cloud servers, edge servers and heterogeneous mobile devices.
It improves the generalization capabilities across tasks and devices, optimizes resource allocation, reduces energy consumption and delays, and improves the efficiency and flexibility of industrial Internet systems.
Smart Images

Figure CN120281766A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mobile communication technologies, and particularly to a method and device for scheduling and allocating multi-layer heterogeneous computing power resources in an industrial Internet. Background Art
[0002] With the progress of technology, reinforcement learning has gradually moved from early theoretical exploration to modern practical applications, and both its theoretical foundation and algorithm practice have made great progress. In recent years, the rise of meta-reinforcement learning and the DDPG (Deep Deterministic Policy Gradient) algorithm has injected new vitality into this field.
[0003] Meta-reinforcement learning (Meta-RL), as a cutting-edge direction of reinforcement learning, integrates the idea of meta-learning into it, aiming to improve the fast adaptation and learning ability of agents when facing new tasks or environments. Meta-reinforcement learning not only focuses on the learning of the current task, but also on how to optimize learning itself through learning, that is, "learning how to learn". It usually involves two levels of learning: the meta-learning level is responsible for optimizing the learning algorithm or policy, and the task-learning level is responsible for executing these algorithms or policies on specific tasks. Through meta-reinforcement learning, agents can more effectively explore and utilize historical experience, thus accelerating the learning process in new environments.
[0004] The DDPG algorithm is a reinforcement learning algorithm for continuous action spaces. It combines the advantages of deep learning and deterministic policy gradients to achieve efficient learning and decision-making in complex environments. The specific process of the algorithm is as Figure 1 shown. The DDPG algorithm approximates the state-action value function (Q-function) and the policy function through a deep neural network. The policy function is responsible for generating deterministic actions based on the current state, while the Q-function is responsible for evaluating the value of this action. This structure enables DDPG to achieve efficient policy search and optimization in continuous action spaces. By iteratively updating the Q value and policy parameters, the DDPG algorithm can guide agents to select optimal actions in different states, thereby maximizing the long-term reward. Since the DDPG algorithm was proposed, it has become a benchmark algorithm in the field of reinforcement learning for continuous action spaces and is widely used in multiple fields such as robot control, autonomous driving, and game AI. Its powerful learning ability and adaptability enable agents to make precise and efficient decisions in complex and changing environments.
[0005] In the field of industrial Internet, although edge computing reinforcement learning technology has brought many advantages to application scenarios such as intelligent manufacturing, intelligent energy management, and intelligent logistics, such as reducing data transmission latency, improving system response speed, and reducing cloud pressure, there are still a series of technical challenges and deficiencies, which limit the wide application and in-depth development of this technology. Summary of the Invention
[0006] To solve the technical problems existing in the prior art of how to systematically solve the optimization problem under resource constraints, improve the generalization ability across tasks and devices, and the traditional reinforcement learning algorithms are often trained for specific tasks and devices, lacking the generalization ability for multiple heterogeneous tasks and different devices, the embodiments of the present invention provide a method and device for scheduling and allocating multi-layer heterogeneous computing power resources in industrial Internet. The technical solutions are as follows:
[0007] On the one hand, a method for scheduling and allocating multi-layer heterogeneous computing power resources in industrial Internet is provided. This method is implemented by a scheduling and allocation device, and the method includes:
[0008] S1. Construct an industrial system model; the industrial system model includes a cloud server, multiple edge servers, and multiple heterogeneous mobile devices; the multiple edge servers include a coordination server and other edge servers.
[0009] S2. Multiple heterogeneous mobile devices generate a heterogeneous task set; the heterogeneous task set includes mobile device computing tasks and remote server computing tasks; the remote server computing tasks are offloaded to the edge server or the cloud server through the coordination server.
[0010] S3. Construct a local computing model and a remote computing model according to the heterogeneous task set, and obtain the total energy consumption and latency of the industrial system model based on the local computing model and the remote computing model.
[0011] S4. Based on the meta-learning method, the mobile device performs inner-loop optimization according to the total energy consumption and latency of the industrial system model, the meta-learning strategy generated by the edge server, and the deep deterministic policy gradient algorithm to obtain a task-specific strategy; the edge server performs outer-loop optimization according to the task-specific strategy and the reinforcement learning method to generate a new meta-learning strategy; and then obtain a task offloading scheme based on the dynamic channel state and queue length.
[0012] Optionally, any one mobile device r among the multiple heterogeneous mobile devices is provided with a mobile device task buffer, and the number of bits of task data backlog in the mobile device task buffer at time slot t is updated dynamically according to the queue of the following formula (1):
[0013]
[0014] In the formula, Denotes the number of task data backlogs of the mobile device r at time slot t + 1, μ r (t) represents the amount of computing task data leaving the task buffer of the mobile device at time slot t, and i represents a heterogeneous task. Denotes the task data generated by the mobile device r at time slot t.
[0015] Any one of the multiple edge servers, edge server n, is provided with an edge server task buffer, and the number of task data backlogs of the edge server task buffer at time slot t Is dynamically updated according to the following formula (2):
[0016]
[0017] In the formula, Denotes the number of task data backlogs of the edge server task buffer at time slot t + 1, Denotes the amount of task data calculated by the edge server n at time slot t, Denotes the amount of task data unloaded from the mobile device to the edge server n at time slot t.
[0018] Optionally, the local computing model in S3 includes:
[0019] The energy consumption required for the mobile device to complete the task data volume of the current time slot is as shown in the following formula (3):
[0020]
[0021] In the formula, Denotes the energy consumption required for the mobile device r to complete the task data volume of time slot t, and ξ represents the effective capacitance switching coefficient, Denotes the processing frequency of the mobile device r at time slot t.
[0022] Optionally, the remote computing model in S3 includes:
[0023] The energy consumption on the edge server is as shown in the following formula (4):
[0024]
[0025] In the formula, Denotes the energy consumption required for the edge server n to complete the task data volume of time slot t, Denotes the amount of task data unloaded to the remote device, v k (t) represents the transmit power of the mobile device r transmitting data in time slot t, R r,n (t) represents the rate at which the mobile device r uploads data and instructions to the edge server at time slot t, ξ represents the effective capacitance switching coefficient, and τ represents the duration of time slot t. Indicates the processing frequency of edge server n.
[0026] The latency on the edge server is as shown in Equation (5) below:
[0027]
[0028] In the formula, Indicates the execution time of task i on the edge server, Indicates the task data size, α i Indicates the proportion of heterogeneous task i processed on the mobile device, R r,n (t) indicates the rate at which mobile device r uploads data and instructions to the edge server at time slot t, ε indicates the number of CPU cycles required to execute 1 bit of task, I indicates the set of heterogeneous tasks, Indicates the set of heterogeneous mobile devices, Indicates the set of edge servers.
[0029] Optionally, the remote computing model in S3 further includes:
[0030] The energy consumption on the cloud server is as shown in Equation (6) below:
[0031]
[0032] In the formula, Indicates the energy consumption required for the cloud server to complete the task data volume at time slot t, Indicates the set of heterogeneous mobile devices, i indicates the heterogeneous task, β i Indicates the offloading vector, R r,n (t) indicates the rate at which mobile device r uploads data and instructions to the edge server at time slot t.
[0033] The latency on the cloud server is as shown in Equation (7) below:
[0034]
[0035] In the formula, Indicates the execution time of task i on the cloud server, R n,c (t) indicates the rate at which edge server n uploads data and instructions to the cloud server at time slot t.
[0036] Optionally, the goal of the inner loop optimization in S4 is to minimize the total energy consumption of the industrial system model under the task latency constraint, as shown in Equations (8)-(15) below:
[0037]
[0038] 0 ≤ v r (t) ≤ v(t) r,max(10)
[0039]
[0040] α i ∈[0,1] (12)
[0041] β i ∈{0,1} (13)
[0042]
[0043] Wherein, T represents the task deadline, and E total (t) represents the total energy consumption of the industrial system model, represents the set of mobile devices, w r,n (t) represents the channel bandwidth allocated by the system when mobile device r offloads tasks to edge server n at time slot t, w n represents the channel bandwidth of edge server n. Equation (9) indicates that the channel bandwidth of the mobile device is less than that of the edge server, v r (t) represents the transmit power of mobile device r for transmitting data at time slot t, v(t) r,max represents the maximum transmit power of mobile device r for transmitting data at time slot t, μ r (t) represents the amount of computational task data leaving the task buffer of the mobile device at time slot t. c represents the number of CPU cycles required to execute 1 bit of task, represents the processing frequency of edge server n. τ represents the duration of time slot t, α i represents the proportion of heterogeneous task i processed on the mobile device, β i represents the offloading vector, represents the number of bits of task data backlog in the task buffer of the mobile device at time slot t, represents the set of edge servers, represents the number of bits of task data backlog in the task buffer of the edge server at time slot t.
[0044] Optionally, the outer loop optimization in S4 includes:
[0045] Making a decision on the task data routing path according to the routing topology relationship between edge servers; wherein, the state space is the edge server node where the task is located, the action space is all edge servers and the cloud server, and a corresponding reward function is set according to the routing topology relationship.
[0046] Taking the transmission power of the channel at time slot t, the channel bandwidth, the queue length of the mobile device task buffer, and the queue length of the edge server task buffer as the state space of the meta-reinforcement learning algorithm, and the computing frequencies of heterogeneous mobile devices and edge servers as the action space, optimal computing frequencies and offloading decision parameters are assigned to each edge server.
[0047] On the other hand, a scheduling and allocation device for multi-layer heterogeneous computing power resources in the industrial Internet is provided. The device is applied to the scheduling and allocation method of multi-layer heterogeneous computing power resources in the industrial Internet. The device includes:
[0048] A construction module for constructing an industrial system model; the industrial system model includes a cloud server, multiple edge servers, and multiple heterogeneous mobile devices; the multiple edge servers include a coordination server and other edge servers.
[0049] A task generation module for generating a heterogeneous task set by multiple heterogeneous mobile devices; the heterogeneous task set includes mobile device computing tasks and remote server computing tasks; the remote server computing tasks are offloaded to the edge server or the cloud server through the coordination server.
[0050] A total energy consumption and delay calculation module for constructing a local computing model and a remote computing model according to the heterogeneous task set, and obtaining the total energy consumption and delay of the industrial system model according to the local computing model and the remote computing model.
[0051] An output module for, based on the meta-learning method, the mobile device performing inner-loop optimization according to the total energy consumption and delay of the industrial system model, the meta-learning strategy generated by the edge server, and the deep deterministic policy gradient algorithm to obtain a task-specific policy; the edge server performing outer-loop optimization according to the task-specific policy and the reinforcement learning method to generate a new meta-learning strategy; and further obtaining a task offloading scheme based on the dynamic channel state and queue length.
[0052] Optionally, any one mobile device r among the multiple heterogeneous mobile devices is provided with a mobile device task buffer, and the number of task data backlogs in the mobile device task buffer at time slot t is updated dynamically according to the following queue formula (1):
[0053]
[0054] where represents the number of task data backlogs of mobile device r at time slot t + 1, μ r (t) represents the amount of computing task data leaving the mobile device task buffer at time slot t, i represents the heterogeneous task, represents the task data generated by mobile device r at time slot t.
[0055] Any one of the multiple edge servers, edge server n, is provided with an edge server task buffer, and the number of backlogged task data bits in the edge server task buffer at time slot t is dynamically updated according to the following formula (2):
[0056]
[0057] wherein, represents the number of backlogged task data bits in the edge server task buffer at time slot t + 1, represents the amount of task data calculated by edge server n at time slot t, represents the amount of task data unloaded from the mobile device to edge server n at time slot t.
[0058] Optionally, the local computing model includes:
[0059] The energy consumption required for the mobile device to complete the amount of task data in this time slot is as shown in the following formula (3):
[0060]
[0061] wherein, represents the energy consumption required for mobile device r to complete the amount of task data at time slot t, and ξ represents the effective capacitance switching coefficient, represents the processing frequency of mobile device r at time slot t.
[0062] Optionally, the remote computing model includes:
[0063] The energy consumption on the edge server is as shown in the following formula (4):
[0064]
[0065] wherein, represents the energy consumption required for edge server n to complete the amount of task data at time slot t, represents the amount of task data unloaded to the remote device, v k (t) represents the transmission power of mobile device r for transmitting data in time slot t, R r,n (t) represents the rate at which mobile device r uploads data and instructions to the edge server at time slot t, ξ represents the effective capacitance switching coefficient, τ represents the duration of time slot t, represents the processing frequency of edge server n.
[0066] The delay on the edge server is as shown in the following formula (5):
[0067]
[0068] wherein, Denote the execution time of task \(i\) on the edge server. Denote the task data size, \(\alpha\). i Denote the proportion of heterogeneous task \(i\) processed on the mobile device, \(R\). r,n \(\varepsilon_{r}(t)\) represents the rate at which mobile device \(r\) uploads data and instructions to the edge server at time slot \(t\), \(\varepsilon\) represents the number of CPU cycles required to execute 1 bit of task, and \(I\) represents the set of heterogeneous tasks. Denote the set of heterogeneous mobile devices. Denote the set of edge servers.
[0069] Optionally, the remote computing model further includes:
[0070] The energy consumption on the cloud server is as shown in the following formula (6):
[0071]
[0072] In the formula, Denote the energy consumption required for the cloud server to complete the task data volume at time slot \(t\). Denote the set of heterogeneous mobile devices, \(i\) represents a heterogeneous task, \(\beta\). i Denote the offloading vector, \(R\). r,n \(\varepsilon_{r}(t)\) represents the rate at which mobile device \(r\) uploads data and instructions to the edge server at time slot \(t\).
[0073] The latency on the cloud server is as shown in the following formula (7):
[0074]
[0075] In the formula, Denote the execution time of task \(i\) on the cloud server, \(R\). n,c \(\varepsilon_{n}(t)\) represents the rate at which edge server \(n\) uploads data and instructions to the cloud server at time slot \(t\).
[0076] Optionally, the objective of the inner loop optimization is to minimize the total energy consumption of the industrial system model under the task latency constraint, as shown in the following formulas (8)-(15):
[0077]
[0078] 0 ≤ \(v_{r}(t)\) r ≤ \(v(t)\) (10) r,max (10)
[0079]
[0080] \(\alpha\) i ∈ [0, 1] (12)
[0081] \(\beta\) i ∈ {0, 1} (13)
[0082]
[0083]
[0084] In the formula, T represents the task deadline, and E total (t) represents the total energy consumption of the industrial system model, represents the set of mobile devices, and w r,n (t) represents the channel bandwidth allocated by the system when mobile device r offloads tasks to edge server n at time slot t, and w n represents the channel bandwidth of edge server n. Equation (9) indicates that the channel bandwidth of the mobile device is less than that of the edge server, and v r (t) represents the transmission power of mobile device r for transmitting data at time slot t, and v(t) r,max represents the maximum transmission power of mobile device r for transmitting data at time slot t, and μ r (t) represents the amount of computational task data leaving the task buffer of the mobile device at time slot t. c represents the number of CPU cycles required to execute 1 bit of task, represents the processing frequency of edge server n. τ represents the duration of time slot t, and α i represents the proportion of heterogeneous task i processed on the mobile device, and β i represents the offloading vector, represents the number of bits of task data backlog in the task buffer of the mobile device at time slot t, represents the set of edge servers, represents the number of bits of task data backlog in the task buffer of the edge server at time slot t.
[0085] Optionally, the outer loop optimization includes:
[0086] Making a decision on the task data routing path according to the routing topology relationship between edge servers; among them, the state space is the edge server node where the task is located, the action space is all edge servers and the cloud server, and the corresponding reward function is set according to the routing topology relationship.
[0087] Taking the transmission power, channel bandwidth, queue length of the mobile device task buffer, and queue length of the edge server task buffer of the time slot t as the state space of the meta-reinforcement learning algorithm, and the heterogeneous mobile device and edge server computing frequencies as the action space, and allocating the optimal computing frequency and offloading decision parameters for each edge server.
[0088] On the other hand, a scheduling and allocation device is provided, which includes: a processor; a memory storing computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the methods for scheduling and allocating multi-layer heterogeneous computing power resources in the industrial Internet as described above is implemented.
[0089] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any one of the methods for scheduling and allocating multi-layer heterogeneous computing power resources in the industrial Internet as described above.
[0090] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0091] In the embodiments of the present invention, research on a joint optimization mechanism for task offloading and communication computing resources is carried out for heterogeneous industrial scenarios, multiple heterogeneous industrial service models and a three-layer cloud-edge-end collaborative transmission architecture are established, and training is performed for multiple heterogeneous tasks and different devices, which is a reinforcement learning algorithm with strong generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0093] Figure 1 It is a flowchart of the DDPG algorithm provided by the embodiments of the present invention;
[0094] Figure 2 It is a flowchart of a method for scheduling and allocating multi-layer heterogeneous computing power resources in the industrial Internet provided by the embodiments of the present invention;
[0095] Figure 3 It is an overall architecture diagram of the solution provided by the embodiments of the present invention;
[0096] Figure 4 It is the pseudocode of the generalization algorithm provided by the embodiments of the present invention;
[0097] Figure 5 It is a block diagram of a device for scheduling and allocating multi-layer heterogeneous computing power resources in the industrial Internet provided by the embodiments of the present invention;
[0098] Figure 6 It is a schematic structural diagram of a scheduling and allocation device provided by the embodiments of the present invention. DETAILED DESCRIPTION
[0099] The technical solutions in the present invention will be described below with reference to the accompanying drawings.
[0100] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Specifically, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either of the two can be selected.
[0101] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.
[0102] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.
[0103] To make the technical problems to be solved, technical solutions and advantages of the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0104] The embodiments of the present invention provide a method for scheduling and allocating multi-layer heterogeneous computing power resources in an industrial Internet. This method can be implemented by a scheduling and allocation device, and the scheduling and allocation device can be a terminal or a server. As Figure 1 shown in the flowchart of the method for scheduling and allocating multi-layer heterogeneous computing power resources in the industrial Internet, the processing flow of this method can include the following steps:
[0105] S1. Build an industrial system model; the industrial system model includes a cloud server, multiple edge servers and multiple heterogeneous mobile devices; the multiple edge servers include a coordination server and other edge servers.
[0106] In a feasible implementation manner, the present invention sets a cloud server, multiple edge server devices and multiple local devices for the heterogeneous industrial business scenarios in the industrial Internet. For the time delay and energy consumption requirements of this industrial system model, a task offloading scheme is studied, and the allocation mechanisms of computing resources and communication resources are determined. The overall architecture of the scheme is as Figure 3 shown.
[0107] Specifically, consider a MEC offloading system consisting of R heterogeneous mobile devices, N edge servers (including an edge coordinator), and a remote cloud server C, defined as Among them, the computing resources of local devices are much smaller than those of cloud servers.
[0108] Furthermore, the heterogeneity of mobile devices is reflected in computing power and battery power. Edge servers can be a wireless access point, a small cell base station, or a 4G / 5G macro cell base station. The available computing resources and bandwidth resources of edge servers are limited.
[0109] In the computing task set D, the characteristics of each task are represented by four parameters to express, represents the task data size, represents the task arrival slot time, represents the task deadline, represents the task type, and its value set is Obviously, when the task's resource requirement type is CPU resources, otherwise it is GPU resources and needs to be offloaded to the edge server or cloud device for execution. Assume that each time slot t classifies and sorts tasks according to the resource type and deadline of the current time calculation task. The shorter the deadline, the more forward the task is located in the task buffer.
[0110] The computing task data studied in the present invention is splittable. Assume that the proportion α i of computing task i processed locally, i its value range is 0 ≤ α i ≤ 1. Correspondingly, the proportion of task data offloaded remotely is 1 - α i . The coordination server decides whether to offload the task to the edge or the cloud, and the corresponding offloading vector is β i When β i = 1, the task is offloaded to the edge server. When β
[0111] = 0, the task is offloaded to the cloud server.
[0112] In a feasible implementation manner, at the beginning of time slot t, the local device (heterogeneous device) generates a computing task and stores it in the local dataset. Due to the limitation of the local device's computing resources, the computing task is offloaded to the remote server. Since the local device and the edge server cannot immediately process the arriving task data, a local task buffer and an edge task buffer are set up to store the task data, assuming that the buffer has sufficient capacity.
[0113] Let the amount of computing task data executed by the local device r in time slot t be The amount of computing task data offloaded to the remote device is The number of bits of task data backlog of the local device r in time slot t is denoted as It is updated dynamically according to the following queue:
[0114]
[0115] In the formula, represents the number of bits of task data backlog of the mobile device r in time slot t + 1, represents the amount of computing task data leaving the task buffer of the mobile device in time slot t, including the amount of task data completed by the local device during time slot t and the amount of task data offloaded from the local device to the edge server. i represents the heterogeneous task, represents the task data generated by the mobile device r in time slot t.
[0116] The number of bits of task data backlog of the edge server n in time slot t is denoted as Its dynamic update of the queue length is as follows:
[0117]
[0118] In the formula, represents the number of bits of task data backlog of the task buffer of the edge server in time slot t + 1, represents the amount of task data calculated by the edge server n in time slot t, represents the amount of task data offloaded from the mobile device to the edge server n in time slot t.
[0119] Furthermore, according to the definition of the average rate stability of the independent queue, if all computing tasks satisfy the following constraint conditions, then the queue is stable:
[0120]
[0121] S3. Construct a local computing model and a remote computing model according to the heterogeneous task set, and obtain the total energy consumption and latency of the industrial system model based on the local computing model and the remote computing model.
[0122] In a feasible implementation, in the local computing model, task processing does not involve the uplink transmission of data. At this time, during time slot t, the local device executes task data and unloads it to the remote device. The energy consumption required to process this computing task is calculated according to the following mathematical model. First, assume that executing 1 bit of task requires a fixed number of CPU cycles ε (CPU Cycles / bit), and the local device has a processing frequency of (Cycles / s) at time slot t.
[0123] The computing task D i,r has an execution time at the local as follows:
[0124]
[0125] The amount of task data that the device can process in time slot t is:
[0126]
[0127] where τ is the duration of time slot t. The energy consumption required for the device to complete the task data volume in this time slot is:
[0128]
[0129] In the formula, represents the energy consumption required for mobile device r to complete the task data volume in time slot t, ξ represents the effective capacitance switching coefficient, and its value is usually determined by the chip structure. Here, it takes 10 -27 , represents the processing frequency of mobile device r in time slot t.
[0130] Furthermore, when the computing task is offloaded to the edge server, it is necessary to determine the transmission rate when data is transmitted between local device r and edge server n. According to Shannon's second theorem, the rate at which local device r uploads data and instructions to the edge server in time slot t is:
[0131]
[0132] where w r,n (t) represents the channel bandwidth allocated by the system when mobile device r offloads tasks to edge server n in time slot t; N0 is the background Gaussian additive white noise on the wireless channel during task transmission; v r (t) is the transmission power of device r for transmitting data in time slot t; g r (t) is the channel gain of the wireless channel during task transmission, and this value follows an exponential distribution with a mean of , γ is the path loss constant, d r,jDenote the distance between the local device r and the edge server j; ∑ k∈R,k≠r v k (t)g k (t) is the interference noise from other devices except device r on the wireless channel. Considering the heterogeneity of local devices, the transmission rates of each device on the wireless channel are different, and the transmission rate is affected by the distance between the local device and the edge server and the transmission power.
[0133] Furthermore, after the coordinator receives the offloading request sent by the local device, it allocates the edge server to which the local device is to be offloaded to improve the rate of computing result delivery. Set the offloading vector of task i on the edge server to be β i β i When β = 1, the task is executed on the edge server. When β i = 0, the task is offloaded to the cloud server for execution.
[0134] At time slot t, the amount of task data offloaded from the local device to the remote device is denoted as:
[0135]
[0136] The amount of task data received by the edge server n at time slot t is:
[0137]
[0138] The energy consumption on the edge server is divided into three parts. The first part is the energy consumption for data uplink wireless transmission, the second part is the energy consumption for task data calculation, and the third part is the energy consumption for transmitting the data calculation result back. Since the amount of data of the calculation result is small and can be ignored compared with the uplink transmission energy consumption, let (Cycles / s) be the processing frequency of the edge server The energy consumption and delay on the edge server are respectively:
[0139]
[0140] In the formula, represents the energy consumption required for the edge server n to complete the amount of task data at time slot t, represents the amount of task data offloaded to the remote device, v k (t) represents the transmission power of the mobile device r for transmitting data in time slot t, ξ represents the effective capacitance switching coefficient, τ represents the duration of time slot t, represents the processing frequency of the edge server n.
[0141] Similarly, if a task execution requires a high computing frequency, the coordination server will offload the task to the cloud server for execution. Since the computing frequency of the cloud server is very high, its computing energy consumption and downlink transmission energy consumption can be ignored. The energy consumption and latency required by the cloud service are as follows:
[0142]
[0143] In the formula, represents the energy consumption required for the cloud server to complete the task data volume in time slot t, represents the set of heterogeneous mobile devices, i represents the heterogeneous task, and β i represents the offloading vector.
[0144] Therefore, the total energy consumption in the system is:
[0145]
[0146] is the execution time of task i on the edge server, is the execution time of the task on the cloud server. The specific mathematical formula is as follows:
[0147] The latency on the edge server is shown in the following formula (14):
[0148]
[0149] In the formula, represents the execution time of task i on the edge server, represents the task data size, and α i represents the proportion of heterogeneous task i processed by the mobile device, and R r,n (t) represents the rate at which mobile device r uploads data and instructions to the edge server in time slot t, ε represents the number of CPU cycles required to execute 1 bit of the task, I represents the set of heterogeneous tasks, represents the set of heterogeneous mobile devices, represents the set of edge servers.
[0150] The latency on the cloud server is shown in the following formula (15):
[0151]
[0152] In the formula, represents the execution time of task i on the cloud server, and R n,c (t) represents the rate at which edge server n uploads data and instructions to the cloud server in time slot t.
[0153] The time for remotely executing task D i,r is:
[0154]
[0155] S4. Based on the meta - learning method, the mobile device performs inner - loop optimization according to the total energy consumption and latency of the industrial system model, the meta - learning policy generated by the edge server, and the deep deterministic policy gradient algorithm to obtain a task - specific policy; the edge server performs outer - loop optimization according to the task - specific policy and the reinforcement learning method to generate a new meta - learning policy; and then a task offloading scheme based on the dynamic channel state and queue length is obtained.
[0156] In a feasible implementation manner, based on the above - mentioned industrial service environment, the overall implementation process of the present invention is as Figure 4 shown, and the scheme process is described below. First, different types of heterogeneous tasks are randomly generated on the local terminal device. Since the computing resources of each local device and the edge server are different, the task offloading strategy and resource allocation scheme are trained. The optimization goal of the present invention is to minimize the scheme training time on the premise of obtaining the optimal resource allocation scheme. The meta - learning method is used for training, and the inner - loop method is DDPG to optimize the system execution energy consumption.
[0157] Optionally, the goal of the inner - loop optimization in S4 is to minimize the total energy consumption of the industrial system model under the premise of meeting the task latency constraint, as shown in the following formulas (17)-(24):
[0158]
[0159] 0≤v r (t)≤v(t) r,max (19)
[0160]
[0161] α i ∈[0,1] (21)
[0162] β i ∈{0,1} (22)
[0163]
[0164] In the formula, T represents the task deadline, E total (t) represents the total energy consumption of the industrial system model, represents the set of mobile devices, w r,n (t) represents the channel bandwidth allocated by the system when the mobile device r offloads tasks to the edge server n at time slot t, w n represents the channel bandwidth of the edge server n. Formula (18) indicates that the channel bandwidth of the mobile device is less than that of the edge server, v r(t) represents the transmit power of the mobile device r for transmitting data in time slot t, v(t) r,max represents the maximum transmit power of the mobile device r for transmitting data in time slot t, μ r (t) represents the amount of computing task data leaving the task buffer of the mobile device in time slot t, c represents the number of CPU cycles required to execute 1 bit of task, generally c = 740 (Cycle / bit), represents the processing frequency of the edge server n, τ represents the duration of time slot t, α i represents the proportion of the heterogeneous task i processed on the mobile device, β i represents the offloading vector.
[0165] Among them, constraint (18) represents the bandwidth allocation constraint of the base station. (19) represents the transmit power constraint. Constraint condition (20) means that the amount of tasks processed by the edge server is μ r (t) The amount of computing resources does not exceed the available computing resources. (21) represents the partial offloading decision of the task on the local device, ensuring that the proportion of the divisible computing task executed locally and on the edge server is non - negative. (22) represents the 0, 1 offloading decision of the task on the edge server, ensuring the execution efficiency of the task when the task buffer queue of the edge server is long. Constraints (23) and (24) ensure the stability of all task queues.
[0166] Optionally, the outer - loop optimization in S4 includes:
[0167] Making a decision on the task data routing path according to the routing topology relationship between edge servers; among them, the state space is the edge server node where the task is located, the action space is all edge servers and the cloud server, and the corresponding reward function is set according to the routing topology relationship.
[0168] Taking the transmit power of the channel, channel bandwidth, length of the mobile device task buffer queue and length of the edge server task buffer queue in time slot t as the state space of the meta - reinforcement learning algorithm, and the computing frequencies of heterogeneous mobile devices and edge servers as the action space, to allocate the optimal computing frequency and offloading decision parameters for each edge server.
[0169] In a feasible implementation manner, the optimization objective of the outer loop: the optimal generalization learning strategy: use the reinforcement learning method to make decisions on the task data routing path according to the routing topology relationship between edge servers, where the state space is the device node where the task is located, the action space is all edge server devices and cloud server devices, and the corresponding reward function is set according to the path relationship. At the same time, the transmission power of the channel, the channel bandwidth, the local and edge server task buffer queue lengths at time slot t are used as the state space of the meta-reinforcement learning algorithm, and the device computing frequency is used as the action space to allocate optimal computing frequency, offloading decision and other parameters for each edge server. The pseudo-code of the optimization algorithm is as Figure 4 .
[0170] First, the local device downloads the parameters of the meta-learning strategy from the edge server and performs "inner loop" training on each local device based on the meta-policy and local data to obtain the task-specific policy. Then, after receiving the policy parameters uploaded by all local devices, the edge server performs "outer loop" training of meta-learning to generate a new meta-policy and starts a new round of training. Once a stable meta-policy is obtained, it can be used to quickly learn the task-specific policy of the new UE through "inner loop" training. The inner loop training only requires a small number of training steps and a small amount of data. Through this algorithm, the energy consumption of the industrial park can be minimized while ensuring the quality of service, and efficient resource allocation and fast task processing can be achieved.
[0171] The present invention proposes a DDPG double-layer optimization framework based on MAML (Model-Agnostic Meta-Learning), which optimizes the task-specific policy in the inner loop and updates the meta-policy in the outer loop. Based on a three-layer cloud-edge-end collaborative architecture, a dynamic resource allocation mechanism for local devices, edge servers, and cloud servers is proposed. System stability is ensured through Lyapunov optimization.
[0172] In the embodiment of the present invention, research is carried out on the joint optimization mechanism of task offloading and communication computing resources for heterogeneous industrial scenarios, a variety of heterogeneous industrial service models and a three-layer cloud-edge-end collaborative transmission architecture are established, and training is carried out for a variety of heterogeneous tasks and different devices, which is a reinforcement learning algorithm with strong generalization ability.
[0173] Figure 5 It is a block diagram of a scheduling and allocation device for multi-layer heterogeneous computing power resources in an industrial Internet shown according to an exemplary embodiment. This device is used for the scheduling and allocation method of multi-layer heterogeneous computing power resources in the industrial Internet. Refer to Figure 5 , this device includes a construction module 310, a task generation module 320, a total energy consumption and delay calculation module 330, and an output module 340. Among them:
[0174] A building module 310 for building an industrial system model; the industrial system model includes a cloud server, multiple edge servers, and multiple heterogeneous mobile devices; the multiple edge servers include a coordination server and other edge servers.
[0175] A task generation module 320 for generating a heterogeneous task set for multiple heterogeneous mobile devices; the heterogeneous task set includes mobile device computing tasks and remote server computing tasks; the remote server computing tasks are offloaded to edge servers or cloud servers through the coordination server.
[0176] A total energy consumption and latency calculation module 330 for constructing a local computing model and a remote computing model based on the heterogeneous task set, and obtaining the total energy consumption and latency of the industrial system model according to the local computing model and the remote computing model.
[0177] An output module 340 for, based on the meta-learning method, the mobile devices perform inner-loop optimization according to the total energy consumption and latency of the industrial system model, the meta-learning strategy generated by the edge servers, and the deep deterministic policy gradient algorithm to obtain a task-specific policy; the edge servers perform outer-loop optimization according to the task-specific policy and the reinforcement learning method to generate a new meta-learning strategy; and then obtain a task offloading scheme based on the dynamic channel state and queue length.
[0178] In the embodiments of the present invention, research on a task offloading and communication computing resource joint optimization mechanism is carried out for heterogeneous industrial scenarios, multiple heterogeneous industrial service models and a three-layer cloud-edge-end collaborative transmission architecture are established, and training is performed for multiple heterogeneous tasks and different devices, which is a reinforcement learning algorithm with strong generalization ability.
[0179] Figure 6 It is a schematic structural diagram of a scheduling and allocation device provided by an embodiment of the present invention. As Figure 6 shown, the scheduling and allocation device may include the above-mentioned Figure 5 scheduling and allocation device for industrial Internet multi-layer heterogeneous computing power resources shown. Optionally, the scheduling and allocation device 410 may include a first processor 2001.
[0180] Optionally, the scheduling and allocation device 410 may further include a memory 2002 and a transceiver 2003.
[0181] Wherein, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, such as through a communication bus.
[0182] Next, in combination with Figure 6 each component of the scheduling and allocation device 410 will be specifically introduced:
[0183] Among them, the first processor 2001 is the control center of the scheduling and allocation device 410, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0184] Optionally, the first processor 2001 can execute various functions of the scheduling and allocation device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0185] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 6 CPU0 and CPU1 shown in
[0186] In a specific implementation, as an embodiment, the scheduling and allocation device 410 may also include multiple processors, such as Figure 6 the first processor 2001 and the second processor 2004 shown in. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0187] Among them, the memory 2002 is used to store software programs for executing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0188] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 6 not shown) of the scheduling and allocation device 410. The embodiments of the present invention do not make specific limitations thereto.
[0189] A transceiver 2003 is configured to communicate with a network device or communicate with a terminal device.
[0190] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 6 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0191] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 6 not shown) of the scheduling and allocation device 410. The embodiments of the present invention do not make specific limitations thereto.
[0192] It should be noted that Figure 6 the structure of the scheduling and allocation device 410 shown does not constitute a limitation to the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0193] In addition, the technical effects of the scheduling and allocation device 410 may refer to the technical effects of the method for scheduling and allocating industrial Internet multi-layer heterogeneous computing power resources described in the above method embodiments, and will not be elaborated herein.
[0194] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0195] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0196] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more sets of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0197] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally indicates an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0198] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or plural.
[0199] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0200] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0201] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0202] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0203] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0204] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0205] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0206] As described above, the above are only specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A scheduling and allocation method for multi-layer heterogeneous computing power resources in industrial Internet, characterized in that The method includes: S1. Construct an industrial system model. The industrial system model includes a cloud server, multiple edge servers, and multiple heterogeneous mobile devices. The multiple edge servers include a coordination server and other edge servers. S2. The multiple heterogeneous mobile devices generate a heterogeneous task set. The heterogeneous task set includes mobile device computing tasks and remote server computing tasks. The remote server computing tasks are offloaded to the edge server or the cloud server through the coordination server. S3. Construct a local computing model and a remote computing model according to the heterogeneous task set, and obtain the total energy consumption and delay of the industrial system model according to the local computing model and the remote computing model. S4. Based on the meta-learning method, the mobile device performs inner-loop optimization according to the total energy consumption and delay of the industrial system model, the meta-learning strategy generated by the edge server, and the deep deterministic policy gradient algorithm to obtain a task-specific policy. The edge server performs outer-loop optimization according to the task-specific policy and the reinforcement learning method to generate a new meta-learning strategy, thereby obtaining a task offloading scheme based on the dynamic channel state and queue length.
2. The scheduling and allocation method of industrial Internet multi-layer heterogeneous computing power resources according to claim 1, characterized in that Any one of the multiple heterogeneous mobile devices, denoted as mobile device r, is equipped with a mobile device task buffer, and the number of queued task data bits in the mobile device task buffer at time slot t is dynamically updated according to the following queue formula (1): In the formula, represents the number of task data backlogs of the mobile device r at time slot t + 1, and μ r (t) represents the amount of computing task data leaving the task buffer of the mobile device at time slot t. i represents a heterogeneous task, represents the task data generated by the mobile device r at time slot t; Any one of the multiple edge servers, edge server n, is provided with an edge server task buffer, and the number of bits of task data backlog in the edge server task buffer at time slot t is dynamically updated according to the queue in the following formula (2): In the formula, represents the number of queued task data bits in the task buffer of the edge server at time slot t + 1, represents the amount of task data calculated by edge server n at time slot t, represents the amount of task data unloaded from the mobile device to edge server n at time slot t.
3. The scheduling and allocation method of industrial Internet multi-layer heterogeneous computing power resources according to claim 1, wherein, The local computing model in S3 includes: The energy consumption required for the mobile device to complete the task data volume in the current time slot is shown in the following formula (3): wherein, represents the energy consumption required for the mobile device r to complete the task data volume in time slot t, and ξ represents the effective capacitance switching coefficient, represents the processing frequency of the mobile device r in time slot t.
4. The scheduling and allocation method of the industrial Internet multi-layer heterogeneous computing power resources according to claim 1, wherein The remote computing model in S3 includes: The energy consumption on the edge server is shown in the following formula (4): Wherein, represents the energy consumption required for the edge server n to complete the task data volume in time slot t, represents the task data volume unloaded to the remote device, and v k (t) represents the transmit power of the mobile device r to transmit data in time slot t, and R r,n (t) represents the rate at which the mobile device r uploads data and instructions to the edge server in time slot t, ξ represents the effective capacitance switching coefficient, τ represents the duration of time slot t, represents the processing frequency of the edge server n; The delay on the edge server is shown in the following formula (5): In the formula, represents the execution time of task i on the edge server, represents the task data size, and α i represents the proportion of heterogeneous task i processed on the mobile device, and R r,n (t) represents the rate at which mobile device r uploads data and instructions to the edge server at time slot t, ε represents the number of CPU cycles required to execute 1 bit of task, I represents the set of heterogeneous tasks, represents the set of heterogeneous mobile devices, represents the set of edge servers.
5. The scheduling and allocation method of the industrial Internet multi-layer heterogeneous computing power resources according to claim 1, characterized in that, The remote computing model in S3 further includes: The energy consumption on the cloud server is shown in the following formula (6): In the formula, represents the energy consumption required for the cloud server to complete the task data volume in time slot t, represents the set of heterogeneous mobile devices, i represents heterogeneous tasks, and β i represents the offloading vector, and R r,n (t) represents the rate at which mobile device r uploads data and instructions to the edge server in time slot t; The delay on the cloud server is shown in the following formula (7): wherein, represents the execution time of task i on the cloud server, and R n,c (t) represents the rate at which edge server n uploads data and instructions to the cloud server at time slot t.
6. The scheduling and allocation method for multi-layer heterogeneous computing power resources of the industrial Internet according to claim 1, characterized in that, The goal of the inner-loop optimization in S4 is to minimize the total energy consumption of the industrial system model under the task delay constraint, as shown in the following formulas (8)-(15): 0≤v r (t)≤v(t) r,max (10) α i ∈[0,1] (12) β i ∈{0,1} (13) where T represents the task deadline, and E total (t) represents the total energy consumption of the industrial system model, represents the set of mobile devices, and w r,n (t) represents the channel bandwidth allocated by the system when mobile device r offloads tasks to edge server n at time slot t, and w n represents the channel bandwidth of edge server n. Equation (9) indicates that the channel bandwidth of the mobile device is less than that of the edge server, and v r (t) represents the transmit power of mobile device r for transmitting data at time slot t, and v(t) r,max represents the maximum transmit power of mobile device r for transmitting data at time slot t, and μ r (t) represents the amount of computational task data leaving the task buffer of the mobile device at time slot t. c represents the number of CPU cycles required to execute 1 bit of the task, represents the processing frequency of edge server n. τ represents the duration of time slot t, and α i represents the proportion of heterogeneous task i processed by the mobile device, and β i represents the offloading vector, represents the number of bits of task data backlog in the task buffer of the mobile device at time slot t, represents the set of edge servers, represents the number of bits of task data backlog in the task buffer of the edge server at time slot t.
7. The scheduling and allocation method for multi-layer heterogeneous computing power resources in the industrial Internet according to claim 1, wherein The outer-loop optimization in S4 includes: Making a decision on the task data routing path according to the routing topology relationship between edge servers. Among them, the state space is the edge server node where the task is located, the action space is all edge servers and the cloud server, and a corresponding reward function is set according to the routing topology relationship. Taking the transmission power, channel bandwidth, mobile device task buffer queue length, and edge server task buffer queue length of the channel at time slot t as the state space of the meta-reinforcement learning algorithm, and the computing frequencies of heterogeneous mobile devices and edge servers as the action space, and allocating optimal computing frequencies and offloading decision parameters for each edge server.
8. A scheduling and allocation device for multi-layer heterogeneous computing power resources in an industrial Internet, the scheduling and allocation device for multi-layer heterogeneous computing power resources in the industrial Internet is used to implement the scheduling and allocation method for multi-layer heterogeneous computing power resources in the industrial Internet as described in any one of claims 1-7, characterized in that, The device includes: A construction module for constructing an industrial system model. The industrial system model includes a cloud server, multiple edge servers, and multiple heterogeneous mobile devices. The multiple edge servers include a coordination server and other edge servers. A task generation module for the multiple heterogeneous mobile devices to generate a heterogeneous task set. The heterogeneous task set includes mobile device computing tasks and remote server computing tasks. The remote server computing tasks are offloaded to the edge server or the cloud server through the coordination server. The total energy consumption and latency calculation module is used to construct a local computing model and a remote computing model according to the heterogeneous task set, and obtain the total energy consumption and latency of the industrial system model according to the local computing model and the remote computing model; The output module is used to, based on the meta-learning method, the mobile device performs inner-loop optimization according to the total energy consumption and latency of the industrial system model, the meta-learning strategy generated by the edge server, and the deep deterministic policy gradient algorithm to obtain a task-specific policy; the edge server performs outer-loop optimization according to the task-specific policy and the reinforcement learning method to generate a new meta-learning strategy; and then obtains a task offloading scheme based on the dynamic channel state and queue length.
9. A scheduling and allocation device, characterized in that The scheduling and allocation device includes: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method described in any one of claims 1 to 7.