Method for unloading computing task and related product
Patent Information
- Application Number
- CN202510400289.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
Smart Images

Figure CN120256122A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the technical field of serverless computing in the industrial Internet of Things. More specifically, this disclosure relates to a method and a computer-readable storage medium for offloading computing tasks. Further, this disclosure also relates to a serverless edge computing system. Background Art
[0002] The cold start problem in serverless computing, that is, the latency of function instance instantiation, mainly stems from the need to restart the runtime environment after the function is first requested or the instance is destroyed. Common solutions include two strategies: lightweight functions and pre-warmed function instances. The lightweight method shortens the instance startup time through technical means (such as WebAssembly), but cannot truly avoid the occurrence of cold starts; the pre-warming and reuse strategy maintains the instance in an active state in advance by predicting the function request pattern, but has a large error in multi-task types and dynamic environments, easily leading to resource waste. Existing commercial serverless platforms usually set the function instance survival time to 5 to 15 minutes, and the refined management of different task types remains a challenge.
[0003] Offloading the computing task orchestration to an already active instance to avoid restarting a new instance and avoid the occurrence of cold starts is an effective method. Existing methods for making decisions on computing task offloading actions are mainly divided into two categories: traditional methods and reinforcement learning methods. Traditional methods (such as game theory and heuristic algorithms) are difficult to meet the requirements of large-scale computing task offloading in complex dynamic environments, while the offloading action decision based on deep reinforcement learning can dynamically learn the optimal decision and achieve the balanced optimization of service latency and resource cost in multi-objective scenarios. However, existing methods based on deep reinforcement learning mainly focus on the optimization of stateful computing tasks, lack the consideration of the cold start latency of computing tasks in serverless edge computing, and do not fully utilize the distributed nature of edge computing, resulting in low efficiency of serverless function computing, seriously affecting the service quality and user experience.
[0004] In view of this, there is an urgent need to provide a solution for offloading computing tasks to reduce the latency of serverless computing tasks, improve the efficiency of serverless function computing, and enhance the service quality and user experience. Summary of the Invention
[0005] To solve at least one or more of the above-mentioned technical problems, this disclosure proposes a solution for offloading computing tasks in the following aspects.
[0006] In a first aspect, the present disclosure provides a method for offloading computing tasks, which is applied to an agent node deployed on an Internet of Things (IoT) device. The IoT device is a node in a serverless edge computing environment, and the nodes in the serverless edge computing environment further include an edge server and a cloud server. The method includes: when receiving a computing task request generated by the IoT device, obtaining the current environmental state of the serverless edge computing environment; the computing task request includes task attributes of the computing task, and the task attributes include task data volume, required CPU cycles, and maximum tolerable latency; the current environmental state includes the computing resource status and function code pool status of each node in the serverless edge computing environment, and the network link status; inputting the current environmental state into a deep reinforcement learning model for decision-making calculations to output an offloading action decision, where the offloading action decision indicates a target node for executing the computing task; allocating network resources and computing resources for the computing task according to the task attributes and the offloading action decision; and based on the network resources and computing resources, sending the computing task to the target node so that the target node executes the computing task.
[0007] In some embodiments, after the target node finishes executing the computing task, the method further includes: receiving the execution result fed back by the target node, and counting the system cost of the target node for completing the execution of the computing task; determining a reward value using a preset reward function based on the execution result and the system cost; obtaining the next environmental state of the serverless edge computing environment, and taking the current environmental state, the offloading action decision, the reward value, and the next environmental state as offloading experiences; assigning priorities to the offloading experiences, and uploading the offloading experiences and the priorities to an experience pool in the cloud server.
[0008] In some embodiments, the deep reinforcement learning model is a double Dueling DQN neural network model; after uploading the offloading experiences and the priorities to the experience pool in the cloud server, the cloud server performs the following operations to implement training of the double Dueling DQN neural network model: selecting target offloading experiences from the experience pool using a preset sampling probability, and calculating the importance weights of the target offloading experiences; and taking the target offloading experiences and the importance weights as training data and inputting them into the double Dueling DQN neural network model for training.
[0009] In some embodiments, the dual Dueling DQN neural network model includes an online network and a target network; inputting the target offloading experience and the importance weight as training data into the dual Dueling DQN neural network model for training includes: inputting the current environmental state and offloading action decision in the target offloading experience into the online network for value calculation to output the first value of the offloading action decision; based on the next environmental state in the target offloading experience, using the target network to calculate the second value of each offloading action decision in the offloading action decision set to output a plurality of second values; selecting the largest second value from the plurality of second values, and using the reward value and the largest second value in the target offloading experience to calculate the target value; determining the loss value based on the first value, the target value, and the importance weight; updating the parameters of the online network based on the loss value, and periodically synchronizing the updated parameters of the online network to the target network.
[0010] In some embodiments, the method further includes: distributing the updated parameters of the online network to the intelligent agent nodes.
[0011] In some embodiments, the preset reward function can be represented in the following form:
[0012]
[0013] where n represents the nth Internet of Things device in the service-free edge computing environment, i represents the ith computing task, and r n_i (t) represents the reward value determined based on the execution result and the system cost; represents the positive reward value when the execution result is successful execution, represents the negative reward value when the execution result is failed execution; Sys_cost n_i represents the negative reward value brought by the system cost of executing the ith computing task of the nth Internet of Things device; e and f are balance parameters.
[0014] In some embodiments, the system cost includes delay, energy consumption, and monetary cost; the delay includes one or more of waiting transmission delay, transmission delay, waiting resource allocation delay, function function instance restart delay, and computing delay; the energy consumption includes at least one of computing task transmission energy consumption and computing task execution energy consumption; the monetary cost is the cost of offloading to the edge server or the cost of offloading to the cloud server.
[0015] In some embodiments, the cost of offloading to the edge server includes the cost of requesting edge computing and the cost generated by the service duration of the edge server; the cost of offloading to the cloud server includes the cost of requesting cloud computing and the cost generated by the service duration of the cloud server.
[0016] In a second aspect, the present disclosure provides a serverless edge computing system, including: an Internet of Things device, an edge server, a cloud server, and a communication module, wherein an agent proxy node is deployed on the Internet of Things device; the Internet of Things device is configured to generate a computing task request; the agent proxy node is configured to execute the method and its multiple embodiments described in the foregoing first aspect so as to offload the computing task to a target node; the edge server maintains a function code pool and receives and executes the computing task sent by the Internet of Things device; the cloud server is configured to train a deep reinforcement learning model, and a global code repository and an experience pool are stored thereon; the communication module is configured to implement data transmission between the Internet of Things device, the edge server, and the cloud server.
[0017] In a third aspect, the present disclosure provides a computer-readable storage medium, on which program instructions for offloading a computing task are stored. When the program instructions are executed by a processor, the method and its multiple embodiments described in the foregoing first aspect are implemented.
[0018] Through the solution for offloading a computing task provided as above, when the agent proxy node receives a computing task request generated by the Internet of Things device, it inputs the task attributes of the computing task and the current environmental state of the serverless edge computing environment into the deep reinforcement learning model for decision-making calculation, so as to output an offloading action decision. Since the task attributes of the computing task include the task data volume, the number of required CPU cycles, and the maximum tolerable delay, and the current environmental state includes the computing resource states and the function code pool states of each node in the serverless edge computing environment, as well as the network link state. Therefore, the offloading action decision output by the model fully considers the function code pool states of each node, and can preferentially select active function instances to execute the computing task, realizing the hot start of the function. This deep reinforcement learning decision-making method with cold start awareness can avoid the delay caused by cold start, thereby improving the efficiency of serverless function computing, and enhancing the service quality and user experience.
[0019] Furthermore, through the solution of the present disclosure, after the deep reinforcement learning model outputs an offloading action decision, network resources and computing resources can be allocated for the computing task according to the offloading action decision. That is to say, different offloading action decisions result in different allocated network resources and computing resources, thereby realizing dynamic resource allocation, which helps to improve the resource utilization rate of network resources and computing resources. In addition, the agent nodes are deployed on physical network devices, enabling each physical network device in the serverless edge computing environment to make independent decisions, making full use of the distributed nature of edge computing, and being able to avoid the communication delay and single point of failure of cloud centralized processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0021] Figure 1 FIG. shows an exemplary structural diagram of a serverless edge computing system according to an embodiment of the present disclosure;
[0022] Figure 2 FIG. shows an exemplary flowchart of a method for offloading a computing task according to an embodiment of the present disclosure;
[0023] Figure 3 FIG. shows an exemplary flowchart of a process for obtaining offloading experience according to an embodiment of the present disclosure;
[0024] Figure 4 FIG. shows an exemplary flowchart of a process for training a deep reinforcement learning model according to an embodiment of the present disclosure;
[0025] Figure 5 FIG. shows an exemplary schematic diagram of an interaction process between an agent node and a cloud server according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0027] It should be understood that the terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0028] It should also be understood that the terms used in this disclosure specification are for the purpose of describing particular embodiments only and are not intended to limit this disclosure. As used in this disclosure specification and claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should be further understood that the term "and / or" as used in this disclosure specification and claims refers to any combination and all possible combinations of one or more of the associated listed items and includes these combinations.
[0029] As used in this specification and claims, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0030] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.
[0031] Figure 1 An exemplary structural diagram of a serverless edge computing system according to an embodiment of this disclosure is shown. As Figure 1 shown, the serverless edge computing system of this disclosure may include Internet of Things (IoT) devices (which may also be referred to as IoT terminal devices) 101, an edge server 102, a cloud server 103, and a communication module 104 (such as a wireless communication module). The IoT devices 101, the edge server 102, and the cloud server 103 may be regarded as nodes (which may also be referred to as computing nodes) of the serverless edge computing system, and the communication module is used to implement data transmission between the IoT devices, the edge server, and the cloud server. It can be understood that in actual application scenarios, the number of IoT devices and the number of edge servers in the serverless edge computing system may be any specific values selected according to actual needs, and this disclosure does not limit this. A smart agent node may be deployed on each IoT device in the serverless edge computing system, and the smart agent node may be used to execute the following in combination with Figures 2 to 4The described method for offloading computing tasks. In addition, a function code pool (i.e., SF code pool) is maintained on the cloud server, each Internet of Things device, and each edge server in the serverless edge computing system.
[0032] The SF code pool of the cloud server is a remote core code repository, which is a global code repository that stores complete serverless function code images and provides a source for code pulling for edge servers and Internet of Things devices. The SF code pools of edge servers and Internet of Things devices pull frequently used function code from the cloud server based on historical computing tasks and pre-load it locally. For example, the edge server stores common function codes such as industrial monitoring and data processing in its local SF code pool, and the Internet of Things device also caches some lightweight function codes as needed. When executing a computing task, the loaded function instance is preferentially obtained from the SF code pool of the local or neighboring node (such as an edge server) to avoid cold start. If the required function is missing from the local code pool, the code is pulled from the cloud for instantiation, forming an optimized link for code distribution and caching of "cloud - edge - device" to improve task execution efficiency.
[0033] Assume that in the cloud-edge-serverless edge computing environment of this disclosure, N Internet of Things devices N = {1, 2, 3,..., N} and M edge servers M = {1, 2, 3,..., M} are equipped. Then the time is divided into equally spaced discrete time slots {1, 2, 3,..., T}, and there is a fixed time interval τ in each time slot t. The Internet of Things device loT generates an application request with a certain probability at the beginning of each time slot, and the client of the Internet of Things device decomposes the application request into i independent computing task requests of function functions SF. The computing task request can be expressed as Here represents the i-th computing task request of Internet of Things device n in time slot t and is a computing task request belonging to function function f, is the task attribute of the computing task. Specifically, represents the task data volume, represents completing the computing task the number of CPU cycles required, represents the maximum tolerable delay for completing the computing task, represents the type set of serverless function functions.
[0034] After generating the computing task request , the Internet of Things device n will send the computing task request to the intelligent agent node deployed on it. After the intelligent agent node receives the computing task request , it needs to make a decision on the offloading action of the computing task where a nn , anm1 , a nm2 , …, a nmm , a nc ∈(0, 1), and a nn + a nm1 + a nm2 + … + a nmm + a nc =1. For example, when a nn =1, the rest are 0, indicating that this computing task will be executed on the local Internet of Things device (i.e., the Internet of Things device n that issues the computing task request). When a nm1 (t)=1, the rest are 0, and this computing task will be offloaded to the edge server numbered 1 for execution, and so on. a nc =1 means that this computing task will be offloaded to the cloud server for execution. That is to say, in the embodiments of the present disclosure, offloading refers to the intelligent agent node of the Internet of Things device deciding whether to leave the computing task for local processing or send it to the edge server or the cloud server for execution according to the current environmental state. Its essence is a cloud-edge-end collaborative resource scheduling strategy. To facilitate understanding of how the intelligent agent node makes the offloading action decision, the method 200 for offloading computing tasks will be described in detail later in conjunction with Figure 2 and will not be elaborated here.
[0035] In the embodiments of the present disclosure, an experience pool is stored on the cloud server. The offloading experiences and their priorities in the experience pool can be used to train a deep reinforcement learning model. To facilitate understanding of how to obtain offloading experiences and how to train the deep reinforcement learning model, the process 300 of obtaining offloading experiences and the process 400 of training the deep reinforcement learning model will be described in detail later in conjunction with Figure 3 and Figure 4 respectively and will not be elaborated here.
[0036] Figure 2 FIG. shows an exemplary flowchart of the method 200 for offloading computing tasks in the embodiments of the present disclosure. The method 200 is applied to the intelligent agent node deployed on the Internet of Things device. The Internet of Things device is a node in a serverless edge computing environment. The nodes in the serverless edge computing environment may also include edge servers and cloud servers. As described above, the number of Internet of Things devices and the number of edge servers in the serverless edge computing system can be any specific values selected according to actual needs, and the steps of the method 200 executed by the intelligent agent node deployed on each Internet of Things device are the same. For the sake of brevity, here only the intelligent agent node deployed on one Internet of Things device is taken as an example to describe the specific steps of the method 200.
[0037] Based on this, as Figure 2As shown, at step S201, when the method 200 receives a computing task request generated by an Internet of Things device, it obtains the current environmental state of the serverless edge computing environment. Here, the computing task request includes the task attributes of the computing task, which may include the amount of task data, the number of required CPU cycles, and the maximum tolerable latency. The current environmental state includes the computing resource status and the functional function code pool status of each node in the serverless edge computing environment, as well as the network link status. Here, the computing resource status may include, but is not limited to, the CPU utilization rate and the memory occupancy rate. The functional function code pool status may include, but is not limited to, the list of functional functions stored in the code pool, the instance survival status (active or inactive) corresponding to each functional function, the instance creation time, the number of available warm-start instances, the list of functions to be loaded for cold start, and the remaining storage space. The network link status may include, but is not limited to, the remaining bandwidth of the wireless channel, the signal strength, and the network latency.
[0038] Next, at step S202, the method 200 may input the current environmental state into a deep reinforcement learning model for decision-making calculations to output an offloading action decision. The offloading action decision disclosed herein represents the offloading action to be performed for the calculation, such as executed by the Internet of Things device, the edge server, or the cloud server. In other words, the offloading action decision indicates the target node for executing the computing task. Then, at step S203, the method 200 may allocate network resources and computing resources for the computing task according to the task attributes and the offloading action decision. Finally, at step S204, the method 200 may send the computing task to the target node based on the network resources and the computing resources so that the target node executes the computing task.
[0039] At the aforementioned step S202, the deep reinforcement learning model may be a double Dueling DQN (Deep Q-Network) neural network model. The main idea of Dueling DQN is to decompose the Q-value function into two parts: the state value function and the advantage function, which can better estimate the contribution of different offloading actions to the environmental state and improve the learning efficiency. The double Dueling DQN neural network model includes an online network and a target network. The training process is usually carried out on a cloud server. When training the online network, its own parameters can be optimized through gradient descent, and the parameters of the online network are copied to the target network (hard synchronization) regularly (such as every 1000 steps), or the parameters of the target network are gradually updated through a soft update mechanism. This operation serves the training stability and reduces the variance of Q-value estimation. When the cloud server completes a round of training (such as reaching a preset number of iterations and the loss function converges), the optimized parameters of the online network can be pushed to all intelligent agent nodes to ensure that the intelligent agent nodes use the latest online network for decision-making calculations.
[0040] In actual operation, due to the high complexity of the cloud-edge collaborative serverless edge computing environment (such as multi-node, multi-device, dynamic network), it is impossible to fully observe all environmental state changes (such as sudden fluctuations in node resources, instantaneous congestion of network links). Therefore, it can be abstracted into a computable probability system through the Markov Decision Process (MDP) so as to use reinforcement learning (such as Double Dueling DQN) for optimal decision-making. An MDP consists of a five-tuple M = {S, A, ρ, R, γ}, where S represents the state space, A represents the action space, ρ represents the state transition probability, R represents the reward function, and γ represents the discount factor for the importance of future rewards.
[0041] Based on this, the processing process of the foregoing method 200 will be further described as a whole below. The intelligent industrial Internet of Things device generates a computing task and sends a computing task request containing task attributes to the intelligent agent node. After receiving the computing task request, the intelligent agent node needs to perceive the current environmental state S(t) at this time slot, S(t) = {R(t), F n (t), F m (t), F c (t), B(t)}. Among them, R(t) is the profile information of the function to be executed (such as the computational complexity and real-time requirements of the function, computational resource requirements, code pool status), F n (t) represents the computing resources of all Internet of Things devices, F m (t) and F c (t) respectively represent the computing resources of all edge servers and cloud servers. B(t) represents the network link state at the current time t (such as the remaining bandwidth of the wireless channel).
[0042] Next, for the computing task the intelligent agent node can use the deep learning model to output an offloading action decision a ni (t) = {a nn , a nm1 , a nm2 , …, a nmm , a nc} according to the task attributes and the current environmental state. Here, a ni (t) represents the offloading action decided to be executed for the computing task , such as being executed by the local Internet of Things device (i.e., the intelligent industrial Internet of Things device that sends the computing task request), being executed by the edge server, or being executed by the cloud server.
[0043] As described above in combination with Figure 2A method for offloading computing tasks is described. When the intelligent agent node receives a computing task request generated by an Internet of Things device, it makes a decision calculation by inputting the task attributes of the computing task and the current environmental state of the serverless edge computing environment into a deep reinforcement learning model, so as to output an offloading action decision. Since the task attributes of the computing task include the task data volume, the number of required CPU cycles, and the maximum tolerable delay, and the current environmental state includes the computing resource state and the function code pool state of each node in the serverless edge computing environment, as well as the network link state. Therefore, the offloading action decision output by the model fully considers the function code pool state of each node, and can preferentially select active function instances to execute computing tasks, realizing the hot start of function functions. This cold-start-aware deep reinforcement learning decision-making method can avoid the delay caused by cold start, thereby improving the efficiency of serverless function computing, enhancing service quality and user experience.
[0044] In order to improve the decision-making ability and decision-making quality of the deep reinforcement learning model, in some implementation scenarios, after the target node finishes executing the computing task, it can also collect the execution result feedback by the target node, as well as the environmental state (which can also be called the next environmental state) of the serverless edge computing environment after the computing task is completed, as offloading experience for further optimization of the deep reinforcement learning model. Based on this, the present disclosure further provides a process 300 for obtaining offloading experience. Figure 3 Fig. shows an exemplary flowchart of the process 300 for obtaining offloading experience according to an embodiment of the present disclosure.
[0045] As Figure 3 shown, at step S301, the execution result feedback by the target node can be received, and the system cost of the target node completing the computing task can be counted. At step S302, based on the execution result and the system cost, a preset reward function can be used to determine the reward value. Then, at step S303, the next environmental state of the serverless edge computing environment can be obtained, and the current environmental state, the offloading action decision, the reward value, and the next environmental state can be used as offloading experience. Finally, at step S304, a priority can be assigned to the offloading experience, and the offloading experience and the priority can be uploaded to the experience pool in the cloud server, so that the cloud server can obtain the offloading experience from the experience pool for training the deep reinforcement learning model.
[0046] In the embodiments of the present disclosure, the execution result is either successful execution or failed execution. The system costs include latency, energy consumption, and monetary cost. Further, the latency may include one or more of waiting transmission latency, transmission latency, waiting resource allocation latency, function instance restart latency, and computing latency. The energy consumption may include at least one of computing task transmission energy consumption and computing task execution energy consumption. The monetary cost is the cost of offloading to the edge server or the cost of offloading to the cloud server. Specifically, the cost of offloading to the edge server includes the cost of requesting edge computing and the cost generated by the service duration of the edge server. The cost of offloading to the cloud server includes the cost of requesting cloud computing and the cost generated by the service duration of the cloud server.
[0047] Next, the calculation methods of the aforementioned latency, energy consumption, and monetary cost will be described in detail. N Internet of Things devices in a serverless edge computing environment are all connected to the edge base station through a wireless channel. To reduce data transmission latency, the edge base station is usually deployed close to the Internet of Things devices to provide network connections for the Internet of Things devices. Thus, at time slot t, the computing task of Internet of Things device n The transmission latency to the edge server x can be calculated as follows: where r n_mx (t) represents the fixed uplink rate at which Internet of Things device n transmits to edge server x. At the same time, if r n_c (t) represents the fixed uplink rate at which Internet of Things device n transmits to the cloud server, then the transmission latency of the computing task to the cloud server can be calculated as follows: Further, the transmission energy consumption E (t) of the computing task can be expressed as: E ni_tX (t) = P ni_tX (t)T n (t), ni_tX (t), where P n (t) represents the transmission rate to the cloud server.
[0048] In some implementation scenarios of the present disclosure, if the offloading action decision of the computing task is to offload to the current local Internet of Things device node, then T ni_n (t) is used to represent the total latency calculated when it is offloaded to Internet of Things device n, and it can be specifically expressed as: T ni_n (t) = W ni_n (t) + S ni_n (t) + C ni_n (t), and the calculation method of the energy consumption is: E ni_n (t) = E ni_tX (t) + E_C ni_n (t).
[0049] Here, W ni_n (t) is the delay waiting for resource allocation, and S ni_n (t) represents the startup delay of the function instance, and C ni_n (t) is the computing delay. Among them, FS f is the resource package size of the serverless function f, and HP n_R is the maximum load for the IoT device n to obtain the serverless function code SF code; f n_n (t) is the processing capacity of the IoT device n. In this implementation scenario, the computing task is offloaded to the current local IoT device node, so the transmission energy consumption E ni_tX (t) is 0, and the processing energy consumption of the computing task Among them is the energy consumption per unit cycle, that is, the energy consumed per CPU cycle.
[0050] In some other implementation scenarios of this disclosure, if the computing task offloading action decision is to offload to the edge server node, the total delay for executing the computing task is: T ni_mx (t) = W ni_tmx (t) + T ni_tmx (t) + W ni_pmx (t) + S ni_mx (t) + T ni_pmx (t). Here, W ni_tmx (t) represents the waiting transmission delay, represents the transmission delay, and W ni_pmx (t) represents the delay waiting for computing resources. The startup delay of the function instance Among them, FS f is the resource package size of the serverless function f, and HP mx_R is the maximum load for the edge server x to obtain the serverless function code SF code. represents the computing delay in the edge server x, where f n_mx (t) represents the computing resource size allocated by the edge server x according to the priority of the nth IoT device at the tth time slot, and satisfies F mx represents the total computing capacity of the edge server x.
[0051] In addition, the total energy consumption of the computing task executed on the edge server x is E ni_mx (t), and the total cost is M ni_mx (t). Specifically, E ni_mx (t) = E ni_tmx (t) + E ni_pmx (t) = Pn (t)T ni_tmx (t); M ni_mx (t) = M ni_rmx (t) + M ni_pmx (t) = M ni_rmx (t) + Price mx *T ni_pmx (t). The aforementioned M ni_rmx (t) is the cost of requesting edge computing, and Price mx is the computing price of the edge server per unit time. In particular, the computing task is not executed on the Internet of Things device, and its computing energy consumption E ni_pmx (t) = 0.
[0052] In some other implementation scenarios of this disclosure, if the offloading action decision of the computing task is to offload it to the cloud server node, then the total delay T ni_c (t) is: T ni_c (t) = W ni_tc (t) + T ni_tc (t) + S ni_c (t) + T ni_pc (t), where W ni_tc (t), T ni_tc (t) respectively represent the delay of the computing task waiting to be transmitted to the cloud server c on the Internet of Things device n and the time to transmit to the cloud server c, and S ni_c (t) represents the startup delay of the functional function example on the cloud server, and T ni_pc (t) represents the computing delay on the cloud server. The aforementioned where HP c_R is the maximum load for the cloud server to obtain the service-free functional function code SF code, and f n_c (t) represents the computing resources allocated by the cloud server to the Internet of Things device n, and satisfies f 1_c (t) = f 2_c (t) =... = f n_c (t).
[0053] In addition, the total energy consumption E of the computing task ni_c (t) executed on the cloud server can be expressed as E ni_c (t) = E ni_tc (t) + E ni_pc (t) = P n (t) × T ni_tc (t), and the total cost M ni_c (t) can be expressed as M ni_c (t) = M ni_rc (t) + M ni_pc (t) = M ni_rc(t) + Price c ×T ni_pc (t), where M ni_rc (t) is the cost of requesting cloud computing, and Price c is the price of the cloud server's service per unit time. In particular, the computing task is not executed on the Internet of Things device, and the computing energy consumption E ni_pc (t) = 0.
[0054] In the embodiments of the present disclosure, the foregoing preset reward function can be represented in the following form:
[0055]
[0056] where n represents the nth Internet of Things device in the serverless edge computing environment, i represents the ith computing task, and r n_i (t) represents the reward value determined based on the execution result and the system cost; represents the positive reward value for a successful execution result, represents the negative reward value for a failed execution result; Sys_cost n_i represents the negative reward value brought by the system cost of executing the ith computing task of the nth Internet of Things device; e and f are balance parameters. In practical applications, by adjusting these two parameters e and f and comprehensively evaluating the performance metrics, the optimal solution under the current conditions is determined. Therefore, the cumulative reward of the reinforcement learning is Thus, the joint optimization problem of function function offloading is transformed into calculating the maximum reward Max(Reward).
[0057] Figure 4 shows an exemplary flowchart of the process 400 for training the deep reinforcement learning model in the embodiments of the present disclosure. As Figure 4 shown, at step S401, the target offloading experience can be selected from the experience pool using the preset sampling probability, and the importance weight of the target offloading experience can be calculated. Then, at step S402, the target offloading experience and the importance weight can be input into the double Dueling DQN neural network model as training data to train it.
[0058] In the embodiments of the present disclosure, when multiple intelligent agent nodes run in the serverless edge computing environment, their respective offloading experiences are not public. Therefore, centralized priority replay is adopted, and when sampling offloading experiences, it tends to prioritize higher-priority offloading experiences. Based on this, the preset sampling probability can be where k is the number of offloading experiences in the experience pool, and p tFor the priority of offloading experience. Traditional greedy screening may result in experiences with low priority but high value to the network not being sampled. Therefore, a sampling mechanism with a preset sampling probability P(t) is used to make the decision engine more inclined to high-priority offloading experiences while retaining the possibility of low-priority offloading experiences being sampled. The aforesaid decides to use the priority magnitude, when it corresponds to the case of uniform randomness, if then select the offloading experience with the highest priority.
[0059] However, high-priority samples introduce preferences and there is a risk of overfitting. Thus, the importance sampling (IS) weight ω t (i.e., the aforesaid importance weight) can be used to correct this bias. Generally, when updating the model, the importance weight ω t is used to adjust the contribution of each sample to the loss function. Specifically, here P(t) represents the preset adoption probability and N represents the size of the experience pool. Generally, at the beginning of training, β takes a small value (such as 0.004) and gradually increases to close to 1 to better correct the bias in the later stage of training.
[0060] At step S402 aforesaid, the target offloading experience and the importance weight are input into the dual Dueling DQN neural network model as training data for training. Specifically, the following operations can be performed: Input the current environmental state and offloading action decision in the target offloading experience into the online network for value calculation to output the first value of this offloading action decision; Based on the next environmental state in the target offloading experience, use the target network to calculate the second value of each offloading action decision in the offloading action decision set to output multiple second values; Select the largest second value from the multiple second values, and based on the reward value and the largest second value in the target offloading experience, use the Bellman equation to calculate the target value; Based on the first value, the target value, and the importance weight, use the mean squared error loss function or the Huber loss function to determine the loss value; Based on the loss value, use backpropagation to update the parameters of the online network, and regularly synchronize the updated parameters of the online network to the target network. Here, the offloading action decision set may include execution by the local Internet of Things device (i.e., the Internet of Things device that sends a computing task request), execution by an edge server (a certain edge server in the edge computing system without services), or execution by a cloud server.
[0061] As a basic deep neural network structure, DQN can output the Q value corresponding to each offloading action decision by taking the environmental state as the input. In the embodiments of this disclosure, in order to improve the accuracy of value Q value estimation and the stability of the learning process, based on the idea of Double DQN, a dual Dueling DQN neural network model is adopted for training and learning. The online network is used to estimate the Q value of the current state, and is used to select offloading action decisions and calculate the loss function. The target network is used to generate the target Q value. By adding a target network to calculate the target Q value, the bias of overestimation is reduced, and the accuracy and stability of the Q value are improved.
[0062] In practical applications, the expressions of the online network and the target network for calculating the state-action value function Q(s,a) are: Q(s,a) = V(s) + (A(s,a) - ∑ a′ A(s,a′) / |A|), where Q(s,a) represents the value estimate of performing the offloading action decision a in the current environmental state s; V(s) represents the state value function, which is used to measure the inherent value of the current environmental state s and has nothing to do with specific actions; a′ represents any offloading action decision in the offloading action decision set; |A| represents the size of the action space, that is, the total number of offloading action decisions included in the offloading action decision set; A(s,a) represents the advantage function, which measures the relative advantage of selecting the offloading action decision a in the current environmental state s; ∑ a′ A(s,a′) is the average value of the advantage functions of all offloading action decisions, which is used to centralize the advantage function. During model training, the online network iteratively updates its parameters according to the target network, while the target network is used to calculate the Q values of all offloading action decisions in the next environmental state and further determine the target Q value, thereby improving the stability and convergence speed in the model training process and effectively solving the problem of overestimation of actions caused by the online network.
[0063] To facilitate a deeper understanding of the solutions of this disclosure, next, in combination with Figure 5 the process of obtaining offloading experience in this disclosure and the process of training the deep reinforcement learning model, the interaction process between the intelligent agent node and the cloud server will be described.
[0064] As Figure 5 shown, the online network in the dual Dueling DQN neural network model is deployed on each intelligent agent node, and an experience pool, as well as the online network and the target network in the dual Dueling DQN neural network model, are deployed on the cloud server. In actual operation, each intelligent agent node calculates offloading action decisions by using the online network deployed on it, and the cloud server uses the offloading experience in the experience pool to train the online network and the target network in the dual Dueling DQN neural network model.
[0065] Specifically, after the execution of the computing task is completed, the intelligent agent node can upload the collected offloading experience and priority to the experience pool in the cloud server. The cloud server then uses a preset sampling probability to select target offloading experience from the experience pool to train the online network and the target network. During the training process, the parameters of the online network can be updated based on the loss value in each iteration, and the optimized online network parameters can be synchronously updated to the target network regularly. When the cloud server completes a round of training (such as reaching the preset number of iterations or the loss function converges), it can push the optimized online network parameters to each intelligent agent node to ensure that each intelligent agent node uses the latest online network for decision-making calculations.
[0066] The offloading experience includes diverse information such as the current environmental state, offloading action decision, reward value, and the next environmental state, providing diverse and structured training data for the double Dueling DQN neural network model. Through learning, the Dueling DQN neural network model has better environmental adaptability, significantly improving the decision-making ability and decision-making quality in the double dynamic environment.
[0067] According to the above description with reference to the drawings, those skilled in the art can also understand that the embodiments of the present application can also be implemented through software programs. Thus, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores program instructions for training the stroke recurrence risk prediction model and / or for predicting the stroke recurrence risk, and the program instructions can be used to implement the method for offloading computing tasks described in the present application in combination with Figures 2 to 4 what is described.
[0068] It should be noted that although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can be changed in the order of execution. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0069] Although multiple embodiments of the present application have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Many changes, variations, and alternative methods can be thought of by those skilled in the art without departing from the spirit and scope of the present application. It should be understood that various alternative solutions to the embodiments of the present application described herein can be adopted in the practice of the present application. The appended claims are intended to define the scope of protection of the present application and thus cover equivalents or alternative solutions within the scope of these claims.
Claims
1. A method for offloading computing tasks, which is applied to an intelligent agent node deployed on an Internet of Things device. The Internet of Things device is a node in a serverless edge computing environment, and the nodes in the serverless edge computing environment further include an edge server and a cloud server. The method includes: When receiving a computing task request generated by the Internet of Things device, obtaining the current environment state of the serverless edge computing environment; The computing task request includes task attributes of the computing task, and the task attributes include the amount of task data, the number of required CPU cycles, and the maximum tolerable delay. The current environment state includes the computing resource states and functional function code pool states of the nodes in the serverless edge computing environment, and the network link state; Inputting the current environment state into a deep reinforcement learning model for decision-making calculation to output an offloading action decision, where the offloading action decision indicates a target node for executing the computing task; Allocating network resources and computing resources for the computing task according to the task attributes and the offloading action decision; Based on the network resources and computing resources, sending the computing task to the target node so that the target node executes the computing task.
2. According to the method described in claim 1, after the target node completes the execution of the computing task, the method further includes: Receiving the execution result feedback by the target node and counting the system cost for the target node to complete the execution of the computing task; Based on the execution result and the system cost, determining a reward value using a preset reward function; Obtaining the next environment state of the serverless edge computing environment, and taking the current environment state, the offloading action decision, the reward value, and the next environment state as offloading experience; Assigning a priority to the offloading experience and uploading the offloading experience and the priority to the experience pool in the cloud server.
3. According to the method described in claim 2, where the deep reinforcement learning model is a Double Dueling DQN neural network model; after uploading the offloading experience and the priority to the experience pool in the cloud server, the cloud server performs the following operations to implement the training of the Double Dueling DQN neural network model: Selecting target offloading experience from the experience pool using a preset sampling probability and calculating the importance weight of the target offloading experience; Taking the target offloading experience and the importance weight as training data and inputting them into the Double Dueling DQN neural network model to train it.
4. The method according to claim 3, wherein, The Double Dueling DQN neural network model includes an online network and a target network. Taking the target offloading experience and the importance weight as training data and inputting them into the Double Dueling DQN neural network model to train it includes: Inputting the current environment state and the offloading action decision in the target offloading experience into the online network for value calculation to output the first value of the offloading action decision; Based on the next environmental state in the target offloading experience, use the target network to calculate the second value of each offloading action decision in the offloading action decision set, so as to output a plurality of second values; Select the largest second value from the plurality of second values, and use the reward value in the target offloading experience and the largest second value to calculate the target value; Determine the loss value based on the first value, the target value, and the importance weight; Update the parameters of the online network based on the loss value, and regularly synchronize the updated parameters of the online network to the target network.
5. The method according to claim 4 further comprises: Distribute the updated parameters of the online network to the agent proxy nodes.
6. The method according to claim 2, wherein the preset reward function can be represented in the following form: Among them, n represents the nth Internet of Things device in the service - less edge computing environment, i represents the ith computing task, r n_i (t) represents the reward value determined based on the execution result and the system cost; represents the positive reward value when the execution result is successful execution, represents the negative reward value when the execution result is failed execution; Sys_cost n_i represents the negative reward value brought by the system cost of executing the ith computing task of the nth Internet of Things device; e and f are balance parameters.
7. The method according to claim 6, wherein The system cost includes latency, energy consumption, and monetary cost; the latency includes one or more of waiting transmission latency, transmission latency, waiting resource allocation latency, function instance restart latency, and computing latency; the energy consumption includes at least one of computing task transmission energy consumption and computing task execution energy consumption; the monetary cost is the cost of offloading to the edge server or the cost of offloading to the cloud server.
8. The method according to claim 7, wherein, The cost of offloading to the edge server includes the cost of requesting edge computing and the cost generated by the service duration of the edge server; the cost of offloading to the cloud server includes the cost of requesting cloud computing and the cost generated by the service duration of the cloud server.
9. A serverless edge computing system, comprising an Internet of Things device, an edge server, a cloud server, and a communication module, wherein an agent proxy node is deployed on the Internet of Things device; The Internet of Things device is used to generate a computing task request; The agent proxy node is used to execute the method according to any one of claims 1-8, so as to offload the computing task to a target node; The edge server maintains a function code pool and receives and executes the computing tasks sent by the Internet of Things device; The cloud server is used to train a deep reinforcement learning model, and a global code repository and an experience pool are stored thereon; The communication module is used to realize data transmission between the Internet of Things device, the edge server, and the cloud server.
10. A computer-readable storage medium, on which program instructions for offloading computing tasks are stored, and when the program instructions are executed by a processor, the method according to any one of claims 1-8 is implemented.