Task scheduling method and device for computing power network and electronic equipment

By updating the forget gate and input gate of the computing power network, combining historical information and loss values, the meta-strategy parameters of the task are determined, and the problem of low flexibility in task scheduling of computing power network is solved, achieving more efficient resource utilization and delay optimization.

CN120335948APending Publication Date: 2025-07-18CHINA MOBILE GROUP SHANDONG +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510343114.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The task scheduling flexibility of computing power networks is low, making it difficult to quickly adapt to task characteristics with significant differences, resulting in low resource utilization and high latency.

Method used

The forgotten gate and input gate are updated through the trained task allocation model, combined with historical cell state and loss values, the meta-strategy parameters of each task are determined, and task scheduling is performed based on the new policy parameters.

Benefits of technology

It improves the flexibility of computing power network task scheduling, can adapt to tasks in a variety of environments, dynamically adjust policy parameters, improve resource utilization and reduce delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335948A_ABST
    Figure CN120335948A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a task scheduling method and device for a computing power network and electronic equipment, belongs to the technical field of computer networks, and can improve the flexibility of task scheduling in the computing power network. Comprising the steps of obtaining task data of a plurality of tasks related to computing power network scheduling, inputting a trained task allocation model, outputting loss values of the tasks, and matching a forgetting gate, an input gate and an output hidden layer feature when target equipment is matched; updating the forgetting door through a parameter updating module of the task allocation model to obtain an updated forgetting door; updating the input gate to obtain an updated input gate; according to the historical cell state, updating a forgetting gate, updating an input gate and a loss value, outputting a new cell state of the task, and determining a meta-strategy parameter of each task based on the new cell state; and inputting the meta-strategy parameter of each task into the task allocation model, outputting a new strategy parameter shared by the task, and performing task scheduling on all the tasks based on the new strategy parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer network technologies, and particularly to a task scheduling method, apparatus, and electronic device for a computing power network. Background Art

[0002] A computing power network is a new type of information infrastructure. As a key link in the computing power network, task scheduling is crucial for resource utilization. By allocating tasks to the most suitable computing nodes, latency can be effectively reduced, resource utilization can be improved, and the quality of service of various services can be ensured.

[0003] In related technologies, in the field of task scheduling for computing power networks, a meta-reinforcement learning algorithm is used to endow the model with the ability to quickly adapt to dynamic environments. During its training phase, it mainly learns a general initial policy, but the method for updating the policy parameters of each task lacks adaptability. As a result, when facing tasks with significantly different characteristics, the model's adjustment ability is limited, making it difficult to achieve fast and effective policy fine-tuning, leading to low flexibility in task scheduling in the computing power network. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a task scheduling method, apparatus, electronic device, and storage medium for a computing power network, so as to solve the problem of low flexibility in task scheduling for the computing power network.

[0005] To solve the above technical problems, the embodiments of this application are implemented as follows: In a first aspect, the embodiments of this application provide a task scheduling method for a computing power network, including: obtaining task data of multiple tasks related to computing power network scheduling, inputting the task data into a trained task allocation model, and outputting the loss value of the initialization policy parameters of the task on the task, as well as the forget gate, input gate, and output hidden layer features when the task matches the target device; through the parameter update module of the task allocation model, performing an update process on the forget gate to obtain an updated forget gate, where the updated forget gate includes: the output hidden layer features, the historical cell state updated when the task matches the previous device, and the historical forget gate; and performing an update process on the input gate to obtain an updated input gate, where the updated input gate includes: the output hidden layer features, the historical cell state updated when the task matches the previous device, and the historical input gate; according to the historical cell state, the updated forget gate, the updated input gate, and the loss value, outputting a new cell state of the task through the parameter update module, so as to determine the meta-policy parameters of each task based on the new cell state of each task; inputting the meta-policy parameters of each task into the task allocation model, and outputting new shared policy parameters of the task, so as to perform task scheduling for all tasks based on the new policy parameters.

[0006] In a second aspect, an embodiment of the present application provides a task scheduling device for a computing power network, including: an acquisition module, configured to acquire task data of a plurality of tasks related to computing power network scheduling, input the task data into a trained task allocation model, and output a loss value of the initialization policy parameters of the task on the task, as well as a forgetting gate, an input gate, and output hidden layer features when the task matches the target device; a parameter update module, configured to perform an update process on the forgetting gate through the parameter update module of the task allocation model to obtain an updated forgetting gate, where the updated forgetting gate includes: the output hidden layer features, the historical cell state and the historical forgetting gate updated when the task matches the previous device; and perform an update process on the input gate to obtain an updated input gate, where the updated input gate includes: the output hidden layer features, the historical cell state and the historical input gate updated when the task matches the previous device; a determination module, configured to determine a new cell state of the task through the parameter update module according to the historical cell state, the updated forgetting gate, the updated input gate, and the loss value, and determine meta-policy parameters of each task based on the new cell state of each task; an output module, configured to input the meta-policy parameters of each task into the task allocation model and output new policy parameters shared by the tasks, so as to perform task scheduling on all the tasks based on the new policy parameters.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory electrically connected to the processor, where the memory stores a computer program, and the processor is configured to call and execute the computer program from the memory to implement the above-mentioned task scheduling method for a computing power network.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium for storing a computer program, where the computer program can be executed by a processor to implement the above-mentioned task scheduling method for a computing power network.

[0009] In a fifth aspect, an embodiment of the present application provides a chip, where the chip includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is configured to run a program or an instruction to implement the above-mentioned task scheduling method for a computing power network.

[0010] In a sixth aspect, an embodiment of the present application provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned task scheduling method for a computing power network.

[0011] Adopt the technical solution of the embodiment of the present application, obtain the task data of multiple tasks related to computing power network scheduling, input the task data of each task into the trained task allocation model, and output the loss value of the initialization policy parameters of the task on the task, as well as the forget gate, input gate and output hidden layer features when the task matches the target device; through the parameter update module of the task allocation model, update the forget gate to obtain the updated forget gate, and the updated forget gate includes: the output hidden layer features, the historical cell state updated when the task matches the previous device, and the historical forget gate; and update the input gate to obtain the updated input gate, and the updated input gate includes: the output hidden layer features, the historical cell state updated when the task matches the previous device, and the historical input gate; according to the historical cell state, the updated forget gate, the updated input gate and the loss value, output the new cell state of the task through the parameter update module, so as to determine the meta-policy parameters of each task based on the new cell state of each task. Among them, by updating the forget gate and input gate of each task through the trained task allocation model, it is possible to combine the historical cell state, historical forget gate and historical input gate obtained in the step of each task matching the previous device to determine the updated forget gate and updated input gate, output the new cell state of each task, and then determine the meta-policy parameters of each task. Each task can update the internal parameters of the task by combining the context information data of the task, making the finally determined meta-policy parameters more accurate.

[0012] Input the meta-policy parameters of each task into the task allocation model, and output the new policy parameters shared by the tasks, so as to perform task scheduling on all tasks based on the new policy parameters. It can be seen that by inputting the meta-policy parameters of each task into the task allocation model, new policy parameters shared by all tasks can be obtained. The new policy parameters synthesize the information of the meta-policy parameters of each task, and each meta-policy parameter also synthesizes the change data in the step of each task matching the target device. Therefore, each meta-policy parameter can be dynamically adjusted, and the new policy parameters are adjusted according to the changes of the meta-policy parameters, which can adapt to tasks in a variety of different environments and solve the problem of low flexibility in task scheduling of the computing power network. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in one or more embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in one or more embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0014] Figure 1 It is a schematic flowchart of a task scheduling method for a computing power network according to an embodiment of the present application; Figure 2 is a flowchart of computing power network task scheduling according to an embodiment of the present application; Figure 3 is an internal structure diagram of a feature extraction module in a task allocation model according to an embodiment of the present application; Figure 4 is a training flowchart of meta-reinforcement learning according to an embodiment of the present application; Figure 5 is a schematic structural diagram of a task allocation model according to an embodiment of the present application; Figure 6 is a schematic flowchart of a task scheduling method for a computing power network according to another embodiment of the present application; Figure 7 is a schematic block diagram of a task scheduling device for a computing power network according to an embodiment of the present application; Figure 8 is a schematic hardware structure diagram of a task scheduling device for a computing power network according to an embodiment of the present application. Detailed implementation manners

[0015] Embodiments of the present application provide a task scheduling method, device, and electronic device for a computing power network to solve the problem of low flexibility in task scheduling of the computing power network.

[0016] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0017] The task scheduling method for a computing power network provided by the embodiments of the present application can be executed by an electronic device or by software installed in the electronic device. Specifically, the electronic device can be a terminal device or a server device. Among them, the terminal device can include a smart phone, a notebook computer, a smart wearable device, a vehicle-mounted terminal, etc., and the server device can include an independent physical server, a server cluster composed of multiple servers, or a cloud server capable of performing cloud computing.

[0018] The following will, with reference to the accompanying drawings, describe in detail a task scheduling method for a computing power network provided by the embodiments of the present application through specific embodiments and their application scenarios.

[0019] Figure 1 A schematic flowchart of a task scheduling method for a computing power network provided by an embodiment of the present invention is shown. The method includes the following steps: S102. Obtain the task data of multiple tasks related to computing power network scheduling, input the task data into the trained task allocation model, and output the loss value of the initialized policy parameters of the task on the task, as well as the forget gate, input gate, and output hidden layer features when the task matches the target device.

[0020] The computing power network is a new type of information infrastructure.

[0021] Task data refers to the data generated during task scheduling related to computing power network scheduling, such as data on delay and energy consumption generated when different environmental samples and task samples are scheduled.

[0022] The task allocation model includes: a preprocessing module, a feature extraction module, a parameter update module, etc. It calculates the meta-policy parameters corresponding to the task through the task data, and is used to determine the optimal solution for allocating all tasks to the target device. The task allocation model generally adopts a Long Short Term Memory Network (LSTM) model improved based on meta-reinforcement learning. Among them, meta-reinforcement learning refers to improving the learning efficiency of the agent on new tasks through the trained task allocation model and quickly adapting to new tasks. The LSTM model can effectively save and update task information, and the agent can be a scheduling system or an electronic device, etc.

[0023] The target device includes: the device matched by the task at the current moment or current step.

[0024] Policy parameters refer to the parameters used to guide the parameter update of the task allocation model during the meta-reinforcement learning process. The initialized policy parameters refer to initializing the policy parameters of the task and optimizing the initialized policy parameters on this basis to obtain the meta-policy parameters of the task.

[0025] Since the task allocation model includes the cell state and gating mechanism in the LSTM model, the cell state is used to store the long-term information data of the task, and the gating mechanism includes: a forget gate, an input gate, and an output gate. Among them, the forget gate: determines the data to be discarded from the cell state. The input gate: determines the data to be added to the cell state. The cell state: is a memory state, which can discard useless old data and add new useful data by combining the output data of the forget gate and the input gate. The output gate: based on the current cell state and the data of the input gate, determines the output hidden features of the hidden layer. The output hidden features include important data input previously and are used for prediction or performing other tasks.

[0026] As an example, the forget gate can be expressed as formula 1: Formula 1 Among them, Represents the value of the forget gate of the task extracted by the feature extraction module of the task assignment model. Represents the sigmoid activation function. Represents the weight matrix of the forget gate. Represents, in the feature extraction module, the output hidden layer features of the previous step when the task matches the previous device. Represents the loss value corresponding to the task after normalization processing. Represents the bias term of the forget gate. It can be seen that t represents the update step of the parameters extracted by the task in the feature extraction module. The value of the forget gate is used to determine whether to discard this data in the cell state. The feature extraction module can be used as the first layer for the task assignment model to process data.

[0027] The input gate can be expressed as Formula 2: Formula 2 Where Represents the value of the input gate of the task extracted by the feature extraction module of the task assignment model. Represents the sigmoid activation function. Represents the weight matrix of the input gate. Represents, in the feature extraction module, the output hidden layer features of the previous step when the task matches the previous device. Represents the loss value. Represents the bias term of the input gate. It can be seen that t also represents the update step of the parameters extracted by the task in the feature extraction module. The value of the input gate is used to determine the information data added to the cell state. The feature extraction module is also used as the first layer for the task assignment model to process data.

[0028] The candidate cell state can be expressed as Formula 3: Formula 3 Where Represents the candidate cell state, which is used to store the candidate information of the current update step. Represents the activation function. Represents the weight matrix of the candidate cell state. Represents, in the feature extraction module, the output hidden layer features of the previous step when the task matches the previous device. Represents the loss value. Represents the bias term of the input gate. The candidate cell state is also used to determine the information data added to the cell state.

[0029] The cell state can be expressed as Formula 4: Formula 4 Where Represents the cell state of the task calculated by the feature extraction module in the first layer of the task assignment model. Represents the value of the forget gate. Represents the cell state of the task when the task matches the previous device in the feature extraction module, i.e., the cell state of the previous step. Represents the value of the input gate. Represents the candidate cell state. Represents element-wise multiplication. The cell state is also used to determine the finally stored information data.

[0030] The output gate can be expressed as Equation 5: Equation 5 Where, Represents the value of the output gate of the task extracted by the feature extraction module of the task assignment model. Represents the sigmoid activation function. Represents the weight matrix of the input gate. Represents the output hidden layer features when the task matches the previous device in the feature extraction module, i.e., the output hidden layer features of the previous step. Represents the loss value. Represents the bias term of the input gate. The output gate is used to determine the finally output information data.

[0031] The output hidden layer features can be expressed as Equation 6: Equation 6 Where, Represents the output hidden layer features. Represents the value of the output gate. Represents the activation function. Represents the cell state corresponding to the task in the feature extraction module. The output hidden layer features are also used to determine the finally output information data. Taking the data determined by the feature extraction module in the task assignment model above as the first layer data, Also represents the operation of element-wise multiplication.

[0032] By obtaining the task data of multiple tasks related to the computing power network scheduling, after inputting the task data into the trained task assignment model, the policy parameters of the tasks can be initialized to obtain the initial policy parameters, and the loss value of the initial policy parameters on the tasks can be calculated. Calculate and output the forget gate, input gate, and output hidden layer features when the task matches the target device.

[0033] S104. Update the forget gate through the parameter update module of the task allocation model to obtain an updated forget gate, where the updated forget gate includes: the output hidden layer features, the historical cell state updated when the task matches the previous device, and the historical forget gate; and update the input gate to obtain an updated input gate, where the updated input gate includes: the output hidden layer features, the historical cell state updated when the task matches the previous device, and the historical input gate.

[0034] Specifically, update the forget gate and the input gate of the feature extraction module of the task allocation model through the parameter update module of the task allocation model to obtain the updated forget gate and the updated input gate. Among them, the process of the feature extraction module extracting the forget gate, the input gate, and the output hidden layer features of the task, etc., can be regarded as the first layer of the task allocation model, and the process of the parameter update module corresponding to updating the parameters can be regarded as the second layer of the task allocation model.

[0035] The historical cell state is the cell state obtained by the parameter update module of the task allocation model when the task matches the previous device.

[0036] The historical forget gate is the forget gate obtained by the parameter update module of the task allocation model when the task matches the previous device.

[0037] As an example, the updated forget gate can be expressed as Equation 7: Equation 7 Where, represents the value of the updated forget gate in the trained task allocation model, represents the activation function, represents the weight matrix of the updated forget gate, represents the output hidden layer features of the task in the feature extraction module of the first layer, represents the historical cell state of the parameter update module of the second layer in the step of matching the previous device, represents the value of the historical forget gate of the parameter update module of the second layer in the step of matching the previous device, represents the bias term of the updated forget gate. It can be seen that t represents the update step of the parameters corresponding to the task, 1 in corresponds to the data in the feature extraction module of the first layer, 2 in corresponds to the data in the parameter update module of the second layer.

[0038] The updated input gate can be expressed as Equation 8: Equation 8 Where, represents the value of the updated input gate in the trained task allocation model, represents the activation function, Represents the weight matrix for updating the forget gate, Represents the output hidden layer features of the task in the feature extraction module of the first layer, Represents the historical cell state when the parameter update module of the second layer matches the previous device step, Represents the value of the historical input gate when the parameter update module of the second layer matches the previous device step, Represents the bias term for updating the input gate. It can be seen that t also represents the update step of the parameters corresponding to the task, 1 also corresponds to the first layer feature extraction module, and 2 also corresponds to the second layer parameter update module.

[0039] S106. According to the historical cell state, the updated forget gate, the updated input gate, and the loss value, the parameter update module outputs the new cell state of the task, and based on the new cell state of each task, determines the meta-policy parameters of each task.

[0040] The meta-policy parameters include: the new cell state when the task matches the last device.

[0041] Continuing with the example of S104 above, the historical cell state is represented as The updated forget gate is represented as The updated input gate is represented as The loss value of the initial policy parameters of the task on the task is represented as , It can also be represented as the loss data after normalization processing. The new cell state of the task is represented as Then, the new cell state can be expressed by Equation 9 as: Equation 9 Represents element-wise multiplication. That is, the values of the updated forget gate and the historical cell state are subjected to element-wise multiplication operation to obtain the remaining data filtered by the updated forget gate, which is used as the first cell state. The values of the updated input gate and the loss value are subjected to element-wise multiplication operation to obtain the input data, which is used as the second cell state. The first cell state and the second cell state are added together to obtain the new cell state.

[0042] It can be seen that the new cell state of each task changes with the change of the update step t. Each update step corresponds to a new cell state. When it is the last update step, the new cell state corresponding to the last update step is determined as the meta-policy parameter of the task. For example, the update step t includes 1, 2......N, that is, N steps of update are performed, and the new cell state The parameter is used as the meta-policy parameter corresponding to the task.

[0043] It should be noted that when the value of the updated forget gate and when the value of the input gate is, Formula 9 becomes the standard gradient descent form, where the input gate is equivalent to the learning rate of the gradient descent algorithm, and the forget gate is equivalent to the product term of the old policy parameters. Compared with the method of fixing the gradient descent within the task, at each task update step, it is possible to comprehensively consider the latest input data and historical information data to calculate the update step of the current parameters and the decay factor for the parameters after the update in the previous step to dynamically determine the update method of each parameter within the task during multi-step updates. At the same time, compared with the meta-learning-based trading method, the task allocation model not only gives the initial policy corresponding to the meta-initial policy parameters of each task, but also dynamically determines the update method of each task's parameters.

[0044] S108: Input the meta-policy parameters of each task into the task allocation model, and output the new shared policy parameters of the tasks, so as to perform task scheduling for all tasks based on the new policy parameters.

[0045] Determine the meta-policy parameters of each task according to S106, input the meta-policy parameters of each task into the task allocation model, determine the new shared policy parameters of the tasks through the task allocation model, and perform personalized scheduling policies for all tasks through the new policy parameters.

[0046] Adopting the technical solution of the embodiment of the present application, obtaining the task data of multiple tasks related to computing power network scheduling, inputting the task data of each task into the trained task allocation model, outputting the loss value of the initialized policy parameters of the task on the task, as well as the forget gate, input gate, and output hidden layer features when the task matches the target device; through the parameter update module of the task allocation model, performing update processing on the forget gate to obtain the updated forget gate, where the updated forget gate includes: output hidden layer features, the updated historical cell state and historical forget gate when the task matches the previous device; and performing update processing on the input gate to obtain the updated input gate, where the updated input gate includes: output hidden layer features, the updated historical cell state and historical input gate when the task matches the previous device; according to the historical cell state, updated forget gate, updated input gate, and loss value, outputting the new cell state of the task through the parameter update module, so as to determine the meta-policy parameters of each task based on the new cell state of each task. Among them, by updating the forget gate and input gate of each task through the trained task allocation model, it is possible to combine the historical cell state, historical forget gate, and historical input gate obtained in the step of each task matching the previous device to determine the updated forget gate and updated input gate, output the new cell state of each task, and further determine the meta-policy parameters of each task. Each task can update the internal parameters of the task by combining the context information data of the task, making the finally determined meta-policy parameters more accurate.

[0047] Input the meta - policy parameters of each task into the task allocation model, and output the new policy parameters shared by tasks, so as to perform task scheduling for all tasks based on the new policy parameters. It can be seen that by inputting the meta - policy parameters of each task into the task allocation model, new policy parameters shared by all tasks can be obtained. The new policy parameters integrate the information of the meta - policy parameters of each task, and each meta - policy parameter also integrates the change data of each step when each task matches the target device. Therefore, each meta - policy parameter can be dynamically regulated, and the new policy parameters are adjusted according to the changes of the meta - policy parameters, which can adapt to tasks in a variety of different environments and solve the problem of low flexibility in task scheduling of the computing power network.

[0048] In one embodiment, to obtain the task data of multiple tasks related to computing power network scheduling (i.e., S102), the following steps A1 - A7 can be executed: Step A1, in response to multiple task requests, add multiple tasks to the scheduling queue, obtain the queuing waiting time for each task to join the scheduling queue, and obtain the scheduling delay time for scheduling tasks through the controller.

[0049] The task request includes: a request to allocate a task to a target device.

[0050] The agent receives and responds to the task requests of multiple tasks through the controller, and adds multiple task requests to the queue to be scheduled. Since there are many tasks, during the process of tasks joining the scheduling queue, they need to be added in turn through queuing. Among them, the time from the start of queuing to joining the scheduling queue is the queuing waiting time , which depends on the length of the scheduling queue, system load, and scheduling algorithm efficiency.

[0051] The controller makes a scheduling decision based on the current system resource status, allocates tasks to appropriate devices to minimize delay and energy consumption, and the time delay existing in the process of scheduling tasks is the scheduling delay time , which depends on the number of tasks and the complexity of the system.

[0052] Among them, the computing power network consists of a regulator and computing power nodes, and the cluster composed of all computing power nodes can be expressed as: , device i is represented as a tuple , is the computing speed of the central processing unit (CPU) of device i, is the computing speed of the graphic processing unit (GPU) of device i, and both are usually measured by the number of floating - point operations per second executed, represents the network transmission speed, Represents the power consumption of the device. The obtained task j is represented as , where is the instruction length of task j, representing the complexity of the task; is the amount of data required for task execution; is the priority of the task. Tasks with greater impact from latency have higher priorities.

[0053] Step A2: Transmit the task to the edge device through the controller to obtain the transmission delay time and transmission energy consumption; by executing the task, obtain the execution time and execution energy consumption of the task.

[0054] The transmission delay time refers to the time delay generated when the controller transmits the task to the edge device, which is mainly determined by the amount of task data and the network transmission rate of the device.

[0055] The transmission energy consumption refers to the energy consumption generated when the controller transmits the task to the edge device. Mainly because data transmission consumes network bandwidth, corresponding energy consumption will also be generated during the process of transmitting the task.

[0056] Specifically, the transmission delay time is represented as , and can be expressed by Formula 10 as: Formula 10 The transmission energy consumption is represented as , and can be expressed by Formula 11 as: Formula 11 Among them, the transmission power consumption of device i is represented as , then the data transmission power consumption corresponding to the task is represented as , the transmission time is , the transmission energy consumption is , and the interpretations of the remaining letters are the same as in Step A1.

[0057] After some data of the task is transmitted to the target device, the target device executes the task. The time for executing the task is the execution time, and the energy consumption generated during the execution of the task is the execution energy consumption. The execution energy consumption is related to the operating power consumption of the target device.

[0058] Specifically, the execution time is represented as , and can be expressed by Formula 12 as: Formula 12 The execution energy consumption is represented as , and can be expressed by Formula 13 as: Formula 13 Among them, Denote the computing power of device i. The computing processing time is also the time to execute a task. Denote the operating power consumption of the device. After the task calculation is completed, the execution result is transmitted back to the controller through the network. Since the data volume of the execution result is usually small, the transmission delay can be ignored. The interpretations of the remaining letters are the same as those in step A1.

[0059] Step A3: Determine the total delay of the task based on one or more of the queuing waiting time, scheduling delay time, transmission delay time, and execution time; determine the total energy consumption of the task based on one or more of the transmission energy consumption and execution energy consumption.

[0060] Following the interpretations in the formulas of step A1 and step A2, the total delay of the task is expressed as and can be expressed by formula 14 as: Formula 14 The total energy consumption of the task is expressed as and can be expressed by formula 15 as: Formula 15 It should be noted that in the process of task scheduling in the computing power network, its goal is to minimize the weighted sum of delay and energy consumption, and it is balanced by setting weight factors. Among them, the weight factor of the total delay is and the weight factor of the total energy consumption is , reflecting the degree of emphasis of the system on delay and energy consumption. R is the objective function to be optimized as shown in formula 16: Formula 16 The scheduling system continuously adjusts the strategy in the dynamic environment to make the delay and energy consumption in each stage reach the optimal, so as to improve the overall performance. The task scheduling of the computing power network is to minimize the total delay and total energy consumption in task scheduling. Denote the situation where task j matches the target device. The optimization objective function is expressed as formula 17: Formula 17 In summary, the task scheduling process of the computing power network is as shown in Figure 2As shown, multiple task requests are received, and the tasks are requested to join the scheduling queue. First, the process of task reception and queuing is carried out, including: receiving tasks, analyzing task requirements, i.e., task data, adding tasks to the scheduling queue, and calculating the queuing waiting time during the queuing process; making scheduling decisions and resource allocations for the tasks added to the scheduling queue, including: evaluating the device status, sorting the task priorities, determining the scheduling decisions, and calculating the scheduling delay time during the process; after determining the scheduling decisions, data transmission of the tasks is carried out, including: transmitting data, calculating the transmission delay time, and calculating the transmission energy consumption when the data is transmitted to the computing node; after the data transmission ends, task execution is carried out, including: executing the tasks, returning the task execution results after the tasks are executed, and calculating the execution time and execution energy consumption; finally, the total delay and total energy consumption are calculated, and the scheduling process is completed.

[0061] Step A4, through the meta-reinforcement learning method, obtain the state information of the task, where the state information includes: the device list and status list of the task.

[0062] Through the meta-reinforcement learning method, the information of the task can also be determined, such as the state information of the task and the device list including: the CPU computing speed of device i corresponding to the task , the GPU computing speed of device i , the network transmission speed , the power consumption of the device etc., and the status list including the instruction length of task j , the amount of data required for task execution , the priority of the task etc. The state information can be expressed as , and can be represented by formula 18: Formula 18 The state information, representing the state space, comprehensively covers various types of information related to task scheduling, including the device list of the device operating conditions and the status list related to task requests. Specifically, the state space can directly read the hardware parameters and operating status of the devices through the interface provided by the cluster management platform to obtain the current environment information.

[0063] The device list includes the CPU / GPU computing speeds, network transmission rates, and power consumption of all nodes, while the status list contains the resource requirements and priorities of the tasks.

[0064] Step A5, obtain the set of various allocation situations where the task is allocated to the target device, and obtain the spatial information corresponding to the task.

[0065] The spatial information represents the action set of all possible strategies that the agent can execute. In the computing power network, the action involves allocating the task to a specific computing device, i.e., the target device.

[0066] For N tasks j to be processed and M computing devices i, Denoting the specific situation where tasks are assigned to target devices, the spatial information A can be expressed by Equation 19: Equation 19 Step A6, determine the reward function according to the negative value of the weighted sum of the total latency and the total energy consumption; the reward function is used to determine the priority of task assignment to target devices.

[0067] The reward function is guided by the optimization objective of task scheduling. The optimization objective of task scheduling is as shown in Equation 17 above, and the goal is to reduce the latency and energy consumption of the system. The reward function is defined as the negative value of the weighted sum of latency and energy consumption: Equation 20 The letter interpretations of the formulas in Step A4 and Step A5 are the same as those of the formula in Step A3 above, and will not be elaborated here. By designing the reward function, the agent will preferentially select those actions that can effectively reduce latency and energy consumption when optimizing the scheduling strategy, so as to achieve efficient task scheduling and be able to determine the priority of task assignment to target devices. Denotes collecting a specific number of location parameters.

[0068] Step A7, determine the state information, spatial information, and reward function as the task data of the task.

[0069] All the obtained data such as the state information, spatial information, and reward function corresponding to each task can be determined as the task data of the task.

[0070] It should be noted that during the task execution process, the scheduling system will continuously monitor feedback information such as the execution time of the task and changes in the device status to adjust the scheduling strategy in real time. When the task environment changes, the task assignment model can use a small number of samples for rapid fine-tuning to ensure the effectiveness and robustness of the scheduling strategy in a dynamic environment and enhance the scheduling performance of the scheduling system under different task requirements. That is, the task data can be adjusted according to the feedback information of the scheduling system, and the task data in the feedback information can be sampled to facilitate subsequent adjustment of the generated new policy parameters and deployment of the new policy.

[0071] In this embodiment, obtaining the task data through task scheduling in the computing power network and also obtaining the task data through meta-reinforcement learning, and determining all the data obtained after calculating the two as the task data of the task can facilitate subsequent task assignment.

[0072] In one embodiment, the task data is input into the trained task allocation model, and the loss value of the initialized policy parameters of the task on the task is output (i.e., S102), and the following steps B1 - B2 can also be executed: Step B1, input the obtained state information, spatial information, and reward function into the task allocation model, and calculate the loss function and the loss function gradient of each task through the preprocessing module of the task allocation model.

[0073] In this application, for the task data such as the state information, spatial information, and reward function obtained in step A7 above, the task data is input into the task allocation model. In the task allocation model, the task set is defined as , and each task represents a specific scheduling task requirement. For each task , the policy parameters in the LSTM meta - learner of meta - reinforcement learning can each extract an initialized policy parameter, and the LSTM meta - learner can be the trained task allocation model. That is, before the task allocation starts, the policy parameters are first initialized, and under the condition of the initialized policy parameters, the loss function and the loss function gradient of each task are calculated.

[0074] Specifically, the initial policy parameters can be expressed as , and the loss calculated on the task is expressed as , which is expressed as formula 21: Formula 21 where, represents the loss function of the initial policy parameters on the task, t represents the current step or the current moment, represents the generalized advantage function, represents the trajectory generated during the interaction between the policy corresponding to the initial policy parameters and the environment. represents the state information of the task at the current step; represents the spatial information of the task at the current step, that is, it represents the task when executing the initial policy according to the initial policy parameters, the generated task data. is the logarithmic function.

[0075] Generalized advantage function combines the advantages of the Monte Carlo method and the temporal difference method, and combines a hyperparameter to balance the weights of the two methods, and is defined as shown in formula 22: Formula 22 where, , represents the temporal difference error; is the reward function, is the reward discount factor, which is used to weigh the impact of future rewards on the current decision. represents the state The corresponding state value function indicates the cumulative reward that is expected to be obtained after executing the current policy in the state . is the balance parameter, which controls the weighting of the Monte Carlo method and the temporal difference method. When, the advantage function tends to the Monte Carlo estimate; , then the temporal difference estimate is completely adopted. H represents the length of the corresponding time stage.

[0076] The loss function of the task is obtained through the preprocessing module , and the corresponding gradient of the loss function is .

[0077] Step B2: The loss function and the gradient of the loss function are normalized through the preprocessing module, and the loss value of the initialized policy parameters of the processed task on the task is output.

[0078] Before normalizing the loss function, the loss function in step B1 is expressed as , and the gradient of the loss function is . The loss function and the gradient of the loss function are respectively input into the preprocessing module for normalization. The normalization process is as shown in Equation 23: Equation 23 is the sign function, and its definition is as shown in Equation 24: 24 Specifically, the normalization method is used to separate the input amplitude and sign. When the absolute value of the input value is greater than a certain range, it will first be compressed through the logarithmic function, and the sign direction is given by the sgn(x) function. For a large positive value of the amplitude, the direction sign is +1, and for a large negative value of the amplitude, the direction sign is -1. When the absolute value of the input value is less than the threshold, the value will be enlarged through the exponential function, and the sign is fixed at -1. It can adjust the scale of the loss and gradient calculated in each update step to an appropriate range. Among them, x in Equation 23 can be the input loss function and the gradient of the loss function. After normalization, the loss value is output. represents if, othereise represents otherwise, is the preset threshold, which is an exponential function.

[0079] In this embodiment, by obtaining the task data, using the preprocessing module of the task allocation model, calculating the numerical values of the loss function and the gradient of the loss function of the task, and normalizing the loss function and the gradient of the loss function, the loss value of the initialized policy parameters of the task on the task is obtained. Through the normalization process, the loss value can be within a suitable range, which facilitates subsequent calculations such as the forget gate.

[0080] In one embodiment, inputting the task data into the trained task allocation model, outputting the loss value of the initialized policy parameters of the task on the task, and the forget gate, input gate, and output hidden layer features (i.e., S102) when the task matches the target device, the following steps C1 - C3 can be executed: Step C1, input the task data into the trained task allocation model, input the loss value into the feature extraction module of the task allocation model, and according to the loss value of the task, activation function, output hidden layer features when the task matches the previous device, weight matrix of the forget gate, and bias of the forget gate, obtain the forget gate when the task matches the target device; and according to the loss value of the task, activation function, output hidden layer features when the task matches the previous device, weight matrix of the input gate, and bias of the input gate, obtain the input gate when the task matches the target device.

[0081] The feature extraction module is a basic LSTM unit, which can receive the normalized loss and loss gradient generated by the preprocessing module in the current update step, that is, the normalized loss value, and combine the hidden layer state passed in the previous step to determine the output hidden layer features when the task matches the previous device, that is, in the previous step.

[0082] Specifically, the loss value can be expressed as , represents the sigmoid activation function, represents the output hidden layer features when the task matches the previous device, that is, in the previous step, represents the weight matrix of the forget gate, represents the bias of the forget gate. According to the above formula 1, the forget gate can be determined. The value of the forget gate is limited between 0 and 1, and is used to determine which historical information the cell retains and discards. A forget gate control of 1 means completely retaining the corresponding information, and 0 means completely forgetting this information. Therefore, the forget gate when the task matches the target device can be calculated through the feature extraction module of the task allocation model ,and determine the forgotten data, that is, judge whether to discard the task matching the target device.

[0083] Similarly, for the tensor information of the loss and loss gradient in the current update step ,and the input hidden layer state feature vector Perform splicing, and then obtain the value of the input gate through linear transformation and activation function to determine the information of the current input. According to Formula 2, the input gate when the task matches the target device can be calculated. .

[0084] Step C2: Calculate the candidate cell state of the task according to the forget gate and the input gate, and determine the cell state of the task based on the candidate cell state.

[0085] Calculate the candidate cell state according to the forget gate and the input gate of the task when matching the target device calculated in Step C1. For example, the candidate cell state can be obtained according to the above Formula 3. Determine the cell state according to the candidate cell state. .

[0086] Step C3: Determine the output gate of the task based on the cell state and the activation function, and output the output hidden layer feature of the task when matching the target device according to the output gate and the cell state.

[0087] Obtain the output gate of the task by using the activation function according to the above Formula 5. ; According to the above Formula 6, it can be known that according to the output gate and the cell state, the output hidden layer feature of the task when matching the target device can be determined. The functions of the output hidden layer feature include two aspects. One is as the input of the second-layer parameter update module, and the other is as the output hidden layer feature of the previous step when the first layer is updated in the next step.

[0088] In summary, as Figure 3 shown, it is the internal structure diagram of the feature extraction module in the task allocation model. The feature extraction module is a basic LSTM unit, and the output hidden layer feature when the task matches the previous device, that is, the previous step , and the task calculation loss value After being calculated by the activation function, the forget gate and the input gate are obtained. The forget gate is then subjected to an element-wise multiplication operation with the cell state of the task when matching the previous device, that is, the previous step to obtain the first cell state. And After being calculated by the activation function, the candidate cell state is obtained. The input gate and the candidate cell state are subjected to an element-wise multiplication operation to obtain the second cell state. The first value and the second value are added to obtain the cell state . And After Calculation of the activation function to obtain the output gate, and the output gate and the cell state that have passed through the activation function After performing the element-wise multiplication operation, the output hidden layer features are obtained.

[0089] In this embodiment, according to the calculated loss value of the task, in the feature extraction module part of the task allocation model, the forgetting gate, input gate, cell state, output gate, and output hidden layer features corresponding to the task are determined. The data parameters of the task can be initially calculated, and on this basis, the parameters of the forgetting gate and input gate are changed to obtain a more accurate cell state.

[0090] In one embodiment, the meta-policy parameters of each task are input into the task allocation model, and the new policy parameters shared by the tasks are output, so as to perform task scheduling on all tasks based on the new policy parameters (i.e., S108), and the following steps D1-D2 can be executed: Step D1, input the meta-policy parameters of each task into the task allocation model, and the meta-policy parameters include: the new cell state when each task matches the last device.

[0091] According to the above formula 9, determine the new cell state of each task among all tasks, and use the new cell state of the last update step in the new cell state of each task as the meta-policy parameter. Such as the updated parameter , N is the last update step of the task, and the updated parameter is used as the query set for input into the task allocation model.

[0092] Step D2, subtract the surrogate loss and surrogate loss gradient of the meta-policy parameters of each task from the shared meta-policy parameters in the task allocation model, and output the new policy parameters shared by the tasks; among them, the surrogate loss includes a clipping function.

[0093] The shared meta-policy parameters refer to the policy parameters used in the task allocation model to allocate all tasks, and are the basic parameters of the LSTM model.

[0094] The surrogate loss refers to the loss of all tasks responding to the task request. Considering that the Proximity Policy Optimization (PPO) has high sample efficiency, PPO is generally used to update the shared meta-policy parameters.

[0095] The new policy parameters shared by the tasks can be expressed by formula 25 as: Formula 25 Among them, represents the shared meta-policy parameters, that is, the basic parameters of the LSTM model; K represents the number of all tasks; Denotes in the task The parameters obtained after N-step updates according to the above formula 9 ; Denotes the surrogate loss. The parameters of the task matching model Include the weight and bias parameters for calculating the forget gate, input gate, output gate, and candidate cell state in the basic feature extraction module, the weight and bias parameters in the parameter update module, and the new cell state for each task. Compared with the method of updating under a fixed gradient, it can determine the learning rate for each step update ( ) and the decay coefficient for historical information ( ) according to the dynamics of specific tasks, realizing dynamic parameter updates and having stronger adaptability.

[0096] The surrogate loss of PPO is defined as shown in formula 26: Formula 26 Among them, Is the importance sampling ratio, used to correct the probability distributions of the old and new policies during multi-step updates, Denotes the minimum value. The importance sampling ratio is shown in formula 27: Formula 27 Among them, Denotes the initial policy parameters extracted from the shared meta-policy parameters of the task assignment model that have not been updated since the beginning In , and the policy obtained after N-step fine-tuning by the parameter update module on the task ; while Denotes the shared meta-policy parameters of the updated task assignment model The initial policy extracted After N-step fine-tuning by the parameter update module on the task The obtained policy. The ratio of the two reflects the gap between the probability distributions of the old and new policies and is used to correct the gap between the probability distributions of the policies before and after the update.

[0097] The clip clipping function is defined as follows, used to prevent the failure of importance sampling caused by the excessive gap between the old and new policies during multi-step updates, as shown in formula 28: Formula 28 In Includes All the meta-policy parameters in. Each can obtain the result of the corresponding clipping function according to different ranges, , with values of etc. Among them, Is a specific parameter.

[0098] As shown Figure 4 in the training flowchart of meta-reinforcement learning, where PPO is used to update the shared meta-policy parameters. Specifically, the LSTM policy parameters are extracted , and the meta-policy is calculated by gradually updating each task according to the initial policy parameters. For example, the policy corresponding to task 1 (task1) is the initial policy calculated based on the initial policy parameters , and the meta-policy after updating the initial policy parameters is , and the corresponding loss generated is ; For example, for task task , the corresponding policy is the initial policy calculated based on the initial policy parameters , and the meta-policy after updating the initial policy parameters is , and the corresponding loss generated is ; The losses of all tasks are collected, and the numerical values corresponding to the losses are input into PPO for updating the LSTM meta-policy parameters

[0099] It can be seen that the trajectory generated during the interaction between the initial policy and the environment can be used as the sample data during the policy training process. K represents the total number of tasks. In each task inner loop, the LSTM meta-learner is used to guide the update of the in-task policy parameters. Combining the characteristics of each task, the learning rate for updating the adaptive policy parameters and the scaling term for the old policy parameters are output. Based on Equation 9, multi-step updates are performed. The policy corresponding to the updated meta-policy parameters is applied to this specific task again. The meta-policy parameters of each task are used as the query set, and the loss corresponding to the query set is calculated. Finally, the losses of the query sets of K tasks are accumulated, and multi-step updates are performed on the parameters of the LSTM meta-learner model based on Equation 25 to finally obtain the newly trained policy parameters

[0100] In this embodiment, the calculation formula obtained through the outer loop training during the training of the task allocation model can be assigned to the tasks. By obtaining the meta-policy parameters of each task, the task allocation model can determine the new policy parameters shared by all tasks and allocate all tasks according to the new policy parameters. The new policy parameters can be adapted to all tasks

[0101] In one embodiment, the training of the task allocation model can be specifically executed according to the following steps E1 - E6 Step E1: Obtain the sample task data and the corresponding sample policies of multiple sample tasks, and input the sample task data into the task allocation model to be trained; use the preprocessing module of the task allocation model to be trained to calculate the sample loss value corresponding to the sample task

[0102] The task allocation model to be trained mainly optimizes the meta - learning model based on the strategy parameter update method guided by the LSTM meta - learner. The task allocation model to be trained includes a pre - processing module, a feature extraction module, and a parameter update module. Among them, the pre - processing module is used to normalize the loss function and the gradient of the loss function; the feature extraction module is a basic LSTM cell, which is used for feature extraction and cell state transfer, etc.; the parameter update module is a custom LSTM cell, which is used for updating the policy parameters.

[0103] The training of the task allocation model can be carried out through two loops, including an inner loop and an outer loop. Among them, in the inner loop, the scheduling strategy can be optimized to adapt to specific tasks, and in the outer loop, a shared meta - strategy can be learned on different tasks.

[0104] A sample task refers to a task that responds to a task request.

[0105] Sample task data, the data corresponding to the sample task, includes a large number of collected task samples and environmental samples, etc. The data covers diverse scenarios that may occur in the computing power network. Among them, the large number of collected task samples include the priority, data volume, computational complexity, etc. of the sample task, and the environmental samples include the computing power, network transmission rate, and power consumption of the sample device.

[0106] A sample strategy refers to a matching method that assigns each sample task to a sample target device according to the sample task.

[0107] In each task inner loop, the task allocation model to be trained quickly adapts to each task to achieve a personalized scheduling strategy. Specifically, the sample task is input into the task allocation model to be trained. Through the pre - processing module in the task allocation model to be trained, the loss function and the gradient of the loss function calculated according to the sample task data are normalized to obtain the sample loss value corresponding to the sample task. One sample task corresponds to one sample loss value, and multiple sample tasks correspond to multiple sample loss values.

[0108] Step E2: Input the sample loss value into the feature extraction module of the task allocation model to be trained, output the sample forget gate and the sample input gate, and based on the sample forget gate and the sample input gate, output the sample output gate and the sample output hidden - layer feature of the task.

[0109] Input the loss value output by the pre - processing module of the task allocation model to be trained into the feature extraction module. Through the feature extraction module, calculate the sample forget gate and the sample input gate of the current step. According to the sample forget gate and the sample input gate, calculate the sample candidate cell state to obtain the sample cell state, and output the sample output gate and the sample output hidden - layer feature of the task according to the sample cell state.

[0110] The feature extraction module can determine the corresponding sample forgetting gate, sample input gate, cell state, sample output gate, and sample output hidden layer features according to the loss value of each sample task.

[0111] Step E3: Through the parameter update module of the task assignment model to be trained, perform parameter update processing on the sample forgetting gate and the sample input gate to obtain the sample updated forgetting gate and the sample updated input gate.

[0112] The parameter update module receives the output hidden layer features output by the feature extraction module, and changes the parameters in the sample forgetting gate and the sample input gate output by the feature extraction module to obtain the sample updated forgetting gate and the sample updated input gate.

[0113] Step E4: According to the sample historical cell state, sample updated forgetting gate, sample updated input gate, and sample loss value updated when the sample task matches the previous device, output the sample new cell state.

[0114] The parameter update module can also calculate the new cell state according to the sample historical cell state, sample updated forgetting gate, sample updated input gate, and sample loss value updated when the sample task matches the previous device.

[0115] That is to say, the parameter update module obtains the sample updated forgetting gate and the sample updated input gate by modifying the values of the sample forgetting gate and the sample input gate output by the feature extraction module, and then determines the sample new cell state. The formula corresponding to it after training is the formula used in the above embodiment.

[0116] Step E5: Based on the sample new cell state of each sample task, obtain the sample meta-policy parameters of each sample task, and input the sample meta-policy parameters of each task into the task assignment model to be trained, and output the sample new policy parameters shared by all sample tasks.

[0117] According to the sample new cell state obtained by the inner loop training in step E4, determine the sample meta-policy parameters of each sample task. Perform outer loop training on the task assignment model to be trained. Specifically, perform multi-step updates on the sample shared meta-policy parameters of the task assignment model to be trained, make full use of the experience on multiple sample tasks, determine the sample new policy parameters, and the obtained sample new policy parameters can be shared among all sample tasks. The sample new policy parameters can not only give the initial policy common to all sample tasks, but also give an adaptive parameter update method according to the input information features of each sample task.

[0118] Step E6: Train the task assignment model to be trained according to the sample new policy parameters and the sample policy to obtain the trained task assignment model.

[0119] Determine the allocation method of sample tasks to sample devices according to the sample new policy parameters. By adjusting the parameters in the inner loop and outer loop training processes of the task allocation model to be trained until the sample sharing new policy parameters of the LSTM meta-learner converge, the allocation method and the sample policy are made to meet the expectations, such as being consistent.

[0120] When the allocation method determined according to the sample new policy parameters output by the task allocation model to be trained and the sample policy meet the expectations, the trained task allocation model is obtained. Among them, the parameters and calculation formulas in the trained task allocation model are the formulas used in the above calculation of specific tasks.

[0121] Specifically, as Figure 5 shown, the schematic structural diagram of the task allocation model includes a preprocessing module, a feature extraction module, and a parameter update module. Among them, the loss function of the task is calculated through task data and the gradient of the loss function are input into the preprocessing module for preprocessing, and the loss value after normalization processing is output, that is . The output loss value is input into the feature extraction module and the parameter update module, and the feature extraction module is used as the first layer, and the cell state of the corresponding task matching the previous device in the previous step and the output hidden layer feature of the corresponding task matching the previous device in the previous step are input into the feature extraction module. The feature extraction module calculates and outputs the cell state and the output hidden layer feature corresponding to the current step and inputs them into the parameter update module; the parameter update module receives the data output by the feature extraction module and the data output by the preprocessing module, and also receives the historical update forget gate generated by the parameter update module for the task matching the previous device in the previous step, the historical update input gate and the historical cell state . The parameter update module outputs the update forget gate , the update input gate and the new cell state . The historical cell state corresponds to the policy parameter of the previous step, and the new cell state corresponds to the policy parameter of this step.

[0122] In this embodiment, by training the task allocation model to be trained, the new policy parameters of the samples learned by the task allocation model can quickly adjust to the optimal policy through a small number of samples after receiving a new task. The task allocation model to be trained learns the new policy corresponding to the new policy parameters of the samples with strong generalization ability, which has strong generalization ability and can be efficiently adapted in different task scenarios. When the scheduling system faces changes in tasks and network environments, it does not need to be retrained and can quickly generate a scheduling policy through a small number of samples, realizing the efficient scheduling of the computing power network.

[0123] Figure 6 is a schematic flowchart of a task scheduling method for a computing power network according to another embodiment of the present application. As Figure 6 shown, the method includes the following steps: S601, the system responds to multiple task requests, sequentially adds each request to the scheduling queue, obtains the queuing waiting time for each task to be added to the scheduling queue, and obtains the scheduling delay time for scheduling tasks through the controller.

[0124] S602, the controller transfers the task to the edge device, obtains the transmission delay time and transmission energy consumption, and obtains the execution time and execution energy consumption for executing the task.

[0125] S603, determine the total delay of the task according to the queuing waiting time, scheduling delay time, transmission delay time, and execution time; determine the total energy consumption of the task through the transmission energy consumption and execution energy consumption.

[0126] S604, through the meta-reinforcement learning method, obtain the state information, spatial information, and reward function of the task, and determine the state information, spatial information, and reward function as the task data of the task.

[0127] S605, input the task data into the task allocation model, obtain the loss function and loss function gradient of each task, and input the loss function and loss function gradient into the preprocessing module to obtain the loss value corresponding to the task.

[0128] S606, input the loss value into the feature extraction module of the task allocation model, and obtain the forgetting gate when the task matches the target device according to the loss value of the task, activation function, output hidden layer features when the task matches the previous device, weight matrix of the forgetting gate, and bias of the forgetting gate.

[0129] S607, according to the loss value of the task, activation function, output hidden layer features when the task matches the previous device, weight matrix of the input gate, and bias of the input gate, obtain the input gate when the task matches the target device.

[0130] S608. Calculate the candidate cell state of the task according to the forget gate and the input gate, determine the cell state of the task based on the candidate cell state; determine the output gate of the task based on the cell state and the activation function, and output the output hidden layer features when the task matches the target device according to the output gate and the cell state.

[0131] S609. Through the parameter update module of the task assignment model, update the forget gate according to the output hidden layer features, the historical cell state updated when the task matches the previous device, and the historical forget gate, and obtain the updated forget gate.

[0132] S610. Through the parameter update module of the task assignment model, update the input gate according to the output hidden layer features, the historical cell state updated when the task matches the previous device, and the historical input gate, and obtain the updated input gate.

[0133] S611. According to the historical cell state, the updated forget gate, the updated input gate, and the loss value, output the new cell state of the task through the parameter update module, so as to determine the meta-policy parameters of each task based on the new cell state of each task.

[0134] S612. Input the meta-policy parameters of each task into the task assignment model, and the meta-policy parameters include: the new cell state when each task matches the last device.

[0135] S613. Subtract the surrogate loss and the surrogate loss gradient of the meta-policy parameters of each task from the shared meta-policy parameters in the task assignment model, and output the new policy parameters shared by the tasks. So as to perform task scheduling for all tasks based on the new policy parameters.

[0136] The specific processes of the above S601 to S613 have been described in detail in the above embodiments, and will not be repeated here.

[0137] Adopt the technical solution of the embodiment of the present application, obtain the task data of multiple tasks related to computing power network scheduling, input the task data of each task into the trained task allocation model, and output the loss value of the initialization policy parameters of the task on the task, as well as the forget gate, input gate and output hidden layer features when the task matches the target device; through the parameter update module of the task allocation model, perform update processing on the forget gate to obtain the updated forget gate, and the updated forget gate includes: output hidden layer features, the updated historical cell state and historical forget gate when the task matches the previous device; and perform update processing on the input gate to obtain the updated input gate, and the updated input gate includes: output hidden layer features, the updated historical cell state and historical input gate when the task matches the previous device; according to the historical cell state, updated forget gate, updated input gate and loss value, output the new cell state of the task through the parameter update module, so as to determine the meta-policy parameters of each task based on the new cell state of each task. Among them, by updating the forget gate and input gate of each task through the trained task allocation model, it is possible to combine the historical cell state, historical forget gate and historical input gate obtained in the step of each task matching the previous device to determine the updated forget gate and updated input gate, output the new cell state of each task, and then determine the meta-policy parameters of each task. Each task can update the internal parameters of the task in combination with the context information data of the task, making the finally determined meta-policy parameters more accurate.

[0138] Input the meta-policy parameters of each task into the task allocation model, and output the new policy parameters shared by the tasks, so as to perform task scheduling on all tasks based on the new policy parameters. It can be seen that by inputting the meta-policy parameters of each task into the task allocation model, it is possible to obtain the new policy parameters shared by all tasks. The new policy parameters synthesize the information of the meta-policy parameters of each task, and each meta-policy parameter also synthesizes the change data in the step of each task matching the target device. Therefore, each meta-policy parameter can be dynamically adjusted, and the new policy parameters are adjusted according to the changes of the meta-policy parameters, which can adapt to tasks in a variety of different environments and solve the problem of low flexibility in task scheduling of the computing power network.

[0139] In summary, specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing may be advantageous.

[0140] The above is a task scheduling method for a computing power network provided by an embodiment of the present application. Based on the same idea, an embodiment of the present application also provides a task scheduling device for a computing power network.

[0141] Figure 7 It is a schematic structural diagram of a task scheduling device for a computing power network according to an embodiment of the present invention. As shown in FIG. 7, the task scheduling device for a computing power network includes: an acquisition module 71, a parameter update module 72, a determination module 73, and an output module 74: The acquisition module 71 is configured to acquire task data of multiple tasks related to computing power network scheduling, input the task data into a trained task allocation model, and output the loss value of the initial policy parameters of the task on the task, as well as the forget gate, input gate, and output hidden layer features when the task matches the target device; The parameter update module 72 is configured to perform an update process on the forget gate through the parameter update module of the task allocation model to obtain an updated forget gate, where the updated forget gate includes: output hidden layer features, the updated historical cell state and historical forget gate when the task matches the previous device; and perform an update process on the input gate to obtain an updated input gate, where the updated input gate includes: output hidden layer features, the updated historical cell state and historical input gate when the task matches the previous device; The determination module 73 is configured to output a new cell state of the task through the parameter update module according to the historical cell state, updated forget gate, updated input gate, and loss value, and determine the meta-policy parameters of each task based on the new cell state of each task; The output module 74 is configured to input the meta-policy parameters of each task into the task allocation model and output new policy parameters shared by the tasks, so as to perform task scheduling on all tasks based on the new policy parameters.

[0142] In one embodiment, the acquisition module 71 includes: An extraction unit, configured to input the task data into a trained task allocation model, input the loss value into the feature extraction module of the task allocation model, and obtain the forget gate when the task matches the target device according to the loss value of the task, activation function, output hidden layer features when the task matches the previous device, weight matrix of the forget gate, and bias of the forget gate; and obtain the input gate when the task matches the target device according to the loss value of the task, activation function, output hidden layer features when the task matches the previous device, weight matrix of the input gate, and bias of the input gate; A calculation unit, configured to calculate the candidate cell state of the task according to the forget gate and input gate, and determine the cell state of the task based on the candidate cell state; A first determination unit, configured to determine the output gate of the task based on the cell state and activation function, and output the output hidden layer features when the task matches the target device according to the output gate and cell state.

[0143] In one embodiment, the output module 74 is specifically configured to input the meta-policy parameters of each task into the task assignment model. The meta-policy parameters include: the new cell state when each task matches the last device; subtract the proxy loss and proxy loss gradient of the meta-policy parameters of each task from the shared meta-policy parameters in the task assignment model, and output the new policy parameters shared by the tasks; wherein, the proxy loss includes a clipping function.

[0144] In one embodiment, the acquisition module 71 further includes: A scheduling unit, configured to respond to multiple task requests, add multiple tasks to a scheduling queue, obtain the queuing waiting time for each task to join the scheduling queue, and obtain the scheduling delay time for scheduling tasks through a controller; A transmission unit, configured to transmit tasks to an edge device through a controller, obtain the transmission delay time and transmission energy consumption; obtain the execution time and execution energy consumption for executing the tasks; An execution unit, configured to determine the total delay of a task according to one or more of the queuing waiting time, scheduling delay time, transmission delay time, and execution time; determine the total energy consumption of a task according to one or more of the transmission energy consumption and execution energy consumption; A reinforcement unit, configured to obtain the state information of a task through a meta-reinforcement learning method. The state information includes: the device list and state list of the task; An allocation unit, configured to obtain a set of multiple allocation situations for a task to be allocated to a target device, and obtain the spatial information corresponding to the task; A reward unit, configured to determine a reward function according to the negative value of the weighted sum of the total delay and total energy consumption; the reward function is used to determine the priority for a task to be allocated to a target device; A second determination unit, configured to determine the state information, spatial information, and reward function as the task data of the task.

[0145] In one embodiment, the acquisition module 71 is specifically further configured to input the obtained state information, spatial information, and reward function into the task assignment model, calculate the loss function and loss function gradient of each task through the preprocessing module of the task assignment model; perform normalization processing on the loss function and loss function gradient through the preprocessing module, and output the loss value of the initialized policy parameter of the processed task on the task.

[0146] In one embodiment, the apparatus further includes a training module, which specifically includes: An acquisition unit, configured to acquire the sample task data and corresponding sample policies of multiple sample tasks, and input the sample task data into the task assignment model to be trained; use the preprocessing module of the task assignment model to be trained to calculate the sample loss value corresponding to the sample task; A feature extraction unit, configured to input a sample loss value into a feature extraction module of a task allocation model to be trained, output a sample forgetting gate and a sample input gate, and based on the sample forgetting gate and the sample input gate, output a sample output gate and a sample output hidden layer feature of a task; A parameter update unit, configured to perform parameter update processing on the sample forgetting gate and the sample input gate through a parameter update module of the task allocation model to be trained, and obtain a sample updated forgetting gate and a sample updated input gate; An update unit, configured to output a sample new cell state according to a sample historical cell state, a sample updated forgetting gate, a sample updated input gate, and a sample loss value updated when a sample task matches the previous device; An input unit, configured to obtain a sample meta-policy parameter of each sample task based on the sample new cell state of each sample task, input the sample meta-policy parameter of each task into the task allocation model to be trained, and output a sample new policy parameter shared by all sample tasks; An output unit, configured to train the task allocation model to be trained according to the sample new policy parameter and the sample policy, and obtain a trained task allocation model.

[0147] By adopting the technical solution of the embodiment of the present application, task data of multiple tasks related to computing power network scheduling is obtained, the task data of each task is input into the trained task allocation model, and the loss value of the initialization policy parameter of the task on the task, as well as the forgetting gate, the input gate, and the output hidden layer feature when the task matches the target device are output; through the parameter update module of the task allocation model, the forgetting gate is updated to obtain an updated forgetting gate, and the updated forgetting gate includes: an output hidden layer feature, a historical cell state updated when the task matches the previous device, and a historical forgetting gate; and the input gate is updated to obtain an updated input gate, and the updated input gate includes: an output hidden layer feature, a historical cell state updated when the task matches the previous device, and a historical input gate; according to the historical cell state, the updated forgetting gate, the updated input gate, and the loss value, the new cell state of the task is output through the parameter update module, so as to determine the meta-policy parameter of each task based on the new cell state of each task. Among them, by updating the forgetting gate and the input gate of each task through the trained task allocation model, the historical cell state, the historical forgetting gate, and the historical input gate obtained in the step of each task matching the previous device can be combined to determine the updated forgetting gate and the updated input gate, output the new cell state of each task, and further determine the meta-policy parameter of each task. Each task can update the internal parameter of the task by combining the context information data of the task, so that the finally determined meta-policy parameter is more accurate.

[0148] Input the meta-policy parameters of each task into the task allocation model, and output the new policy parameters shared by the tasks, so as to perform task scheduling for all tasks based on the new policy parameters. It can be seen that by inputting the meta-policy parameters of each task into the task allocation model, new policy parameters shared by all tasks can be obtained. The new policy parameters integrate the information of the meta-policy parameters of each task, and each meta-policy parameter also integrates the change data of each step when each task matches the target device. Therefore, each meta-policy parameter can be dynamically adjusted, and the new policy parameters are adjusted according to the changes of the meta-policy parameters, which can adapt to tasks in a variety of different environments and solve the problem of low flexibility in task scheduling of the computing power network.

[0149] Those skilled in the art should understand that Figure 7 the task scheduling device of the computing power network in can be used to implement the task scheduling method of the computing power network described above. The detailed description therein should be similar to the description in the method part above. To avoid redundancy, it will not be elaborated here.

[0150] Based on the same technical concept, an embodiment of the present application also provides an electronic device, which is used to execute the above-mentioned task scheduling method of the computing power network. Figure 8 FIG. is a schematic structural diagram of an electronic device for implementing various embodiments of the present application. The electronic device may vary greatly due to configuration or performance differences, and may include a processor 810, a communication interface 820, a memory 830, and a communication bus 1140. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call a computer program stored in the memory 830 and running on the processor 810 to execute the following steps: Obtain the task data of multiple tasks related to computing power network scheduling, input the task data into the trained task allocation model, and output the loss value of the initial policy parameters of the task on the task, as well as the forget gate, input gate, and output hidden layer features when the task matches the target device; Through the parameter update module of the task allocation model, perform update processing on the forget gate to obtain an updated forget gate, where the updated forget gate includes: output hidden layer features, the updated historical cell state and historical forget gate when the task matches the previous device; and perform update processing on the input gate to obtain an updated input gate, where the updated input gate includes: output hidden layer features, the updated historical cell state and historical input gate when the task matches the previous device; According to the historical cell state, updated forget gate, updated input gate, and loss value, output the new cell state of the task through the parameter update module, so as to determine the meta-policy parameters of each task based on the new cell state of each task; Input the meta-policy parameters of each task into the task allocation model, and output the new policy parameters shared by the tasks, so as to perform task scheduling for all tasks based on the new policy parameters.

[0151] Adopt the technical solution of the embodiment of the present application, obtain the task data of multiple tasks related to computing power network scheduling, input the task data of each task into the trained task allocation model, and output the loss value of the initialized policy parameters of the task on the task, as well as the forget gate, input gate and output hidden layer features when the task matches the target device; through the parameter update module of the task allocation model, perform update processing on the forget gate to obtain the updated forget gate, and the updated forget gate includes: output hidden layer features, the updated historical cell state and historical forget gate when the task matches the previous device; and perform update processing on the input gate to obtain the updated input gate, and the updated input gate includes: output hidden layer features, the updated historical cell state and historical input gate when the task matches the previous device; according to the historical cell state, updated forget gate, updated input gate and loss value, output the new cell state of the task through the parameter update module, so as to determine the meta-policy parameters of each task based on the new cell state of each task. Among them, by updating the forget gate and input gate of each task through the trained task allocation model, it is possible to combine the historical cell state, historical forget gate and historical input gate obtained in the step of each task matching the previous device to determine the updated forget gate and updated input gate, and output the new cell state of each task, and then determine the meta-policy parameters of each task. Each task can update the internal parameters of the task by combining the context information data of the task, making the finally determined meta-policy parameters more accurate.

[0152] Input the meta-policy parameters of each task into the task allocation model, and output the new policy parameters shared by the tasks, so as to perform task scheduling for all tasks based on the new policy parameters. It can be seen that by inputting the meta-policy parameters of each task into the task allocation model, new policy parameters shared by all tasks can be obtained. The new policy parameters synthesize the information of the meta-policy parameters of each task, and each meta-policy parameter also synthesizes the change data in the step of each task matching the target device. Therefore, each meta-policy parameter can be dynamically adjusted, and the new policy parameters are adjusted according to the changes of the meta-policy parameters, which can adapt to tasks in a variety of different environments and solve the problem of low flexibility in task scheduling of the computing power network.

[0153] The specific execution steps can refer to the steps of the embodiment of the task scheduling method of the above computing power network, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0154] It should be noted that the electronic devices in the embodiments of the present application include: servers, terminals or other devices other than terminals.

[0155] The above structure of the electronic device does not limit the electronic device. The electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. For example, the input unit may include a Graphics Processing Unit (GPU) and a microphone, and the display unit may be configured with a display panel in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit includes at least one of a touch panel and other input devices. The touch panel is also called a touch screen. Other input devices may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0156] The memory can be used to store software programs and various data. The memory mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory can include volatile memory or non-volatile memory, or the memory can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synchlink DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM).

[0157] The processor may include one or more processing units; optionally, the processor integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor may not be integrated into the processor either.

[0158] The embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the task scheduling method embodiment of the above-mentioned computing power network and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0159] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc, etc.

[0160] The embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement each process of the task scheduling method embodiment of the above-mentioned computing power network and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0161] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0162] The embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the processor is used to run the program or instruction to implement each process of the product recommendation method embodiment of the above-mentioned product and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0163] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising such element. In addition, it should be pointed out that the scope of the methods and apparatuses in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in a different order from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0164] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0165] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit and scope protected by the claims of the present application, can still make many forms, all of which fall within the protection scope of the present application.

Claims

1. A task scheduling method for a computing power network, characterized in that, The method includes: Obtain the task data of multiple tasks related to computing power network scheduling, input the task data into the trained task allocation model, and output the loss value of the initialized policy parameters of the task on the task, as well as the forget gate, input gate, and output hidden layer features when the task matches the target device; Through the parameter update module of the task allocation model, update the forget gate to obtain an updated forget gate, where the updated forget gate includes: the output hidden layer features, the historical cell state updated when the task matches the previous device, and the historical forget gate; and update the input gate to obtain an updated input gate, where the updated input gate includes: the output hidden layer features, the historical cell state updated when the task matches the previous device, and the historical input gate; According to the historical cell state, the updated forget gate, the updated input gate, and the loss value, output the new cell state of the task through the parameter update module, and determine the meta-policy parameters of each task based on the new cell state of each task; Input the meta-policy parameters of each task into the task allocation model, and output the new policy parameters shared by the tasks, so as to perform task scheduling on all the tasks based on the new policy parameters.

2. The method according to claim 1, wherein Inputting the task data into the trained task allocation model and outputting the forget gate, input gate, and output hidden layer features when the task matches the target device includes: Input the task data into the trained task allocation model, input the loss value into the feature extraction module of the task allocation model, and obtain the forget gate when the task matches the target device according to the loss value of the task, the activation function, the output hidden layer features when the task matches the previous device, the weight matrix of the forget gate, and the bias of the forget gate; and obtain the input gate when the task matches the target device according to the loss value of the task, the activation function, the output hidden layer features when the task matches the previous device, the weight matrix of the input gate, and the bias of the input gate; Calculate the candidate cell state of the task according to the forget gate and the input gate, and determine the cell state of the task based on the candidate cell state; Determine the output gate of the task based on the cell state and the activation function, and output the output hidden layer features when the task matches the target device according to the output gate and the cell state.

3. The method according to claim 1, characterized in that The step of inputting the meta-policy parameters of each task into the task allocation model, outputting the new policy parameters shared by the tasks, and performing task scheduling on all the tasks based on the new policy parameters includes: Input the meta-policy parameters of each task into the task allocation model, where the meta-policy parameters include: the new cell state when each task matches the last device; Subtract the proxy loss and proxy loss gradient of the meta-policy parameters of each task from the shared meta-policy parameters in the task allocation model, and output the new policy parameters shared by the tasks; where the proxy loss includes a clipping function.

4. The method according to claim 1, wherein Obtaining task data of multiple tasks related to computing power network scheduling includes: Responding to multiple task requests, adding the multiple tasks to a scheduling queue, obtaining the queuing waiting time for each task to be added to the scheduling queue, and obtaining the scheduling delay time for scheduling the task through a controller; Transmitting the task to an edge device through the controller to obtain the transmission delay time and transmission energy consumption; by executing the task, obtaining the execution time and execution energy consumption for executing the task; Determining the total delay of the task according to one or more of the queuing waiting time, the scheduling delay time, the transmission delay time, and the execution time; determining the total energy consumption of the task according to one or more of the transmission energy consumption and the execution energy consumption; Obtaining the state information of the task through a meta-reinforcement learning method, where the state information includes: the device list and state list of the task; Obtaining a set of multiple allocation situations for the task to be assigned to a target device, and obtaining the spatial information corresponding to the task; Determining a reward function according to the negative value of the weighted sum of the total delay and the total energy consumption; the reward function is used to determine the priority of the task to be assigned to the target device; Determining the state information, the spatial information, and the reward function as the task data of the task.

5. The method according to claim 4, wherein Inputting the task data into a trained task allocation model, and outputting the loss value of the initialized policy parameter of the task on the task, including: Inputting the obtained state information, the spatial information, and the reward function into the task allocation model, and calculating the loss function and the loss function gradient of each task through the preprocessing module of the task allocation model; Performing normalization processing on the loss function and the loss function gradient through the preprocessing module, and outputting the loss value of the initialized policy parameter of the task on the task after processing.

6. The method according to claim 1, characterized in that, The training of the task allocation model includes: Obtaining the sample task data and the corresponding sample policy of multiple sample tasks, and inputting the sample task data into the task allocation model to be trained; using the preprocessing module of the task allocation model to be trained to calculate the sample loss value corresponding to the sample task; Inputting the sample loss value into the feature extraction module of the task allocation model to be trained, outputting a sample forget gate and a sample input gate, and based on the sample forget gate and the sample input gate, outputting a sample output gate and a sample output hidden layer feature of the task; Updating the parameters of the sample forget gate and the sample input gate through the parameter update module of the task allocation model to be trained to obtain a sample updated forget gate and a sample updated input gate; Outputting a sample new cell state according to the sample historical cell state updated when the sample task matches the previous device, the sample updated forget gate, the sample updated input gate, and the sample loss value; Based on the sample new cell state of each of the sample tasks, obtain the sample meta-policy parameters of each of the sample tasks, and input the sample meta-policy parameters of each of the tasks into the task allocation model to be trained, and output the sample new policy parameters shared by all the sample tasks; Train the task allocation model to be trained according to the sample new policy parameters and the sample policy to obtain the trained task allocation model.

7. A task scheduling device for a computing power network, characterized in that It includes: An acquisition module, configured to acquire the task data of multiple tasks related to computing power network scheduling, input the task data into the trained task allocation model, and output the loss value of the initialized policy parameters of the task on the task, as well as the forget gate, input gate, and output hidden layer features when the task matches the target device; A parameter update module, configured to perform an update process on the forget gate through the parameter update module of the task allocation model to obtain an updated forget gate, where the updated forget gate includes: the output hidden layer features, the updated historical cell state and historical forget gate when the task matches the previous device; and perform an update process on the input gate to obtain an updated input gate, where the updated input gate includes: the output hidden layer features, the updated historical cell state and historical input gate when the task matches the previous device; A determination module, configured to determine the new cell state of the task through the parameter update module according to the historical cell state, the updated forget gate, the updated input gate, and the loss value, so as to determine the meta-policy parameters of each task based on the new cell state of each task; An output module, configured to input the meta-policy parameters of each task into the task allocation model, and output the new policy parameters shared by the tasks, so as to perform task scheduling on all the tasks based on the new policy parameters.

8. An electronic device, characterized in that, It includes a processor and a memory electrically connected to the processor, where the memory stores a computer program, and the processor is configured to call and execute the computer program from the memory to implement a task scheduling method for a computing power network according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium is used to store a computer program, and the computer program can be executed by a processor to implement a task scheduling method for a computing power network according to any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a processor, it implements a task scheduling method for a computing power network according to any one of claims 1 to 6.

Citation Information

Cited By

  • LSTM reasoning method based on channel cutting and gating configuration and electronic equipment

    CN121212207A