Task scheduling method and device, computing power network fusion platform and storage medium

By connecting the computing power network fusion platform with the cloud management platform and the network management platform, and by using reinforcement learning and Markov decision process models to optimize the allocation of cloud center resources, the complex problem of computing power network scheduling is solved, and efficient computing network scheduling and personalized services are achieved.

CN120994337APending Publication Date: 2025-11-21CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511094915.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

The scheduling process of computing power networks is complex, and there is considerable room for improvement in scheduling efficiency. It is difficult to consider multiple factors such as equipment status, task requirements, and network conditions in a short period of time, which makes it difficult to guarantee service quality and efficiency.

Method used

By connecting the computing power network fusion platform with the cloud management platform and the network management platform, the trained target policy network is used for task scheduling. The Markov decision process model and the actor-critic architecture are combined for reinforcement learning to generate scheduling decision groups and optimize the allocation of cloud center resources.

Benefits of technology

It enables personalized computing network scheduling services, improves the efficiency and accuracy of computing network scheduling, meets user needs, reduces operating costs, and improves service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994337A_ABST
    Figure CN120994337A_ABST
Patent Text Reader

Abstract

The invention provides a task scheduling method and device, a computing power network fusion platform and a storage medium, and relates to the technical field of computing power networks. The method comprises the following steps: acquiring a target computing task of a target user, target scheduling preference information of the target user, a current link network state between the target user and each cloud center and current computing power data corresponding to each cloud center; matching the target computing task with the current computing power data corresponding to each cloud center and the current link network state to generate a corresponding scheduling adaptation group; inputting all the scheduling adaptation groups and the target scheduling preference information into a target strategy network trained to be converged, and generating corresponding scheduling decision groups; the scheduling decision group comprises a target cloud center and a network link between the target user and the target cloud center; and scheduling a resource calculation target calculation task corresponding to the target cloud center according to the scheduling decision group. According to the task scheduling method, the efficiency of computing power network scheduling is improved in a reinforcement learning mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computing power network technology, specifically relating to a task scheduling method, device, computing power network fusion platform, and storage medium. Background Technology

[0002] Computing power scheduling is a new capability system for resolving the contradiction between computing power supply and demand, computing power network transmission issues, and the problem of universal access to computing power resources. Based on the supply capacity of computing power resource providers and the dynamic resource needs of application users, computing power scheduling integrates multi-dimensional resources such as computing, storage, and network at the underlying level of computing power infrastructure within a region. Based on a computing power scheduling platform, it performs consistent management, integrated orchestration, and unified scheduling of computing power resources, achieving collaborative linkage and precise matching of computing power resources across industries, regions, and levels.

[0003] In related technologies, computing demands and equipment status may change rapidly. Multiple factors such as equipment status, task requirements, and network conditions need to be considered, and decisions need to be made in a short period of time to ensure service quality and efficiency. This makes the computing power network scheduling process very complex, and there is considerable room for improvement in scheduling efficiency.

[0004] Therefore, a new computing power network scheduling method is needed in related technologies to improve the efficiency of computing power network scheduling. Summary of the Invention

[0005] The technical problem to be solved by this application is to address the above-mentioned shortcomings of the existing technology by providing a task scheduling method, apparatus, computing power network fusion platform and storage medium. Using this task scheduling method can improve the efficiency of computing power network scheduling.

[0006] In a first aspect, embodiments of this application provide a task scheduling method, wherein a computing power network fusion platform is connected to a cloud management platform and a network management platform respectively, and the computing power network fusion platform schedules cloud center resources through the cloud management platform and the network management platform. The method is applied to the computing power network fusion platform and includes:

[0007] Acquire the target computing tasks of the target user, the target scheduling preference information of the target user, the current link network status between the target user and each cloud center, and the current computing capacity data of each cloud center;

[0008] Match the target computing task with the current computing capacity data and current link network status of each cloud center to generate the corresponding scheduling adaptation group;

[0009] All scheduling adaptation groups and target scheduling preference information are input into the convergent target policy network to generate corresponding scheduling decision groups; the scheduling decision groups include the target cloud center and the network links between the target user and the target cloud center.

[0010] The scheduling decision group schedules the resource calculation tasks for the corresponding target cloud center.

[0011] In some implementations of the first aspect, obtaining the current link network status between the target user and each cloud center includes:

[0012] Obtain the network status of three links between the target user and each cloud center up to the current time. The first link network status corresponds to the time one moment before the current time, the second link network status corresponds to the time two moments before the current time, and the third link network status corresponds to the time three moments before the current time. Each link network status includes transmission latency, bandwidth utilization, and packet loss rate.

[0013] The first, second, and third link network states are input into a convergent gated recurrent unit to generate the current link network state between the target user and each cloud center.

[0014] In some implementations of the first aspect, the target computing task includes: the number of processor clock cycles required by the task, memory size, disk size, and device time; the current computing power data includes: processor storage capacity, memory capacity, disk storage capacity, computing power index, price per processor, price per memory, and price per disk; and there are multiple scheduling adaptation groups.

[0015] The target computing task is matched with the current computing capacity data and current link network status of each cloud center to generate corresponding scheduling adaptation groups, including:

[0016] Based on the target computing task and the current computing capacity data, all candidate cloud centers capable of handling the target computing task are determined from each cloud center.

[0017] Each candidate cloud center and its corresponding network link are combined to generate multiple scheduling adaptation groups.

[0018] In some implementations of the first aspect, before inputting all scheduling adaptation groups and target scheduling preference information into a converged target policy network to generate the corresponding scheduling decision group, the method further includes:

[0019] The five-tuple of the Markov decision process model is defined based on historical data to establish the Markov decision process model; wherein, the historical data includes the historical computing task pool corresponding to the target user, the historical scheduling preference information corresponding to the target user, the historical link network status between the target user and each cloud center, and the historical computing capacity data pool corresponding to each cloud center.

[0020] Based on the Actor-Critic architecture, the preset policy network and preset value network are iteratively trained according to the training adaptation group, generating the training decision group and the decision score corresponding to the training decision group for each iteration; the training adaptation group corresponds to the current state space in the quintuple; the training decision group corresponds to the action space in the quintuple.

[0021] The immediate reward and rating score are determined based on the historical scheduling preference information and decision scores corresponding to the training and adaptation groups; the rating score is the score corresponding to the advantage function.

[0022] An experience pool is built based on training adaptation groups, training decision groups, instant rewards, and rating scores.

[0023] The preset policy network and preset value network are updated based on the experience pool, and the updated preset policy network and preset value network are iteratively trained until the convergence condition is met.

[0024] In some implementations of the first aspect, historical scheduling preference information includes four weights corresponding to execution energy consumption, equipment redundancy, price, and network service quality, respectively.

[0025] The five-tuples of a Markov decision process model, defined based on historical data, include:

[0026] The historical computing power data at the target time and the historical link network status corresponding to the target time are used as the current state space; the historical computing power data are the computing power data in the historical computing power data pool that matches the historical computing tasks; the historical computing tasks are the computing tasks at the target time in the historical computing task pool.

[0027] The training decision group corresponding to the historical computing task is used as the action space; the training decision group includes the scheduling cloud center corresponding to the historical computing task and the network link between the target user and the scheduling cloud center.

[0028] The total cost of historical computing tasks is determined based on the historical link network status at the target time, the four weights, and the historical computing capacity data of the scheduling cloud center.

[0029] The reciprocal of the total cost of the task is used as the immediate reward;

[0030] Define a discount factor and determine the next state space based on the current state space and the training decision group.

[0031] In some implementations of the first aspect, a preset policy network and a preset value network are iteratively trained according to a training adaptation group to generate a training decision group and a decision score corresponding to each iteration of training, including:

[0032] Match historical computing tasks with historical computing capability data in the historical computing capability data pool and the corresponding historical link network status to generate corresponding training adaptation groups.

[0033] Input all training adaptation groups and their corresponding historical scheduling preference information into the preset policy network to generate the probability of each training adaptation group.

[0034] The probabilities of each training adaptation group are sampled to generate the corresponding training decision group;

[0035] Input the training decision group and the training adaptation group into the preset value network to generate the decision score corresponding to the training decision group.

[0036] In some implementations of the first aspect, updating the preset policy network and the preset value network based on the experience pool includes:

[0037] The data in the experience pool are sampled using both priority sampling and random sampling methods;

[0038] The preset strategy network and preset value network are updated based on the obtained data.

[0039] Based on the same inventive concept, in a second aspect, embodiments of this application also provide a task scheduling device. A computing power network fusion platform is connected to both a cloud management platform and a network management platform. The computing power network fusion platform schedules cloud center resources through the cloud management platform and the network management platform. The device is located on the computing power network fusion platform and includes:

[0040] The acquisition module is used to acquire the target computing task of the target user, the target scheduling preference information of the target user, the current link network status between the target user and each cloud center, and the current computing capacity data of each cloud center.

[0041] The matching module is used to match the target computing task with the current computing capacity data and current link network status of each cloud center to generate the corresponding scheduling adaptation group.

[0042] The generation module is used to input all scheduling adaptation groups and target scheduling preference information into the target policy network trained to convergence, and generate corresponding scheduling decision groups; the scheduling decision groups include the target cloud center and the network link between the target user and the target cloud center.

[0043] The computing module is used to perform target computing tasks according to the resource allocation of the corresponding target cloud center by the scheduling decision group.

[0044] In some implementations of the second aspect, when the acquisition module acquires the current link network status between the target user and each cloud center, it is specifically used for:

[0045] Obtain the three link network states between the target user and each cloud center up to the current time. The first link network state corresponds to the time one moment before the current time, the second link network state corresponds to the time two moments before the current time, and the third link network state corresponds to the time three moments before the current time. Each link network state includes transmission latency, bandwidth utilization, and packet loss rate. Input the first, second, and third link network states into a gated recurrent unit trained to convergence to generate the current link network states between the target user and each cloud center.

[0046] In some implementations of the second aspect, the target computing task includes: the number of processor clock cycles required by the task, memory size, disk size, and device time; the current computing capability data includes: processor storage capacity, memory capacity, disk storage capacity, computing power indicators, price per processor, price per memory, and price per disk; and there are multiple scheduling adaptation groups.

[0047] The matching module is specifically used for:

[0048] Based on the target computing task and the current computing capacity data, all candidate cloud centers capable of handling the target computing task are determined from each cloud center; each candidate cloud center and its corresponding network link are combined to generate multiple scheduling adaptation groups.

[0049] In some embodiments of the second aspect, the apparatus further includes:

[0050] The training module defines the quintuples of the Markov decision process model based on historical data to establish the model. Historical data includes the historical computing task pool corresponding to the target user, historical scheduling preference information for the target user, historical link network status between the target user and each cloud center, and historical computing capacity data pools for each cloud center. Based on the Actor-Critic architecture, the module iteratively trains the preset policy network and preset value network according to the training adaptation group, generating the training decision group and decision score corresponding to each iteration. The training adaptation group corresponds to the current state space in the quintuple; the training decision group corresponds to the action space in the quintuple. The module determines the immediate reward and rating score based on the historical scheduling preference information and decision score corresponding to the training adaptation group; the rating score is the score corresponding to the dominance function. An experience pool is built based on the training adaptation group, training decision group, immediate reward, and rating score. The preset policy network and preset value network are updated based on the experience pool, and iterative training continues on the updated preset policy network and preset value network until convergence conditions are met.

[0051] In some implementations of the second aspect, historical scheduling preference information includes four weights corresponding to execution energy consumption, equipment redundancy, price, and network service quality, respectively.

[0052] The training module is specifically used to define the quintuples of a Markov decision process model based on historical data:

[0053] The current state space is defined as the historical computing power data at the target time and the historical link network state corresponding to the target time. The historical computing power data is the computing power data in the historical computing power data pool that matches the historical computing tasks. The historical computing tasks are the computing tasks at the target time in the historical computing task pool. The training decision group corresponding to the historical computing tasks is defined as the action space. The training decision group includes the scheduling cloud center corresponding to the historical computing tasks and the network link between the target user and the scheduling cloud center. The total task consumption cost corresponding to the historical computing tasks is determined based on the historical link network state at the target time, the four weights, and the historical computing power data corresponding to the scheduling cloud center. The reciprocal of the total task consumption cost is determined as the immediate reward. A discount factor is defined, and the next state space is determined based on the current state space and the training decision group.

[0054] In some implementations of the second aspect, when the training module iteratively trains the preset policy network and the preset value network according to the training adaptation group, and generates the training decision group and the decision score corresponding to each iteration of training, it is specifically used for:

[0055] The historical computing tasks are matched with the historical computing capacity data in the historical computing capacity data pool and the corresponding historical link network status to generate corresponding training adaptation groups; all training adaptation groups and their corresponding historical scheduling preference information are input into a preset policy network to generate the probability of each training adaptation group; the probability of each training adaptation group is sampled to generate the corresponding training decision group; the training decision group and training adaptation group are input into a preset value network to generate the decision score corresponding to the training decision group.

[0056] In some implementations of the second aspect, when the training module updates the preset policy network and the preset value network based on the experience pool, it is specifically used for:

[0057] Data in the experience pool is sampled using priority sampling and random sampling methods; the preset policy network and preset value network are updated based on the obtained data.

[0058] Based on the same inventive concept, in a third aspect, embodiments of this application also provide a computing power network fusion platform, which includes:

[0059] Memory and processor;

[0060] The memory stores instructions that the computer executes;

[0061] The processor executes computer execution instructions stored in memory to implement a task scheduling method as described in any of the first aspects.

[0062] Based on the same inventive concept, in a fourth aspect, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the task scheduling method as described in any of the first aspects.

[0063] According to the task scheduling method, apparatus, computing power network fusion platform, and storage medium provided in this application, target computing tasks are matched with the current computing capacity data and current link network status of each cloud center to generate corresponding scheduling adaptation groups. All scheduling adaptation groups and target scheduling preference information are input into a converged target policy network to generate corresponding scheduling decision groups, thereby integrating user needs into the scheduling decision-making process and realizing personalized computing network scheduling services. Simultaneously, resource computing target tasks are scheduled according to the corresponding target cloud center according to the scheduling decision groups, thereby improving the efficiency of computing power network scheduling through reinforcement learning, i.e., making decisions through the target policy network. Attached Figure Description

[0064] Figure 1 This illustration shows a flowchart of a task scheduling method provided in an embodiment of this application;

[0065] Figure 2 This illustration shows another flowchart of the task scheduling method provided in an embodiment of this application;

[0066] Figure 3 This diagram illustrates the scheduling architecture of the task scheduling method provided in an embodiment of this application.

[0067] Figure 4 This diagram illustrates the signal characteristics of the gated loop unit provided in an embodiment of this application.

[0068] Figure 5 This illustration shows a training process diagram provided in an embodiment of this application;

[0069] Figure 6 This illustration shows a policy network diagram provided in an embodiment of this application.

[0070] Figure 7 This illustration shows a value network diagram provided in an embodiment of this application;

[0071] Figure 8 This is a schematic diagram of the structure of the task scheduling device provided in an embodiment of this application. Detailed Implementation

[0072] To enable those skilled in the art to better understand the technical solutions of this application, the application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0073] The features and exemplary embodiments of various aspects of this application will now be described in detail. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only configured to explain this application and are not configured to limit this application. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples of this application.

[0074] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0075] As a key productive force in the digital economy era, computing power is experiencing explosive growth in demand across various industries. This has brought new challenges due to the mismatch between computing power demand and supply. Computing networks connect computing resources distributed across cloud, edge, and terminal devices in various locations, forming a distributed, dynamic pool of computing resources. This pool enables unified management and scheduling of all computing resources, thereby improving resource utilization, reducing operating costs, and enhancing service quality.

[0076] The computing demands and equipment status in related technologies may change rapidly. Multiple factors such as equipment status, task requirements, and network conditions need to be considered, and decisions need to be made in a short period of time to ensure the quality and efficiency of services. This makes the computing power network scheduling process very complex.

[0077] Therefore, a new computing power network scheduling method is needed in related technologies to improve the efficiency of computing power network scheduling.

[0078] Example 1

[0079] The task scheduling method provided in this application can be executed by a task scheduling device and a computing power-network convergence platform. The following description uses the execution of this task scheduling method by a computing power-network convergence platform as an example. The computing power-network convergence platform is connected to both a cloud management platform and a network management platform, and schedules cloud center resources through these platforms. The overall scheduling architecture is divided into two levels. Level 1 scheduling uses the computing power-network convergence platform to manage the cloud management and network management platforms, enabling cloud and routing operations. Level 2 scheduling uses technologies such as Kubernetes (or k8s, an open-source container orchestration platform) to offload computing tasks to virtual machines or bare metal machines, achieving load balancing.

[0080] like Figure 1 As shown, the task scheduling method provided in this application embodiment may include steps S101 to S104.

[0081] S101. Obtain the target computing task of the target user, the target scheduling preference information of the target user, the current link network status between the target user and each cloud center, and the current computing capacity data of each cloud center.

[0082] For example, the target user can be any user, which can refer to a single user or a group of users. The target computing task can include multiple subtasks, each with its corresponding task requirements, including the number of CPU (processor) clock cycles required to execute task i, the amount of memory required to execute task i, the disk capacity required to execute task i, and the device time required to execute task i, etc.

[0083] For example, target scheduling preference information is related to scheduling policies, which may favor one or more of the following: execution energy consumption, device redundancy, price, and network service quality. Users can rate or select weights for these four factors, allowing the policy network to choose a more suitable scheduling policy based on user needs, thereby improving the user experience.

[0084] For example, the current link network state can be either no network connection or a network connection state, and the state of the network connection can be represented by transmission latency, bandwidth utilization, and packet loss rate.

[0085] For example, the current computing power data corresponding to the cloud center refers to the computing power data that can currently be used for computing tasks. This computing power data includes processor storage capacity, memory capacity, disk storage capacity, computing power indicators, price per processor, price per memory, price per disk, etc.

[0086] In some implementations, the process of obtaining the current link network status between the target user and each cloud center can be as follows:

[0087] Obtain the network states of three links between the target user and each cloud center up to the current time. The first link network state corresponds to the time one moment before the current time, the second link network state corresponds to the time two moments before the current time, and the third link network state corresponds to the time three moments before the current time. Each link network state includes transmission latency, bandwidth utilization, and packet loss rate.

[0088] The first, second, and third link network states are input into a convergent gated recurrent unit to generate the current link network state between the target user and each cloud center.

[0089] By employing Gated Recurrent Units (GRUs) to train and predict the state of network links, the current state of the links can be accurately captured, and the state at the time of future task distribution can be predicted. This improves the accuracy of network link state prediction, and consequently, the accuracy of task scheduling.

[0090] For example, assuming the current time is t, the first link network state corresponds to time t-1, the second link network state corresponds to time t-2, and the third link network state corresponds to time t-3.

[0091] If the current time t is the initial time after the device is started, then the first link network state, the second link network state, and the third link network state can be preset default values.

[0092] S102. Match the target computing task with the current computing capacity data and current link network status of each cloud center to generate the corresponding scheduling adaptation group.

[0093] For example, the scheduling adaptation group consists of cloud centers that, after filtering, can support the target computing task, as well as the network links between the target user and those cloud centers. The ability to support the target computing task indicates that the current computing capacity of the cloud center meets the requirements of the target computing task.

[0094] In some implementations, the target computing task includes: the number of processor clock cycles required by the task, memory size, disk size, and device time. Current computing capability data includes: processor storage capacity, memory capacity, disk storage capacity, computing power metrics, price per processor, price per memory module, and price per disk. Furthermore, there are multiple scheduling adaptation groups.

[0095] S102 can be specifically as follows:

[0096] Based on the target computing task and the current computing capacity data, all candidate cloud centers capable of handling the target computing task are identified from each cloud center.

[0097] Each candidate cloud center and its corresponding network link are combined to generate multiple scheduling adaptation groups.

[0098] For example, if cloud centers b and d can meet the task requirements of the target computing task, then cloud centers b and d are candidate cloud centers. Cloud center b has two network links, namely link a and link b. Cloud center d has two network links, namely link c and link d.

[0099] The scheduling adaptation groups are (cloud center b, link a), (cloud center b, link b), (cloud center d, link c), and (cloud center d, link d).

[0100] S103. Input all scheduling adaptation groups and target scheduling preference information into the converged target policy network to generate the corresponding scheduling decision group. The scheduling decision group includes the target cloud center and the network link between the target user and the target cloud center.

[0101] For example, the target cloud center is the cloud center that is ultimately chosen to be invoked, through which the target computing task is computed. The network link between the target user and the target cloud center is the network link that is ultimately used, and this network link may be the one with the best or highest network quality between the target user and the target cloud center. The target user communicates with the target cloud center and unloads the task through this network link.

[0102] S104. Calculate the target computing task according to the resource calculation of the corresponding target cloud center as scheduled by the scheduling decision group.

[0103] According to the task scheduling method provided in this application, a corresponding scheduling adaptation group is generated by matching the target computing task with the current computing capacity data and current link network status of each cloud center. All scheduling adaptation groups and target scheduling preference information are input into a converged target policy network to generate a corresponding scheduling decision group, thereby integrating user needs into the scheduling decision process and realizing personalized computing network scheduling services. Simultaneously, the resource computing target tasks of the corresponding target cloud center are scheduled according to the scheduling decision group, thus improving the efficiency of computing network scheduling through reinforcement learning, i.e., making decisions through the target policy network.

[0104] Example 2

[0105] like Figure 2As shown, the task scheduling method provided in this application embodiment, based on the task scheduling method provided in embodiment 1 of this application, further explains the network training process and may include steps S201 to S205.

[0106] S201. Define the quintuple of the Markov decision process model based on historical data to establish the Markov decision process model. The historical data includes the historical computing task pool corresponding to the target user, the historical scheduling preference information corresponding to the target user, the historical link network status between the target user and each cloud center, and the historical computing capacity data pool corresponding to each cloud center.

[0107] For example, the historical computing task pool includes the historical computing tasks of the target user. In this embodiment, the target user can be multiple different users, or it can correspond to the target user in Embodiment 1. The historical computing capacity data pool includes the historical computing capacity data of each cloud center.

[0108] In some implementations, historical scheduling preference information includes four weights corresponding to execution energy consumption, equipment redundancy, price, and network service quality, respectively.

[0109] The specific process for defining the five-tuple of a Markov decision process model based on historical data is as follows:

[0110] The current state space is defined by the historical computing power data at the target time and the corresponding historical link network status. The historical computing power data is the computing power data in the historical computing power data pool that matches the historical computing tasks. The historical computing tasks are the computing tasks at the target time in the historical computing task pool.

[0111] The training decision group corresponding to the historical computing tasks is used as the action space. The training decision group includes the scheduling cloud center corresponding to the historical computing tasks and the network link between the target user and the scheduling cloud center.

[0112] The total cost of historical computing tasks is determined based on the historical link network status at the target time, the four weights, and the historical computing capacity data of the scheduling cloud center.

[0113] The reciprocal of the total cost of the task is used as the immediate reward.

[0114] Define a discount factor and determine the next state space based on the current state space and the training decision group.

[0115] For example, the target time refers to the current time used during training. Based on the target time, the corresponding historical computing power data and historical link network status are used as the current state space, and the corresponding historical computing tasks are used as the training tasks. Moreover, the historical computing power data and historical link network status are related to the historical computing tasks. For example, the historical computing power data can meet the requirements of the historical computing tasks, and the historical link network status corresponds to the cloud center that meets the requirements of the historical computing tasks.

[0116] For example, the training decision group is similar to the scheduling decision group, where the training decision group is the decision group generated during training. The scheduling cloud center is the cloud center determined during training for computing historical computing tasks.

[0117] For example, the total cost of historical computing tasks can be determined using a general calculation method, such as calculating the relationship between the task requirements and prices of historical computing tasks, and simultaneously weighting the execution energy consumption, equipment redundancy, price and network service quality to obtain the final total cost of the task.

[0118] For example, suppose there are k clouds in the network, C = {C1, C2, C3, ..., C...} k} represents the set of computing cloud nodes, where C j Let j represent the j-th cloud computing node, where j∈1,2,...k. The computing power of a cloud node is represented as... Among them U j Indicates CPU storage capacity, M j D represents the amount of memory. j P represents the disk storage capacity. cj This indicates the calculation of power indicators. This indicates the price per CPU. Indicates the price per unit of memory. This indicates the price of a single disk.

[0119] There are s links between user m and cloud j, R = {R1, R2, R3 ... R...} S}

[0120] Among them, R s = [TD,PL,BW], where TD represents transmission delay, PL represents packet loss rate, and BW represents bandwidth utilization.

[0121] The set of computational tasks A submitted by the user can contain i tasks, i.e., A = {A1, A2, A3, ..., A...} i}, where i = 1, 2, ..., n, and each task is described as A i =[u i ,m i ,d i ,τi ].

[0122] Where u i m represents the number of CPU clock cycles required to execute task i. i d represents the memory size required to execute task i. i τ represents the disk required to execute task i. i This indicates the device time required to execute task i. When a user initiates a computing task, a resource scheduling decision is made based on the user's computing task requirements, the network status of the link to the user, and the computing capabilities of each cloud center node.

[0123] Decision-making uses a set of Boolean variables B = {B1, B2, B3, ..., B} k Let} represent the number of cloud center nodes, where k represents the number of nodes in the cloud center, and B k It is a Boolean variable, 0 or 1, where 1 indicates that the task is executed on the cloud central node.

[0124] For subtask A i When B j When = 1, the computation task is executed in cloud center j, and the computing node in that cloud center processes subtask A. i The time consumed is τ i Cloud center for sub-task A i The energy consumption during processing is:

[0125]

[0126] P cj Let j be the computing power of the cloud center device. Task A i The cost required to execute j in the cloud center is:

[0127]

[0128] Where 0≤α1, α2, α3≤1, α1, α2, α3 represent the weighting factors of execution energy consumption, equipment redundancy, and price, respectively.

[0129] Network link cost is predicted using a gated recurrent unit. Input sample X t The sample set of state information of the link from user m to cloud j required to predict the network link state at time t:

[0130] X t ={[R1(t-3),R1(t-2),R1(t-1)],

[0131] [R2(t-3),R2(t-2),R2(t-1)]……

[0132] [R s (t-3 ),R s (t- 2 ),R s (t- 1 )]}(3)

[0133] Among them, R s = [TD(t),PL(t),BW(t)], where TD(t) represents the transmission delay at sampling time t, PL(t) represents the packet loss rate at sampling time t, and BW(t) represents the bandwidth utilization at sampling time t.

[0134] The sample information required at each time step includes the network link state information from the previous three time steps. In GRU network training, a time-series dataset is used. To update the r, z gate and hidden layer information, we finally obtain the required network link state prediction information:

[0135] Y t ={R1(t),R2(t),R3(t)……R s (t)} (4)

[0136] Based on the transmission delay, bandwidth utilization, and packet loss rate obtained at that moment, task A can be derived. i The network cost of one of the network links s that is offloaded to cloud center j:

[0137]

[0138] Where α4 is the network service quality index factor, and α4 < 1.

[0139] If subtask i can only choose one of many cloud computing centers j and selects one network link s for computation offloading, then the total cost incurred by the user to complete a certain task is:

[0140]

[0141] The immediate reward is the reciprocal of the total cost incurred by the task.

[0142] For example, by scheduling computing tasks to a specific cloud center and network link combination through a scheduling algorithm, and removing the remaining resources (remaining cloud centers and network links) after removing the cloud center and network link combination, the next state space is obtained.

[0143] S202. Based on the Actor-Critic architecture, iteratively train the preset policy network and preset value network according to the training adaptation group, generating the training decision group and the decision score corresponding to each iteration. The training adaptation group corresponds to the current state space in the quintuple. The training decision group corresponds to the action space in the quintuple.

[0144] For example, the Actor-Critic architecture refers to a policy network that makes decisions based on the state, and a value network that evaluates the decisions, thereby optimizing the policy network to produce better decisions.

[0145] For example, the training decision group is generated by the policy network, and the decision score is generated by the value network. The input of the value network can be the training decision group and the training adaptation group.

[0146] In some implementations, S202 may be specifically as follows:

[0147] Historical computing tasks are matched with historical computing capability data in the historical computing capability data pool and the corresponding historical link network status to generate corresponding training adaptation groups.

[0148] All training adaptation groups and their corresponding historical scheduling preference information are input into a preset policy network to generate the probability of each training adaptation group.

[0149] The probabilities of each training adaptation group are sampled to generate the corresponding training decision group.

[0150] Input the training decision group and the training adaptation group into the preset value network to generate the decision score corresponding to the training decision group.

[0151] For example, sampling can be based on the highest probability sample or random sampling. Decision scores can be used to guide the updating and iteration of a pre-defined policy network.

[0152] S203. Determine the immediate reward and rating score based on the historical scheduling preference information and decision scores corresponding to the training and adaptation groups. The rating score is the score corresponding to the dominance function.

[0153] For example, the immediate reward can be calculated based on historical scheduling preference information and the aforementioned algorithm for immediate rewards. The rating score, which corresponds to the advantage function, can be calculated based on the decision score.

[0154] S204. Build an experience pool based on training adaptation groups, training decision groups, instant rewards, and rating scores.

[0155] For example, the experience pool is used for the experience replay mechanism, which allows historical samples to be reused, avoiding over-reliance on the latest acquired data. Each time training is performed, the training adaptation group, training decision group, immediate reward, and rating score obtained from that training are combined, and when it is determined that the data obtained from the training is of good quality, it is stored in the experience pool.

[0156] S205. Update the preset policy network and preset value network according to the experience pool, and continue to iterate the training of the updated preset policy network and preset value network until the convergence condition is met.

[0157] For example, the importance of empirical data is assessed, it is incorporated into an experience pool, and network parameters are optimized based on priority and a random sampling and replay mechanism.

[0158] For example, the convergence condition could be that the losses of the policy network and the value network meet the loss conditions, or that the number of iterations reaches a certain threshold.

[0159] In some implementations, the process of updating the preset policy network and the preset value network based on the experience pool can be as follows:

[0160] The data in the experience pool are sampled using priority sampling and random sampling methods.

[0161] The preset strategy network and preset value network are updated based on the obtained data.

[0162] During parameter updates, a target network mechanism can be incorporated. This mechanism employs a dual neural network approach, using identical structures for the value network and target value network, as well as the policy network and target policy network. This addresses the feedback loop problem during training, improving algorithm stability and training efficiency. Specifically, the policy network and value network can adopt different update rates. A faster update rate for the target value network helps reduce overfitting and facilitates faster learning of the policy network, as it continuously learns from the latest data. More frequent value network updates may help stabilize the learning process, adapt more quickly to policy changes, and provide more timely value estimates.

[0163] This embodiment employs a two-level scheduling method and four scheduling strategies to address the issues of limited scheduling strategies, inconsistent metrics, and poor user matching in computing networks. Furthermore, it utilizes a GRU network to resolve the problem of unpredictable perception of network state information.

[0164] The task scheduling method in this embodiment transforms the computation-network fusion scheduling problem into a Markovian solution process. Experience data is graded and placed into an experience pool, and network parameters are updated and learned using an experience replay mechanism based on priority. The policy network and value network update their parameters at different rates. During training, the output of the policy network is modified; instead of directly outputting action values, the probability of outputting action values ​​is sampled, introducing randomness to improve search efficiency and avoid getting trapped in local optima.

[0165] To better understand the task scheduling method provided in the embodiments of this application, a specific application implementation method will be described below.

[0166] The task scheduling method in this embodiment uses a computing power network fusion platform as the execution entity, as detailed below:

[0167] 1. Network Scenarios

[0168] like Figure 3 As shown, this embodiment proposes a two-level centralized computing network scheduling method. The first-level scheduling implements cloud selection and routing operations through a computing power network fusion platform that integrates cloud management platform and network management platform. Only cloud data center A and cloud data center N are used as examples for illustration. The second-level scheduling uses technologies such as Kubernetes to offload computing tasks to virtual machines (e.g., virtual machine A) or bare metal (e.g., bare metal A) to achieve load balancing. Figure 3 In this context, VPC (Virtual Private Cloud) and EIP (Elastic IP Address, Internet Protocol) are used.

[0169] Among them, the cloud management platform performs (1) computing power resource perception, the network management platform performs (2) network resource status perception, the computing power network fusion platform performs (3) intelligent processing of computing network status information through the cloud management platform, and performs (4) calculation of the optimal path based on computing power and network status through the network management platform. Thus, it performs (5) unloading the computing to a virtual machine or bare metal.

[0170] The characteristics of Level 1 scheduling are cross-regional, long-distance, and multi-node operations, which can meet the "separation of storage and computation" requirements of large models. Level 2 scheduling is relatively shorter and more refined. This level of scheduling belongs to the traditional cloud service side and has a relatively complete research foundation and technical system.

[0171] Task cost:

[0172] The example and algorithm used are the same as those in (1) to (2) above. There are k clouds in the network, C = {C1, C2, C3 ... C2}. k} represents the collection of computing cloud nodes; other details will not be elaborated here.

[0173] Network links:

[0174] Because the state of network links is constantly changing due to various factors such as environmental changes, data transmission load, and sampling timing, this embodiment employs a gated recurrent unit (GRU) to train and predict the state of network links in order to accurately capture the current state of the links and predict their state during future task distribution. The input-output structure of the GRU is the same as that of a regular RNN (Recurrent Neural Network). There is a current input X. tAnd the hidden state passed down from the previous node. t-1 This hidden state contains information about previous nodes. Combined with X t and h t-1 GRU will obtain the output y of the currently hidden node. t and the hidden state h passed to the next node t .

[0175] like Figure 4 As shown, firstly, based on the state h transmitted from the previous transmission... t-1 and the input x of the current node t To obtain two gate states, r is the gate that controls the reset (reset gate) and z is the gate that controls the update (update gate).

[0176] After receiving the gating signal, use the reset gating to obtain the "reset" data h. t-1' =h t-1 ×r, where × represents the element-wise multiplication of the operation matrices, requiring that the two multiplied matrices be of the same type. Then h t-1' With input x t The data is concatenated and then scaled to the range of -1 to 1 using a tanh activation function. Here, h' primarily contains the current input x. t Data. Therefore, selectively adding h′ to the current hidden state is equivalent to “memorizing the state at the current moment”.

[0177] in This represents matrix addition. During the "update memory" phase, both forgetting and remembering steps occur simultaneously. Furthermore, the previously obtained update gate z is used, with the expression: h. t =(1-z)h t-1 +zh'. The range of the gating signal z is 0-1. The closer the gating signal is to 1, the more data is remembered, while the closer it is to 0, the more data is forgotten.

[0178] Finally, we obtain the contents of the aforementioned algorithms (3) to (5). These will not be elaborated upon further here.

[0179] Total cost of the computation task:

[0180] If subtask i can only choose one of many cloud computing centers j and selects one network link s for computation offloading, then the total cost incurred by the user to complete a certain task is:

[0181]

[0182] Considering the resource scheduling problem of minimizing combined energy consumption and price costs while improving equipment redundancy and network service quality, and with the objective of minimizing system cost, the problem is modeled as follows:

[0183] min COST a ll(i,j,s)

[0184] stC1: Set B = {B1, B2, ..., B} k In}, there is one and only one B k =1, the rest are 0.

[0185] C2: α1+α2+α3+α4=1

[0186] The optimization objective is to minimize the weighted system cost of latency, energy consumption, and price for users processing computational tasks. Constraint C1 restricts the task scheduling location, and constraint C2 limits the magnitude of the weighting factors.

[0187] Since the various factors and indicators in the integrated scheduling of computing networks often have different dimensions and units, this can affect the results of data analysis. To eliminate the influence of dimensions between indicators, data normalization and standardization are required to compare different data indicators. After data normalization / standardization, the indicators are all on the same order of magnitude, enabling comprehensive comparison and evaluation. Therefore, MIN-MAX normalization is used in the data preprocessing stage to scale all indicator factor values ​​to between 0 and 1. The MIN-MAX standardization formula is:

[0188] x'=(x-min) / (max-min) (6)

[0189] Where x represents the original data, x' represents the standardized data, min represents the minimum value in the dataset, and max represents the maximum value in the dataset.

[0190] The converged computing and network scheduling strategy includes four categories: device redundancy priority-aware scheduling strategy, energy consumption priority-aware scheduling strategy, price priority-aware scheduling strategy, and network service quality priority-aware scheduling strategy. The differences in the scheduling strategy model are reflected in the different values ​​of the indicator factors α1, α2, α3, and α4. The weights of these indicator factors affect the final result of the algorithm's execution decision, and the weight values ​​of these indicator factors mainly depend on the user's ratings for the four strategies. Before scheduling, the user needs to submit ratings for the four strategies: Judge = [α1, α2, α3, α4].

[0191] After n scheduling iterations, a user historical scheduling score dataset JUDGE will be generated:

[0192]

[0193] Next, the entropy weight method is used to obtain the historical data weight scores. First, standardization is performed according to the following formula to obtain the standardized matrix (i represents the weight type, j represents the index):

[0194]

[0195] The deviation matrix P is constructed according to the following formula.

[0196]

[0197] Calculate the standard information entropy using the following formula:

[0198]

[0199] Then calculate the information utility value using the following formula:

[0200] d j =1-e j (14)

[0201] Finally, by normalizing the information utility value, the entropy weight of each indicator can be obtained.

[0202]

[0203] The final weighted combination is W = {W1, W2, W3, W4}. This is then combined with the user's current interest score regarding the scheduling strategy, judge = [α]. t1 ,α t2 ,α t3 ,α t4 Finally, the α value is incorporated into the scheduling algorithm. final,i =0.1×W i +0.9×α ti .

[0204] The training steps for computer-network converged scheduling, and the overall training process are as follows: Figure 5 As shown:

[0205] A. Step 1: Establish a Markov decision process model

[0206] (1) Calculate the state space S t (Corresponding to State in the diagram)

[0207] Computational task resource scheduling first goes through a pre-matching phase. The system state at the current time t consists of two parts: the computing power status of all cloud devices adapted to the computational task. Network conditions Then the state space

[0208] (2) Calculate the action space A (corresponding to Action in the diagram).

[0209] At the current time t, the total set of scheduling decisions for user computation task A can be represented as:

[0210] a t =[a1,a2,……a i (16)

[0211] Among them, a i ={C j ,R s} indicates that the service request will be scheduled to cloud node C. j Process and select the link number R between the user and the cloud center. s Perform Level 1 scheduling.

[0212] (3) Calculate the reward function R (corresponding to Reward in the figure).

[0213] At time t, the environmental state S t The action taken a t Denote as a set of state-action pairs (s t ,a t Each (s) t ,a t Each has a reward value, represented by R(s). t ,a t The optimization objective is to minimize system cost by reducing execution energy consumption and price, and improving device redundancy and network service quality. The reward function is set as the reciprocal of the system cost, i.e.:

[0214]

[0215] Among them, COST(s) t ,a t The reward function R(s) represents the system cost in the current state. t ,a t The larger the value of , the lower the current system cost. The long-term reward value throughout the process can be expressed as:

[0216]

[0217] Where 0≤γ≤1 represents the discount factor.

[0218] (4) Calculate the state space S t+1 (Corresponding to the Next-state in the diagram)

[0219] By scheduling computational tasks to a specific cloud center and network link combination using a scheduling algorithm, and then removing the remaining resources after removing the cloud center and network link combination, the state space S is obtained. t+1 .

[0220] Step 2: Adapt tasks to the task pool and cloud network resource pool.

[0221] Step 3: The matching groups are fed into the policy network, and the matching group probability is selected.

[0222] Step 4: Sampling to generate decision groups (corresponding to...) Figure 5 In this context, Expect-Action refers to predicting actions.

[0223] Step 5: The fitting group and the decision-making group are fed into the value network for decision scoring (corresponding to...). Figure 5 (Value). Also, it can be like... Figure 5 As shown, the adaptation group is input into the value network for training. Figure 5 In this implementation, Policy-loss represents the policy loss corresponding to the policy network, and Value-loss represents the value loss corresponding to the value network. The parameters of the value network and policy network are updated using these two losses. Additionally, this embodiment introduces a target network mechanism, which also softly updates the parameters of the target value network and the target policy network.

[0224] Step 6: Calculate and build the experience pool based on the results.

[0225] Step 7: Update the results based on the experience pool.

[0226] Scheduling algorithm design:

[0227] The scheduling algorithm designed in this embodiment follows the idea of ​​the DDPG (Deep Deterministic Policy Gradient) algorithm, adopting an Actor-Critic architecture, introducing an experience replay mechanism, and a target network mechanism. The Actor-Critic architecture refers to the policy network making decisions based on the state, and the value network evaluating these decisions, thereby optimizing the policy network to produce better decisions. The experience replay mechanism allows for the reuse of historical samples, avoiding over-reliance on recently acquired data. The target network mechanism optimization refers to using a dual neural network approach, i.e., setting up identical value network and target value network, and policy network and target policy network structures. This solves the feedback loop problem during training, improving the algorithm's stability and training efficiency. (For example...) Figure 5The target policy network generates the Next-Action (the next decision group), inputs the target value network (Target-Value, Expect-Value, predicted value, and loss determination), and stores the Next-Reward, Next-Grade, and Next-Next-State in the experience pool.

[0228] In the decision-making phase, this embodiment first adapts the task and cloud network resources, then sends all adapted groups into the policy network to output the probability of selecting each adapted group, and samples according to the data distribution to generate decision groups. Next, all adapted groups and the selected decision groups are fed into the value network as input, outputting a decision score. Based on the output decision group and scheduling policy, its reward (instant reward) is calculated, and then a grade is assigned according to the reward, ultimately generating a State-Action-Reward-Next_state-grade combination which is stored in the storage pool. In the network parameter update phase, storage content is first selected through both priority sampling and random sampling, and then, combined with the target network, a two-parameter soft update method is used to update the network parameters.

[0229] The value network structure mentioned above is as follows: Figure 6 The policy network structure shown is as follows Figure 7 As shown, both networks have basic structures containing several 256-dimensional hidden layers. The policy network has an input layer with a state dimension of 9 and an output layer with an action dimension of 1. The input layer dimension refers to the dimension of a single adaptation group generated during the resource adaptation phase. This dimension includes the sum of the dimensions of cloud, network resources, and task resources. The output layer represents the probability of selecting this group for scheduling, and is 1-dimensional. It also contains multiple hidden layers (hidden layers 1 to 3 in the figure) and activation functions ReLU and tanh. The value network's input includes the dimension of the adaptation group plus a dimension of 9+1 (adapter group number), and its output represents the value score, with a dimension of 1. Besides the differences in input / output content and dimensions, the policy network adds a tanh operation before the output.

[0230] The scheduling method in this embodiment evaluates the importance of empirical data, such as... Figure 5As shown, the network parameters are incorporated into an experience pool and optimized based on priority order and a random sampling and replay mechanism. During parameter updates, the policy network and value network adopt differentiated update rates. Allowing the target value network to update faster helps reduce overfitting and facilitates faster learning of the policy network, as it continuously learns from the latest data. More frequent updates to the value network may help stabilize the learning process, adapt more quickly to policy changes, and provide more timely value estimates. Furthermore, the algorithm innovatively adjusts the output of the policy network, not directly determining actions but outputting the probability distribution of action values ​​for sampling. This increases the exploratory nature of decision-making and helps the algorithm seek the globally optimal policy.

[0231] The construction of the computing-network converged resource scheduling platform can be based on the two-level converged scheduling model constructed in this embodiment. When sensing or cloud management and network management report computing and network indicators, the unified evaluation and measurement system for computing-network resources established in this embodiment can be used to uniformly abstract and describe parameters of different dimensions such as multi-dimensional cloud resources and networks, thereby comprehensively evaluating the real-time status of distributed resources in the computing power network to serve computing-network scheduling decisions.

[0232] Based on the GRU network awareness method proposed in this embodiment, information prediction and network pre-inspection functions can be achieved according to historical line conditions. Combined with the scheduling algorithm and scheduling strategy proposed in this embodiment, personalized network scheduling services can be provided.

[0233] Example 3

[0234] like Figure 8 As shown in the embodiment of this application, the task scheduling device 400 is connected to the computing power network fusion platform, which is connected to the cloud management platform and the network management platform respectively. The computing power network fusion platform schedules cloud center resources through the cloud management platform and the network management platform. The task scheduling device 400 is located on the computing power network fusion platform and may include:

[0235] The acquisition module 401 is used to acquire the target computing task of the target user, the target scheduling preference information of the target user, the current link network status between the target user and each cloud center, and the current computing capacity data of each cloud center.

[0236] The matching module 402 is used to match the target computing task with the current computing capacity data and current link network status of each cloud center to generate a corresponding scheduling adaptation group.

[0237] The generation module 403 is used to input all scheduling adaptation groups and target scheduling preference information into the converged target policy network to generate corresponding scheduling decision groups. The scheduling decision groups include the target cloud center and the network links between the target user and the target cloud center.

[0238] The computing module 404 is used to compute target computing tasks according to the resource computing of the corresponding target cloud center as scheduled by the scheduling decision group.

[0239] In some implementations, when acquiring the current link network status between the target user and each cloud center, the acquisition module 401 is specifically used for:

[0240] Obtain the network states of three links between the target user and each cloud center up to the current time. The first link network state corresponds to the time one moment before the current time, the second to the time two moments before the current time, and the third to the time three moments before the current time. Each link network state includes transmission latency, bandwidth utilization, and packet loss rate. Input the first, second, and third link network states into a convergent gated recurrent unit to generate the current link network states between the target user and each cloud center.

[0241] In some implementations, the target computing task includes: the number of processor clock cycles required by the task, memory size, disk size, and device time. Current computing capability data includes: processor storage capacity, memory capacity, disk storage capacity, computing power metrics, price per processor, price per memory unit, and price per disk unit. Multiple scheduling adaptation groups are available.

[0242] Matching module 402 is specifically used for:

[0243] Based on the target computing task and the current computing capacity data, all candidate cloud centers capable of handling the target computing task are identified from each cloud center. Each candidate cloud center and its corresponding network link are combined to generate multiple scheduling adaptation groups.

[0244] In some embodiments, the task scheduling device 400 further includes:

[0245] The training module defines the quintuples of a Markov decision process model based on historical data to establish the model. Historical data includes the historical computing task pool for the target user, historical scheduling preference information for the target user, historical link network status between the target user and each cloud center, and historical computing capacity data pools for each cloud center. Based on an actor-critic architecture, the pre-defined policy network and pre-defined value network are iteratively trained according to the training adaptation group, generating a training decision group and a decision score for each iteration. The training adaptation group corresponds to the current state space in the quintuple. The training decision group corresponds to the action space in the quintuple. The immediate reward and rating score are determined based on the historical scheduling preference information and decision score corresponding to the training adaptation group. The rating score is the score corresponding to the dominance function. An experience pool is built based on the training adaptation group, training decision group, immediate reward, and rating score. The pre-defined policy network and pre-defined value network are updated based on the experience pool, and iterative training continues on the updated network until convergence is met.

[0246] In some implementations, historical scheduling preference information includes four weights corresponding to execution energy consumption, equipment redundancy, price, and network service quality, respectively.

[0247] The training module is specifically used to define the quintuples of a Markov decision process model based on historical data:

[0248] The current state space is defined by the historical computing power data at the target time and the corresponding historical link network state. The historical computing power data is the computing power data from the historical computing power data pool that matches the historical computing tasks. The historical computing tasks are the computing tasks at the target time from the historical computing task pool. The training decision group corresponding to the historical computing tasks is defined as the action space. The training decision group includes the scheduling cloud center corresponding to the historical computing tasks and the network link between the target user and the scheduling cloud center. Based on the historical link network state at the target time, the four weights, and the historical computing power data corresponding to the scheduling cloud center, the total task consumption cost corresponding to the historical computing tasks is determined. The reciprocal of the total task consumption cost is determined as the immediate reward. A discount factor is defined, and the next state space is determined based on the current state space and the training decision group.

[0249] In some implementations, when the training module iteratively trains the preset policy network and the preset value network according to the training adaptation group, and generates the training decision group and the decision score corresponding to each iteration of training, it is specifically used for:

[0250] Historical computing tasks are matched with historical computing capacity data and corresponding historical link network states in the historical computing capacity data pool to generate corresponding training adaptation groups. All training adaptation groups and their corresponding historical scheduling preference information are input into a preset policy network to generate the probability of each training adaptation group. The probabilities of each training adaptation group are sampled to generate corresponding training decision groups. The training decision groups and training adaptation groups are input into a preset value network to generate decision scores corresponding to the training decision groups.

[0251] In some implementations, when the training module updates the preset policy network and the preset value network based on the experience pool, it specifically performs the following:

[0252] Data in the experience pool is sampled using both priority sampling and random sampling methods. The preset policy network and preset value network are then updated based on the obtained data.

[0253] The task scheduling device provided in this application has the beneficial effects and implementation methods of the task scheduling methods provided in Embodiments 1 and 2 of this application. For details, please refer to the specific descriptions of the task scheduling methods in Embodiments 1 and 2 above. This embodiment will not repeat them here.

[0254] Example 4

[0255] This application embodiment also provides a computing power network fusion platform, which includes:

[0256] Memory and processor.

[0257] The memory stores the instructions that the computer executes.

[0258] The processor executes computer execution instructions stored in memory to implement the task scheduling methods as described in Embodiments 1 and 2.

[0259] The computing power network fusion platform provided in this application has the beneficial effects and implementation methods of the task scheduling methods in Embodiments 1 and 2 of this application. For details, please refer to the specific descriptions of the task scheduling methods in Embodiments 1 and 2 above. This embodiment will not repeat them here.

[0260] Example 5:

[0261] This embodiment also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the task scheduling method in Embodiment 1 or Embodiment 2 above.

[0262] The computer-readable storage medium provided in this application has the beneficial effects and implementation methods of the task scheduling methods in Embodiments 1 and 2 of this application. For details, please refer to the specific descriptions of the task scheduling methods in Embodiments 1 and 2 above. This embodiment will not repeat them here.

[0263] Other embodiments of the present application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the embodiments of this application that follow the general principles of the embodiments of this application and include common knowledge or customary techniques in the art not disclosed in the embodiments of this application.

[0264] It is understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of this application, and this application is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and substance of this application, and these modifications and improvements are also considered to be within the scope of protection of this application.

Claims

1. A task scheduling method, characterized in that, The computing power network fusion platform is connected to both a cloud management platform and a network management platform. The computing power network fusion platform schedules cloud center resources through the cloud management platform and the network management platform. The method is applied to the computing power network fusion platform, and the method includes: Acquire the target computing task of the target user, the target scheduling preference information of the target user, the current link network status between the target user and each cloud center, and the current computing capacity data corresponding to each cloud center; The target computing task is matched with the current computing capacity data of each cloud center and the current link network status to generate a corresponding scheduling adaptation group; All the aforementioned scheduling adaptation groups and the target scheduling preference information are input into a converged target policy network to generate a corresponding scheduling decision group; the scheduling decision group includes the target cloud center and the network link between the target user and the target cloud center. The target computing task is calculated by scheduling the resources of the corresponding target cloud center according to the scheduling decision group.

2. The method according to claim 1, characterized in that, Obtain the current network status of the link between the target user and each cloud center, including: Obtain the network status of three links between the target user and each cloud center up to the current time. The first link network status corresponds to the time one moment before the current time, the second link network status corresponds to the time two moments before the current time, and the third link network status corresponds to the time three moments before the current time. Each link network status includes transmission latency, bandwidth utilization, and packet loss rate. The first link network state, the second link network state, and the third link network state are input into a convergent gated recurrent unit to generate the current link network state between the target user and each cloud center.

3. The method according to claim 1, characterized in that, The target computing task includes: the number of processor clock cycles required for the task, memory size, disk size, and device time; the current computing capability data includes: processor storage capacity, memory capacity, disk storage capacity, computing power index, price per processor, price per memory, and price per disk; the scheduling adaptation group can be multiple; The step of matching the target computing task with the current computing capacity data of each cloud center and the current link network status to generate a corresponding scheduling adaptation group includes: Based on the target computing task and the current computing capacity data, all candidate cloud centers capable of processing the target computing task are determined from each cloud center. Each candidate cloud center and its corresponding network link are combined to generate multiple scheduling adaptation groups.

4. The method according to any one of claims 1 to 3, characterized in that, Before inputting all the scheduling adaptation groups and the target scheduling preference information into a converged target policy network to generate the corresponding scheduling decision group, the method further includes: The five-tuple of the Markov decision process model is defined based on historical data to establish the Markov decision process model; wherein, the historical data includes the historical computing task pool corresponding to the target user, the historical scheduling preference information corresponding to the target user, the historical link network status between the target user and each cloud center, and the historical computing capacity data pool corresponding to each cloud center. Based on the Actor-Critic architecture, the preset policy network and preset value network are iteratively trained according to the training adaptation group, generating the training decision group and the decision score corresponding to the training decision group for each iteration; the training adaptation group corresponds to the current state space in the five-tuple; the training decision group corresponds to the action space in the five-tuple. The immediate reward and rating score are determined based on the historical scheduling preference information corresponding to the training adaptation group and the decision score; the rating score is the score corresponding to the dominance function. An experience pool is built based on the training adaptation group, the training decision group, the instant reward, and the rating score; The preset policy network and preset value network are updated according to the experience pool, and the updated preset policy network and preset value network are iteratively trained until the convergence condition is met.

5. The method according to claim 4, characterized in that, The historical scheduling preference information includes four weights corresponding to execution energy consumption, equipment redundancy, price, and network service quality, respectively. The five-tuples of the Markov decision process model defined based on historical data include: The historical computing power data at the target time and the historical link network status corresponding to the target time are used as the current state space; the historical computing power data is the computing power data in the historical computing power data pool that matches the historical computing tasks; the historical computing tasks are the computing tasks at the target time in the historical computing task pool. The training decision group corresponding to the historical computing task is used as the action space; the training decision group includes the scheduling cloud center that calculates the historical computing task and the network link between the target user and the scheduling cloud center. The total cost of the historical computing task is determined based on the historical link network status corresponding to the target time, the four weights, and the historical computing capacity data corresponding to the scheduling cloud center. The reciprocal of the total cost of the task is determined as the immediate reward; Define a discount factor and determine the next state space based on the current state space and the training decision group.

6. The method according to claim 5, characterized in that, The step of iteratively training the preset policy network and the preset value network according to the training adaptation group, and generating the training decision group and the decision score corresponding to the training decision group for each iteration, includes: The historical computing tasks are matched with the historical computing capability data in the historical computing capability data pool and the corresponding historical link network status to generate corresponding training adaptation groups. All the training adaptation groups and the historical scheduling preference information corresponding to the training adaptation groups are input into a preset policy network to generate the probability of each training adaptation group. The probabilities of each training adaptation group are sampled to generate corresponding training decision groups; The training decision group and the training adaptation group are input into a preset value network to generate a decision score corresponding to the training decision group.

7. The method according to claim 4, characterized in that, The step of updating the preset policy network and preset value network according to the experience pool includes: The data in the experience pool are sampled using both priority sampling and random sampling methods; The preset strategy network and preset value network are updated based on the obtained data.

8. A task scheduling device, characterized in that, The computing power network convergence platform is connected to both a cloud management platform and a network management platform. The computing power network convergence platform schedules cloud center resources through the cloud management platform and the network management platform. The device is located on the computing power network convergence platform and includes: The acquisition module is used to acquire the target computing task of the target user, the target scheduling preference information of the target user, the current link network status between the target user and each cloud center, and the current computing capacity data corresponding to each cloud center. The matching module is used to match the target computing task with the current computing capacity data and the current link network status of each cloud center to generate a corresponding scheduling adaptation group. The generation module is used to input all the scheduling adaptation groups and the target scheduling preference information into the convergent target policy network to generate a corresponding scheduling decision group; the scheduling decision group includes the target cloud center and the network link between the target user and the target cloud center. The computing module is used to compute the target computing task according to the resources of the corresponding target cloud center scheduled by the scheduling decision group.

9. A computing power network fusion platform, characterized in that, include: Memory and processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the task scheduling method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the task scheduling method as described in any one of claims 1 to 7.