Computing power network resource scheduling method and device, storage medium and electronic equipment

By obtaining task delay and computing power costs in the computing power network, building optimization goals, and using the MADDPG algorithm for cross-domain resource scheduling, the problems of low resource scheduling efficiency and insufficient resource utilization in the computing power network are solved, and efficient and low-latency resource allocation is achieved.

CN120066704APending Publication Date: 2025-05-30STATE GRID BEIJING ELECTRIC POWER CO +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510018031.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Cross-domain resource scheduling in computing power networks has problems such as high latency, low resource utilization and insufficient scheduling efficiency, which is difficult to meet the needs of business applications for efficient, low latency and resource intensiveness.

Method used

By obtaining the total task delay of multiple computing tasks and the total computing cost of computing power service nodes in multiple domains, we build optimization goals, and use the multi-agent deep deterministic strategy gradient algorithm (MADDPG) to optimize cross-domain resource scheduling, dynamic allocation and scheduling resources.

Benefits of technology

It has achieved the improvement of computing resource scheduling efficiency and resource utilization, reduced the execution delay and computing cost of tasks, and met the needs of business applications for efficient, low latency and resource intensiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066704A_ABST
    Figure CN120066704A_ABST
Patent Text Reader

Abstract

The invention discloses a computing power network resource scheduling method and device, a storage medium and electronic equipment. The method comprises the following steps: acquiring the total task time delay of a plurality of calculation tasks; the total computing cost of computing power service nodes included in multiple domains is obtained, the multiple domains are obtained by dividing the computing power network, and the computing power service nodes are used for indicating devices providing computing resources; constructing an optimization target based on the total calculation cost of the computing power service nodes included in the plurality of domains and the total task delay of the plurality of calculation tasks; and based on the optimization target, performing cross-domain resource scheduling optimization on the plurality of computing tasks to obtain a resource scheduling result, the resource scheduling result being at least used for indicating computing resources respectively allocated to the plurality of computing tasks. The technical problem of low computing power resource utilization rate caused by low scheduling accuracy of a computing power resource scheduling method in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart grids, and in particular, to a computing power network resource scheduling method, device, storage medium, and electronic device. Background Art

[0002] In the context of the rapid development of cloud computing, edge computing, and Internet of Things technologies, the Computing Power Network (CPN), as a new generation of information infrastructure that integrates cloud, network, and edge computing resources, is facing huge challenges in resource scheduling complexity and dynamics. Especially in cross-domain resource scheduling, resource scheduling methods in related technologies, such as scheduling schemes based on historical records and centralized scheduling strategies, are difficult to cope with the sudden load changes and heterogeneous resource environments frequently occurring in the computing power network. These problems are mainly reflected in: lag and efficiency issues: The scheduling mechanism relying on historical records has obvious lag when dealing with real-time changing loads and is difficult to respond quickly. In large networks, the high computational complexity of centralized scheduling affects the system's response speed and real-time performance. Insufficient resource utilization: The resource distribution in each domain of the computing power network is uneven, and the resource types are diverse, including Central Processing Unit (CPU), Graphics Processing Unit (GPU), Field-Programmable Gate Array (FPGA), etc. The scheduling methods in related technologies fail to effectively utilize these heterogeneous resources, resulting in low utilization rate of computing power resources and a large number of idle resources. High latency and energy consumption: Data transmission and resource allocation in the computing power network are often accompanied by high latency and energy consumption. Especially in cross-domain scheduling, it is necessary to consider the channel state of network connections between different domains, making latency and energy consumption key issues.

[0003] In summary, there are problems such as high latency, low resource utilization rate, and insufficient scheduling efficiency in cross-domain resource scheduling of the computing power network in related technologies, which are difficult to meet the requirements of current business applications for high efficiency, low latency, and resource intensification.

[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present invention provide a computing power network resource scheduling method, device, storage medium, and electronic device to at least solve the technical problem of low scheduling accuracy in the computing power resource scheduling method in related technologies, which in turn leads to low utilization rate of computing power resources.

[0006] According to one aspect of an embodiment of the present invention, a computing power network resource scheduling method is provided, including: obtaining the total task latency of multiple computing tasks; obtaining the total computing cost of computing power service nodes included in multiple domains, where the multiple domains are obtained by partitioning the computing power network, and the computing power service nodes are used to indicate devices providing computing resources; constructing an optimization objective based on the total computing cost of the computing power service nodes included in the multiple domains and the total task latency of the multiple computing tasks; and performing cross-domain resource scheduling optimization on the multiple computing tasks based on the optimization objective to obtain a resource scheduling result, where the resource scheduling result is at least used to indicate the computing resources respectively allocated to the multiple computing tasks.

[0007] According to another aspect of an embodiment of the present invention, a computing power network resource scheduling device is further provided, including: a first obtaining module, configured to obtain the total task latency of multiple computing tasks; a second obtaining module, configured to obtain the total computing cost of computing power service nodes included in multiple domains, where the multiple domains are obtained by partitioning the computing power network, and the computing power service nodes are used to indicate devices providing computing resources; a target constructing module, configured to construct an optimization objective based on the total computing cost of the computing power service nodes included in the multiple domains and the total task latency of the multiple computing tasks; and a scheduling optimization module, configured to perform cross-domain resource scheduling optimization on the multiple computing tasks based on the optimization objective to obtain a resource scheduling result, where the resource scheduling result is at least used to indicate the computing resources respectively allocated to the multiple computing tasks.

[0008] According to another aspect of an embodiment of the present invention, a non-volatile storage medium is further provided. The non-volatile storage medium stores multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform any one of the computing power network resource scheduling methods.

[0009] According to another aspect of an embodiment of the present invention, an electronic device is further provided, including one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement any one of the computing power network resource scheduling methods.

[0010] In an embodiment of the present invention, the total task latency of multiple computing tasks is obtained; the total computing cost of computing power service nodes included in multiple domains is obtained, where the multiple domains are obtained by partitioning the computing power network, and the computing power service nodes are used to indicate devices providing computing resources; an optimization objective is constructed based on the total computing cost of the computing power service nodes included in the multiple domains and the total task latency of the multiple computing tasks; based on the optimization objective, cross-domain resource scheduling optimization is performed on the multiple computing tasks to obtain a resource scheduling result, where the resource scheduling result is at least used to indicate the computing resources respectively allocated to the multiple computing tasks, achieving the purpose of realizing cross-domain resource scheduling optimization of computing tasks by comprehensively considering task latency and computing cost and utilizing the multi-domain structure of the computing power network, thereby achieving the technical effect of improving the efficiency of computing power resource scheduling and resource utilization rate, and further solving the technical problem that the computing power resource scheduling method in the related art has low scheduling accuracy, resulting in low computing power resource utilization rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and the illustrative embodiments and descriptions of the present invention are used to explain the present invention, and do not constitute an improper limitation of the present invention. In the drawings:

[0012] Figure 1 is a schematic diagram of an optional computing power network architecture for resource scheduling according to an embodiment of the present invention;

[0013] Figure 2 is a flowchart of a computing power network resource scheduling method according to an embodiment of the present invention;

[0014] Figure 3 is a schematic diagram of an optional algorithm architecture based on similarity-assisted experience replay according to an embodiment of the present invention;

[0015] Figure 4 is a schematic diagram of an optional CPN controller module structure according to an embodiment of the present invention;

[0016] Figure 5 is a flowchart of an optional resource perception module according to an embodiment of the present invention;

[0017] Figure 6 is a schematic diagram of an optional resource scheduling according to an embodiment of the present invention;

[0018] Figure 7 is a schematic diagram of an optional comparison result according to an embodiment of the present invention;

[0019] Figure 8 is a schematic diagram of a computing power network resource scheduling device according to an embodiment of the present invention. Detailed implementation manners

[0020] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0022] First of all, for the convenience of understanding the embodiments of the present invention, some terms or nouns involved in the present invention will be explained below:

[0023] Multi-Agent Deep Deterministic Policy Gradient (MADDPG), a deep reinforcement learning algorithm for collaborative multi-agent systems. This algorithm is an extension based on the Deterministic Policy Gradient algorithm (DDPG) and can be effectively applied to scenarios where there are cooperation or competition relationships among multiple agents.

[0024] In view of the above problems, an embodiment of the present invention provides an embodiment of a method for scheduling computing power network resources. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that here.

[0025] As an optional embodiment, the computing power network resource scheduling method provided in this embodiment can be applied to a computing power network architecture for resource scheduling, Figure 1 is a schematic diagram of an optional computing power network architecture for resource scheduling according to an embodiment of the present invention, as Figure 1As shown in the figure, the computing power network includes an application service layer, a perception management layer, and an infrastructure layer, where:

[0026] The infrastructure layer is the bottom layer of the computing power network and is composed of ubiquitous heterogeneous computing resource devices and network infrastructure in the network. The computing resource devices include a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Field-Programmable Gate Array (FPGA), etc. The network infrastructure refers to the physical network that connects the actual computing resources. In the infrastructure layer, the physical network of the computing power network is divided into multiple domains according to gateways, so that each domain contains a subnet managed by the same router, and the regional computing power network provides data to the perception management layer in the form of computing power service nodes.

[0027] The perception management layer is the middle layer of the computing power network and is responsible for resource allocation, computing task scheduling and forwarding, etc. in the computing power service nodes. A CPN controller is set in each domain to be responsible for the computing power services within its own domain, exchange information with other controllers, and perform resource modeling, etc.

[0028] The application service layer is the top layer of the computing power network and faces actual business. With the development of actual business, the application service scenarios will also be diverse. These scenario requirements do not need to obtain the details of the computing network, but only need to send requests to the computing network.

[0029] Based on the above computing power network, the following network modeling is carried out:

[0030] The computing power network G is divided into multiple (such as N) regions {G 1 ,G 2 ,…,G N}, and a CPN controller C i ∈{C 1 ,C 2 ,…,C N} is deployed in each domain to realize the management of resources within the domain. The i-th domain can be modeled as a directed graph G i =(V i ,E i ). V i represents all the nodes in this domain, which are composed of two types. One type is the computing power providing nodes p i ={p i1 ,p i2 ,…}, where the number of computing power providing nodes in the region V i is |p i |, and p ij represents the j-th computing power providing node in the i-th domain. The other type is the user nodes ui = {u i1 ,u i2 ,…},V i =p i ∪u i .

[0031] In terms of computing power node modeling (i.e., computing power service node), consider that each round of scheduling tasks consists of a set of time slots T = {1, ..., t}, each time slot lasts for τ, and the computing power resources of the computing power providing node are re-perceived every τ period. The computing power providing node p at time slot t is recorded as ij The idle computing resources are They represent the idle CPU resources, idle GPU resources, idle storage capacity, CPU computing cost coefficient, GPU computing cost coefficient and storage cost coefficient of the node respectively. This embodiment considers the situation that the resource scale will not change during the task execution process, and once it is assigned a computing task, it cannot be assigned other tasks.

[0032] In terms of task modeling, consider the case where a random task arrives at CPN and arrives at region G at time slot t. i Mission Will be stored in its C i Task queue in Represents the jth computing task request in the queue. i Computing power is allocated to tasks according to the first-in-first-out principle. The specific information of the task can be further represented by an octet. They represent the arrival time of the computing task, the maximum tolerable delay, and the C i The bandwidth of the task, the size of the task input data and the size of the output result, the CPU and GPU computing required for the task, and the set of pre-tasks for the task. The team leader task Before allocating computing power, it is necessary to calculate the queuing delay based on the current time a like The task is discarded.

[0033] Use binary variables to represent C i Decisions about tasks. Representation Task Whether to perform cross-domain resource scheduling. Indicates that it needs to be performed, otherwise it is 0. Indicates area G i Mission Whether to schedule area G k Resources, Indicates that scheduling is required, otherwise it is 0. Representation Task Deployed in region G k Node p kq If it is, the value is 1; otherwise, it is 0. According to the task process, when the CPN controller returns decision information to the user, it will send configuration information to the computing power service node. The decision of the configuration information mainly includes the relevant information of the computing power tasks assigned to the corresponding nodes for execution.

[0034] Figure 2 It is a flowchart of a computing power network resource scheduling method according to an embodiment of the present invention. As Figure 2 shown, the method includes the following steps:

[0035] Step S102, obtain the total task latency of multiple computing tasks;

[0036] Optionally, the total task latency is the latency information of all the to-be-processed computing tasks collected, which may include, but is not limited to, transmission latency, transfer latency, and queuing latency, etc. In a computing power network, latency is an important indicator to measure service quality, directly affecting the user experience and the task completion efficiency.

[0037] In an optional embodiment, obtaining the total task latency of multiple computing tasks includes: obtaining the task latency of any one of the multiple computing tasks in the following way: obtaining the scheduling region selection vector and the set of prerequisite tasks of any one of the computing tasks, where the scheduling region selection vector is used to indicate whether cross-domain resource scheduling is performed for the corresponding computing task, and the set of prerequisite tasks is used to indicate the set of other tasks or services that need to be completed before executing the corresponding task; determining the target latency calculation method from multiple latency calculation methods based on the scheduling region selection vector and the set of prerequisite tasks; obtaining the task latency of any one of the computing tasks based on the target latency calculation method; using the method of obtaining the task latency of any one of the computing tasks to obtain the task latencies of multiple computing tasks, where the task latency includes task transmission latency and task transfer latency; performing a summation operation on the task latencies of multiple computing tasks to obtain the total task latency.

[0038] It can be understood that by defining a specific delay calculation method, the delay of each computing task in different scheduling scenarios can be accurately quantified. The delay calculation based on the task's scheduling area selection vector and the set of prerequisite tasks not only takes into account the requirements of the task itself but also fully considers the dependencies between tasks, thereby ensuring that the scheduling policy can optimize the delay performance of the entire network while satisfying the constraints of sequential task execution. By clarifying that the delay can be divided into task sending delay and task transmission delay, this helps to more meticulously analyze and optimize the delay issues in the network. The sending delay involves the time for the task to be sent from the user to the network, while the transmission delay involves the time for data to be transmitted in the network, especially in the case of cross-domain scheduling. By determining the target delay calculation method from multiple delay calculation methods, the most appropriate delay calculation strategy can be selected according to whether the task requires cross-domain resource scheduling and the characteristics of its set of prerequisite tasks. This flexibility helps to optimize resource scheduling in different scenarios and improve the efficiency of the computing power network. In summary, the above method not only improves the accuracy of resource scheduling in the computing power network but also enhances the flexibility and optimization ability of the strategy by precisely defining the delay calculation process, considering the characteristics of task dependencies and cross-domain scheduling, and differentiating delay types, which plays an important role in improving the response speed of the computing power network, reducing energy consumption, and optimizing service quality.

[0039] Optionally, this embodiment considers the resource decision-making process for user computing tasks, without considering the execution situation of the computing tasks, and the queuing delay has been calculated and included in the constraints before the initial decision. Therefore, only the sending delay and the transmission delay are considered here. There can be 4 types of multiple delay calculation methods, and the task delay is calculated through 4 cases, including: ① The task does not require cross-domain resource scheduling and the set of prerequisite tasks is an empty set. ② The task does not require cross-domain resource scheduling and its set of prerequisite tasks is not an empty set. ③ The task requires cross-domain resource scheduling and its set of prerequisite tasks is an empty set. ④ The task requires cross-domain resource scheduling and its set of prerequisite tasks is not an empty set.

[0040] For case ①, the task 's scheduling area selection vector if and only if i = k. At the same time, the set of prerequisite vectors At this time, the sending delay only needs to consider the task's, and the transmission delay only needs to consider C i Transmitting the task to the computing power node p i in G iq The delay is as follows:

[0041]

[0042] Among them, represents the CPN controller C of the i-th domain i and the transmission rate between the p kq nodes within the k-th domain.

[0043] For case ②, the task of scheduling area selection vector if and only if i = k. At the same time, the set of precursor vectors of At this time, the transmission delay not only needs to consider the task's, but also needs to consider the transmission delay and propagation delay of its precursor tasks also not only needs to consider C i transmitting the task to the computing node in G i The delay, but also needs to consider the transmission delay of the computing result of its precursor task to the selected node. It is stipulated that the precursor task of The corresponding cross-domain resource scheduling decision, area selection decision, and task deployment decision are respectively At this time, its delay is calculated as follows.

[0044] If the node where the precursor task is deployed and the node where k′q′ is deployed are in the same area, that is, and k = k ′ i . Then the result can be directly transmitted within the domain.

[0045]

[0046] Among them, is the CPN controller C of the i-th domain i and the transmission rate between the p kq nodes within the k-th domain, is the transmission rate between node p k′q′ and node p kq ; represents the size of the task output result of the j-th task request in area G i ; represents the bandwidth of the area where the task is located, which is used to calculate the transmission delay.

[0047] If the nodes where the precursor task is deployed and the deployed node are not in the same area, then the CPN controller needs to be used to transfer the result between domains.

[0048]

[0049] Among them, represents the prerequisite task, represents the prerequisite task Whether to schedule the resources of area G k′ (a Boolean variable). Here, 'is used to distinguish the prerequisite task from the current task. represents the prerequisite task Whether it is deployed on the node p k′ in area G k′q′ is also a Boolean variable. In the summation formula, when it is 1, that is, the node where it is deployed participates in the calculation of the formula. represents the transmission rate between the controllers C k′ and C k of the two domains.

[0050] For case ③, the task needs to perform cross-domain resource scheduling and its set of prerequisite tasks is an empty set.

[0051]

[0052] For case ④, the task needs to perform cross-domain resource scheduling and its set of prerequisite tasks is not an empty set. Similarly to ②, it is just the issue of whether to cross domains. If the prerequisite task is deployed on the node p k′q′ and the deployed node is in the same area, that is and k = k ′ . Then the result can be directly transmitted within the domain.

[0053]

[0054] If the node where the prerequisite task is deployed and the deployed node are not in the same area, then the CPN controller needs to be used to perform inter-domain result transfer.

[0055]

[0056] The total delay T of all tasks in the entire CPN within time slot t t is:

[0057]

[0058] Step S104, obtain the total computing cost of the computing power service nodes included in multiple domains, where the multiple domains are obtained by partitioning the computing power network, and the computing power service nodes are used to indicate the devices providing computing resources;

[0059] Optionally, the computing power network is physically divided into multiple domains, and each domain contains several computing power service nodes, which may include, but are not limited to, CPUs, GPUs, FPGAs, etc. The total computing cost can be obtained based on the CPU computing costs, GPU computing costs, and storage costs of all computing power service nodes in multiple domains.

[0060] Optionally, the CPU computing costs, GPU computing costs, and storage costs corresponding to the computing power service nodes included in multiple domains can be obtained first. The total CPU computing cost is obtained by summing up the CPU computing costs corresponding to the computing power service nodes included in multiple domains; the total GPU computing cost is obtained by summing up the GPU computing costs corresponding to the computing power service nodes included in multiple domains; the total storage cost is obtained by summing up the storage costs corresponding to the computing power service nodes included in multiple domains; and the total computing cost is obtained by summing up the total CPU computing cost, the total GPU computing cost, and the total storage cost.

[0061] Optionally, for the CPU computing cost. The computing cost of the CPU can mainly consider the cost generated by the time when the computing task occupies the virtual machine of the computing power service node. This part of the cost is given a fixed value by the computing power service node itself according to the actual energy consumption operating cost. For the task assuming that the allocated computing power service node is the q-th node p in region G k of kq . In this time slot, the CPU computing cost CA i of all computing tasks in G t is:

[0062]

[0063] For the GPU computing cost. For the task assuming that the allocated computing power service node is the q-th node p in region G k of kq . In this time slot, the GPU computing cost CB t of all computing tasks is:

[0064]

[0065] For the storage cost. The cost generated by the data it receives and the communication cost between the computing power service node and the user are mainly considered. In this time slot, the storage cost CC t of all computing tasks is:

[0066]

[0067] The total cost C t of the tasks in this time slot is obtained in the following way.

[0068] C t = CA t + CB t + CC t

[0069] Step S106, construct an optimization objective based on the total computing cost of the computing power service nodes included in multiple domains and the total task latency of multiple computing tasks;

[0070] Optionally, the optimization objective is to combine latency and cost to achieve a balance, ensuring that while meeting the task latency requirements, the computing cost is reduced as much as possible.

[0071] In an alternative embodiment, constructing an optimization objective based on the total computing cost of the computing power service nodes included in multiple domains and the total task latency of multiple computing tasks includes: determining a first weight value corresponding to the total computing cost and a second weight value corresponding to the total task latency; performing a weighted calculation on the total computing cost and the total task latency based on the first weight value and the second weight value to obtain a weighted calculation result; constructing an optimization objective based on the weighted calculation result, where the optimization objective is used to indicate that the weighted calculation result is minimized.

[0072] It can be understood that the first weight is used to reflect the importance of the computing cost in the optimization objective, and the second weight emphasizes the importance of the task latency in the optimization objective. By adjusting the first weight value and the second weight value, a balance between cost control and task processing efficiency can be achieved. For example, in an environment with limited resources but cost-sensitive, the first weight value can be set higher than the second weight value, prompting the scheduling strategy to prefer resources with lower costs, even if this means sacrificing some processing speed. The introduction of the weight value provides flexibility to the optimization objective, enabling the algorithm to dynamically adapt to the requirements of different scenarios. For example, the weight value can be dynamically adjusted at different times of the day to cope with different resource demands during peak and off-peak periods. During peak periods, more emphasis may be placed on reducing task latency, while during off-peak periods, more attention may be paid to cost control. The weight value allows for more refined control of the resource scheduling strategy, enabling the algorithm to be adjusted according to specific business requirements. This not only helps improve resource utilization but also ensures that the service quality meets user needs.

[0073] In this embodiment, the comprehensive optimization objective is the overall system latency and cost, but there is no strong correlation between latency and cost. Therefore, when considering the comprehensive optimization objective, a weighted calculation can be performed on latency and cost. Here, the latency weight value is μ, and the cost weight value is θ. The comprehensive optimization objective is as follows:

[0074]

[0075] where the constraint respectively represent for the tasks the computing power resources of the nodes where it is deployed cannot exceed its available resources, constraint represents cross-domain resource scheduling decision-making, region selection decision-making, and task deployment decision-making is in binary representation.

[0076] Step S108, based on the optimization objective, perform cross-domain resource scheduling optimization on multiple computing tasks to obtain a resource scheduling result, where the resource scheduling result is at least used to indicate the computing resources allocated for multiple computing tasks respectively.

[0077] Optionally, during the resource scheduling process, the computing power cost and energy consumption are considered simultaneously. Through the optimization strategy, the operating cost of the entire network can be minimized, including the operating cost of computing power nodes and the cost of data transmission between nodes. Through cross-domain resource scheduling, the computing power resources in the entire network can be utilized more effectively, including CPU, GPU, and storage resources, dynamically allocate and schedule resources, avoid resource idleness and waste, and improve the overall resource utilization rate.

[0078] In an optional embodiment, based on the optimization objective, perform cross-domain resource scheduling optimization on multiple computing tasks to obtain a resource scheduling result, including: based on the optimization objective, use the communication common network controllers corresponding to multiple domains as the corresponding multiple agents, and adopt the multi-agent deep deterministic policy gradient algorithm to perform cross-domain resource scheduling optimization on multiple computing tasks to obtain a resource scheduling result.

[0079] It should be noted that reinforcement learning enables the agent to optimize its decision-making strategy by continuously exploring and learning from feedback through the interaction between the agent and the environment. Most of the reinforcement learning methods in the related technologies focus on the interaction between a single agent and the environment. However, cross-domain resource scheduling involves resource allocation and cooperation between multiple domains. A single agent cannot comprehensively perceive and manage all the dynamic changes and requirements inside and outside all domains, which makes simple strategies difficult to adapt to the complexity and uncertainty in cross-domain resource scheduling. Multi-agent reinforcement learning (MARL) allows multiple agents to cooperate and compete in different domains. The agents can make local decisions in different regions based on their respective observations and behaviors, and at the same time share global information to achieve the optimal scheduling of cross-domain resources. This embodiment adopts the architecture of multi-agent deep deterministic policy gradient algorithm (Multi-Agent Deep Deterministic Policy Gradient, MADDPG) with centralized training and decentralized execution, allowing each agent to share global information during the training phase, thereby improving the accuracy of computing power resource scheduling and further improving the utilization rate of computing power network resources.

[0080] In an alternative embodiment, when the multi-agent deep deterministic policy gradient algorithm is constructed based on an action generation network and a value evaluation network, based on the optimization objective, using the communication common network controllers corresponding to multiple domains as the corresponding multiple agents, the multi-agent deep deterministic policy gradient algorithm is used to optimize the cross-domain resource scheduling of multiple computing tasks to obtain a resource scheduling result, including: constructing a reward function based on the optimization objective; for each agent among the multiple agents, obtaining the current state within the corresponding domain through each agent, where the current state includes computing power resource information and task information within the corresponding domain; based on the current state, through each agent, selecting an action and executing the action using the corresponding action generation network to obtain an execution result; calculating a reward value using the reward function based on the execution result; evaluating the value of the action through each agent using the corresponding value evaluation network to obtain an evaluation result; updating the action generation network and the value evaluation network corresponding to each agent based on the reward value and the evaluation result; repeating the above operations until a predetermined termination condition is reached; determining the resource scheduling result based on the action output by the value evaluation network at the predetermined termination condition.

[0081] Optionally, the optimization goal is achieved by reducing task execution latency and computational cost. The reward function, as the cornerstone of reinforcement learning, is used to guide the agent (i.e., the CPN controller) to learn the optimal policy. Each agent (CPN controller) needs to obtain the current state from its corresponding domain, including computing power resource information (such as idle resources and cost coefficients of CPUs and GPUs) and task information (such as task arrival time, required resource types and quantities, etc.). This is the basis for the agent's decision-making, ensuring that the agent can respond based on real-time resource and task conditions. Based on the obtained current state, the agent selects an action (i.e., a decision to allocate computing power resources for a certain task) through its dedicated action generation network and executes it. After executing the action, the system will feedback the execution result. For example, the execution result may include information such as whether the task is successfully completed, the actual latency, and the actual cost. According to the execution result, the agent calculates the reward value. The reward value is the key for the agent to learn and adjust the policy, and it reflects the effect of the action in a specific environment. The higher the reward value, the better the effect of reducing latency and cost brought by the action. The agent uses its value evaluation network to evaluate the value of the selected action. This step helps the agent understand the impact of this action on the future reward expectation in the current state. The evaluation result reflects the long-term benefit of the action, rather than just the immediate reward. The agent updates its action generation network and value evaluation network according to the reward value and the evaluation result. This is the core of the learning process. By adjusting the network parameters, the agent can more accurately predict and evaluate the potential effects of actions in future decisions, thus continuously optimizing its policy. The agent repeats the above steps until a predetermined termination condition is met, which may include reaching a certain number of iterations, the performance metric reaching a preset threshold, or other termination criteria. This iterative process ensures that the agent can gradually approach the optimal policy through continuous exploration and learning. Finally, when the termination condition is reached, the value evaluation network of the agent will output an optimal action, and this action is the decision result of resource scheduling. This result will be used to guide the actual allocation of computing power resources to minimize the task execution latency and computational cost, thereby improving the efficiency and performance of the entire computing power network. In the above way, it is ensured that the agent can efficiently learn and optimize the resource scheduling policy based on the real-time in-domain state and task requirements through the MADDPG algorithm, and finally achieve the goal of reducing latency and cost, and improving the resource utilization efficiency and overall performance of the computing power network.

[0082] Optionally, the computing power resource information may include, but is not limited to, the number of computing tasks in the corresponding domain, the number of computing power providing nodes, the idle computing power resources of the computing power providing nodes, and the task information may include, but is not limited to, the arrival time of the tasks in the corresponding domain, the maximum tolerable latency, the bandwidth of the domain where the task is located, the sizes of the task input data and output results, the CPU and GPU computing amounts required by the task, and the set of prerequisite tasks of the task, etc.

[0083] Optionally, the states and actions of each agent can be represented by a state space and an action space respectively. Specifically, each CPN controller corresponds to an agent in a domain, still denoted as C i , and the problem can be represented as a Markov process. This process generally considers the state space S, action space A, transition probability P, and reward function R of the problem. Among them:

[0084] The state space (S) is used to record the various indicators of the computing power network during the process and describe the state changes throughout the process. For the i-th domain at time slot t, the selected reference indicators are:

[0085] The number of computing tasks generated in domain i at the current time slot;

[0086] |p i |: The number of computing power service nodes in domain i;

[0087] The idle computing power resources of the computing power providing node p ij at time slot t, respectively representing the idle CPU resources, idle GPU resources, idle storage capacity, CPU computing power cost coefficient, GPU computing power cost coefficient, and storage cost coefficient of the node.

[0088] The arrival time, maximum tolerable delay of task j at the current time slot, the bandwidth of C i where the task is located, the size of the task input data and the size of the output result, the CPU and GPU computing amounts required by the task, and the set of prerequisite tasks of the task.

[0089] Then the state space of the i-th domain at time slot t is:

[0090]

[0091] For the action space (A), considering making decisions for each task, we will give the decision A of allocating nodes to each task. The action space of the i-th domain at time slot t is:

[0092]

[0093] It represents the node information allocated to the j-th computing task in domain i. Here, a vector representation is adopted, that is, an N-dimensional vector. Only the unique 1 in the vector indicates being allocated to the corresponding computing power service node; the others are all 0, indicating not being allocated to the corresponding computing power service node.

[0094] In the algorithm, a method of making decisions on tasks sequentially is adopted, which greatly reduces the size of the action space and avoids the occurrence of exponential explosion.

[0095] In an alternative embodiment, the reward function includes a positive reward and a negative reward. Among them, the positive reward is generated when the corresponding computing task is completed according to the predetermined time, and the negative reward is determined based on the computing cost and task delay of the computing power service nodes included in the corresponding domain.

[0096] Optionally, the reward function is set to include two parts: a positive reward and a negative reward. The positive reward encourages the agent (i.e., the CPN controller) to complete the computing task within the predetermined time. By setting the positive reward, the agent will give priority to the strategies that can complete the task on time during the learning process, which helps to reduce the queuing time of the task and lower the overall transmission and execution delay. The negative reward is determined based on the computing cost and task delay of the computing power service nodes in the corresponding domain, which is intended to punish those decisions that result in high costs and high delays. By introducing the negative reward, those strategies that can complete the task but have too high costs or too large delays can be identified and avoided, thus guiding the agent to make more economical and time-sensitive decisions. By combining the positive reward and the negative reward, the economy of resource allocation and the dynamic changes of the environment can be considered, so as to find a balance point in multi-objective optimization. The design of this reward mechanism helps the agent to learn the optimal strategy that can not only meet the task requirements but also control the cost and delay in a complex computing power network environment.

[0097] The reward function is mainly composed of two parts: the first part is a positive reward, which is the reward obtained when the task is completed within the specified time slot, and this part of the reward is set as a constant CompletionReward; the second part is a negative reward. The longer the task execution time and the greater the computing cost, the less the reward.

[0098] The design of the reward function considers two goals, one is to reduce the execution time of all tasks, and the other is to obtain the least computing cost. The specific reward function is as follows, where T ij,t represents the time required for the j-th computing task in domain i, and Cost ij,t represents the computing cost required for the j-th computing task in domain i.

[0099]

[0100] Among them, r i t represents the reward value.

[0101] In an alternative embodiment, based on the reward value and the evaluation result, the action generation network and the value evaluation network corresponding to each agent are updated, including: constructing a first experience tuple based on the current state, the reward value, the action, and the next state, and storing the first experience tuple in a shared experience pool, where the next state is used to indicate the state of each agent after the current state; selecting a second experience tuple from the shared experience values whose similarity to the first experience tuple is less than a preset similarity threshold; and updating the action generation network and the value evaluation network corresponding to each agent based on the evaluation result and the second experience tuple.

[0102] Optionally, the constructed first experience tuple is stored in the experience pool to record each decision of the agent, including its action selection, immediate reward, and state transition information in a specific state, which provides a data basis for subsequent policy optimization. The mechanism of similarity-assisted experience replay is introduced, and its core idea is to improve the learning efficiency and enhance the generalization ability of the policy by avoiding repeated learning of similar experiences. Through similarity threshold filtering, it is ensured that the agent only replays other experience tuples whose similarity to the current experience tuple is lower than the preset threshold during training, which means that the agent can learn decisions in different scenarios more diversely, avoiding the waste of resources caused by repeated learning, and at the same time promoting policy generalization, making the decisions of the agent more reasonable in different scenarios. Updating based on the evaluation result and the second experience tuple means that when the agent uses the selected second experience tuple for learning, it will adjust its action generation network and value evaluation network according to the actual effect of its action (evaluation result) to optimize future decisions. This process not only utilizes the immediate reward but also considers the long-term reward in the selected experience tuple after similarity threshold filtering, thus improving the learning efficiency and quality. It should be noted that since in MADDPG, the mutual dependence between agents will affect the generation of the policy, the mechanism of similarity-assisted experience replay is proposed to help the agent optimize based on similar experiences and policies, so that the excellent policies in the scenario can affect the global system faster to improve the learning efficiency of the multi-agent system.

[0103] Optionally, the MADDPG algorithm is a variant of the Deep Deterministic Policy Gradient (DDPG) algorithm applied in Multi-Agent Reinforcement Learning (MARL). The algorithm is based on the dual-action generation network Actor - dual-value evaluation network Critic. Each agent implements a DDPG algorithm, which consists of four DNN networks, namely the Actor network, the target action generation Target Actor network, the Critic network, and the target value evaluation Target Critic network. The Actor network is responsible for generating deterministic actions in a given state and is responsible for policy learning. The Critic network is responsible for the value of the generated actions and guides the optimization of the Actor network. During the execution process, the agents are trained in a centralized manner and executed in a decentralized manner. That is, in the training phase, the Critic network of each agent not only considers its own actions and states but also can access the action and state information of other agents. However, in the execution phase, each agent decides its actions only based on its own state, maintaining a decentralized execution mode. For the resource scheduling task of the computing power network, similar scheduling tasks may exist in different domains. If the experience of similar tasks is repeatedly replayed, it may lead to weak generalization ability of the model.

[0104] Furthermore, manage the information interaction between agents through a shared experience pool, and propose a similarity-assisted experience replay mechanism so that each agent can replay experiences with lower similarity to itself with a higher probability during the training phase to improve the model's sensitivity to different tasks. And evaluate the scheduling situation of each domain, use the dominant strategy region to guide policy updates, and replay the experiences in the dominant strategy region more to improve the efficiency and quality of policy learning. Figure 3 It is a schematic diagram of an optional algorithm architecture based on similarity-assisted experience replay according to an embodiment of the present invention. The overall architecture of the similarity-assisted experience replay MADDPG algorithm is as Figure 3 shown.

[0105] A shared experience pool is used between agents to store the joint state and joint action information geometry of each agent at different time slots. Denote the joint experience {S t , A t , R t , S ′t} at time slot t, where represents the observation vector of all agents on the current environmental state (i.e., the current state), represents the action information of all agents, represents the reward values obtained by all agents respectively. Represents the observation vector of the agent for the next state of the environment (i.e., the next state).

[0106] Considering three elements: the similarity of task resource requirements, the similarity of state space, and the similarity of action space, to measure the similarity degree between agents, quantified using cosine similarity and accumulated and averaged by time slot. Taking agent Agent i and agent Agent j as examples, there is the following quantification method.

[0107] ① Similarity of task resource requirements: Each scheduling task may contain different resource requirements, such as CPU, bandwidth, storage, etc. Denote the total resource requirements of all scheduling tasks of Agent i within time slot t as Then the similarity of task resource requirements between Agent i and Agent j at time slot t is

[0108] ② Similarity of state space: In MADDPG, the state space of each agent contains information such as the resource state and task queue state within the current domain. Denote the state vector of Agent i within time slot t as Then the similarity of state space between Agent i and Agent j at time slot t is

[0109] ③ Similarity of action space: The actions taken by agents in different states (such as which resources to allocate, scheduling order, etc.) also reflect the similarity of tasks. Therefore, the similarity of tasks can be evaluated by comparing the actions taken by agents in similar states. Denote the action vector of Agent i within time slot t as Then the similarity of action space between Agent i and Agent j at time slot t is

[0110] Weight the cosine similarities of the three elements to obtain the similarity between agents at time slot t:

[0111]

[0112] During experience replay, the agent only needs to exclude the experience groups with higher similarity according to the calculated similarity, and selectively extract a series of experiences from the shared experience pool for replay training.

[0113] Through the above steps S102 to S108, it is possible to achieve the purpose of optimizing the cross-domain resource scheduling of computing tasks by comprehensively considering task latency and computing costs and utilizing the multi-domain structure of the computing power network, thereby achieving the technical effect of improving the scheduling efficiency and resource utilization rate of computing power resources, and further solving the technical problem of low scheduling accuracy in the computing power resource scheduling method in the related art, which in turn leads to low utilization rate of computing power resources.

[0114] It should be noted that this embodiment can effectively improve the utilization rate of computing power resources. By optimizing the allocation and scheduling path of computing power nodes, it realizes the inter-domain utilization of resources, significantly reduces the transmission latency and execution latency of tasks, and reduces the idle and waste of resources. By improving the experience replay mechanism of the MADDPG algorithm and selectively replaying the in-domain experiences with relatively low similarity to the agent, the generalization of the scheduling algorithm is improved, the reward value of the model is maximized, and more reasonable resource scheduling is achieved.

[0115] Based on the above embodiments and alternative embodiments, the present invention proposes an alternative implementation manner. Figure 4 It is a schematic structural diagram of an optional CPN controller module according to an embodiment of the present invention, as Figure 4 shown. This CPN controller can be used for any of the above computing power network resource scheduling methods.

[0116] In the computing power network, the CPN controller is responsible for computing power network control. Although centralized control can improve the control ability of the computing network, the high latency it brings and the massive concurrency requirements for the core network are difficult to meet the actual needs. Therefore, this embodiment proposes a solution to set a CPN controller in each domain to implement the architecture and algorithm of this embodiment. Reduce the latency and overall energy consumption of the computing route, so as to meet the optimization goals and service requirements.

[0117] The resource awareness module senses the resource information and network connection information of the computing power service nodes in the domain, processes and models them. These information and the requests from the application service layer together constitute the input of the resource scheduling module, that is, the environmental information of the deep reinforcement learning algorithm. The resource scheduling module can return the allocated computing power service nodes to the user and will notify the client of the allocated task information and configuration information. In addition, the CPN controllers can communicate with each other. The purpose is that when the domain where the CPN controller is located has sufficient idle resources, this domain can provide some computing power resources for the neighboring domains, thereby improving the overall performance.

[0118] For the two parts of the CPN controller, generally the C / S architecture form is adopted, that is, the computing power service node is used as the client, and the CPN controller is used as the server side. The resource awareness module and the resource scheduling module adopt a solution based on the Telemetry technology. The following is a specific introduction to the two modules.

[0119] The resource awareness module is as above Figure 4 as shown in the left half of the figure above. The resource awareness module in the CPN controller mainly consists of two parts: data storage and resource modeling, and client awareness configuration. The data storage and resource modeling module stores node resource information, models network and node resources, and uses the model as the input of the resource scheduling module.

[0120] Telemetry technology is based on another new generation network device configuration data modeling language (Yet Another Next Generation, YANG) model for data collection. Therefore, before the specific implementation of the method in this embodiment, it is necessary to define the data sources (Yang) to be collected. The information selected for collection in this embodiment includes CPU frequency, GPU frequency, CPU energy consumption, GPU energy consumption, storage cost, the number of idle CPUs, the number of idle GPUs. For network switching devices, the information selected for collection includes in / out port traffic information, routing information, etc.

[0121] Figure 5 is an optional working flowchart of the resource awareness module according to an embodiment of the present invention, as Figure 5 shown. Taking the CPN controller (server side) and the computing power service node (client side) as examples, the working process of the resource awareness module is introduced. The server side of the resource awareness module is in the CPN controller, and the client side of the awareness module is in the computing power service node. For a more detailed understanding, here the client side of the awareness module is regarded as a middleware between the server side of the resource awareness module of the CPN controller and the computing power service node. The resource awareness is completed through the interaction requests among the three.

[0122] The resource scheduling module is as Figure 4 shown on the right side of the figure above. To meet the needs of massive data access, a solution based on the UDP network layer protocol using the C / S architecture is used. The user's local program constructs a specific file format and sends a request to the server side of the CPN controller through an external interface. After obtaining the decision information sent by the CPN controller from the external interface, it will send task data to the computing power service node. After the user obtains the calculation result, the result is returned to the application program. Figure 6 is an optional resource scheduling schematic diagram according to an embodiment of the present invention.

[0123] The MADDPG mechanism of similarity-assisted experience replay in this embodiment, the MADDPG mechanism and the Multi-Agent Proximal Policy Optimization (MAPPO) mechanism are compared in an environment of 100 computing power nodes under different numbers of tasks Figure 7 is an optional comparison result graph according to an embodiment of the present invention, and the following is obtainedFigure 7 The results shown. It can be found that the MADDPG mechanism (improved MADDPG algorithm, i.e., improved_MADDPG) of similarity-assisted experience replay in this embodiment has low energy consumption and low latency, and performs excellently.

[0124] In this embodiment, a computing power network resource scheduling device is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the terms "module" and "device" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0125] According to an embodiment of the present invention, an apparatus embodiment for implementing the above-mentioned computing power network resource scheduling method is also provided. Figure 8 It is a schematic structural diagram of a computing power network resource scheduling device according to an embodiment of the present invention, as Figure 8 shown, the above-mentioned computing power network resource scheduling device includes: a first acquisition module 200, a second acquisition module 202, a target construction module 204, and a scheduling optimization module 206, wherein:

[0126] The first acquisition module 200 is used to acquire the total task latency of multiple computing tasks.

[0127] The second acquisition module 202 is connected to the first acquisition module 200 and is used to acquire the total computing cost of computing power service nodes included in multiple domains. Among them, the multiple domains are obtained by dividing the computing power network, and the computing power service nodes are used to indicate devices that provide computing resources.

[0128] The target construction module 204 is connected to the second acquisition module 202 and is used to construct an optimization target based on the total computing cost of computing power service nodes included in multiple domains and the total task latency of multiple computing tasks.

[0129] The scheduling optimization module 206 is connected to the target construction module 204 and is used to perform cross-domain resource scheduling optimization on multiple computing tasks based on the optimization target to obtain a resource scheduling result, where the resource scheduling result is at least used to indicate the computing resources respectively allocated to multiple computing tasks.

[0130] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following manner: the above-mentioned various modules can be located in the same processor; or, the above-mentioned various modules are located in different processors in any combination.

[0131] It should be noted that the above first acquisition module 200, second acquisition module 202, target construction module 204, and scheduling optimization module 206 correspond to steps S102 to S108 in the embodiment. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment. It should be noted that the above modules, as part of the device, can run in a computer terminal.

[0132] It should be noted that the optional or preferred implementation manners of this embodiment can be referred to the relevant descriptions in the embodiment, and will not be repeated here.

[0133] The above computing power network resource scheduling device may further include a processor and a memory. The above first acquisition module 200, second acquisition module 202, target construction module 204, scheduling optimization module 206, etc. are all stored in the memory as program modules, and the processor executes the above program modules stored in the memory to implement corresponding functions.

[0134] The processor includes a kernel, and the kernel retrieves the corresponding program module from the memory. One or more kernels can be set. The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM (flash RAM). The memory includes at least one storage chip.

[0135] According to an embodiment of the present application, an embodiment of a non-volatile storage medium is also provided. Optionally, in this embodiment, the above non-volatile storage medium includes a stored program, wherein when the above program runs, it controls the device where the non-volatile storage medium is located to execute any one of the above computing power network resource scheduling methods.

[0136] Optionally, in this embodiment, the above non-volatile storage medium can be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the mobile terminals in a mobile terminal group. The above non-volatile storage medium includes a stored program.

[0137] Optionally, when the program runs, it controls the device where the non-volatile storage medium is located to execute the following functions: obtaining the total task delay of multiple computing tasks; obtaining the total computing cost of computing power service nodes included in multiple domains, where the multiple domains are obtained by dividing a computing power network, and the computing power service nodes are used to indicate devices providing computing resources; constructing an optimization target based on the total computing cost of computing power service nodes included in multiple domains and the total task delay of multiple computing tasks; and performing cross-domain resource scheduling optimization on multiple computing tasks based on the optimization target to obtain a resource scheduling result, where the resource scheduling result is at least used to indicate the computing resources respectively allocated to multiple computing tasks.

[0138] According to an embodiment of the present application, an embodiment of a processor is further provided. Optionally, in this embodiment, the above-mentioned processor is used to run a program, and when the program runs, it executes any one of the above-mentioned computing power network resource scheduling methods.

[0139] According to an embodiment of the present application, an embodiment of a computer program product is further provided. When executed on a data processing device, it is adapted to execute a program initialized with the steps of any one of the above-mentioned computing power network resource scheduling methods.

[0140] Optionally, when the above-mentioned computer program product is executed on a data processing device, it is adapted to execute a program initialized with the following method steps: obtaining the total task latency of multiple computing tasks; obtaining the total computing cost of computing power service nodes included in multiple domains, where the multiple domains are obtained by partitioning a computing power network, and the computing power service nodes are used to indicate devices providing computing resources; constructing an optimization objective based on the total computing cost of computing power service nodes included in multiple domains and the total task latency of multiple computing tasks; performing cross-domain resource scheduling optimization on multiple computing tasks based on the optimization objective to obtain a resource scheduling result, where the resource scheduling result is at least used to indicate the computing resources respectively allocated to multiple computing tasks.

[0141] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented: obtaining the total task latency of multiple computing tasks; obtaining the total computing cost of computing power service nodes included in multiple domains, where the multiple domains are obtained by partitioning a computing power network, and the computing power service nodes are used to indicate devices providing computing resources; constructing an optimization objective based on the total computing cost of computing power service nodes included in multiple domains and the total task latency of multiple computing tasks; performing cross-domain resource scheduling optimization on multiple computing tasks based on the optimization objective to obtain a resource scheduling result, where the resource scheduling result is at least used to indicate the computing resources respectively allocated to multiple computing tasks.

[0142] The order of the above embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments.

[0143] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0144] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the above-mentioned module division can be a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of modules or modules can be in electrical or other forms.

[0145] The modules described above as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0147] If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned non-volatile storage media include: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.

[0148] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A computing power network resource scheduling method, characterized in that: include: Get the total task latency of multiple computing tasks; Obtaining a total computing cost of computing power service nodes included in a plurality of domains, wherein the plurality of domains are obtained by dividing the computing power network, and the computing power service nodes are used to indicate devices that provide computing resources; Constructing an optimization target based on the total computing cost of the computing service nodes included in the multiple domains and the total task delay of the multiple computing tasks; Based on the optimization target, cross-domain resource scheduling optimization is performed on the multiple computing tasks to obtain a resource scheduling result, wherein the resource scheduling result is at least used to indicate the computing resources respectively allocated to the multiple computing tasks.

2. The method according to claim 1, characterized in that The constructing of the optimization target based on the total computing cost of the computing service nodes included in the multiple domains and the total task delay of the multiple computing tasks includes: Determine a first weight value corresponding to the total computing cost and a second weight value corresponding to the total task delay; Based on the first weight value and the second weight value, weighted calculation is performed on the total computing cost and the total task delay to obtain a weighted calculation result; Based on the weighted calculation result, the optimization target is constructed, wherein the optimization target is used to indicate that the weighted calculation result is minimum.

3. The method according to claim 1, characterized in that The obtaining of the total task delay of the plurality of computing tasks comprises: The task delay of any computing task among the multiple computing tasks is obtained by: Obtaining a scheduling region selection vector and a set of predecessor tasks of any of the computing tasks, wherein the scheduling region selection vector is used to indicate whether the corresponding computing task performs cross-domain resource scheduling, and the set of predecessor tasks is used to indicate a set of other tasks or services that need to be completed before executing the corresponding task; Determining a target delay calculation method from a plurality of delay calculation methods based on the scheduling region selection vector and the set of predecessor tasks; Based on the target delay calculation method, obtaining the task delay of any computing task; Obtaining the task delays of the plurality of computing tasks in a manner of obtaining the task delay of any of the computing tasks, wherein the task delays include task sending delays and task transmission delays; The task delays of the multiple computing tasks are summed to obtain the total task delay.

4. The method according to claim 1, characterized in that: The cross-domain resource scheduling optimization is performed on the multiple computing tasks based on the optimization target to obtain a resource scheduling result, including: Based on the optimization goal, the communication public network controllers corresponding to the multiple domains are used as the corresponding multiple intelligent agents, and a multi-agent deep deterministic policy gradient algorithm is used to perform cross-domain resource scheduling optimization on the multiple computing tasks to obtain the resource scheduling result.

5. The method according to claim 4, characterized in that In the case where the multi-agent deep deterministic policy gradient algorithm is constructed based on an action generation network and a value evaluation network, based on the optimization objective, the communication public network controllers corresponding to the multiple domains are used as the corresponding multiple agents, and the multi-agent deep deterministic policy gradient algorithm is used to perform cross-domain resource scheduling optimization on the multiple computing tasks to obtain the resource scheduling result, including: Constructing a reward function based on the optimization objective; For each of the multiple agents, obtain a current state in a corresponding domain through each agent, wherein the current state includes computing power resource information and task information in the corresponding domain; Based on the current state, each agent uses a corresponding action generation network to select an action and execute the action to obtain an execution result; Based on the execution result, a reward value is calculated using a reward function; By using each of the intelligent agents, using a corresponding value evaluation network to evaluate the value of the action to obtain an evaluation result; Based on the reward value and the evaluation result, updating the action generation network and the value evaluation network corresponding to each intelligent agent; Repeat the above operations until a predetermined termination condition is reached; The resource scheduling result is determined based on the action output by the value evaluation network when the predetermined termination condition occurs.

6. The method according to claim 5, characterized in that Based on the reward value and the evaluation result, The action generation network and the value evaluation network corresponding to each intelligent agent are updated, including: Based on the current state, the reward value, the action and the next state, a first experience tuple is constructed, and the first experience tuple is stored in a shared experience pool, wherein the next state is used to indicate the state of each agent after the current state; Selecting a second experience tuple from the shared experience values, the second experience tuple having a similarity with the first experience tuple less than a preset similarity threshold; Based on the evaluation result and the second experience tuple, the action generation network and the value evaluation network corresponding to each intelligent agent are updated.

7. The method according to claim 5, characterized in that The reward function includes positive rewards and negative rewards, wherein the positive reward is generated when the corresponding computing task is completed within the scheduled time, and the negative reward is determined based on the computing cost and task delay of the computing power service nodes included in the corresponding domain.

8. A computing power network resource scheduling device, characterized in that: include: A first acquisition module is used to acquire the total task delay of multiple computing tasks; A second acquisition module is used to obtain a total computing cost of computing service nodes included in a plurality of domains, wherein the plurality of domains are obtained by dividing the computing network, and the computing service nodes are used to indicate devices that provide computing resources; A target construction module, configured to construct an optimization target based on a total computing cost of the computing service nodes included in the multiple domains and a total task delay of the multiple computing tasks; The scheduling optimization module is used to perform cross-domain resource scheduling optimization on the multiple computing tasks based on the optimization target to obtain a resource scheduling result, wherein the resource scheduling result is at least used to indicate the computing resources respectively allocated to the multiple computing tasks.

9. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the computing power network resource scheduling method described in any one of claims 1 to 7.

10. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the computing power network resource scheduling method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Network resource allocation method and device, electronic equipment and storage medium

    CN121419018A