Multi-target coupling scheduling system for computing power resources and electric power resources
The multi-objective coupled scheduling system enables rapid and accurate resource matching for power business tasks in the new power system, solving the problems of resource matching difficulties and multi-objective decision imbalance, and improving the reliability and economy of system operation.
Patent Information
- Application Number
- CN202610162810.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-19
AI Technical Summary
How to quickly and accurately allocate appropriate computing nodes and power nodes for each power business task in a new power system, and ensure that the scheduling results simultaneously meet multiple objectives such as task performance, grid security, economic cost and low carbon emissions, is a challenge that existing technologies face with difficulties in resource matching, lack of awareness of dynamic constraints, imbalance in multi-objective decision-making, and lack of integrated closed-loop verification.
A multi-objective coupled scheduling system for computing and power resources is adopted. Through components such as a global scheduling server, a multi-objective optimization engine, and a global state-aware cluster, task attributes are acquired and evaluated in real time. Combined with an SDN controller and a power control cloud gateway, accurate matching and scheduling of cross-domain resources are achieved. Power grid security and section verification are introduced to ensure the reliability and security of the scheduling scheme.
It significantly shortens the overall time from the submission of power business tasks to the generation of target scheduling schemes, improves the timeliness and continuity of the scheduling process, avoids resource waste, ensures the reliability and green economy of system operation, and realizes the spatiotemporal matching of clean energy and computing load.
Smart Images

Figure CN122068568A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power technology, and in particular to a multi-objective coupled scheduling system for computing resources and power resources. Background Technology
[0002] With the rapid advancement of the construction of new power systems, power business tasks are characterized by high frequency, high concurrency, and high complexity, leading to an explosive growth in the demand for computing resources. The new power system has evolved into a multi-resource collaborative system integrating "user terminals, power nodes, computing nodes, and network nodes": users submit power business tasks through user terminals, computing nodes execute these tasks, power nodes provide power to the computing nodes, and network nodes ensure efficient and reliable data transmission between the nodes.
[0003] However, in actual operation, how to allocate appropriate computing power nodes and power nodes for each power business task has become a core technical challenge that restricts the system's operating efficiency and security. Summary of the Invention
[0004] In view of the above problems, this application provides a multi-objective coupled scheduling system for computing power resources and power resources, so as to allocate appropriate computing power nodes and power nodes for each power service task. The specific solution is as follows:
[0005] The first aspect of this application provides a multi-objective coupled scheduling system for computing resources and power resources, comprising:
[0006] The user terminal is used to obtain power business tasks and task attributes. The task attributes include at least one of the following: computing power required to execute the power business task, maximum tolerable latency, expected completion time, budgeted electricity cost, carbon emission requirements, security requirements, and task type.
[0007] A global scheduling server is used to obtain a complete task label based on the task attributes. The complete task label includes a basic sub-label representing the processing capability of the required computing power node, a power sub-label representing the economic environment requirements of the required power node, and a network sub-label representing the communication capability of the required network node.
[0008] An SDN controller cluster is used to obtain a network communication path composed of network nodes that satisfy the network sub-label from multiple network nodes;
[0009] A power control cloud gateway is used to filter out multiple candidate power sets that satisfy the power sub-label from multiple power nodes, wherein the candidate power set includes one or more power nodes;
[0010] A global scheduling server is used to filter multiple candidate computing nodes that meet the basic sub-labels from multiple computing power nodes;
[0011] A multi-objective optimization engine is used to obtain multiple candidate computing power sets, each candidate computing power set including one or more computing power nodes; combine any one of the candidate computing power sets, any one of the network communication paths, and any one of the candidate power sets to obtain a candidate scheduling scheme; and determine multiple optimal scheduling schemes from the multiple candidate scheduling schemes.
[0012] The global scheduling server is also used to determine a target scheduling scheme that satisfies N-1 security check and cross-sectional power flow check from multiple optimal scheduling schemes.
[0013] One possible implementation also includes:
[0014] A global state-aware cluster is used to acquire real-time computing resource information of multiple computing nodes under various evaluation indicators, communication information of multiple network nodes, economic environment information of multiple power nodes, and static hardware capabilities of the multiple computing nodes under various evaluation indicators. Based on the real-time computing resource information and static hardware capabilities of the multiple computing nodes under various evaluation indicators, a weight vector corresponding to each evaluation indicator and a normalized value corresponding to each evaluation indicator are obtained. Based on the weight vector, real-time computing resource information and static hardware capabilities, the performance level indicators corresponding to the multiple computing nodes are determined.
[0015] In one possible implementation, the multi-objective optimization engine, when performing the acquisition of multiple candidate computing power sets, specifically includes:
[0016] Based on the performance level indicators corresponding to multiple candidate computing power nodes, multiple candidate computing power sets are obtained, and the candidate computing power sets include one or more computing power nodes.
[0017] In one possible implementation, when the global state-aware cluster performs the steps of obtaining the weight vectors corresponding to each evaluation index and the normalized values of the multiple computing nodes for each evaluation index based on the real-time computing resource information of the multiple computing nodes under each evaluation index and the static hardware capabilities of the multiple computing nodes under each evaluation index, the steps include:
[0018] Based on the real-time computing power resource information of multiple computing power nodes under various evaluation indicators, the first weight vector corresponding to each evaluation indicator is obtained.
[0019] Based on the static hardware capabilities of multiple computing nodes under various evaluation indicators, the second weight vector corresponding to each evaluation indicator is obtained.
[0020] Based on the first weight vector corresponding to each evaluation index and the real-time computing power resource information of multiple computing power nodes under each evaluation index, the real-time weighting matrix corresponding to multiple computing power nodes is obtained.
[0021] Based on the second weight vector corresponding to each of the evaluation indicators and the static hardware capabilities of multiple computing nodes under each evaluation indicator, obtain the static weighting matrix corresponding to each of the multiple computing nodes.
[0022] Based on the real-time weighted matrix and the static weighted matrix corresponding to each of the multiple computing power nodes, the weighted comprehensive matrix corresponding to each of the multiple computing power nodes is obtained.
[0023] Based on the weighted comprehensive matrix corresponding to each of the multiple computing power nodes, the performance level labels corresponding to each of the multiple computing power nodes are determined.
[0024] In one possible implementation, the number of evaluation indicators is n, and the number of multiple computing nodes is m. The step of obtaining the first weight vector corresponding to each evaluation indicator based on the real-time computing resource information of multiple computing nodes under each evaluation indicator includes:
[0025] Obtain real-time computing resource information of multiple computing nodes under various evaluation indicators. ,in, It refers to the first average value of real-time computing resource information of the i-th computing node under the current period of the j-th evaluation index;
[0026] If the j-th evaluation metric of the i-th computing node is a positive metric, then by formula... The first standardized value is calculated. The positive indicator is positively correlated with the capability of the computing node. It refers to the maximum value of the j-th evaluation index for the i-th computing power node. It refers to the minimum value of the j-th evaluation index of the i-th computing power node;
[0027] If the j-th evaluation metric of the i-th computing power node is negative, then by formula... The first standardized value is calculated. The negative indicator is negatively correlated with the capability of the computing node.
[0028] Through formula The first weight vector of the j-th evaluation index is calculated.
[0029] In one possible implementation, the number of evaluation metrics is n, and the number of computing nodes is m. Based on the static hardware capabilities of the computing nodes under each evaluation metric, the second weight vector corresponding to each evaluation metric is obtained as follows:
[0030] Obtain the static hardware capabilities of multiple computing nodes under various evaluation metrics. ,in, It refers to the second average value of the static hardware capability of the i-th computing node under the current period of the j-th evaluation index;
[0031] If the j-th evaluation metric of the i-th computing node is a positive metric, then by formula... The second standardized value was calculated. The positive indicator is positively correlated with the capability of the computing node. This refers to the maximum value of the j-th second preset evaluation index for the i-th computing power node. It refers to the minimum value of the j-th evaluation index of the i-th computing power node;
[0032] If the j-th evaluation metric of the i-th computing power node is negative, then by formula... The second standardized value was calculated. The negative indicator is negatively correlated with the capability of the computing node.
[0033] Through formula The second weight vector of the j-th evaluation index is calculated.
[0034] In one possible implementation, the step of obtaining the weighted composite matrix corresponding to each of the multiple computing power nodes based on the real-time weighted matrix and the static weighted matrix corresponding to each of the multiple computing power nodes includes:
[0035] For each computing power node, determine the real-time weighting matrix of that computing power node. The static weighting matrix of the computing power nodes The sum of these is the weighted composite matrix of the computing power nodes.
[0036] In one possible implementation, the step of determining the performance level labels corresponding to the multiple computing power nodes based on the weighted comprehensive matrix corresponding to the multiple computing power nodes includes:
[0037] The weighted comprehensive matrix corresponding to multiple computing power nodes is input into the pre-constructed performance level prediction model, and the performance level prediction model outputs the performance level labels of multiple computing power nodes.
[0038] The performance level prediction model is trained by using the weighted comprehensive matrix of the sample computing power nodes as input and the labeled performance level tags of the sample computing power nodes as the training target.
[0039] In one possible implementation, after determining the target scheduling scheme, the global scheduling server is further configured to:
[0040] Generate scheduling instructions for the power service task;
[0041] The scheduling instructions are sent to the computing nodes, network nodes, and power nodes included in the target scheduling scheme.
[0042] In one possible implementation, the multi-objective optimization engine, when performing the step of determining multiple optimal scheduling schemes from multiple candidate scheduling schemes, specifically uses:
[0043] Multiple candidate scheduling schemes were identified as the initial population;
[0044] The initial population is input into the NSGA-II algorithm, with the optimization objectives of minimizing total execution cost, total task latency, and total task carbon emissions, resulting in multiple undetermined scheduling schemes. The total execution cost is the sum of the costs of computing nodes, power nodes, and network nodes. The total task latency includes the latency of computing nodes and the transmission latency of network communication paths. The total task carbon emissions refer to the total amount of carbon emissions during the execution of the power service task.
[0045] Using the first preset weight of the total execution cost, the second preset weight of the total task delay, and the third preset weight of the total task carbon emissions as reference points in a three-dimensional space, and taking multiple undetermined scheduling schemes as an initial population, multiple optimal scheduling schemes are obtained through NSGA-Ⅲ; the three-dimensional coordinates of the three-dimensional space are the total execution cost, the total task delay, and the total task carbon emissions, respectively.
[0046] In one possible implementation, when the global scheduling server performs the step of determining a target scheduling scheme that satisfies the N-1 security check and cross-sectional power flow check from a plurality of said optimal scheduling schemes, it includes:
[0047] From the correspondence between preset decision rules and calculation formulas, determine the target calculation formula corresponding to the preset decision rules that the power business task satisfies;
[0048] Based on the target calculation formula, scores corresponding to multiple optimal scheduling schemes are obtained respectively;
[0049] Obtain the unverified scheduling scheme from the multiple optimal scheduling schemes sorted by the score from high to low;
[0050] If the scheduling scheme under test satisfies the N-1 security check and the cross-sectional power flow check, the scheduling scheme under test is determined to be the target scheduling scheme.
[0051] If the proposed scheduling scheme does not meet the N-1 security check or the cross-sectional power flow check, return an unchecked proposed scheduling scheme from the multiple optimal scheduling schemes sorted by the scores from high to low.
[0052] In one possible implementation, when the global scheduling server retrieves the complete task tag based on the task attribute, it includes:
[0053] If the task type in the task attributes is an emergency task, the priority with the lowest total delay is determined to be higher than the priority with the lowest total execution cost and higher than the priority with the lowest total carbon emissions; the total execution cost is the sum of the cost of the computing node, the cost of the power node, and the cost of the network node; the total delay includes the delay of the computing node and the transmission delay of the network communication path; the total carbon emissions refer to the total amount of carbon emissions during the execution of the power service task;
[0054] If the task type in the task attributes is a non-urgent task and the carbon emission requirement in the task attributes is not marked as green electricity, the priority of the task with the lowest total execution cost is determined to be higher than the priority of the task with the lowest total delay and higher than the priority of the task with the lowest total carbon emissions.
[0055] If the task type in the task attributes is a non-urgent task and the carbon emission requirement in the task attributes is marked as green electricity, the priority of the task with the lowest total carbon emissions is determined to be higher than the priority with the lowest total execution cost and higher than the priority with the lowest total delay.
[0056] By employing the aforementioned technical solutions, this application provides a multi-objective coupled scheduling system for computing and power resources. By transforming the task attributes of power business tasks into multi-dimensional constraints, it achieves synchronous screening and matching of computing nodes, power nodes, and network nodes, significantly shortening the overall time from power business task submission to the generation of the target scheduling scheme, and improving the timeliness and consistency of the scheduling process. Relying on the integrated processing of the processing capabilities of multiple computing nodes, the communication capabilities of multiple network nodes, and the economic and environmental requirements of multiple power nodes, the resource assessment process avoids subjective biases caused by manual intervention, making the evaluation results more stable, consistent, and repeatable. The parallel processing mechanism of network and power constraints reduces repeated calculations caused by traditional serial screening, making the solution process more efficient and avoiding waste of resources and time. Introducing power grid safety and section verification as the final control link ensures that the selected target scheduling scheme has sufficient safety margin during operation, moving risk prevention forward and enhancing the reliability of system operation. The overall process unifies the modeling and collaborative scheduling of task attributes with the states of computing nodes, network nodes, and power nodes, promoting the spatiotemporal matching of clean energy and computing load, and helping to achieve green and economical operation of the power grid and computing system. Attached Figure Description
[0057] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0058] Figure 1 A schematic diagram of the system architecture of a multi-objective coupled scheduling system for computing power resources and power resources provided in this application;
[0059] Figure 2 A flowchart illustrating an implementation of a multi-objective coupled scheduling system for computing and power resources provided in this application embodiment;
[0060] Figure 3 This is a flowchart illustrating another implementation of a multi-objective coupled scheduling system for computing and power resources provided in this application. Detailed Implementation
[0061] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0062] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0063] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0064] How to allocate appropriate computing power nodes and power nodes for each power business task has become a core technical challenge restricting the system's operational efficiency and security, for the following reasons:
[0065] Reason 1: Heterogeneous and dispersed resources lead to matching difficulties.
[0066] Computing nodes, such as CPUs (Central Processing Units), GPUs (Graphics Processing Units), and NPUs (Neural Processing Units), vary significantly in terms of architecture, performance, power consumption, and geographical location; power nodes exhibit dynamic changes in attributes such as power supply type (thermal power, wind power, photovoltaic), carbon emission intensity, real-time electricity price, and available capacity; and network nodes also exhibit high uncertainty in terms of bandwidth, latency, packet loss rate, and topology path.
[0067] In related technologies, the allocation of computing nodes, network nodes, and power nodes for power service tasks often adopts a static capability matching and manual experience-based decision-making approach. This means that computing nodes, network nodes, and power nodes are manually allocated to power service tasks based on their static hardware capabilities. This method struggles to achieve accurate adaptation of heterogeneous resources across domains, often resulting in mismatches such as assigning high-capacity nodes to high-carbon power sources or assigning low-latency tasks to high-cost paths, leading to resource waste and increased operating costs.
[0068] Static hardware capabilities refer to the physical specifications of a node that are determined at the time of manufacture or deployment and are independent of real-time load.
[0069] Reason 2: Lack of effective sensing and closed-loop management of dynamic constraints on the power side.
[0070] The operation mode of the new power system is constantly changing. For example, the cross-sectional power flow limit may be tightened instantaneously due to N-1 faults; the carbon emission intensity of power nodes changes drastically with the fluctuation of wind power and photovoltaic output; and the real-time electricity price can vary by several times during peak and off-peak periods.
[0071] In the process of allocating computing nodes, network nodes, and power nodes for power business tasks, the aforementioned dynamic constraints on the power side cannot be obtained and internalized in real time. This results in the power business tasks that have been issued facing unpredictable risks such as electricity price jumps, carbon emission exceeding limits, and power flow overload during the execution phase. This not only increases the pressure on the power grid operation but also weakens the credibility of enterprises' green and low-carbon commitments.
[0072] Reason 3: Decision-making imbalance under multiple conflicting objectives
[0073] Different power business tasks have significantly different preferences for "performance-cost-carbon emissions-safety". For example, real-time control tasks require millisecond-level latency and can accept higher costs; batch analysis tasks pursue the lowest cost and can tolerate longer completion times; while policy-related tasks prioritize the use of green electricity.
[0074] In the process of allocating computing nodes, network nodes, and power nodes for power business tasks, related technologies generally adopt single-objective or fixed-weight multi-objective optimization, which makes it difficult to weigh conflicting objectives and provide an interpretable and auditable scheduling scheme within seconds. Ultimately, it can only rely on manual offline adjustments, which is inefficient and prone to problems such as priority inversion and high-priority tasks starving.
[0075] Reason 4: Lack of an integrated closed-loop verification mechanism for "computing power nodes - power nodes - network nodes"
[0076] In related technologies, the interface between the computing platform and the power dispatching system (EMS, DMS) is closed. Power consumption adjustment commands issued by the computing nodes cannot be verified and executed by the power grid in real time. Conversely, power grid power flow exceeding limits and insufficient reserves cannot be promptly transmitted to the computing dispatching end. Whether the resource allocation results truly meet the rigid constraints of the power grid's N-1 safety, cross-sectional power flow, and carbon emission quotas lacks online closed-loop verification methods, leading to a disconnect between "dispatch-execution-feedback" and insufficient overall system reliability.
[0077] In summary, how to quickly and accurately allocate suitable computing nodes and power nodes in a heterogeneous, distributed, and real-time changing resource pool of "computing nodes-power nodes-network nodes" for each power business task, and ensure that the allocation results simultaneously meet multiple objectives such as task performance, grid security, economic cost, and low carbon emissions, has become a core technical challenge that urgently needs to be overcome by those skilled in the art.
[0078] Based on this, this application provides a multi-objective coupled scheduling system for computing power resources and power resources (hereinafter referred to as the system). The hardware architecture involved in the system is described below.
[0079] See Figure 1 , Figure 1 A schematic diagram of a multi-objective coupled scheduling system for computing and power resources is shown. The system may include one or more data centers, each of which includes, but is not limited to: user terminal 110, application server 120, global scheduling server 130, multi-objective optimization engine 140, global state-aware cluster 150, power control cloud gateway 160, SDN controller cluster 170, computing nodes 180, network nodes 190, power nodes 200, scheduling and management terminal 210, and resource database 220.
[0080] User terminal 110 allows various users to submit power service tasks, monitor dispatch status, and obtain results. Users can access the user terminal via the network to submit power service tasks. After submitting a power service task, users can use the user terminal 110 to view the queuing status, execution progress, and final result of the power service task in real time, as well as view historical task records and resource consumption reports.
[0081] The dispatch and control terminal 210 is a professional interface for administrators to perform advanced management and intervention. The dispatch and control terminal 210 provides administrators with system management functions, including: approving user-submitted power business tasks, manually adjusting and optimizing automatically generated dispatch strategies, defining and managing global resource dispatch strategies, setting system operating parameters and constraints, monitoring system health status, handling alarm events, and performing system configuration and maintenance operations. All management commands are issued to the global dispatch server 130 through the dispatch and control terminal 210.
[0082] "Approving user-submitted power service tasks" refers to manually reviewing power service requests submitted by administrators through user terminals to confirm their rationality, security, and resource feasibility, and to prevent unreasonable power service tasks from entering the system.
[0083] "Dynamically adjust and optimize automatically generated scheduling strategies" means that although the system can automatically generate target scheduling plans, administrators can manually intervene based on experience or unexpected situations to fine-tune or re-optimize the strategies in order to improve effectiveness or deal with anomalies.
[0084] "Defining and managing global resource scheduling strategies" refers to formulating system-level scheduling rules, such as strategy templates like "low-carbon priority," "cost priority," and "high reliability priority," for use by the multi-objective optimization engine to achieve consistency and configurability of scheduling objectives.
[0085] "Setting system operating parameters and constraints" refers to configuring key operating parameters (such as queue timeout, number of retries, and maximum parallelism) and hard constraints (such as latency limits, carbon emission limits, and security domain limits) to ensure that the system operates within the allowable range.
[0086] "Monitoring the health status of the system" refers to viewing the operating metrics (CPU, memory, network, response time) of each component (server, node, network, resource database) in real time to understand the overall health of the system and identify potential risks in advance.
[0087] "Handling alarm events" refers to manually confirming, classifying, assigning tasks, or ignoring various alarm events (such as node failures, network jitter, and power outages) reported by the system to ensure that anomalies are responded to and handled in a timely manner.
[0088] "Performing system configuration and maintenance operations" refers to routine operation and maintenance tasks such as user permission management, model parameter adjustment, version upgrade, log cleanup, backup and recovery, to ensure the long-term stability, controllability and maintainability of the system.
[0089] The global scheduling server is the core decision-making hub of the system, deployed in the central cloud or main data center. It receives power business tasks from user terminals and management instructions from scheduling and control terminals, and embeds modules such as policy management, resource registration, and constraint management. The global scheduling server interacts closely with the multi-objective optimization engine, the global state awareness cluster, the power control cloud gateway, the SDN controller cluster, and the resource database, aggregating real-time and non-real-time data from the entire network and outputting global resource collaborative allocation instructions and scheduling strategies.
[0090] The application server (general-purpose server) interacts with core components such as the global scheduling server, resource database, and global status awareness cluster through the internal network, providing rich business function support for user terminals while ensuring the high availability and scalability of the system.
[0091] The multi-objective optimization engine, deployed in conjunction with the global scheduling server, forms its core computing component. This engine receives optimization objectives and constraints from the global scheduling server and, based on built-in optimization algorithms such as reinforcement learning and genetic algorithms, performs large-scale, multi-objective joint solutions (e.g., computing power performance, electricity cost, carbon efficiency, network latency and bandwidth). It outputs the optimal or near-optimal integrated computing power-electricity-network resource allocation scheme and sends the scheme back to the global scheduling server for decision-making.
[0092] The Global Status Awareness Cluster 150 is an independent distributed monitoring system responsible for collecting and aggregating real-time information from all resources (computing nodes, network nodes, and power nodes in all data centers) across regions and levels, such as real-time computing resource information of computing nodes, communication information of network nodes, and economic environment information of power nodes.
[0093] The global state-aware cluster 150 continuously collects real-time computing resource information (such as CPU / GPU utilization, load, power consumption, and temperature) from computing nodes 180 in different data centers, economic and environmental information (such as real-time electricity price, power supply carbon intensity, available capacity, and line load) from power nodes 200 in different data centers, and communication information (such as real-time link bandwidth, port throughput, network latency, packet loss rate, and jitter) from network nodes and SDN controller cluster 170 in different data centers. After cleaning and aggregating this data, the global state-aware cluster 150 pushes it to the global scheduling server 130 in real time to support decision-making, and also persists it in the resource database 220 for historical analysis and big data optimization.
[0094] The Power Control Cloud Gateway 160 serves as a secure bridge connecting multi-objective coupled dispatching systems and power control systems (such as Energy Management Systems, EMS). It is responsible for acquiring rigid security constraints (such as cross-sectional power flow limits, reserve capacity requirements, and frequency stability ranges) from the power control system and converting them into a model understandable by the multi-objective coupled dispatching system. Simultaneously, it securely and reliably transmits dispatching instructions for power business tasks issued by the multi-objective coupled dispatching system to power nodes for execution, ensuring that all dispatching activities are conducted within the power grid security framework.
[0095] The SDN controller cluster 170 is the brain of network resources, responsible for the centralized control and flexible scheduling of underlying network resources in a multi-objective coupled scheduling system. The SDN controller cluster 170 communicates with all network nodes (such as SDN switches and routers) through interfaces (such as OpenFlow) to implement network policies issued by the global scheduling server 130 (such as dynamically creating / tearing up SRv6 tunnels across data centers, allocating guaranteed bandwidth for critical scheduling command flows, and performing load balancing). Simultaneously, the SDN controller cluster 170 provides real-time network topology and traffic information to the global scheduling server 130 and the global state-aware cluster 150, which is crucial for achieving "network programmability" and supporting efficient collaboration between computing power and power flow.
[0096] Resource Database 220 consists of multiple highly available storage servers, providing a data foundation for the multi-objective coupled scheduling system. Resource Database 220 stores metadata for all registered computing nodes (such as model, architecture, and performance indicators), metadata for power nodes (such as geographical location, power supply type, capacity, and carbon emission coefficient), metadata for network nodes (such as device model, port speed, link capacity, and supported protocols), historical status data, a scheduling policy library, and policy execution logs, supporting the system's analysis, decision-making, backtracking, and optimization.
[0097] The computing node 180 is equipped with a computing power agent module, which is a standardized interface for the computing node. The computing power agent module is responsible for collecting real-time computing power resource information of the computing node and receiving and executing resource allocation instructions from the global scheduling server 130 (such as starting / hibernating the computing node, allocating computing tasks, and adjusting the power consumption limit).
[0098] Power node 200 is equipped with a power agent module. For example, power nodes include, but are not limited to, distributed photovoltaic power stations, energy storage power stations, and adjustable load centers. The power agent module is the standardized interface of the power node. Each power agent module is responsible for collecting economic environment information of the power nodes within its jurisdiction and receiving and executing power regulation instructions (such as increasing / decreasing generation power, charging / discharging, and switching loads) from the global scheduling server.
[0099] Network node 190 is equipped with a network proxy module. For example, network nodes include, but are not limited to, various types of network nodes, such as core, aggregation, and access switches and routers. The network proxy module is responsible for collecting communication information from network nodes, such as fine-grained data like port-level traffic, status, and error counts, and for receiving and executing instructions from the SDN controller cluster, such as flow table distribution, policy configuration, and path adjustment. It is the foundation for realizing network resource awareness and controllability.
[0100] The following description Figure 1 The product form of the middle user terminal 110;
[0101] The user terminal 110 in this application embodiment can be a mobile phone, tablet computer, wearable device, vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., and this application embodiment does not impose any restrictions on it.
[0102] User terminal 110 may include a radio frequency unit, memory, input unit, display unit, camera (optional), audio circuitry (optional), speaker (optional), microphone (optional), headphone jack (optional), processor, external interface, power supply, and other components. Those skilled in the art will understand that the above components are merely examples and do not constitute a limitation on the terminal or multifunctional device; it may include more or fewer components, or a combination of certain components, or different components.
[0103] The input unit can be used to receive input numeric or character information, and to generate key signal inputs related to user settings and function control of the portable multi-functional device. Specifically, the input unit may include a touchscreen (optional) and / or other input devices. Other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0104] Among them, the input device can receive input data, etc.
[0105] The display unit can be used to display information input by the user or information provided to the user, various menus of the terminal, interactive interfaces, file display, and / or playback of any multimedia file. In the embodiments of this application, the display unit can be used to display interfaces, processing results, etc.
[0106] This radio frequency unit (optional) can be used to receive and send signals during information transmission or calls.
[0107] In this embodiment of the application, the radio frequency unit can send data to the server and receive the processing results sent by the server.
[0108] It should be understood that this radio frequency unit is optional and can be replaced with other communication interfaces, such as a network port.
[0109] User terminal 110 also includes a power source (such as a battery) that supplies power to the various components.
[0110] User terminal 110 also includes an external interface, which can be a standard Micro USB interface or a multi-pin connector, which can be used to connect user terminal 110 to other devices for communication, or to connect a charger to charge user terminal 110.
[0111] For example, any of the devices among application server 120, global scheduling server 130, multi-objective optimization engine 140, global state-aware cluster 150, and SDN controller cluster 170 can be a server or a server cluster. The server includes a bus, processor, communication interface, and memory. The processor, memory, and communication interface communicate with each other via the bus.
[0112] The following is combined Figure 1 The specific implementation process of the multi-objective coupled scheduling system for computing power resources and power resources in the embodiments of this application is described.
[0113] Reference Figure 2 , Figure 2This is a flowchart illustrating an implementation of a multi-objective coupled scheduling system for computing and power resources provided in this application embodiment. The process may include steps S201 to S209. These steps are described in detail below.
[0114] Step S201: The user terminal obtains the power service task and task attributes.
[0115] The task attributes include at least one of the following: computing power required to execute the power business task, maximum tolerable latency, expected completion time, budgeted electricity cost, carbon emission requirements, security requirements, and task type.
[0116] The task type represents the urgency of the power business task, and the security requirement represents the geographical location where the data related to the power business task is allowed to be located.
[0117] The power business task is a quantifiable, dispatchable, and closed-loop power system calculation or control operation that must be completed collaboratively by computing nodes, power nodes, and network nodes with the goal of ensuring the safe, economical, and low-carbon operation of the power grid.
[0118] The following are examples of power service tasks. Power service tasks can include real-time fault diagnosis at the 20ms level, load clustering and profiling of 1 million users, assessment of dispatchable energy storage capacity and cost, and detection of abnormal user electricity consumption behavior.
[0119] The following explains the task attributes.
[0120] Computing power scale refers to the total theoretical computing capacity required to complete power business tasks, which can be further divided into CPU, GPU, NPU, memory, storage capacity and bandwidth requirements.
[0121] Maximum tolerable delay refers to the maximum allowed delay from the submission of a power business task to the return of the final result. This delay includes the entire process of queuing, calculation, and network transmission.
[0122] The expected completion time refers to the latest date on which the power business task can be completed.
[0123] Budgeted electricity charges refer to the maximum electricity charges that users are willing to pay for electricity business tasks, including computing, storage, and network energy consumption costs.
[0124] Carbon emission demand refers to the maximum amount of carbon emissions allowed to be generated from the start to the end of an electricity business operation.
[0125] Task type indicates the urgency of power operations.
[0126] The following example illustrates the security requirements.
[0127] Assume the power business task is to complete the 0–72 hour wind and solar power forecast for the entire province before 08:00 AM on June 26, 2025. The task attributes are: expected completion time: 08:00 AM tomorrow; input: 3GB meteorological data; budgeted electricity cost ≤ 300 yuan; carbon emission requirement ≤ 150g / kWh; results must remain within Province AA. Therefore, the security requirement is that the data related to the power business task cannot leave Province AA. In other words, the computing nodes executing the power business task and the network nodes transmitting the relevant data should all be within Province AA.
[0128] Step S202: The global scheduling server obtains the complete task tag based on the task attributes.
[0129] The complete task label includes a basic sub-label representing the capabilities of the required computing power nodes, a power sub-label representing the economic environment requirements of the required power nodes, and a network sub-label representing the communication capabilities of the required network nodes.
[0130] The basic sub-tags include computing power scale, task execution period, potential geographical location of the computing power node executing the power service task, and task type; the power sub-tags represent the carbon emission and electricity price requirements of the power service task during execution; the network sub-tags represent the hard performance limits and path requirements for network transmission met for the power service task.
[0131] Understandably, power service tasks need to be standardized and tagged after submission. Power service task requests submitted by user terminals pass through a load balancer. The load balancer parses the HTTP / WebSocket request body of the power service task and transforms it into an internal system description containing basic requirements. Subsequently, a tag generation microservice (deployed on a global scheduling server) calls a text classification machine learning model to determine the task type and adds relevant task tags based on a rule base.
[0132] Understandably, based on the static hardware capabilities of computing nodes, the required computing power scale and security requirements of power business tasks, qualified computing nodes can be selected; combined with the task execution period, computing nodes that meet the task execution period can be further selected; the geographical location of qualified computing nodes is the potential geographical location.
[0133] Assume the power business task is to complete the 0–72 h wind and solar power forecast for the entire province before 08:00 on June 26, 2025. The task attributes are: expected completion time: 08:00 tomorrow morning, input 3 GB of meteorological data, budget electricity cost ≤ 300 yuan, carbon emission requirement ≤ 150 g / kWh, and the results must remain in AA province. The basic sub-tags include: computing power scale of 512 CPU cores + 1TB memory (batch processing parallel); task execution deadline of 2025-06-26T08:00:00; potential geographical location of [Hangzhou-Xiasha Cloud Valley, Ningbo-Meishan, Jiaxing-Wuzhen]; the urgency level represented by the task type is high, i.e., priority=HIGH; the power sub-tags include: budgeted electricity cost max_cost=300 yuan; carbon emission requirement of umax_carbon=150g / kWh; the network sub-tags include: maximum tolerable latency max_latency=5 min (3GB of data to be uploaded within 30 seconds), minimum bandwidth min_bandwidth=1 Gbps, and avoid_geo= [provinces other than AA province].
[0134] Hard performance limits refer to the minimum or maximum allowable thresholds given in absolute numerical form.
[0135] Step S203: The SDN controller cluster obtains the network communication path composed of network nodes that satisfy the network sub-label from multiple network nodes.
[0136] The network communication path refers to the sequence of links through which data is transmitted in network nodes during the execution of the power service task.
[0137] For example, the power control cloud gateway 160 can be invoked via RPC (Remote Procedure Call) to obtain parameters such as carbon emission intensity and dynamic electricity price of the current regional power nodes; at the same time, the SDN controller cluster 170 calculates the network communication path that meets the latency requirements based on the basic sub-label, power sub-label, network sub-label, and network constraints.
[0138] For example, network constraints include, but are not limited to: maximum tolerable latency for the task, minimum required bandwidth, maximum packet loss rate, and path preference. For example, network constraints can be uploaded by the user through their terminal.
[0139] Taking the aforementioned power business task of completing the 0–72 h wind and solar power forecast for the entire province before 08:00 AM on June 26, 2025 as an example, max_latency=5min indicates that the end-to-end transmission time + queuing time + processing time ≤ 300s; min_bandwidth=1Gbps indicates that the effective throughput ≥ 1 Gbps (3GB of data transmitted within 30s); avoid_geo = [provinces other than AA province] indicates that any path passing through routers, POPs, or backbone nodes in provinces other than AA province will be discarded. Multiple network communication paths can be obtained using the K-shortest path algorithm for network nodes that meet the above "hard performance limits and path requirements".
[0140] It is understandable that the bottleneck bandwidth of each network communication path is ≥1Gbps, the data transmission time of each network communication path is ≤30s, and the network nodes in each network communication path do not belong to Jiangsu Province.
[0141] Step S204: The power control cloud gateway filters out multiple candidate power sets that meet the power sub-label from multiple power nodes.
[0142] The candidate power set includes one or more power nodes.
[0143] The following example illustrates step S204.
[0144] Assume the power sub-label is: max_cost = 300 yuan, max_carbon = 150g / kWh.
[0145] Assume multiple power nodes are as follows:
[0146] The real-time electricity price of the Hangzhou-Xiasha Yung Valley power node when performing power business tasks is 0.65 yuan / kWh, and the carbon emission intensity is 380 gCO2 / kWh. Due to exceeding the carbon emission and budget, the high computing power node was eliminated.
[0147] The Ningbo-Meishan offshore wind power dedicated line node has a real-time electricity price of 0.32 yuan / kWh and a carbon emission intensity of 45gCO2 / kWh when performing power business tasks. Since both of these meet the standards, the power node is classified into the candidate power set.
[0148] Step S205: The global scheduling server selects multiple candidate computing nodes from multiple computing nodes that meet the basic sub-labels.
[0149] Step S206: The multi-objective optimization engine obtains multiple candidate computing power sets, wherein the candidate computing power sets include one or more computing power nodes.
[0150] Step S207: The multi-objective optimization engine combines any of the candidate computing power sets, any of the network communication paths, and any of the candidate power sets to obtain a candidate scheduling scheme.
[0151] Step S208: The multi-objective optimization engine determines multiple optimal scheduling schemes from the multiple candidate scheduling schemes.
[0152] Step S209: The global scheduling server determines the target scheduling scheme that satisfies the N-1 security check and cross-sectional power flow check from multiple optimal scheduling schemes.
[0153] After obtaining multiple optimal scheduling schemes output by the multi-objective optimization engine, the global scheduling server 130 will initiate the power grid security constraint solving process to ensure that the selected target scheduling scheme not only optimizes performance but also meets the rigid security requirements of power grid operation. The target scheduling scheme that best aligns with user preferences will be selected.
[0154] N-1 safety check refers to the process by which the power grid, in the event of the loss of any one component (N-1), does not cause overload of other components, grid disconnection, or voltage collapse.
[0155] In cross-sectional power flow calculation, a "cross-section" refers to a set of critical lines in a power grid, whose total transmission power directly affects the stability of the entire power grid. "Power flow calculation" is a type of steady-state calculation for power systems, used to determine the power, current, voltage, and other states of all power grid components.
[0156] This application provides a multi-objective coupled scheduling system for computing and power resources. By transforming the task attributes of power business tasks into multi-dimensional constraints, it achieves synchronous screening and matching of computing nodes, power nodes, and network nodes, significantly shortening the overall time from power business task submission to target scheduling scheme generation, and improving the timeliness and consistency of the scheduling process. Relying on the integrated processing of the processing capabilities of multiple computing nodes, the communication capabilities of multiple network nodes, and the economic and environmental requirements of multiple power nodes, the resource assessment process avoids subjective biases caused by human intervention, making the evaluation results more stable, consistent, and repeatable. The parallel processing mechanism of network and power constraints reduces repeated calculations caused by traditional serial screening, making the solution process more efficient and avoiding waste of resources and time. Introducing grid security and section verification as the final control link ensures that the selected target scheduling scheme has sufficient safety margin during operation, moving risk prevention forward and enhancing the reliability of system operation. The overall process unifies the modeling and collaborative scheduling of task attributes with the states of computing nodes, network nodes, and power nodes, promoting the spatiotemporal matching of clean energy and computing load, and helping to achieve green and economical operation of the power grid and computing system.
[0157] like Figure 3The diagram shown illustrates another implementation of a multi-objective coupled scheduling system for computing and power resources provided in this application. The process may include steps S301 to S308, which are described in detail below.
[0158] Step S301: User terminal 110 obtains the power service task and task attributes.
[0159] Step S302: The global state-aware cluster 150 acquires real-time computing resource information of multiple computing nodes under various evaluation indicators, communication information of multiple network nodes, economic environment information of multiple power nodes, and static hardware capabilities of the multiple computing nodes under various evaluation indicators; based on the real-time computing resource information and static hardware capabilities of the multiple computing nodes under various evaluation indicators, it obtains the weight vectors corresponding to each evaluation indicator and the normalized values of the multiple computing nodes under each evaluation indicator; based on the weight vectors, real-time computing resource information and static hardware capabilities of the multiple computing nodes under various evaluation indicators, it determines the performance level indicators corresponding to the multiple computing nodes.
[0160] Understandably, computing nodes, power nodes, and network nodes can register with the global state-aware cluster 150, which can then obtain the static hardware capabilities of each computing node, each power node, and each network node.
[0161] For example, the static hardware capabilities of a computing node include, but are not limited to: manufacturer, model, architecture, number of CPU cores / GPU stream processors, theoretical peak performance, memory capacity, serial number, and physical location (data center-server room-rack).
[0162] For example, the static hardware capabilities of a power node include, but are not limited to: device model, port type / number, maximum speed per port, supported protocols, and topology endpoints.
[0163] For example, the static hardware capabilities of a network node include, but are not limited to: power supply type (mains / wind / solar), maximum sustainable power supply capacity, and theoretical carbon intensity coefficient.
[0164] Understandably, after acquiring the aforementioned data, the global state-aware cluster 150 will perform normalization processing. For example, it will standardize the units of measurement, such as unifying memory units to GB, power units to kW, and theoretical carbon emission intensity coefficient units to gCO2 / kWh. It will map the unstructured theoretical carbon emission intensity coefficient into structured key-value pairs according to a predefined JSON Schema / Protobuf model. It will perform a SHA-256 hash on the "manufacturer + model + serial number" to generate a globally unique resource ID, ensuring that only one identity exists for the same node. This will obtain a static resource registry and store it in the resource database.
[0165] The resource registration service is a software microservice deployed on the global scheduling server 130. It listens for registration requests from computing nodes, network nodes, and power nodes, executes data standardization and processing logic, and interacts with the resource database 220 to store the data. When the resource registration service starts, it loads the code for the computing resource description model, such as {"node_id": "string","gpu_model":"string","total_flops":"number","unit":"TFLOPS","memory_capacity": "number","unit": "GB", ...}.
[0166] The evaluation metrics are explained below. It is understood that the evaluation metrics for computing nodes, network nodes, and power nodes may be the same or different. For example, evaluation metrics for computing nodes include, but are not limited to: real-time computing utilization (CPU / GPU / NPU), current idle computing power, number of running and queued tasks, instantaneous power consumption (W), core temperature (°C), and VRAM / RAM utilization; evaluation metrics for network nodes include, but are not limited to: link real-time bandwidth utilization (%), port forwarding latency (ms), packet loss rate (%), and jitter (ms); evaluation metrics for power nodes include, but are not limited to: real-time dynamic electricity price (yuan / kWh), current actual carbon emission intensity (gCO2 / kWh), and real-time load (kW).
[0167] Understandably, based on static registration, the global state-aware cluster 150 captures real-time resource status through continuous state awareness. Similarly, relying on network proxy modules in network nodes, computing power proxy modules in computing power nodes, and power proxy modules in power nodes, they periodically poll or subscribe to events at high frequency (seconds to milliseconds) to obtain dynamic operational metrics, namely, real-time computing power resource information of computing power nodes, communication information of network nodes, and economic environment information of power nodes.
[0168] Understandably, massive amounts of real-time data, such as real-time computing resource information of computing nodes, communication information of network nodes, and economic environment information of power nodes, are reported to the global state awareness cluster 150. The global state awareness cluster 150 first performs data cleaning (removing outliers) and verification. Subsequently, by querying the static resource registry in the resource database 220, it correlates and integrates dynamic operating indicators with static hardware capabilities to form a complete resource profile for each computing node, covering its own computing power, dependent network conditions, and power supply attributes.
[0169] For example, the entropy weight method can be used to objectively calculate the weight vector of each evaluation index based on the data dispersion, and machine learning algorithms such as CART decision tree can be used to obtain the performance level index of each computing node.
[0170] For example, the higher the performance level index, the higher the computing power node should be selected, and the lower the performance level index, the lower the computing power node should be selected.
[0171] For example, the real-time status, performance score and alarm events after the power business task is processed are immediately pushed to the global scheduling server 130 to support immediate decision-making, and are stored in the "historical status database" of the resource database 220 in time sequence for trend analysis, capacity planning and backtracking.
[0172] Step S303: The global scheduling server 130 obtains the complete task tag based on the task attributes.
[0173] Step S304: The SDN controller cluster 170 obtains the network communication path composed of network nodes that satisfy the network sub-label from multiple network nodes.
[0174] Step S305: The power control cloud gateway 160 filters out multiple candidate power sets that satisfy the power sub-label from multiple power nodes.
[0175] Step S306: The global scheduling server 130 selects multiple candidate computing nodes that meet the basic sub-labels from the multiple computing nodes.
[0176] For example, the global scheduling server 130 can obtain pending power business tasks and complete task tags for power business tasks from a queue in the resource database.
[0177] The global scheduling server 130 first verifies the integrity of the power business task and confirms the digital signature. After successful verification, the global scheduling server 130 parses the complete task tags, extracts the dimensional information of each sub-tag, and constructs a structured "task profile" object. Simultaneously, the global scheduling server 130 pulls the latest global resource snapshot from the resource database 220. Based on the computing power requirements in the "task profile," it quickly filters out all computing power nodes whose static hardware capabilities meet the requirements from the resource snapshot. This is a coarse-grained filtering process that does not involve complex optimizations and only performs rapid "yes / no" judgments.
[0178] If a power service task must be completed by a single computing node (i.e., "allow_distributed": false in the complete task label), the static hardware capabilities of the single computing node must be sufficient to meet the computing power requirements. If a power service task can be executed in parallel across multiple computing nodes (i.e., "allow_distributed": true in the complete task label), the selection criteria are relaxed. It is not required that a single computing node meets the computing power requirements; only that the static architecture of the computing nodes supports parallel computing and that the real-time idle computing power of each computing node is greater than zero.
[0179] Step S307: The multi-objective optimization engine 140 obtains multiple candidate computing power sets based on the performance level indicators corresponding to multiple candidate computing power nodes, wherein each candidate computing power set includes one or more computing power nodes; it combines any one of the candidate computing power sets, any one of the network communication paths, and any one of the candidate power sets to obtain a candidate scheduling scheme; each candidate scheduling scheme includes a candidate computing power set, a network communication path, and a candidate power set; and it determines multiple optimal scheduling schemes from the multiple candidate scheduling schemes.
[0180] The multi-objective optimization engine can determine which computing nodes to combine and how much computing power to allocate to each node to jointly meet the needs of the entire power business task. This forms a "multiple candidate computing power set".
[0181] "Obtaining multiple candidate computing power sets based on the performance level indicators corresponding to multiple candidate computing power nodes" means that the higher the performance level indicator, the greater the probability that the candidate computing power node will be selected.
[0182] A candidate computing power set, a network communication path, and a candidate power set can be randomly selected to obtain a candidate scheduling scheme.
[0183] Step S308: The global scheduling server 130 determines a target scheduling scheme that satisfies N-1 security verification and cross-sectional power flow verification from multiple optimal scheduling schemes; generates a scheduling instruction for the power service task; and sends the scheduling instruction to the computing power nodes, network nodes, and power nodes included in the target scheduling scheme.
[0184] The target scheduling scheme is transformed into executable scheduling instructions and securely and reliably distributed to various resource execution nodes across the entire domain. This step is a crucial bridge connecting decision-making and execution, and its core objective is to ensure the confidentiality, integrity, and non-repudiation of the instructions, and to establish a preliminary execution feedback mechanism.
[0185] The global scheduling server 130 decomposes the final target scheduling scheme into a set of atomic operation instructions (i.e., scheduling instructions) and encapsulates them into a structured instruction set object. This object includes: computing power instructions, sent to specific computing power nodes, which include tasks loading, resource allocation, power consumption adjustment, etc.; network instructions, sent to the SDN controller cluster 170 and then to network nodes through the SDN controller cluster 170, which include path establishment / teardown, flow table distribution, bandwidth reservation, etc.; and power instructions, forwarded to power nodes through the power control cloud gateway 160, which include power generation adjustment, load switching, etc.
[0186] After encrypting and digitally signing the scheduling instructions using encryption algorithms, the instructions are distributed through a secure tunnel. Computing power instructions are directly sent point-to-point to the corresponding computing power nodes; network instructions are sent to the SDN controller cluster, which converts them into device-level configuration commands through an interface (such as OpenFlow) and then sends them to the corresponding network nodes; power instructions are converted into protocols through the power control cloud gateway 160 and then securely transmitted to the corresponding power nodes or power grid control system (EMS) for execution using power industry standard protocols (such as IEC 104 and IEC 61850).
[0187] Upon receiving the instruction, each node verifies the digital signature to confirm its legitimacy. If everything is correct, it executes local operations (such as starting a container, configuring flow tables, or adjusting power). Upon success or failure, it immediately sends an instruction execution acknowledgment (ACK / NACK) to the global state-aware cluster 150 and the global scheduling server 130. This triggers real-time monitoring of the execution status, allowing the system to monitor the actual execution effect of the scheduling instruction, thus forming a preliminary closed loop from "issuance" to "awareness."
[0188] In one optional implementation, when the global state-aware cluster performs the steps of obtaining the weight vectors corresponding to each evaluation index and the normalized values corresponding to each evaluation index based on the real-time computing power resource information of multiple computing power nodes under each evaluation index and the static hardware capabilities of multiple computing power nodes under each evaluation index, the specific steps include steps B1 to B6.
[0189] Step B1: Based on the real-time computing power resource information of multiple computing power nodes under various evaluation indicators, obtain the first weight vector corresponding to each of the evaluation indicators.
[0190] For example, the number of each evaluation index is n, the number of multiple computing nodes is m, and step B1 includes steps B11 to B14. Wherein, n and m are both positive integers greater than or equal to 1.
[0191] Step B11: Obtain real-time computing resource information for multiple computing nodes under various evaluation metrics. ,in, It refers to the first average value of real-time computing resource information of the i-th computing node under the current period of the j-th evaluation index;
[0192] For example, the cycle can be determined based on the actual situation; for instance, the cycle can be 15 minutes.
[0193] Step B12: If the j-th evaluation metric of the i-th computing node is a positive metric, then use the formula... The first standardized value is calculated. The positive indicator is positively correlated with the capability of the computing node. It refers to the maximum value of the j-th evaluation index for the i-th computing power node. It refers to the minimum value of the j-th evaluation index of the i-th computing power node;
[0194] A positive indicator means that the larger the indicator value, the better the computing power node; a negative indicator means that the larger the indicator value, the worse the computing power node.
[0195] Step B13: If the j-th evaluation metric of the i-th computing node is a negative metric, then use the formula... The first standardized value is calculated. The negative indicator is negatively correlated with the capability of the computing node.
[0196] It is understandable that the first standardized value is obtained. Calculate the proportion of the i-th computing power node under the j-th evaluation index. Then, the information entropy of the j-th evaluation index is calculated. , is used to measure the degree of uncertainty or data disorder of the j-th evaluation indicator. It is the difference coefficient. Obviously, the smaller the information entropy, the larger the difference coefficient, which means that the evaluation index provides more information and should be given a higher weight.
[0197] Finally, the difference coefficient and weight of the j-th evaluation indicator are calculated, resulting in the weight vector. The calculation formula.
[0198] Step B14: Using the formula The first weight vector of the j-th evaluation index is calculated.
[0199] The entropy weighting method is used to objectively calculate the weight vector of each evaluation indicator based on its dispersion across computing nodes. The greater the dispersion of an evaluation indicator, the richer its information content, and the higher its weight vector. This step is entirely driven by the data itself, eliminating subjective judgment and ensuring the fairness and scientific rigor of the weight vector.
[0200] Step B2: Based on the static hardware capabilities of multiple computing nodes under various evaluation indicators, obtain the second weight vector corresponding to each of the evaluation indicators.
[0201] For example, the number of each evaluation metric is n, and the number of multiple computing nodes is m. Step B2 includes steps B21 to B24.
[0202] Step B21: Obtain the static hardware capabilities of multiple computing nodes under various evaluation metrics. ,in, It refers to the second average value of the static hardware capability of the i-th computing node under the current period of the j-th evaluation index;
[0203] Step B22: If the j-th evaluation metric of the i-th computing node is a positive metric, then use the formula... The second standardized value was calculated. The positive indicator is positively correlated with the capability of the computing node. This refers to the maximum value of the j-th second preset evaluation index for the i-th computing power node. It refers to the minimum value of the j-th evaluation index of the i-th computing power node;
[0204] Step B23: If the j-th evaluation metric of the i-th computing node is a negative metric, then use the formula... The second standardized value was calculated. The negative indicator is negatively correlated with the capability of the computing node.
[0205] Step B24: Using the formula The second weight vector of the j-th evaluation index is calculated.
[0206] Step B3: Based on the first weight vector corresponding to each of the evaluation indicators and the real-time computing power resource information of multiple computing power nodes under each evaluation indicator, obtain the real-time weighted matrix corresponding to each of the multiple computing power nodes.
[0207] Step B4: Based on the second weight vector corresponding to each of the evaluation indicators and the static hardware capabilities of multiple computing nodes under each evaluation indicator, obtain the static weighting matrix corresponding to each of the multiple computing nodes.
[0208] Step B5: Based on the real-time weighted matrix and the static weighted matrix corresponding to each of the multiple computing power nodes, obtain the weighted comprehensive matrix corresponding to each of the multiple computing power nodes.
[0209] For example, for each computing power node, a real-time weighting matrix for the computing power node is determined. The static weighting matrix of the computing power nodes The sum of these is the weighted composite matrix of the computing power nodes.
[0210] Step B6: Based on the weighted comprehensive matrix corresponding to each of the multiple computing power nodes, determine the performance level labels corresponding to each of the multiple computing power nodes.
[0211] For example, step B6 includes the following step B61.
[0212] Step B61: Input the weighted comprehensive matrix corresponding to the multiple computing power nodes into the pre-constructed performance level prediction model, and output the performance level labels of the multiple computing power nodes through the performance level prediction model.
[0213] The performance level prediction model is trained by using the weighted comprehensive matrix of the sample computing power nodes as input and the labeled performance level tags of the sample computing power nodes as the training target.
[0214] The training process of the performance level prediction model is explained below.
[0215] After obtaining the weight vectors of each evaluation indicator, the standardized eigenvalues are weighted using the calculated weight vectors of each evaluation indicator to generate a new weighted composite matrix. This operation is equivalent to amplifying the influence of highly discriminative and important evaluation indicators in subsequent models. Among them, At the same time, the system calculates a weighted comprehensive score for each computing node, i.e. , that is, the sum of all weighted eigenvalues in the weighted comprehensive matrix. According to the distribution of the scores of all computing power nodes, the three - percentile method or the clustering algorithm (k = 3) is used to divide all computing power nodes into three categories. If S ,
[0218] , ,
[0219] ≥ the upper percentile, the performance level label Y i = "high"; if the lower percentile ≤ S i < the upper percentile, the performance level label Y i = "medium"; S i < the lower percentile, the performance level label Y i = "low"), and finally the training data D is obtained. The training data D contains multiple weighted comprehensive matrices and the performance level label set .
[0216] Taking multiple weighted comprehensive matrices as input features and the performance level label set as the training target, the CART decision tree model is trained to obtain the performance level prediction model.
[0217] It can be understood that the process of training the CART decision tree model is a recursive binary splitting process: starting from the root node, the algorithm greedily searches for all possible features and all possible splitting points on the current node, and calculates the "Gini impurity" reduction brought by each splitting method. Finally, it selects the feature and splitting point that can maximize the "purity" of the child nodes (that is, making the samples in the child nodes belong to the same category as much as possible), and splits the current node into two. This process is recursively carried out on each newly generated child node until the preset stopping conditions are met (such as too few samples in the node, the purity has reached the standard, or the tree depth has reached the limit, then stop). After the model training and optimization are completed, its final output is a series of clear if - then classification rules. Each rule corresponds to a path from the root node to a leaf node, for example, "IF weighted GPU computing power feature > 0.8 AND weighted network bandwidth feature > 0.6 THEN performance level label = high". These rules give the CART decision tree model strong interpretability, and the operation and maintenance personnel can clearly know the basis of the model decision without understanding the complex algorithm black box.
[0218] When a new computing power node needs to be evaluated, the system will apply exactly the same pre - processing process to it: obtain the weighted comprehensive matrix of this computing power node. Then input the weighted comprehensive matrix of this computing power node into the CART decision tree model, and the CART decision tree model will automatically traverse the decision tree according to a series of preset rule conditions, and finally output the performance level label of "high, medium, low" for this computing power node. This set of processes runs periodically, enabling the system to adapt to the dynamic changes of the resource pool and realizing intelligent operation and maintenance in the true sense.
[0219] This application couples the entropy weight method with the CART decision tree model, using the objective weighting results of the entropy weight method as the core input feature of the CART decision tree model. This makes the classification process dependent on both the distribution of the data itself and the importance of each evaluation indicator. The key to this method is that its evaluation dimensions cover computing power nodes, power nodes, and network nodes. The performance level labels output by the CART decision tree model are a quantitative evaluation of the comprehensive service capabilities of resource entities, completely replacing the traditional method that relies on subjective scoring based on expert experience. This enables fully automated performance evaluation of large-scale, heterogeneous computing power nodes, eliminating human bias, and ensuring that the evaluation results are reproducible and auditable. The entire evaluation system can be automatically re-executed periodically (e.g., daily) or triggered automatically (e.g., when the resource pool changes), with the weight vector and CART decision tree model updated accordingly to adapt to the dynamic changes in the resource pool. The generated performance level labels provide a direct and efficient decision-making basis for resource pre-screening and multi-objective optimization in subsequent scheduling processes.
[0220] In one alternative implementation, the multi-objective optimization engine includes steps C1 to C3 when performing the step of determining multiple optimal scheduling schemes from multiple candidate scheduling schemes.
[0221] Step C1: Determine multiple candidate scheduling schemes as the initial population.
[0222] Step C2: Input the initial population into the NSGA-II algorithm, with the optimization objectives of minimizing total execution cost, minimum total task latency, and minimum total task carbon emissions, to obtain multiple undetermined scheduling schemes; the total execution cost is the sum of the cost of computing nodes, power nodes, and network nodes; the total task latency includes the latency of computing nodes and the transmission latency of network communication paths; the total task carbon emissions refer to the total amount of carbon emissions during the execution of the power service task.
[0223] For example, Where fcost(x) is the total execution cost, fdelay(x) is the total task delay, fcarbon(x) is the total carbon emissions of the task, and W i W(t) (including Wcost(t), Wdelay(t), and Wcarbon(t)) is a dynamic weight function. It is not fixed but changes with time t (task execution time) and system state. The initial values of the dynamic weights are not arbitrarily set but are automatically generated by the system based on a set of predefined business rule bases and strategies. These rule bases are pre-configured by domain experts (such as power grid dispatchers and system administrators) to ensure the rationality and scientific nature of the initial weight values.
[0224] "The optimization objective is to minimize the total execution cost, the total task delay, and the total task carbon emissions" means to minimize the minimum of F(x,t).
[0225] This application introduces the concept of a dynamic weighting function, constructing it as a dynamic multi-objective optimization problem. The core innovation lies in the fact that the weights of the objective function are no longer fixed values, but rather dynamically adjusted functions based on time, grid conditions, and task priorities. For example, during peak electricity consumption periods, the weight vectors of total carbon emissions and total execution cost automatically increase, guiding the system to prioritize green and economical resources. The decision variable (x) represents a complete undetermined scheduling scheme, including not only resource and network path selection but also start time and power consumption patterns. Ultimately, the multi-objective optimization engine models the scheduling problem as a multi-objective optimization problem.
[0226] Step C3: Using the first preset weight of the total execution cost, the second preset weight of the total task delay, and the third preset weight of the total task carbon emissions as reference points in the three-dimensional space, and using multiple pending scheduling schemes as the initial population, multiple optimal scheduling schemes are obtained through NSGA-Ⅲ; the three-dimensional coordinates of the three-dimensional space are the total execution cost, the total task delay, and the total task carbon emissions, respectively.
[0227] For example, the first preset weight, the second preset weight, and the third preset weight can be determined based on the actual situation, and are not limited here.
[0228] The system injects knowledge from areas such as current time, current task priority, and power grid status into the NSGA-Ⅲ algorithm in the form of dynamic weights and reference points, and uses the NSGA-Ⅲ mechanism to guide the population to evolve towards the preferred region that the decision-maker is most interested in. Furthermore, NSGA-Ⅲ can better maintain the diversity of the population in the preferred region.
[0229] According to the dynamic weighting function W i(t) Define a set of reference points in the target space. These points indicate the preferred regions of the power service tasks. The multiple pending scheduling schemes obtained in the first phase are used as the initial population of NSGA-III. NSGA-III uses a reference point-based selection mechanism to replace the congestion metric in NSGA-II, thus better maintaining the diversity of the population in preference directions. Similarly, each individual is evaluated, and the calculated objective function value is normalized. For each normalized individual, its vertical distance to each reference line (the line connecting the origin (0,0,0) and a reference point) is calculated. The reference line with the shortest vertical distance is found, and the individual is associated with the reference point corresponding to this reference line. When constructing the next generation, individuals associated with sparse reference points (reference points with fewer associated individuals) are preferentially retained to ensure that the population evolves uniformly in all preference directions.
[0230] After each generation of evolution, the top Q individuals (Q elite solutions) are selected from the highest-ranking non-dominant individuals. For each selected elite solution x, its "neighborhood" needs to be defined, which is the set of all new solutions that can be obtained through small, reasonable modifications, because there are likely better solutions in the vicinity of a good solution. Subsequently, simulated annealing local search is used to reintegrate the high-quality new solutions found by the simulated annealing search into the population of the main evolutionary algorithm. These new solutions enrich the diversity of the population, provide better genes, and guide the entire population to evolve in a more positive direction.
[0231] For example, in addition to providing multiple optimal scheduling schemes, the multi-objective optimization engine also provides a detailed decision insight report for each optimal scheduling scheme. The decision insight report employs quantitative analysis methods, including quantitative trade-off analysis (presenting the specific values and relative advantages / disadvantages of each solution in various objectives such as total execution cost, total task latency, and total task carbon emissions in the form of radar charts or tables), sensitivity analysis (pointing out the solution's "vulnerabilities." For example, "This scheme has the lowest total execution cost, but it is very sensitive to network latency; if the latency increases by 2ms, the cost will increase sharply by 10%."), compromise cost calculation (clearly explaining the cost to other objectives that comes with choosing this optimal scheduling scheme. For example, "Choosing this low-carbon scheme requires you to pay 15% more than the lowest-cost scheme, but the latency only increases by 5ms."), and a constraint satisfaction overview (listing the satisfaction of all soft constraints, clearly showing the "compliance" of the scheme). These methods visually demonstrate the specific performance of each optimal scheduling scheme in various objectives, potential risks, and the costs incurred in choosing the optimal scheduling scheme. For example, the report might explicitly state that "choosing this lowest-cost option will result in carbon emissions that are 25% higher than the optimal environmentally friendly option." This transforms the output from cold data points into informative decision support material, greatly assisting subsequent decision rule engines or system administrators in making more informed and reliable choices.
[0232] For example, the resilience of the optimal scheduling scheme to uncertainties can be assessed to identify vulnerabilities, for instance, using perturbation analysis. By fine-tuning key input parameters (e.g., assuming a 5% increase in network latency or a 10% increase in electricity prices) and rerunning the process, the changes in total execution cost, total task latency, and total task carbon emissions can be observed. If (total task latency increases by 5%) THEN (total latency change rate > 10%), it can be concluded that the optimal scheduling scheme is highly sensitive to network latency. This generates the "potential risk points" conclusion in the decision insight report.
[0233] For example, the decision insight report may also include: the cost compromise rate, carbon emission compromise rate, and delay compromise rate of each optimal scheduling scheme;
[0234] The formula for calculating the cost compromise rate of the optimal scheduling scheme is as follows: (Total execution cost of the optimal scheduling scheme - Lowest total execution cost among all optimal scheduling schemes) / Lowest total execution cost among all optimal scheduling schemes.
[0235] The carbon emission compromise rate of the optimal scheduling scheme is calculated as follows: (Total carbon emissions of the task under the optimal scheduling scheme - Lowest total carbon emissions of the task among all optimal scheduling schemes) / Lowest total carbon emissions of the task among all optimal scheduling schemes.
[0236] The formula for calculating the delay compromise rate of the optimal scheduling scheme is as follows: (Total task delay of the optimal scheduling scheme - Lowest total task delay among all optimal scheduling schemes) / Lowest total task delay among all optimal scheduling schemes.
[0237] For example, when the global scheduling server 130 executes the process of determining a target scheduling scheme that satisfies the N-1 security check and cross-sectional power flow check from a plurality of said optimal scheduling schemes, it specifically includes the following steps D1 to D4.
[0238] Step D1: Determine the target calculation formula corresponding to the preset decision rules that the power business task satisfies from the correspondence between preset decision rules and calculation formulas.
[0239] For example, the correspondence between preset decision rules and calculation formulas is as follows:
[0240] Safety First Rule:
[0241] If any optimal scheduling scheme triggers a power grid safety alarm, this is returned by the power control cloud gateway 160 verification.
[0242] THEN The optimal scheduling scheme is immediately discarded and not considered.
[0243] Emergency Mission Rules:
[0244] IF task.priority == "CRITICAL" / / Emergency task, such as troubleshooting;
[0245] THEN Select the optimal scheduling scheme with the lowest total task latency from among several unverified optimal scheduling schemes.
[0246] The formula for calculating the scheme with the lowest total task latency is: Total Score = First Weight × (Total Execution Cost Score) + Second Weight × (Total Task Latency Score) + Third Weight × (Total Task Carbon Emission Score). The optimal scheduling scheme with the lowest total task latency is the optimal scheduling scheme with the highest total score.
[0247] Cost control rules:
[0248] IF task.priority == "LOW" OR time.is_night() == true / / Low-priority tasks or off-peak electricity periods;
[0249] THEN selects the optimal scheduling scheme with the lowest total execution cost from among several unverified optimal scheduling schemes.
[0250] The formula for calculating the solution with the lowest total execution cost is: Total Score = First Weight × (Total Execution Cost Score) + Second Weight × (Total Task Delay Score) + Third Weight × (Total Task Carbon Emission Score). Select the optimal scheduling solution with the highest total score.
[0251] Green priority rule:
[0252] IF task.has_label("power_preference", "green") / / Carbon emission requirement is marked with priority given to green electricity;
[0253] THEN selects the optimal scheduling scheme with the lowest total carbon emissions from the multiple unverified optimal scheduling schemes.
[0254] The formula for calculating the scheme with the lowest total carbon emissions is: Total Score = First Weight × (Total Execution Cost Score) + Second Weight × (Total Task Delay Score) + Third Weight × (Total Task Carbon Emission Score). The optimal scheduling scheme with the highest total score is selected.
[0255] It is understandable that the first weight, second weight, and third weight in the calculation formulas corresponding to different preset decision rules (such as emergency task rules, cost control rules, and green priority rules) are different.
[0256] For example, the total execution cost score = (power consumption of computing nodes × running time × electricity price during that period) + (network bandwidth rental cost) + (other possible costs).
[0257] For example, the total task latency score = (computation latency, based on task volume / node real-time computing power) + (network transmission latency, based on data volume / path available bandwidth).
[0258] For example, the total carbon emission score of a task = (power consumption of computing nodes × running time × carbon emission intensity of the power grid during that period).
[0259] Step D2: Based on the target calculation formula, obtain the scores corresponding to the multiple optimal scheduling schemes respectively.
[0260] Step D3: Obtain the unverified scheduling schemes to be tested from the multiple optimal scheduling schemes sorted from high to low according to the scores.
[0261] Step D4: Verify whether the scheduling scheme under test meets the N-1 security check and cross-sectional power flow check. If it does not meet the requirements, return to step D3. If it does meet the requirements, determine that the scheduling scheme under test is the target scheduling scheme.
[0262] The global scheduling server first performs a weighted scoring of multiple optimal scheduling schemes based on the target calculation formula, selecting the scheduling scheme with the highest comprehensive score to be tested. The target calculation formula can be adjusted according to factors such as task priority, time period strategy, and user preferences. For each scheduling scheme to be tested, the global scheduling server calls the API of the power grid energy management system (EMS) through the power control cloud gateway to perform the following key verifications:
[0263] N-1 Safety Verification: This simulates whether the system can remain stable and free from overload, voltage over-limit, or other problems after any critical component (such as a line or transformer) in the power grid fails and is taken out of service.
[0264] Cross-sectional power flow verification: Verify whether the implementation of the scheme will lead to power flow exceeding the limit in key transmission sections (such as inter-provincial tie lines and important channels).
[0265] If the proposed scheduling scheme passes all security checks, it is marked as a grid-safe feasible solution and used as the target scheduling scheme. If the proposed scheduling scheme fails the checks (e.g., causing cross-section overruns or violating N-1 security checks), the system automatically excludes it and selects the second-best-scoring proposed scheduling scheme for re-verification. This process is iterated until a target scheduling scheme that satisfies both multi-objective optimization and grid-safety checks is found. The final determined target scheduling scheme is converted by the global scheduling server into an executable atomic instruction set for downstream infrastructure (computing power, network, power agent) and enters the instruction issuance process.
[0266] In an optional implementation, the global scheduling server is further configured to: determine the target priority level of the power service task based on the complete task label of the power service task; and store the power service task in the queue corresponding to the target priority level.
[0267] For example, the queue is stored in a resource database.
[0268] For example, a rule base is pre-configured by domain experts (such as grid dispatchers and system administrators): each power business task corresponds to a complete task label, and priority can be determined based on the complete task label. The rule base is the core of priority decision-making.
[0269] The global scheduling server deploys a tag generation microservice, which embeds a rule engine (Drools). The rule engine matches extracted task attributes with a pre-configured rule base. A complete task tag for a power business task may simultaneously match multiple rules in the rule base. For example, if a VIP user submits a real-time fault diagnosis task, which needs to be executed by the dispatch control center, the complete task tag of this real-time fault diagnosis task may simultaneously trigger rule 001 (task.department=="dispatch control center"), rule 003 (task.type=="real-time fault diagnosis"), and rule 005 (user.level == "VIP") in Table 1. There are four ways to determine the priority of this power business task:
[0270] Scenario 1: If the action label.set("priority", "CRITICAL or HIGH") is matched, which means the task type is an emergency task, then the priority of the power business task is directly determined as the first priority.
[0271] Scenario 2: If a rule in the rule base is matched, and only one rule with a trigger condition is matched, then the priority of the power business task is the priority specified by the matched rule.
[0272] Scenario 3: If a rule in the rule base is matched and multiple triggering conditions are matched, then the rule with the highest priority among the multiple rules will be taken as the priority of the power business task.
[0273] Scenario 4: If no rule is matched in the rule base, the default rule is triggered, and the priority of the power business task is set to the third priority.
[0274] For example, the complete task tag of a power business task can be written into the "task tag library" in the resource database for querying by the global scheduling server 130, and can also be pushed to the global scheduling server 130 through a message queue.
[0275] The rule base is explained below in conjunction with the four scenarios mentioned above, as detailed in Table 1.
[0276] Table 1
[0277]
[0278] In Table 1, priorities are sorted from highest to lowest as follows: First priority > Second priority > Third priority > Fourth priority.
[0279] Understandably, the power service task and its complete tag can be sent as a message body to the queue management service via the API (Application Programming Interface) of the global scheduling server. Upon receiving the message body, the queue management service parses the complete task tag and reads the value of the priority field. Based on the parsed priority, the queue management service stores the power service task in the corresponding queue.
[0280] For example, the queues include the Q_CRITICAL queue (which stores power service tasks of the first priority), the Q_HIGH queue (which stores power service tasks of the second priority), the Q_MEDIUM queue (which stores power service tasks of the third priority), and the Q_LOW queue (which stores power service tasks of the fourth priority).
[0281] The global scheduling server continuously attempts to consume power service tasks from the queues. Its consumption strategy is a strict priority round-robin: it first checks the Q_CRITICAL queue, and if a power service task exists, it retrieves it sequentially until the queue is empty. Only when the Q_CRITICAL queue is empty does it check the Q_HIGH queue, and so on, finally processing the Q_LOW queue. This strategy ensures that high-priority power service tasks are always processed before low-priority ones.
[0282] It is understandable that there may be conflicts between minimizing total execution cost, minimizing total latency, minimizing total carbon emissions, and security requirements for power service tasks. Total execution cost is the sum of the costs of computing nodes, power nodes, and network nodes; total latency includes the latency of computing nodes and the transmission latency of network communication paths; total carbon emissions refer to the total amount of carbon emissions during the execution of the power service task. An example is provided below for illustration.
[0283] There is a conflict between minimizing total task latency and minimizing total task carbon emissions. For example, network nodes that meet the minimum latency requirement may be powered by fossil fuels (high carbon emissions); while computing nodes that use green electricity (such as wind power and solar power) may be geographically remote and have high network latency.
[0284] For example, a power service task requires real-time fault diagnosis with a latency of <50ms and the desire to use green electricity as much as possible. Network path A (low latency): 35ms latency, but the carbon emission intensity of the connected power nodes is 400 gCO2 / kWh (mainly thermal power); Network path B (green path): 65ms latency, but the carbon emission intensity of the connected power nodes is 50gCO2 / kWh (mainly green electricity). Network path B cannot meet the latency requirement, and network path A cannot meet the green preference. The two are in direct conflict.
[0285] There is a conflict between minimizing total task latency and minimizing total execution cost. For example, high-performance computing nodes (such as the latest GPU servers) are usually located in core data centers with high electricity prices; while low-cost electricity may be available at night or at edge sites, but their computing performance or network conditions may not meet the low latency requirements.
[0286] For example, a power business task involves training an AI model, requiring a high-performance GPU while controlling costs. Computing node A (lowest total latency): equipped with an A100 GPU, offering strong computing power and fast task completion, but the electricity price in its area is 1.2 yuan / kWh; Computing node B (lowest total execution cost): equipped with an older GPU, offering weaker computing power and slower task completion, but its area operates during off-peak electricity hours at night, with an electricity price of 0.3 yuan / kWh. It is impossible to simultaneously achieve both the lowest total latency and the lowest total execution cost.
[0287] There is a conflict between security requirements and optimal resources. For example, power business tasks require data to remain within the province / city, but there may not be optimal resources within the province / city that meet all performance, cost, and environmental requirements.
[0288] For example, a power service task involving processing user privacy data, such as `label.set("data_security", "high");label.set("zone", "secure")`, must be scheduled to a high-security zone, preventing data from leaving the zone and requiring high computing power. Within the security zone, there is only one computing node, whose GPU is nearing full capacity, and the estimated completion time for the power service task is very long. Outside the security zone, there are ample and idle high-performance computing nodes, but their use is prohibited due to data security rules, creating a conflict between data security and optimal resource availability.
[0289] Therefore, this application needs to resolve the aforementioned conflicts in the process of determining the complete task label for power business tasks. The conflict resolution method is as follows:
[0290] Step A1: If the task type in the task attribute is an emergency task, determine that the priority with the lowest total delay is higher than the priority with the lowest total execution cost and the priority with the lowest total carbon emissions.
[0291] Assuming that the task type in the task attribute is represented by priority, then if priority == "CRITICAL" (urgent task, such as real-time control, fault diagnosis), then the task with the lowest total latency is retained, and the total execution cost and total carbon emissions of the task can be appropriately relaxed.
[0292] Step A2: If the task type in the task attributes is a non-urgent task and the carbon emission requirement in the task attributes does not indicate green electricity, determine that the priority with the lowest total execution cost is higher than the priority with the lowest total delay and higher than the priority with the lowest total carbon emissions.
[0293] For example, non-urgent tasks refer to medium-urgent tasks (priority=="MEDIUM") or low-urgent tasks (priority=="LOW").
[0294] For example, if priority == "LOW" (for non-urgent tasks, such as batch data processing or historical analysis), then the lowest total execution cost is retained, and longer total task latency and higher total task carbon emissions can be accepted.
[0295] For example, if time.now().hour >= 23 && time.now().hour <= 5; and ("cost_preference", "true"), then the lowest total execution cost is retained, which can accept a longer total task delay and a higher total task carbon emission.
[0296] Step A3: If the task type in the task attributes is a non-urgent task and the carbon emission requirement in the task attributes is marked as green electricity, determine that the priority with the lowest total carbon emissions is higher than the priority with the lowest total execution cost and higher than the priority with the lowest total delay.
[0297] For example, if priority == "MEDIUM" (medium urgency task) AND power_preference == "green" (green electricity), the task with the lowest total carbon emissions can be reserved, and the total delay and total execution cost of the task can be appropriately relaxed.
[0298] For example, if time.now().hour (today's time) >= 8 && time.now().hour <= 18; and ("power_preference", "green") and (max_carbon", "≤200 gCO2 / kWh"), then the task with the lowest total carbon emissions can be retained, and the total task delay and total execution cost can be appropriately relaxed.
[0299] Understandably, if the automated rules cannot resolve the conflict, a conflict alarm will be generated, and the power business task will be suspended and submitted to the dispatch and control terminal 210 for manual adjudication and tag adjustment by the administrator.
[0300] In one optional implementation, monitoring agent modules are deployed on computing nodes, network nodes, and power nodes. These modules can collect fine-grained metrics at high frequencies. The computing nodes, network nodes, and power nodes asynchronously push the collected data to a queue stored in the resource database. The streaming data processing unit in the global state-aware cluster 150 consumes the data from the queue in real time, performing rapid aggregation, computation, and anomaly detection. The processed results are stored in the time-series data table of the resource database 220. The monitoring rule engine within the global state-aware cluster 150 determines whether the metrics are abnormal in real time and generates alarm events.
[0301] In one optional implementation, the global scheduling server 130 receives alarm events. The global scheduling server 130 retrieves monitoring data and historical logs, and runs a lightweight diagnostic model (such as a rule-based or machine learning classifier) to determine the type and severity of the anomaly. The results of the determination include node-level failures (such as physical server crashes or power outages), task-level failures (such as a task process crashing due to an internal error, but the node itself is healthy), localized network failures (such as a network path deteriorating in quality, but other paths are normal), and performance failures (such as power service tasks running, but with performance far below expectations, or computing power being preempted), etc. The global scheduling server 130 makes decisions based on the diagnostic results, as follows:
[0302] IF exception type == "node-level failure" THEN decision = "immediate rescheduling";
[0303] IF exception type == "task-level failure" AND retries < 3 THEN decision = "reboot in place";
[0304] IF Exception Type == "Partial Network Failure" THEN Decision = "Trigger Network Repath (Executed by SDN Controller Cluster 170)";
[0305] IF exception type == "performance not up to standard" THEN decision = "rescheduling immediately".
[0306] If the decision is "immediate rescheduling," the global scheduling server 130 will generate a new task object (or modify the original power service task). The new task object inherits the complete task tag of the original power service task. Fault information is injected into the new task object. An avoidance list is appended to the new task object, containing all resource IDs diagnosed as faults (e.g., avoid_nodes: ["Node-123"], avoid_paths: ["Path-A"]). The priority of the rescheduled task is usually automatically increased to ensure the system can respond quickly to faults. If the power service task is stateful (e.g., a simulation training run in progress), the system will attempt to recover task data from the previous checkpoint, and the rescheduling request will carry the storage path information of the latest state data.
[0307] A checkpoint is an action that periodically or strategically saves the current memory state (or disk state) of a power service task to persistent storage (such as a distributed file system or object storage). The resulting file is called a "checkpoint file".
[0308] The rescheduling request is sent back to the entry point of the global scheduling server 130, entering a new round of complete scheduling process. Since the new task object has an avoidance list attached, when filtering candidate schemes (such as pending scheduling schemes, optimal scheduling schemes), the multi-objective optimization engine can avoid failed candidate schemes according to the avoidance list, thereby allocating appropriate healthy resources to the new task object instance; the multi-objective optimization engine also adds "migration cost", that is, try to select healthy nodes that are geographically close to the faulty node to reduce the network overhead of data migration; after the new computing power agent node starts the task, it will register new monitoring items with the global state awareness cluster 150; the global state awareness cluster 150 will switch the monitoring node from the old, faulty node to the new healthy node.
[0309] In one optional implementation, if the power service task is completed, the global state awareness cluster 150 confirms that the task process has exited and that the resource occupancy reported by the relevant nodes has reached zero. Based on this, the global scheduling server 130 generates a resource release command, which is then distributed to each node for execution via the SDN controller cluster 170 and the security gateway. Each node executes the resource release command, releases physical resources, and sends an acknowledgment signal. The global resource pool status in the resource database 220 is updated to "idle." Simultaneously, the system extracts detailed records of the task's entire lifecycle from the resource database 220, generates a final report, stores it in the resource database 220 and distributed storage for archiving, and pushes it to the user terminal 110.
[0310] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0311] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network node, etc.) to execute the methods described in the various embodiments of this application.
[0312] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0313] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A multi-objective coupled scheduling system for computing power resources and power resources, characterized in that, include: The user terminal is used to obtain power business tasks and task attributes. The task attributes include at least one of the following: computing power required to execute the power business task, maximum tolerable latency, expected completion time, budgeted electricity cost, carbon emission requirements, security requirements, and task type. A global scheduling server is used to obtain a complete task label based on the task attributes. The complete task label includes a basic sub-label representing the processing capability of the required computing power node, a power sub-label representing the economic environment requirements of the required power node, and a network sub-label representing the communication capability of the required network node. An SDN controller cluster is used to obtain a network communication path composed of network nodes that satisfy the network sub-label from multiple network nodes; A power control cloud gateway is used to filter out multiple candidate power sets that satisfy the power sub-label from multiple power nodes, wherein the candidate power set includes one or more power nodes; A global scheduling server is used to filter multiple candidate computing nodes that meet the basic sub-labels from multiple computing power nodes; A multi-objective optimization engine is used to obtain multiple candidate computing power sets, wherein the candidate computing power sets include one or more computing power nodes; A candidate scheduling scheme is obtained by combining any of the candidate computing power sets, any of the network communication paths, and any of the candidate power sets; multiple optimal scheduling schemes are determined from the multiple candidate scheduling schemes. The global scheduling server is also used to determine a target scheduling scheme that satisfies N-1 security check and cross-sectional power flow check from multiple optimal scheduling schemes.
2. The multi-objective coupled scheduling system for computing power resources and power resources according to claim 1, characterized in that, Also includes: A global state-aware cluster is used to acquire real-time computing resource information of multiple computing nodes under various evaluation indicators, communication information of multiple network nodes, economic environment information of multiple power nodes, and static hardware capabilities of the multiple computing nodes under various evaluation indicators. Based on the real-time computing resource information and static hardware capabilities of the multiple computing nodes under various evaluation indicators, a weight vector corresponding to each evaluation indicator and a normalized value corresponding to each evaluation indicator are obtained. Based on the weight vector, real-time computing resource information and static hardware capabilities, the performance level indicators corresponding to the multiple computing nodes are determined.
3. The multi-objective coupled scheduling system for computing power resources and power resources according to claim 2, characterized in that, When the multi-objective optimization engine acquires multiple candidate computing power sets, it specifically includes: Based on the performance level indicators corresponding to multiple candidate computing power nodes, multiple candidate computing power sets are obtained, and the candidate computing power sets include one or more computing power nodes.
4. The multi-objective coupled scheduling system for computing power resources and power resources according to claim 2, characterized in that, When the global state-aware cluster executes the steps of obtaining the weight vectors corresponding to each evaluation index and the normalized values of the multiple computing nodes for each evaluation index based on the real-time computing resource information of the multiple computing nodes under each evaluation index and the static hardware capabilities of the multiple computing nodes under each evaluation index, the steps include: Based on the real-time computing power resource information of multiple computing power nodes under various evaluation indicators, the first weight vector corresponding to each evaluation indicator is obtained. Based on the static hardware capabilities of multiple computing nodes under various evaluation indicators, the second weight vector corresponding to each evaluation indicator is obtained. Based on the first weight vector corresponding to each evaluation index and the real-time computing power resource information of multiple computing power nodes under each evaluation index, the real-time weighting matrix corresponding to multiple computing power nodes is obtained. Based on the second weight vector corresponding to each of the evaluation indicators and the static hardware capabilities of multiple computing nodes under each evaluation indicator, obtain the static weighting matrix corresponding to each of the multiple computing nodes. Based on the real-time weighted matrix and the static weighted matrix corresponding to each of the multiple computing power nodes, the weighted comprehensive matrix corresponding to each of the multiple computing power nodes is obtained. Based on the weighted comprehensive matrix corresponding to each of the multiple computing power nodes, the performance level labels corresponding to each of the multiple computing power nodes are determined.
5. The multi-objective coupled scheduling system for computing power resources and power resources according to claim 4, characterized in that, The number of evaluation indicators is n, and the number of computing nodes is m. The step of obtaining the first weight vector corresponding to each evaluation indicator based on the real-time computing resource information of the multiple computing nodes under each evaluation indicator includes: Obtain real-time computing resource information of multiple computing nodes under various evaluation indicators. ,in, It refers to the first average value of real-time computing resource information of the i-th computing node under the current period of the j-th evaluation index; If the j-th evaluation metric of the i-th computing node is a positive metric, then by formula... The first standardized value is calculated. The positive indicator is positively correlated with the capability of the computing node. It refers to the maximum value of the j-th evaluation index for the i-th computing power node. It refers to the minimum value of the j-th evaluation index of the i-th computing power node; If the j-th evaluation metric of the i-th computing power node is negative, then by formula... The first standardized value is calculated. The negative indicator is negatively correlated with the capability of the computing node. Through formula The first weight vector of the j-th evaluation index is calculated.
6. The multi-objective coupled scheduling system for computing power resources and power resources according to claim 4, characterized in that, The number of evaluation indicators is n, and the number of computing nodes is m. Based on the static hardware capabilities of the computing nodes under each evaluation indicator, the second weight vector corresponding to each evaluation indicator is obtained as follows: Obtain the static hardware capabilities of multiple computing nodes under various evaluation metrics. ,in, It refers to the second average value of the static hardware capability of the i-th computing node under the current period of the j-th evaluation index; If the j-th evaluation metric of the i-th computing node is a positive metric, then by formula... The second standardized value was calculated. The positive indicator is positively correlated with the capability of the computing node. This refers to the maximum value of the j-th second preset evaluation index for the i-th computing power node. It refers to the minimum value of the j-th evaluation index of the i-th computing power node; If the j-th evaluation metric of the i-th computing power node is negative, then by formula... The second standardized value was calculated. The negative indicator is negatively correlated with the capability of the computing node. Through formula The second weight vector of the j-th evaluation index is calculated.
7. The multi-objective coupled scheduling system for computing power resources and power resources according to any one of claims 4 to 6, characterized in that, The step of obtaining the weighted composite matrix corresponding to multiple computing power nodes based on the real-time weighted matrix and the static weighted matrix corresponding to multiple computing power nodes includes: For each computing power node, determine the real-time weighting matrix of that computing power node. The static weighting matrix of the computing power nodes The sum of these is the weighted composite matrix of the computing power nodes.
8. The multi-objective coupled scheduling system for computing power resources and power resources according to claim 4, characterized in that, The step of determining the performance level labels corresponding to multiple computing power nodes based on the weighted comprehensive matrix corresponding to each computing power node includes: The weighted comprehensive matrix corresponding to multiple computing power nodes is input into the pre-constructed performance level prediction model, and the performance level prediction model outputs the performance level labels of multiple computing power nodes. The performance level prediction model is trained by using the weighted comprehensive matrix of the sample computing power nodes as input and the labeled performance level tags of the sample computing power nodes as the training target.
9. The multi-objective coupled scheduling system for computing power resources and power resources according to claim 1, characterized in that, After determining the target scheduling scheme, the global scheduling server is further configured to: Generate scheduling instructions for the power service task; The scheduling instructions are sent to the computing nodes, network nodes, and power nodes included in the target scheduling scheme.
10. The multi-objective coupled scheduling system for computing power resources and power resources according to claim 1, characterized in that, When the multi-objective optimization engine performs the step of determining multiple optimal scheduling schemes from multiple candidate scheduling schemes, it is specifically used for: Multiple candidate scheduling schemes were identified as the initial population; The initial population is input into the NSGA-II algorithm, and multiple undetermined scheduling schemes are obtained with the optimization objectives of minimizing total execution cost, minimizing total task latency, and minimizing total task carbon emissions. The total execution cost is the sum of the cost of computing nodes, the cost of power nodes, and the cost of network nodes; the total task latency includes the latency of computing nodes and the transmission latency of network communication paths; the total task carbon emissions refer to the total amount of carbon emissions during the execution of the power service task. Using the first preset weight of the total execution cost, the second preset weight of the total task delay, and the third preset weight of the total task carbon emissions as reference points in a three-dimensional space, and taking multiple undetermined scheduling schemes as an initial population, multiple optimal scheduling schemes are obtained through NSGA-Ⅲ; the three-dimensional coordinates of the three-dimensional space are the total execution cost, the total task delay, and the total task carbon emissions, respectively.
11. The multi-objective coupled scheduling system for computing power resources and power resources according to claim 1, characterized in that, When the global scheduling server performs the step of determining a target scheduling scheme that satisfies N-1 security check and cross-sectional power flow check from multiple optimal scheduling schemes, it includes: From the correspondence between preset decision rules and calculation formulas, determine the target calculation formula corresponding to the preset decision rules that the power business task satisfies; Based on the target calculation formula, scores corresponding to multiple optimal scheduling schemes are obtained respectively; Obtain the unverified scheduling scheme from the multiple optimal scheduling schemes sorted by the score from high to low; If the scheduling scheme under test satisfies the N-1 security check and the cross-sectional power flow check, the scheduling scheme under test is determined to be the target scheduling scheme. If the proposed scheduling scheme does not meet the N-1 security check or the cross-sectional power flow check, return an unchecked proposed scheduling scheme from the multiple optimal scheduling schemes sorted by the scores from high to low.
12. The multi-objective coupled scheduling system for computing power resources and power resources according to claim 1, characterized in that, When the global scheduling server performs the task tag retrieval based on the task attributes, it includes: If the task type in the task attributes is an emergency task, the priority with the lowest total delay is determined to be higher than the priority with the lowest total execution cost and higher than the priority with the lowest total carbon emissions; the total execution cost is the sum of the cost of the computing node, the cost of the power node, and the cost of the network node; the total delay includes the delay of the computing node and the transmission delay of the network communication path; the total carbon emissions refer to the total amount of carbon emissions during the execution of the power service task; If the task type in the task attributes is a non-urgent task and the carbon emission requirement in the task attributes is not marked as green electricity, the priority of the task with the lowest total execution cost is determined to be higher than the priority of the task with the lowest total delay and higher than the priority of the task with the lowest total carbon emissions. If the task type in the task attributes is a non-urgent task and the carbon emission requirement in the task attributes is marked as green electricity, the priority of the task with the lowest total carbon emissions is determined to be higher than the priority with the lowest total execution cost and higher than the priority with the lowest total delay.