A distributed off-site multi-available area computing power scheduling method

By using a dynamic resource view and adaptive cost function under a master-slave architecture, the problems of precise matching and scalability in distributed computing power scheduling are solved, achieving low-latency, high-efficiency cross-domain computing power scheduling to adapt to different task requirements.

CN121542049BActive Publication Date: 2026-05-08GUIZHOU POLYMER COMPUTING SERVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUIZHOU POLYMER COMPUTING SERVICE CO LTD
Filing Date
2026-01-16
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing distributed computing power scheduling technologies struggle to achieve accurate evaluation and optimal matching during cross-domain scheduling, lack scalability, and are ill-suited to cope with ultra-large-scale and dynamically changing computing power environments. Furthermore, there is a lack of mature solutions for balancing bandwidth costs and efficiency in cross-regional scheduling.

Method used

By adopting a master-slave architecture approach, a dynamically updated global resource view is constructed through centralized collaboration and distributed execution. Combined with an adaptive cost function and a comprehensive load index, computing power scheduling with low latency, high availability, and high resource utilization is achieved.

Benefits of technology

It enables real-time and accurate perception of resource utilization and load status in each availability zone, reduces network transmission performance bottlenecks, improves system throughput and resource utilization, enhances the system's ability to cope with node or network failures, and meets the flexibility and adaptability requirements of different tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542049B_ABST
    Figure CN121542049B_ABST
Patent Text Reader

Abstract

This invention discloses a distributed, multi-availability-zone computing power scheduling method, relating to the field of distributed management technology. The method includes the following steps: a management node collects the status information of its local availability zone and reports it to a master service node to construct a dynamic global resource view; the master service node receives computing tasks, parses and breaks them down into subtasks; an adaptive cost function is constructed to calculate the cost of scheduling each subtask to different availability zones, and the lowest cost is selected as the target availability zone; the subtasks are distributed to the management nodes of the target availability zones, which then assign them to computing nodes for execution; the results are aggregated and returned to the master service node; simultaneously, the availability zone status is monitored, and tasks in faulty availability zones are rescheduled. This invention achieves low-latency, high-availability, and high-resource-utilization computing power scheduling by combining centralized coordination with distributed execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed management technology, and specifically to a distributed multi-availability zone computing power scheduling method. Background Technology

[0002] With the full arrival of the digital economy era, the total amount of data has exploded, and computing power has become a core productive force. However, computing resources are extremely unevenly distributed across regions, levels, and entities: the eastern region has abundant data but high computing costs, while the western region has abundant computing resources but insufficient demand. In recent years, distributed computing scheduling technology has developed rapidly, moving from initial intra-cluster scheduling to cross-regional and cross-entity global scheduling. Its core goal is to transform from traditional resource-based supply to task-based service, ultimately achieving ubiquitous accessibility and efficient utilization of computing resources.

[0003] Efficient and intelligent distributed computing power scheduling is of great strategic significance for enhancing digital competitiveness, promoting coordinated development between eastern and western regions, and driving the digital transformation of industries. It is not only key to solving the structural mismatch of computing power resources and reducing the total social computing power cost, but also the digital foundation supporting the development of cutting-edge applications such as AI large-scale model training, telemedicine, and smart cities.

[0004] Despite significant progress in distributed computing power scheduling, existing technologies still face numerous challenges and shortcomings in practical application, primarily in the following aspects: The strong heterogeneity of computing resources among different computing power entities makes it difficult to accurately assess and compare computing power during cross-domain scheduling, hindering optimal matching. Distributed scheduling systems suffer from scalability bottlenecks; each additional node imposes extra load on the master node, restricting the construction of large-scale computing power pools. Existing scheduling algorithms still have room for improvement in their intelligence level when dealing with ultra-large-scale, dynamically changing computing power environments. Furthermore, there is a lack of mature business models and product solutions for balancing economic factors such as bandwidth costs associated with cross-regional scheduling with scheduling efficiency.

[0005] For example, Chinese Patent No. CN117472549B discloses a distributed computing power scheduling system based on AIGC. The system includes a node acquisition module, a task management module, a computing power scheduling module, a monitoring module, and a user service module. The node acquisition module is used to acquire available distributed computing power nodes within the system. The task management module is used to acquire computing power task information and decompose the computing power tasks into multiple sub-tasks for execution. The computing power scheduling module is used to allocate computing power tasks to each computing power node by combining historical and real-time data. The monitoring module is used to monitor the task execution status of the computing power nodes. The user service module is used to provide a user interface to facilitate information interaction between the user and the system. This invention optimizes system resource utilization and response speed by combining real-time and historical data.

[0006] For example, Chinese patent CN115051988B discloses a fusion scheduling system based on distributed computing power, mainly involving the field of computing power scheduling technology. It aims to solve technical problems such as the time-consuming and labor-intensive nature of existing computing power scheduling schemes, data redundancy, idle computing power in some areas, and waste of computing power resources. The system includes: a registration center module for obtaining the list of algorithms uploaded by computing power nodes; a scheduling center module for receiving and allocating computing power requirements and algorithm deployment requests; sequentially distributing computing power requirements to each computing power node to be detected; a weight center module for determining the computing power nodes to be detected; determining the priority of each computing power node to be detected; and a deployment center module for obtaining available computing power nodes or preset computing power nodes; and deploying algorithms on any node among the available computing power nodes or preset computing power nodes. This application, through the above method, achieves overall avoidance of computing power resource waste and more flexible allocation and use of computing power resources. Summary of the Invention

[0007] This invention aims to address the problems of the prior art by providing a distributed, multi-availability zone computing power scheduling method based on a master-slave architecture that supports dynamic perception and intelligent decision-making. This method achieves low latency, high availability, and high resource utilization computing power scheduling through centralized collaboration and distributed execution.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0009] A distributed, multi-availability-zone computing power scheduling method is proposed, running in a system consisting of a master service node, a central database, and N availability zones connected via the Internet. Each availability zone contains a management node and M computing nodes. The method includes the following steps:

[0010] Step S1: The management node of the availability zone periodically collects the status information of the availability zone and reports it to the main service node. Based on this, the main service node builds and maintains a dynamically updated global resource view.

[0011] Step S2: The main service node receives the computing task submitted by the user and parses and splits the task into subtasks.

[0012] Step S3: For each subtask, the main service node calculates the cost of scheduling it to each availability zone based on the global resource view by constructing an adaptive cost function, and selects the availability zone with the lowest cost as the target availability zone.

[0013] Step S4: The main service node distributes the subtasks to the management nodes of their target availability zones, and each management node then assigns the tasks to the compute nodes within its domain for execution.

[0014] Step S5: The compute node returns the results, which are then aggregated by the management node and finally reported to the main service node.

[0015] Step S6: The main service node monitors the status of each availability zone and reschedules the subtasks of the availability zone that has failed.

[0016] Furthermore, in step S1, the global resource view consists of a status tuple for each availability zone, and the status tuple includes: available computing power vector, comprehensive load index, network latency, effective network bandwidth, and status flag.

[0017] The status flags include: online, offline, congested, and error.

[0018] Furthermore, the comprehensive load index is calculated by weighted summation, which includes the following four items: the first item is the product of the average CPU utilization and the CPU weight coefficient; the second item is the product of the average memory utilization and the memory weight coefficient; the third item is the product of the average I / O utilization and the I / O weight coefficient; and the fourth item is the product of the task queue relative saturation and the queue weight coefficient. The task queue relative saturation is obtained by dividing the real-time length of the local task queue on the management node by the maximum queue length.

[0019] Furthermore, in step S3, the adaptive cost function consists of two parts: the first part is the weighted sum of the time estimated cost, network estimated cost, and economic estimated cost multiplied by their respective weight coefficients; the second part is the product of the load penalty factor and the comprehensive load index of the target availability zone.

[0020] Furthermore, the method for calculating the network's estimated cost in the adaptive cost function includes:

[0021] Calculate the input data transmission overhead, which is determined by dividing the subtask input data size by the target availability zone inbound effective bandwidth and then adding the inbound network latency value.

[0022] Calculate the output data transmission overhead, which is determined by dividing the subtask's estimated output data size by the target availability zone's outbound effective bandwidth and then adding the outbound network latency value.

[0023] The estimated network cost is obtained by calculating the sum of the input data transmission overhead and the output data transmission overhead.

[0024] Furthermore, in step S4, the process of the management node assigning the task to a computing node within its domain for execution specifically includes the following steps:

[0025] The management node receives subtasks from the master service node and adds them to its local task queue;

[0026] The local management node calculates a comprehensive mismatch score for each subtask and each compute node, wherein the comprehensive mismatch score includes a resource fit term and a cache affinity term;

[0027] The local management node selects the computing node with the lowest overall mismatch score as the target node and assigns the subtask to the selected target node for execution.

[0028] Furthermore, step S6 specifically includes the following steps:

[0029] The master service node monitors the status of each management node through a heartbeat mechanism;

[0030] If the primary service node loses heartbeat signals from a management node for a consecutive number of times reaching the fault determination threshold, it is determined that the availability zone to which it belongs has failed, and the status of the availability zone is marked as offline, and it is also excluded from the global resource view.

[0031] The primary service node identifies all tasks that have been distributed to the faulty availability zone but have not been confirmed as completed, puts them back into the global task queue, and reschedules them to other healthy availability zones.

[0032] The fault determination threshold is dynamically set based on the heartbeat interval and the maximum tolerable fault recovery time target of the system.

[0033] Furthermore, the management node communicates with the computing nodes within its jurisdiction's availability zones via a remote management connection based on the Secure Shell protocol. This connection is used to execute commands and transmit file data via a secure file transfer protocol. The main service node, the management nodes of each availability zone, and the central database are interconnected via a wide area network (WAN) communication link. This WAN communication link is built on the public Internet or a private dedicated network and uses transport layer security protocols or secure socket layer protocols for communication encryption.

[0034] A storage medium, characterized in that the storage medium stores instructions, which, when read by a computer, cause the computer to execute a distributed multi-availability zone computing power scheduling method.

[0035] An electronic device, characterized in that it includes a processor and a storage medium, wherein the processor executes instructions in the storage medium.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0037] 1. By constructing a dynamically updated global resource view and designing a comprehensive load index, this invention can perceive the resource utilization and load status of each availability zone in real time and accurately, effectively avoiding overload of a single availability zone and improving overall resource utilization and system throughput.

[0038] 2. This invention innovatively quantifies network transmission overhead and incorporates it into the cost model during scheduling decisions, fully considering the size of input and output data and network conditions. This reduces performance bottlenecks caused by cross-network transmission, making it particularly suitable for data-intensive tasks and effectively reducing the overall execution time of the task.

[0039] 3. The cost function weights constructed in this invention can be initially set and dynamically adjusted according to user task priorities and business constraints. This enables the scheduling system to meet both the extreme speed requirements of computationally intensive tasks and the economic requirements of cost-sensitive tasks, providing high flexibility and adaptability.

[0040] 4. This invention continuously monitors the availability zone status. Once an availability zone anomaly is detected, it immediately triggers the rescheduling of affected tasks, seamlessly migrating them to a healthy availability zone for execution. This mechanism greatly enhances the system's ability to cope with node or network failures, ensuring the completion of complex tasks that run for extended periods. Attached Figure Description

[0041] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0042] Figure 1 This is a flowchart illustrating an embodiment of the present invention;

[0043] Figure 2 This is a schematic diagram of the distributed architecture of an embodiment of the present invention. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0045] like Figure 1 As shown, a distributed multi-availability zone computing power scheduling method operates in a system consisting of a master service node, a central database, and N availability zones connected via the Internet. Each availability zone contains one management node and M computing nodes. The method includes the following steps:

[0046] Step S1: The management node of the availability zone periodically collects the status information of the availability zone and reports it to the main service node. Based on this, the main service node builds and maintains a dynamically updated global resource view.

[0047] Step S2: The main service node receives the computing task submitted by the user and parses and splits the task into subtasks.

[0048] Step S3: For each subtask, the main service node calculates the cost of scheduling it to each availability zone based on the global resource view by constructing an adaptive cost function, and selects the availability zone with the lowest cost as the target availability zone.

[0049] Step S4: The main service node distributes the subtasks to the management nodes of their target availability zones, and each management node then assigns the tasks to the compute nodes within its domain for execution.

[0050] Step S5: The compute node returns the results, which are then aggregated by the management node and finally reported to the main service node.

[0051] Step S6: The main service node monitors the status of each availability zone and reschedules the subtasks of the availability zone that has failed.

[0052] like Figure 2 As shown, the system architecture specifically includes:

[0053] Central database: Stores all persistent states, including global resource tables, task metadata tables, historical performance data tables, etc.

[0054] Main service node: The system brain, running the global scheduler module.

[0055] N Availability Zones: Connected via the Internet.

[0056] Each availability zone contains:

[0057] One management node: runs the local scheduler and state agent.

[0058] M computing nodes: providing heterogeneous computing power (CPU, GPU, NPU, etc.).

[0059] In step S1, the global resource view consists of a status tuple for each availability zone, and the status tuple includes: available computing power vector, comprehensive load index, network latency, effective network bandwidth, and status flag.

[0060] The available computing power vector is a weighted aggregation of the available resources of all computing nodes under it; the network latency is the network latency (inbound / outbound direction) from the main service node to availability zone i, which is measured by periodic Ping; the effective network bandwidth is the effective bandwidth (inbound / outbound direction) from the main service node to availability zone i, which is estimated by periodic network probing.

[0061] The status flags include: online, offline, congested, and error.

[0062] The comprehensive load index is calculated by weighted summation, which includes the following four items: the first item is the product of the average CPU utilization and the CPU weight coefficient; the second item is the product of the average memory utilization and the memory weight coefficient; the third item is the product of the average I / O utilization and the I / O weight coefficient; and the fourth item is the product of the task queue relative saturation and the queue weight coefficient. The task queue relative saturation is obtained by dividing the real-time length of the local task queue on the management node by the maximum queue length.

[0063] The formula for calculating the comprehensive load index is:

[0064]

[0065] in, Indicates the availability zone index. Indicates time, This represents the overall load index of availability zone i at time t. This represents the average CPU utilization of all compute nodes in the availability zone. This represents the average memory utilization of all compute nodes in the availability zone. This represents the average I / O utilization of all compute nodes in the availability zone. Indicates the length of the local task queue on the management node. Indicates the maximum length of the queue. , , and These represent the weighting coefficients for CPU utilization, memory utilization, I / O utilization, and queue ratio, respectively.

[0066] The sum of the weighting coefficients for CPU utilization, memory utilization, I / O utilization, and queue ratio is 1. The specific methods for determining these values ​​include setting initial values ​​of 0.4, 0.3, 0.2, and 0.1 for CPU utilization, memory utilization, I / O utilization, and queue ratio, respectively. A configuration interface is provided, allowing administrators to manually adjust the weights based on business characteristics and hardware configuration. For example, for I / O-intensive storage clusters, the weighting coefficient for I / O utilization can be increased, while the other three weighting coefficients can be decreased proportionally.

[0067] Users submit tasks to the master service node, which then formalizes them into a tuple, specifically including:

[0068] Task type, storage location (e.g., URL) and total size of input data, expected storage location and estimated size of output data, task constraints (e.g., maximum completion deadline, minimum computing power requirement), and task priority.

[0069] If a task can be decomposed, it is split into a set of subtasks. Each subtask inherits some attributes of the task and adds its own input data size as a new attribute.

[0070] In step S3, the adaptive cost function consists of two parts: the first part is the weighted sum of the time estimated cost, network estimated cost, and economic estimated cost multiplied by their respective weight coefficients; the second part is the product of the load penalty factor and the comprehensive load index of the target availability zone.

[0071] The formula for calculating the adaptive cost function is:

[0072]

[0073] in, This represents the j-th subtask. This represents the i-th availability zone. This represents the scheduling cost calculated by the primary service node for each subtask and each availability zone. This refers to the estimated cost over time, which is the cost of estimating subtasks based on historical performance data or benchmark models. In the availability zone The runtime on a typical computing node, This indicates the estimated cost of the network. This represents the estimated economic cost, i.e., the cost of performing the sub-task. The financial costs of the required computing and data transmission resources are calculated based on the cloud service provider's pricing model. , and These represent the weighting coefficients for estimated time cost, estimated network cost, and estimated economic cost, respectively. This represents the load penalty factor, a constant coefficient used to increase the cost of high-load availability zones and prevent them from becoming overloaded.

[0074] The sum of the weight coefficients for estimated time cost, estimated network cost, and estimated economic cost is 1. The specific values ​​are determined based on the initial values ​​of the user's preset strategy. When a user submits a task, they can select a preset strategy and map it to a set of preset initial weights, which include: fastest speed: 0.7, 0.2, and 0.1 respectively; lowest cost: 0.1, 0.2, and 0.7 respectively; balanced mode: 0.4, 0.3, and 0.3 respectively.

[0075] The priority and task constraints specified by the user when submitting a task directly determine the initial values ​​of the weights for estimated time cost, estimated network cost, and estimated economic cost.

[0076] Time-sensitive tasks (such as real-time analytics): Initial setup Relatively high.

[0077] Cost-sensitive tasks (such as offline batch processing): Initial setup Relatively high.

[0078] Data-intensive tasks: Initial setup Relatively high.

[0079] The adaptive cost function includes the following methods for calculating the network's estimated cost:

[0080] Calculate the input data transmission overhead, which is determined by dividing the subtask input data size by the target availability zone inbound effective bandwidth and then adding the inbound network latency value.

[0081] Calculate the output data transmission overhead, which is determined by dividing the subtask's estimated output data size by the target availability zone's outbound effective bandwidth and then adding the outbound network latency value.

[0082] The estimated network cost is obtained by calculating the sum of the input data transmission overhead and the output data transmission overhead.

[0083] The formula for calculating the estimated network cost is:

[0084]

[0085] in, Subtasks Input data size, Subtasks Estimated size of output data This indicates the time t when the primary service node reaches the availability zone. Incoming effective bandwidth, Indicates the available area at time t Effective outbound bandwidth to the primary service node This indicates the time t when the primary service node reaches the availability zone. Incoming network latency, Indicates the available area at time t Outbound network latency to the main service node;

[0086] in, To calculate the input data transfer overhead, To mitigate the overhead of data transmission, the estimation of output data size specifically includes: the system sets a general output / input ratio coefficient for all tasks or for major task categories, such as 0.1 or 0.5, and the estimated size of the subtask's output data is obtained by multiplying the input data size of the subtask by the ratio coefficient of that task.

[0087] In step S4, the management node assigns the task to the computing nodes within its domain for execution, specifically including the following steps:

[0088] The management node receives subtasks from the master service node and adds them to its local task queue;

[0089] The local management node calculates a comprehensive mismatch score for each subtask and each compute node, wherein the comprehensive mismatch score includes a resource fit term and a cache affinity term;

[0090] The local management node selects the computing node with the lowest overall mismatch score as the target node and assigns the subtask to the selected target node for execution.

[0091] The formula for calculating the overall mismatch score is:

[0092]

[0093] in, This represents the m-th compute node within the available area. Indicates the subtask to be scheduled. and candidate computing nodes The overall mismatch score, Subtasks The resource demand vector, Represents a computing node The vector of currently available resources, This represents the distance function between two vectors, with cosine distance chosen to measure their dissimilarity. Indicates the attenuation coefficient. Indicates the current time. Represents a computing node The timestamp of the last child task processed from the same parent task. This represents the natural exponential function.

[0094] For environments where cache expires quickly and task data types change frequently, a larger value should be set. A value, such as 0.5; for environments with continuous data processing and high cache hit benefits, a smaller value should be set. Values, such as 0.1.

[0095] Resource fitting terms A task-node specific matching mechanism is introduced; the larger this value, the greater the mismatch between available resources and task requirements. For example, a task requiring a large amount of memory should be assigned to the node with the most currently available memory, even if its CPU load is slightly higher than other nodes. This avoids performance bottlenecks (such as memory overflow, frequent swapping) caused by resource mismatch.

[0096] Cache affinity items Data locality is considered, which is crucial for iterative tasks or tasks that require repeatedly reading the same data. If a node recently processed another subtask of the same user or task, its CPU cache and disk cache are likely to contain residual data or code. Assigning a new task to it can significantly improve cache hit rate, reduce I / O overhead for data loading, and thus accelerate task execution. The larger the time difference, the closer this value is to 0; the smaller the time difference, the closer this value is to 1.

[0097] Step S6 specifically includes the following steps:

[0098] The master service node monitors the status of each management node through a heartbeat mechanism;

[0099] If the primary service node loses heartbeat signals from a management node for a consecutive number of times reaching the fault determination threshold, it is determined that the availability zone to which it belongs has failed, and the status of the availability zone is marked as offline, and it is also excluded from the global resource view.

[0100] The primary service node identifies all tasks that have been distributed to the failed availability zone but have not been confirmed as completed, puts them back into the global task queue, and reschedules them to other healthy availability zones.

[0101] The fault determination threshold is dynamically set based on the heartbeat interval and the system's maximum tolerable fault recovery time target, and the specific formula is as follows:

[0102]

[0103] in, Indicates the fault determination threshold. This represents the maximum tolerable fault recovery time target of the system, that is, the maximum time allowed from the occurrence of a fault to the system detecting and beginning to handle the fault. This indicates the heartbeat interval, specifically how often the management node sends a heartbeat. This represents the function for rounding up.

[0104] The management node communicates with the computing nodes in its assigned availability zones via a remote management connection based on a secure shell protocol. This connection is used to execute commands and transmit file data via a secure file transfer protocol. The main service node, the management nodes of each availability zone, and the central database are interconnected via a wide area network (WAN) communication link. This WAN communication link is built on the public Internet or a private dedicated network and uses transport layer security protocols or secure socket layer protocols for communication encryption.

[0105] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0106] The examples described herein are merely preferred embodiments of the invention and are not intended to limit the concept and scope of the invention. Any modifications and improvements made by those skilled in the art to the technical solutions of the invention without departing from the design concept of the invention should fall within the protection scope of the invention.

Claims

1. A distributed, multi-availability-zone computing power scheduling method, running in a system consisting of one master service node, one central database, and N availability zones connected via the Internet, each availability zone containing one management node and M computing nodes, characterized in that, Includes the following steps: Step S1: The management node of the availability zone periodically collects the status information of the availability zone and reports it to the main service node. Based on this, the main service node builds and maintains a dynamically updated global resource view. Step S2: The main service node receives the computing task submitted by the user and parses and splits the task into subtasks. Step S3: For each subtask, the main service node calculates the cost of scheduling it to each availability zone based on the global resource view by constructing an adaptive cost function, and selects the availability zone with the lowest cost as the target availability zone. Step S4: The main service node distributes the subtasks to the management nodes of their target availability zones, and each management node then assigns the tasks to the compute nodes within its domain for execution. Step S5: The compute node returns the results, which are then aggregated by the management node and finally reported to the main service node. Step S6: The main service node monitors the status of each availability zone and reschedules the sub-tasks of the availability zone that has failed. In step S1, the global resource view consists of a status tuple for each availability zone, and the status tuple includes: available computing power vector, comprehensive load index, network latency, effective network bandwidth, and status flag. The status flags include: online, offline, congested, and error; The comprehensive load index is calculated through a weighted summation, which includes the following four terms: the first term is the product of the average CPU utilization and the CPU weight coefficient; the second term is the product of the average memory utilization and the memory weight coefficient; the third term is the product of the average I / O utilization and the I / O weight coefficient; and the fourth term is the product of the task queue relative saturation and the queue weight coefficient. The task queue relative saturation is obtained by dividing the real-time length of the local task queue on the management node by the maximum queue length, calculated using the following formula: in, Indicates the availability zone index. Indicates time, This represents the overall load index of availability zone i at time t. This represents the average CPU utilization of all compute nodes in the availability zone. This represents the average memory utilization of all compute nodes in the availability zone. This represents the average I / O utilization of all compute nodes in the availability zone. Indicates the length of the local task queue on the management node. Indicates the maximum length of the queue. , , and These represent the weighting coefficients for CPU utilization, memory utilization, I / O utilization, and queue ratio, respectively. In step S3, the adaptive cost function consists of two parts: the first part is the weighted sum of the products of the estimated time cost, estimated network cost, and estimated economic cost with their respective weight coefficients; the second part is the product of the load penalty factor and the comprehensive load index of the target availability zone. The formula for calculating the adaptive cost function is as follows: in, This represents the j-th subtask. This represents the i-th availability zone. This represents the scheduling cost calculated by the primary service node for each subtask and each availability zone. This refers to the estimated cost over time, which is the cost of estimating subtasks based on historical performance data or benchmark models. In the availability zone The runtime on a typical computing node, This indicates the estimated cost of the network. This represents the estimated economic cost, i.e., the cost of performing the sub-task. The financial costs of the required computing and data transmission resources. , and These represent the weighting coefficients for estimated time cost, estimated network cost, and estimated economic cost, respectively. Indicates the load penalty factor; The adaptive cost function includes the following methods for calculating the network's estimated cost: Calculate the input data transmission overhead, which is determined by dividing the subtask input data size by the target availability zone inbound effective bandwidth and then adding the inbound network latency value. Calculate the output data transmission overhead, which is determined by dividing the subtask's estimated output data size by the target availability zone's outbound effective bandwidth and then adding the outbound network latency value. The sum of the input data transmission overhead and the output data transmission overhead is calculated to obtain the estimated network cost; The formula for calculating the estimated network cost is: in, Subtasks Input data size, Subtasks Estimated size of output data This indicates the time t when the primary service node reaches the availability zone. Incoming effective bandwidth, Indicates the available area at time t Effective outbound bandwidth to the primary service node This indicates the time t when the primary service node reaches the availability zone. Incoming network latency, Indicates the available area at time t Outbound network latency to the main service node; In step S4, the management node assigns the task to the computing nodes within its domain for execution, specifically including the following steps: The management node receives subtasks from the master service node and adds them to its local task queue; The local management node calculates a comprehensive mismatch score for each subtask and each compute node, wherein the comprehensive mismatch score includes a resource fit term and a cache affinity term; The local management node selects the computing node with the lowest overall mismatch score as the target node and assigns the subtask to the selected target node for execution. The formula for calculating the overall mismatch score is as follows: in, This represents the m-th compute node within the available area. Indicates the subtask to be scheduled. and candidate computing nodes The overall mismatch score, Subtasks The resource demand vector, Represents a computing node The vector of currently available resources, This represents the distance function between two vectors. Indicates the attenuation coefficient. Indicates the current time. Represents a computing node The timestamp of the last child task processed from the same parent task. This represents the natural exponential function.

2. The method according to claim 1, characterized in that, Step S6 specifically includes the following steps: The master service node monitors the status of each management node through a heartbeat mechanism; If the primary service node loses heartbeat signals from a management node for a consecutive number of times reaching the fault determination threshold, it is determined that the availability zone to which it belongs has failed, and the status of the availability zone is marked as offline, and it is also excluded from the global resource view. The primary service node identifies all tasks that have been distributed to the faulty availability zone but have not been confirmed as completed, puts them back into the global task queue, and reschedules them to other healthy availability zones. The fault determination threshold is dynamically set based on the heartbeat interval and the maximum tolerable fault recovery time target of the system.

3. The method according to claim 2, characterized in that, The management node communicates with the computing nodes in its assigned availability zones via a remote management connection based on a secure shell protocol. This connection is used to execute commands and transmit file data via a secure file transfer protocol. The main service node, the management nodes of each availability zone, and the central database are interconnected via a wide area network (WAN) communication link. This WAN communication link is built on the public Internet or a private dedicated network and uses transport layer security protocols or secure socket layer protocols for communication encryption.

4. A storage medium, characterized in that, The storage medium stores instructions, and when the computer reads the instructions, it executes the distributed multi-availability zone computing power scheduling method according to any one of claims 1-3.

5. An electronic device, characterized in that, It includes a processor and the storage medium of claim 4, wherein the processor executes instructions in the storage medium.

Citation Information

Patent Citations

  • A fusion scheduling system based on distributed computing power

    CN115051988B

  • A distributed computing power scheduling system based on AIGC

    CN117472549B

  • Distributed computing power resource scheduling method and system, storage medium and program product

    CN121144048A