Method and device for healthy scheduling of computing power by intelligent computing center according to computing power node

By evaluating and estimating the health status of computing power nodes and optimizing resource allocation, the problem of resource waste in traditional computing power scheduling methods is solved, and higher stability and reliability are achieved.

CN120560799APending Publication Date: 2025-08-29DATACANVAS LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510660288.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Traditional computing power scheduling methods rely on static indicators, resulting in waste of computing power resources and it is difficult to optimize the operating stability and reliability of computing power nodes.

Method used

By obtaining the health status evaluation results of the computing power node, estimating the health status in the future time period based on the evaluation results, scheduling the computing power node to perform tasks, and optimizing resource allocation using health scores and weights.

Benefits of technology

It improves the stability and reliability of computing power operation, reduces resource waste, increases the task success rate by 8%, extends the life cycle of computing power nodes by 15%, and increases the resource utilization rate by 12%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560799A_ABST
    Figure CN120560799A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for an intelligent computing center to schedule computing power according to computing power node health, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructures, and the method comprises the steps: S1, obtaining an evaluation result of a health state of a computing power node of the intelligent computing center; s2, estimating the health state of the computing power node in a target time period based on the evaluation result to obtain an estimation result; and S3, scheduling the computing power node to execute a computing power operation task based on the estimation result. According to the invention, the health state of the computing power node in the future time period is estimated based on the evaluation result of the health state of the computing power node, and the computing power can be scheduled according to the estimated health state of the computing power node when the computing power operation task is executed, so that the computing power can be fully utilized; the computing power node scheduling is optimized to improve the stability and reliability of computing power operation and reduce the resource waste of computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent computing centers, smart computing centers and computing power infrastructure, and in particular to a method and device for an intelligent computing center to schedule computing power according to the health of computing power nodes. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.

[0003] An "Intelligent Computing Center" is a facility that uses large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.

[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] Since the emergence of intelligent computing centers, fields such as artificial intelligence and high-performance computing have rapidly developed, and computing power faces the challenges of high load, high concurrency, and high availability. However, traditional computing power scheduling methods often rely on static metrics, which can easily lead to wasted computing power resources. Therefore, optimizing computing power node scheduling to improve computing power operation stability and reliability and reduce computing power waste is an urgent problem to be solved. Summary of the Invention

[0008] The present invention provides a method and device for scheduling computing power of an intelligent computing center according to the health of computing power nodes, which is used to solve the problem of how to optimize the scheduling of computing power nodes to reduce resource waste of computing power nodes.

[0009] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:

[0010] In a first aspect, the present invention provides a method for scheduling computing power in an intelligent computing center based on the health of computing power nodes, comprising:

[0011] Step S1: Obtain the evaluation results of the health status of the computing power nodes of the intelligent computing center;

[0012] Step S2: estimating the health status of the computing power node within the target time period based on the evaluation result to obtain an estimation result;

[0013] Step S3: scheduling the computing power node to execute the computing power operation task based on the estimation result.

[0014] Optionally, step S1 includes:

[0015] Step S11: Obtain health indicators of computing nodes in the intelligent computing center within a historical period, wherein the health indicators include at least one of the following: CPU temperature, GPU temperature, CPU usage, GPU usage, memory occupancy, disk input / output load, fan speed, power supply voltage, and historical downtime records;

[0016] Step S12: Evaluate the health status of the computing power node based on the health indicator to obtain the evaluation result.

[0017] Optionally, step S12 includes:

[0018] Step S121: when the number of the health indicators is at least two, obtaining the weight of each health indicator;

[0019] Step S122: Based on each health indicator and the corresponding weight, determine the health level of the computing power node, and the health level is used to characterize the health status of the computing power node.

[0020] Optionally, the evaluation result of the health status of the computing power node includes the health status of the computing power node at N historical moments, where N is an integer greater than 1;

[0021] The step S2 comprises:

[0022] Step S21: Based on the health status at the N historical moments, the health status of the computing power node at the N+1th moment in the future is scored and estimated to obtain the estimation result.

[0023] Optionally, step S3 includes:

[0024] Step S31: When it is determined according to the estimation result that the health status of the computing power node within the target time period meets the preset conditions, the computing power node is scheduled to execute the computing power operation task.

[0025] Optionally, step S3 includes:

[0026] Step S32: determining the scheduling priority of the computing power node based on the estimation result, and scheduling the computing power node to execute the computing power operation task according to the scheduling priority;

[0027] The scheduling priority is determined according to at least one of the following:

[0028] Health score size order;

[0029] The load of the computing power node, the estimated task duration and the health score.

[0030] In a second aspect, the present invention provides a device for scheduling computing power in an intelligent computing center based on the health of computing power nodes, comprising:

[0031] The acquisition module is used to obtain the evaluation results of the health status of the computing power nodes of the intelligent computing center;

[0032] An estimation module, configured to estimate the health status of the computing power node within a target time period based on the evaluation result to obtain an estimation result;

[0033] A scheduling module is used to schedule the computing power node to perform computing power operation tasks based on the estimation result.

[0034] In a third aspect, the present invention provides a server comprising: a processor, a memory, and a program stored in the memory and runnable on the processor. When the program is executed by the processor, the steps of the method for scheduling computing power according to the health of computing power nodes by the intelligent computing center as described in the first aspect above are implemented.

[0035] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method for scheduling computing power of an intelligent computing center according to the health of computing power nodes as described in the first aspect above are implemented.

[0036] In a fifth aspect, the present invention provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of the method for scheduling computing power of an intelligent computing center according to the health of computing power nodes as described in the first aspect above.

[0037] In the present invention, the health status of the computing power node in the future period is estimated based on the evaluation results of the health status of the computing power node. When executing the computing power operation task, the computing power can be scheduled according to the estimated health status of the computing power node. The computing power can be fully utilized and the scheduling of computing power nodes can be optimized to improve the stability and reliability of the computing power operation and reduce the waste of computing power resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0039] Figure 1 A flow chart of a method for scheduling computing power according to the health of computing power nodes in an intelligent computing center according to the present invention;

[0040] Figure 2 This is a schematic diagram of the structure of a device for scheduling computing power according to the health of computing power nodes in the intelligent computing center of the present invention;

[0041] Figure 3 This is a structural diagram of the server of the present invention. DETAILED DESCRIPTION

[0042] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0043] First, the technical terms involved in the present invention are briefly explained below.

[0044] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0045] The "computing power" (CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .

[0046] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capability of the computing power facilities, including the comprehensive capabilities of network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0047] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, security and reliability, and environmental friendliness. It is a comprehensive indicator for measuring a data center's data storage capacity, encompassing both external storage devices such as storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second / TB). Disaster recovery ratio is a key indicator of security and reliability.

[0048] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.

[0049] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0050] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.

[0051] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0052] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing (NLP) and machine vision.

[0053] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.

[0054] The "intelligent computing center" described in this article refers to a facility that provides the computing power, data, and algorithms required for artificial intelligence applications (such as AI deep learning model development, model training, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center encompasses facilities, hardware, and software, and can provide a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0055] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".

[0056] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.

[0057] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0058] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.

[0059] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information computing, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water and electricity.

[0060] The "large language model" mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0061] The "computing power node" mentioned in the present invention refers to the computing resources of the server / container that can process computing tasks.

[0062] The "computing power operation task" mentioned in the present invention refers to: a specific workload or job executed on computing power resources that requires a certain amount of computing power support, usually involving complex data processing, numerical calculations, model training or simulation scenarios.

[0063] See also Figure 1 , Figure 1 The present invention provides a method for intelligent computing center to schedule computing power according to the health of computing power nodes, such as Figure 1 As shown, the method includes:

[0064] Step S1: Obtain the evaluation results of the health status of the computing power nodes of the intelligent computing center;

[0065] Step S2: estimating the health status of the computing power node within the target time period based on the evaluation result to obtain an estimation result;

[0066] Step S3: scheduling the computing power node to execute the computing power operation task based on the estimation result.

[0067] A computing node can be a server or a computing resource within a container that can process computing tasks. The number of such computing nodes can be one or more. When there are multiple computing nodes, the health status assessment results of each computing node are obtained separately, and the assessment results can be used to reflect the health status of the computing nodes.

[0068] In some implementations, operational metrics of computing nodes can be collected in real time, such as CPU temperature, GPU usage, memory usage, disk input / output (I / O) load, network latency, etc. A corresponding weight can be assigned to each metric based on its importance, and a health score for the computing node, i.e., an assessment of the node's health, can be calculated based on the weight.

[0069] In some implementations, the operating indicators of the computing power nodes may be monitored, and the health status of the computing power nodes may be evaluated based on the operating indicators, thereby obtaining an evaluation result of the health status of the computing power nodes.

[0070] The target time period is a time period in the future. For example, if the current time is time t, the target time period can be t+1 to t+T (T is a preset time length, such as 1 hour, 24 hours, etc.).

[0071] When the evaluation results of the health status of the computing power node are obtained, the health status of the computing power node in the future period is estimated based on the evaluation results.

[0072] In some implementations, a sliding window mechanism is used to construct time series data, inputting historical health scores and operating indicators. A long short-term memory (LSTM) model is then used to estimate the health score for the next T time period. For example, if the current time is t and the target time period is t+1 to t+2 (i.e., two time points in the future), the health status of the computing power node is estimated for the next two time points.

[0073] In some embodiments, the health trend of the computing power node is modeled and the health score within a future time T (e.g., 24 hours) is estimated.

[0074] In some implementations, the health status of the node within a future T time period is estimated based on the correlation between multiple operating indicators of the computing power node.

[0075] The estimation results may include information indicating the health status of the computing power node, including health score, health level, etc. The computing power node may be scheduled to perform computing power operation tasks based on the estimation results.

[0076] In some embodiments, when it is determined based on the estimation result that the health score of the computing power node in the future target time period is higher than a preset threshold, the computing power operation task is assigned to the computing power node.

[0077] In some implementations, tasks are assigned to computing nodes with an estimated health score above a preset threshold, and computing nodes with lower loads may be prioritized. For example, if computing node A has an estimated health score of 85 and a current load of 30%, and computing node B has an estimated health score of 75 and a current load of 60%, then tasks are assigned to computing node A first.

[0078] In some embodiments, computing nodes with an estimated health score below a threshold are isolated from resources to prevent tasks from being assigned to these computing nodes.

[0079] In some implementations, multiple computing nodes may be sorted according to the estimated health scores, thereby scheduling the computing nodes according to the order of the estimated health scores.

[0080] Through the above method, the health status of the computing power node in the future period is estimated based on the evaluation results of the health status of the computing power node. When executing tasks, the computing power node can be scheduled according to the estimated health status of the computing power node, which can fully utilize the computing resources of the computing power node and reduce the waste of computing resources corresponding to the computing power node.

[0081] Optionally, step S1 includes:

[0082] Step S11: Obtain health indicators of computing nodes in the intelligent computing center within a historical period, wherein the health indicators include at least one of the following: CPU temperature, GPU temperature, CPU usage, GPU usage, memory occupancy, disk input / output load, fan speed, power supply voltage, and historical downtime records;

[0083] Step S12: Evaluate the health status of the computing power node based on the health indicator to obtain the evaluation result.

[0084] Among them, health indicators can be used to evaluate the operating status of computing nodes, such as one or more of CPU temperature, GPU temperature, usage rate, memory occupancy rate, disk input / output (I / O) load, fan speed, power supply voltage, historical downtime records, etc. These indicators reflect the hardware performance, stability and potential failure risks of the node.

[0085] The historical time period is the time period before the current moment (such as the past 24 hours, 7 days, or 30 days). The operating status of the computing power nodes within this time period is obtained, so that the health trend of the computing power nodes in the future can be analyzed.

[0086] In some implementations, the operating data of the computing nodes is collected in real time and stored as historical data.

[0087] In some implementations, historical downtime records, task failure logs, and other information of computing nodes are obtained.

[0088] Based on the above computing power indicators, evaluate the health status of the computing power nodes.

[0089] In some embodiments, a weighted scoring method is used to calculate the comprehensive score of the computing power node, and the health status of the computing power node is evaluated based on the score.

[0090] In some implementations, the health indicator values ​​are mapped to a set of health levels. For example, when the CPU temperature is ≥85°C, the health level is automatically downgraded to "medium" or "poor."

[0091] In some implementations, a clustering algorithm (e.g., K-means) is used to identify the health level by combining correlation analysis of multiple indicators. For example, if the CPU temperature, GPU usage, and disk I / O load of a computing node all exceed thresholds, the node is marked as "poor."

[0092] In some embodiments, random forest, extreme gradient boosting (XGBoost) or weighted comprehensive scoring method is used, and the output range is 0-100, divided into four levels (excellent, good, medium, poor).

[0093] Through weight distribution, the misjudgment of node health status by a single indicator can be reduced and the reliability of comprehensive evaluation can be improved.

[0094] Optionally, step S12 includes:

[0095] Step S121: when the number of the health indicators is at least two, obtaining the weight of each health indicator;

[0096] Step S122: Based on each health indicator and the corresponding weight, determine the health level of the computing power node, and the health level is used to characterize the health status of the computing power node.

[0097] When multiple health indicators are obtained, the weight of each health indicator is determined based on the contribution ratio of each indicator to health. For example, the weight of CPU temperature is 30%, the weight of GPU usage is 25%, and the weight of memory occupancy is 20%.

[0098] This weight can be adjusted in real time based on the computing power node's operating environment, task requirements, etc. For example, in high-load scenarios, the weight of CPU and GPU usage can be increased, while in low-load scenarios, the weight of power supply voltage can be increased.

[0099] In some embodiments, a weighted comprehensive scoring method is used to calculate the health score = Σ(health index value × weight), and the health level is divided according to the score range. For example, the health level is divided as follows:

[0100] 90-100 points: Excellent (indicates that the computing power node operates stably and can be scheduled with priority);

[0101] 70-89 points: Good (indicates that the computing power node is operating normally and requires regular monitoring);

[0102] 50-69 points: Medium (indicates that the computing power node has potential risks and task allocation needs to be restricted);

[0103] 0-49 points: Poor (indicates that the computing power node needs to be immediately isolated or repaired).

[0104] The health level grading mechanism provides a basis for resource scheduling, such as prioritizing tasks for "excellent" nodes and limiting resource usage for "poor" nodes, thereby optimizing overall system stability and resource utilization. It can also identify potential problem nodes in advance, reducing task interruptions and resource waste.

[0105] Optionally, the evaluation result of the health status of the computing power node includes the health status of the computing power node at N historical moments, where N is an integer greater than 1;

[0106] The step S2 comprises:

[0107] Step S21: Based on the health status at the N historical moments, the health status of the computing power node at the N+1th moment in the future is scored and estimated to obtain the estimation result.

[0108] Among them, the health status of N historical moments represents the health score or health indicator value of the computing power node at N time points (such as once every 5 minutes).

[0109] For example, if N=5, the health status of the computing power node at t-4, t-3, t-2, t-1, and t is recorded.

[0110] Based on the health status at the above N moments, the health status of the computing power node at the N+1 moment can be estimated.

[0111] In some implementations, a deep learning model such as LSTM, Gated Recurrent Unit (GRU), or Transformer is used to estimate the health score of the node for the next T time period.

[0112] In some implementations, a sliding window mechanism is used to construct time series data, where health scores at N historical moments are input and health scores at the N+1th moment in the future are output.

[0113] For example, collect health scores at N historical moments (e.g., scores from time t-4 to time t); construct supervised learning samples with the input being the health scores at [t-4, t-3, t-2, t-1, t] and the output being the health scores at time t+1. Use an LSTM model to train a time series estimation task and output the estimation results. Use an LSTM model to train a time series estimation task and output the estimation results.

[0114] In some implementations, a GRU model is used to model the health trends of nodes and estimate their health status at the N+1th moment in the future. For example, health indicators (such as CPU temperature and GPU usage) at N historical moments are used as input features. The GRU network extracts the dependencies in the time series and outputs a health score or health grade (such as "excellent" or "good") at the N+1th moment in the future.

[0115] In some embodiments, the future health status is estimated based on a sliding window combined with a weighted average method.

[0116] For example, calculate the weighted average of health scores at N historical moments (e.g., recent scores have higher weights), adjust the estimated results based on the historical trend of the scores, and output the health score at the N+1th moment in the future.

[0117] The computing power nodes are estimated based on their health status at N historical moments to avoid accidental errors at a single time point and improve the reliability of future health status estimates.

[0118] Optionally, step S3 includes:

[0119] Step S31: When it is determined according to the estimation result that the health status of the computing power node within the target time period meets the preset conditions, the computing power node is scheduled to execute the computing power operation task.

[0120] After obtaining the estimation results, you can adopt one or more of the following strategies for scheduling:

[0121] No tasks are assigned to nodes whose estimated scores are lower than the set threshold;

[0122] Prioritize scheduling to nodes with high health scores and moderate current loads;

[0123] Combine minimum load priority, task duration estimation, and health score to perform multi-objective optimization scheduling;

[0124] You can match computing nodes with different health levels based on the fault tolerance level of the acquired task;

[0125] Reduce energy consumption while meeting health requirements;

[0126] Build a cross-center health-aware federated scheduling system.

[0127] In some implementations, computing nodes are scheduled based on health score thresholds.

[0128] For example, if the estimation result meets the preset conditions (the health score of the computing power node is ≥80 points), the computing power running task will be assigned to the node; otherwise, the scheduling will be rejected and other nodes will be sought.

[0129] In some implementations, task priority is combined with the health status of computing nodes for scheduling.

[0130] For example, tasks are assigned priorities (high, medium, and low). If the node health score is ≥80 points and the load is low, high-priority tasks are assigned first; if the node health score is ≥60 points and the load is high, medium-priority tasks are assigned.

[0131] In some implementations, health status, load balancing, and task time scheduling are integrated.

[0132] For example, input parameters such as node health score, current load, and expected task execution time. Genetic algorithm or particle swarm optimization algorithm is used to generate the optimal scheduling solution (such as selecting nodes with high health scores and low load).

[0133] In some implementations, the estimation results are combined with historical data to dynamically adjust the scheduling strategy.

[0134] For example, if the estimation results show that the health score of a computing power node continues to decline within the target time period (for example, from 85 to 70 points), its task allocation will be restricted, and nodes with stable health status will be prioritized. If the estimation results show that the node health score is stable (for example, above 80 points), it will be allowed to execute high-load tasks.

[0135] The scheduler integration method can be as follows: Kubernetes plug-in scheduler extender (SchedulerExtender), Slurm scheduler extension module, cloud management platform API Hook to achieve native access.

[0136] Optionally, step S3 includes:

[0137] Step S32: determining the scheduling priority of the computing power node based on the estimation result, and scheduling the computing power node to execute the computing power operation task according to the scheduling priority;

[0138] The scheduling priority is determined according to at least one of the following:

[0139] Health score size order;

[0140] The load of the computing power node, the estimated task duration and the health score.

[0141] The estimation results may include the health status of the computing power node (for example, a health score), as well as the load of the computing power node and the task duration estimation.

[0142] In some implementations, the scheduling priority of a computing node may be determined based on the health score.

[0143] For example, based on the estimated results, a health score is calculated for each node (e.g., 85, 70, 60). Nodes are sorted from high to low by health score, with priority given to scheduling nodes with a health score ≥ 80. If multiple nodes have the same health score, the load is further compared (prioritizing the node with the lower load).

[0144] In some embodiments, the scheduling priority of the computing nodes is determined based on the order of health score size, the load of the computing nodes, and the task duration estimation.

[0145] Input the estimated health score, current load (such as CPU usage), and estimated task duration (such as the expected task execution time). Calculate a comprehensive priority score for each node: Priority score = α × health score + β × load1 + γ × estimated task duration, where α, β, and γ are weighting coefficients that can be adjusted as needed. Schedule tasks from highest to lowest priority score.

[0146] Nodes are prioritized by health scores to reduce the risk of task interruption by assigning tasks to nodes with poor health.

[0147] In order to evaluate the effectiveness of the above invention, simulation tests were conducted based on 100 GPU servers. The test results showed that the task success rate increased by about 8%, the average life cycle of computing nodes was extended by 15%, resource utilization increased by 12%, and the system could avoid about 87% of potential node downtime events in advance.

[0148] The present invention can be applied to: AI training platforms (such as Computer Vision / Large Language Model Training Scheduling, CV / LLM) training scheduling), cloud-native HPC systems, edge computing platforms, and data center AIOps intelligent operation and maintenance systems.

[0149] This paper combines the health scores of computing nodes with an intelligent resource scheduling strategy based on an estimation model to achieve predictive task scheduling. By collecting the operating status of computing nodes in real time, building a health scoring system, and estimating future health status based on time series, the resource allocation strategy can be dynamically adjusted during the scheduling process.

[0150] See also Figure 2 , Figure 2 The present invention provides a device for scheduling computing power in an intelligent computing center based on the health of computing power nodes. The device 200 includes:

[0151] Acquisition module 201, used to obtain the evaluation results of the health status of the computing power nodes of the intelligent computing center;

[0152] An estimation module 202 is configured to estimate the health status of the computing power node within a target time period based on the evaluation result to obtain an estimation result;

[0153] The scheduling module 203 is used to schedule the computing power node to perform the computing power operation task based on the estimation result.

[0154] Optionally, the acquisition module includes:

[0155] An acquisition submodule is used to obtain health indicators of the computing nodes of the intelligent computing center within a historical time period. The health indicators include at least one of the following: CPU temperature, GPU temperature, CPU usage, GPU usage, memory occupancy, disk input / output load, fan speed, power supply voltage, and historical downtime records;

[0156] The evaluation submodule is used to evaluate the health status of the computing power node based on the health indicator to obtain the evaluation result.

[0157] Optionally, the evaluation submodule includes:

[0158] an acquiring unit, configured to acquire a weight of each health indicator when the number of the health indicators is at least two;

[0159] A determination unit is used to determine the health level of the computing power node based on each health indicator and the corresponding weight, and the health level is used to characterize the health status of the computing power node.

[0160] Optionally, the evaluation result of the health status of the computing power node includes the health status of the computing power node at N historical moments, where N is an integer greater than 1;

[0161] The estimation module is specifically used for:

[0162] Based on the health status at the N historical moments, a score is estimated for the health status of the computing power node at the N+1th moment in the future to obtain the estimation result.

[0163] Optionally, the scheduling module is specifically configured to:

[0164] When it is determined according to the estimation result that the health status of the computing power node within the target time period meets a preset condition, the computing power node is scheduled to execute the computing power operation task.

[0165] Optionally, the scheduling module is specifically configured to:

[0166] Determine the scheduling priority of the computing power node based on the estimation result, and schedule the computing power node to perform the computing power operation task according to the scheduling priority;

[0167] The scheduling priority is determined according to at least one of the following:

[0168] Health score size order;

[0169] The load of the computing power node, the estimated task duration and the health score.

[0170] The device provided by the present invention for scheduling computing power of an intelligent computing center according to the health of computing power nodes is capable of realizing the various processes of each embodiment of the method for scheduling computing power of an intelligent computing center according to the health of computing power nodes. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be described here.

[0171] It should be noted that the device for scheduling computing power according to the health of computing power nodes in the intelligent computing center of the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.

[0172] Please refer to Figure 3 The present invention also provides a server 110, including a processor 111, a memory 112, and a computer program stored in the memory 112 and executable on the processor 111. When the computer program is executed by the processor 111, the various processes of the embodiment of the method for scheduling computing power according to the health of computing power nodes in the above-mentioned intelligent computing center are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0173] The present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the computer program implements the various processes of the embodiment of the method for scheduling computing power of the intelligent computing center according to the health of computing power nodes, and can achieve the same technical effect. To avoid repetition, the description is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0174] The present application also provides a computer program product including computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the embodiment of the method for scheduling computing power according to the health of computing power nodes in the intelligent computing center shown can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0175] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0176] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0177] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.

Claims

1. A method for scheduling computing power in an intelligent computing center based on the health of computing power nodes, characterized in that: include: Step S1: Obtain the evaluation results of the health status of the computing power nodes of the intelligent computing center; Step S2: estimating the health status of the computing power node within the target time period based on the evaluation result to obtain an estimation result; Step S3: scheduling the computing power node to execute the computing power operation task based on the estimation result.

2. The method according to claim 1, characterized in that The step S1 comprises: Step S11: Obtain health indicators of computing nodes in the intelligent computing center within a historical period, wherein the health indicators include at least one of the following: CPU temperature, GPU temperature, CPU usage, GPU usage, memory occupancy, disk input / output load, fan speed, power supply voltage, and historical downtime records; Step S12: Evaluate the health status of the computing power node based on the health indicator to obtain the evaluation result.

3. The method according to claim 2, characterized in that The step S12 includes: Step S121: when the number of the health indicators is at least two, obtaining the weight of each health indicator; Step S122: Based on each health indicator and the corresponding weight, determine the health level of the computing power node, and the health level is used to characterize the health status of the computing power node.

4. The method according to claim 1, wherein The evaluation result of the health status of the computing power node includes the health status of the computing power node at N historical moments, where N is an integer greater than 1; The step S2 comprises: Step S21: Based on the health status at the N historical moments, the health status of the computing power node at the N+1th moment in the future is scored and estimated to obtain the estimation result.

5. The method according to claim 1, wherein The step S3 comprises: Step S31: When it is determined according to the estimation result that the health status of the computing power node within the target time period meets the preset conditions, the computing power node is scheduled to execute the computing power operation task.

6. The method according to any one of claims 1 to 5, characterized in that The step S3 comprises: Step S32: determining the scheduling priority of the computing power node based on the estimation result, and scheduling the computing power node to execute the computing power operation task according to the scheduling priority; The scheduling priority is determined according to at least one of the following: Health score size order; The load of the computing power node, the estimated task duration and the health score.

7. A device for scheduling computing power in an intelligent computing center based on the health of computing power nodes, characterized in that: include: The acquisition module is used to obtain the evaluation results of the health status of the computing power nodes of the intelligent computing center; An estimation module, configured to estimate the health status of the computing power node within a target time period based on the evaluation result to obtain an estimation result; A scheduling module is used to schedule the computing power node to perform computing power operation tasks based on the estimation result.

8. A server, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for scheduling computing power according to the health of computing power nodes by an intelligent computing center as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for scheduling computing power according to the health of computing power nodes by an intelligent computing center as described in any one of claims 1 to 6.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the method for scheduling computing power according to the health of computing power nodes by an intelligent computing center as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Graphics card online health assessment and scheduling method and device, electronic equipment and storage medium

    CN121412056A