Data center intelligent scheduling system and method oriented to heterogeneous computing power resources
By using an intelligent scheduling system to monitor in real time and analyze historical data, dynamically adjust resource quotas and migrate containers, the problem of unreasonable scheduling of heterogeneous computing resources in data centers is solved, resource utilization and system stability are improved, and the efficiency of task execution is ensured.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-03
AI Technical Summary
Existing data center computing power scheduling cannot fully and accurately obtain the resource load status of each heterogeneous computing node, and cannot accurately predict the resource gap in the future preset time period, resulting in unreasonable resource allocation, affecting task execution efficiency, and failing to effectively migrate containers when the load is too high, reducing system reliability and availability.
This paper presents an intelligent scheduling system for data centers with heterogeneous computing resources, including a resource awareness module, a queue construction module, and a hierarchical scheduling module. Through real-time monitoring and historical data analysis, it dynamically adjusts resource quotas, migrates containers to low-load nodes when the load exceeds the threshold, and optimizes data storage paths using a distributed storage system.
It enables efficient allocation of heterogeneous computing resources, improves resource utilization, ensures high efficiency in task execution and system stability, reduces the risk of task interruption, and enhances the operational efficiency and service quality of the data center.
Smart Images

Figure CN121785775A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data center scheduling technology, and in particular to a data center intelligent scheduling system and method for heterogeneous computing resources. Background Technology
[0002] With the rapid development of technologies such as big data, artificial intelligence, and cloud computing, the computing tasks undertaken by data centers are becoming increasingly complex and diverse, leading to an explosive growth in the demand for computing power. Against this backdrop, heterogeneous computing resources, due to their ability to integrate the advantages of different types of computing chips (such as CPUs, GPUs, and FPGAs (Field-Programmable Gate Arrays)) to meet the differentiated computing power needs of various application scenarios, have become key to improving computing efficiency and resource utilization in data centers. For example, in artificial intelligence training scenarios, GPUs, with their powerful parallel computing capabilities, can significantly accelerate the model training process; while in traditional transaction processing and data management tasks, CPUs can leverage their strengths in logic control and serial processing. Intelligent scheduling systems and methods for data centers with heterogeneous computing resources are crucial for improving the overall performance and competitiveness of data centers. In the future wave of digital transformation, they have broad application prospects and development potential in scientific research and innovation, enterprise digital operations, and the emerging field of edge computing.
[0003] However, existing data center computing power scheduling methods struggle to comprehensively and accurately acquire the resource load status of each heterogeneous computing node and accurately predict resource gaps within a predetermined timeframe. This leads to inaccurate understanding of resource conditions and an inability to prepare for resource allocation in advance. For task management, the lack of an effective classification mechanism prevents the scientific categorization of tasks within heterogeneous data centers and the creation of corresponding task queues. This results in a lack of targeted task scheduling, failing to fully leverage the advantages of heterogeneous computing resources. In terms of scheduling strategies, the inability to dynamically and reasonably adjust resource quotas for different types of task queues based on the real-time resource load status of each heterogeneous computing node and future resource gap predictions leads to unreasonable resource allocation and impacts task execution efficiency. Furthermore, when heterogeneous computing nodes experience excessive load, current technologies cannot effectively migrate containers running on those nodes to lower-load nodes and adjust the data storage path of the migrated containers based on the multi-path optimization strategies of distributed storage systems. This can not only cause task interruptions but also reduce the reliability and availability of the entire data center.
[0004] Therefore, this invention proposes an intelligent scheduling system and method for data centers with heterogeneous computing resources. Summary of the Invention
[0005] This invention provides an intelligent scheduling system and method for data centers with heterogeneous computing resources. It enables efficient allocation of heterogeneous computing resources, rationally distributing computing resources according to the characteristics and resource requirements of different tasks, avoiding resource waste and idleness, thereby improving resource utilization and reducing operating costs. Simultaneously, this intelligent scheduling system can quickly respond to business changes, ensuring the efficient execution of various tasks and promoting the development of data centers towards intelligent and refined management.
[0006] This invention provides an intelligent scheduling system for data centers with heterogeneous computing resources, comprising: The resource awareness module is used to obtain the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset future period based on multi-dimensional operational data and historical task time data of heterogeneous data centers. The queue building module is used to classify tasks in heterogeneous data centers and build AI training task queues and general computing task queues. The hierarchical scheduling module is used to dynamically adjust the resource quotas of each heterogeneous computing node for the two types of task queues based on the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset period of the future. The fault self-healing module is used to migrate containers running on the corresponding heterogeneous computing nodes to low-load heterogeneous computing nodes when the node load of the heterogeneous computing nodes exceeds a preset load threshold, and adjust the data storage path of the migrated containers based on the multi-path optimization strategy of the distributed storage system.
[0007] Preferably, the resource awareness module includes: The current resource awareness submodule is used to collect multi-dimensional operational data from heterogeneous data centers in real time through hardware monitoring interfaces; The future resource awareness submodule is used to obtain the real-time resource load status of each heterogeneous computing node and the resource gap prediction results within a preset future period based on multi-dimensional operation data and historical task time data of heterogeneous data centers. The resource gap prediction results include resource gap values in multiple dimensions, where each resource gap value is the difference between the task request resource amount and the remaining available resource amount of the node in the corresponding dimension.
[0008] Preferably, the queue building module includes: The demand determination submodule determines the computing power demand type based on the task characteristics of each task in the heterogeneous data center. The task classification submodule is used to build AI training task queues and general computing task queues based on the computing power requirements of all tasks in heterogeneous data centers. The AI training task queue is adapted to tasks with high GPU requirements, while the general computing task queue is adapted to tasks with high CPU requirements.
[0009] Preferably, the hierarchical scheduling module includes: The model building submodule is used to build the scheduling decision model. The hierarchical scheduling submodule is used to dynamically adjust the resource quotas of heterogeneous computing nodes for two types of task queues based on the scheduling decision model with the optimization objectives of maximizing resource utilization and minimizing task latency, combined with the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset period of the future.
[0010] Preferably, the hierarchical scheduling module includes: The history record acquisition submodule is used to acquire historical task execution and scheduling records of heterogeneous data centers within the most recent historical period; The historical record analysis submodule is used to determine the task priority ratio, business density ratio, and historical scheduling completion rate of the two types of task queues at each moment in the most recent historical period based on historical task execution and scheduling records. The linear model building submodule is used to construct a linear model between the priority ratio, business density ratio and scheduling completion rate of the two types of task queues based on the task priority ratio and business density ratio of the two types of task queues at each moment in the most recent historical period, as well as the historical scheduling completion rate of historical task execution and scheduling records at the corresponding moment. The resource quota adjustment submodule is used to dynamically adjust the resource quotas of each heterogeneous computing node for the two types of task queues based on the real-time resource load status of each heterogeneous computing node, the predicted resource gap within a preset time period, and the linear model between the priority ratio, business density ratio and scheduling completion rate of the two types of task queues.
[0011] Preferably, the historical record analysis submodule includes: The priority allocation determination unit is used to determine the task priority distribution set of each type of task queue at each moment in the most recent historical period based on historical task execution and scheduling records, and to calculate the task priority allocation ratio of the two types of task queues at each moment in the most recent historical period based on the task priority distribution set of the two types of task queues at each moment in the most recent historical period. The business density ratio determination unit is used to determine the business density of each type of task queue at each moment in the most recent historical period based on the task submission frequency and resource consumption of the two types of task queues at each moment in the most recent historical period in the historical task execution and scheduling records, and to calculate the business density ratio of the two types of task queues at each moment in the most recent historical period based on the business density of the two types of task queues at each moment in the most recent historical period. The scheduling completion rate analysis unit is used to determine the historical scheduling completion rate of each type of task queue at each moment in the most recent historical period based on historical task execution and scheduling records.
[0012] Preferably, the scheduling completion rate analysis unit includes: The scheduling success rate determination subunit is used to determine the historical scheduling success rate of each type of task queue at each moment in the most recent historical period based on historical task execution and scheduling records. The scheduling effect value determination subunit is used to calculate the historical scheduling effect value of historical task execution and scheduling records at each moment in the most recent historical period based on the supply-side revenue value of each heterogeneous computing node at each moment in the most recent historical period and the demand-side revenue value of the two types of task queues at each moment in the most recent historical period. The scheduling completion rate determination subunit is used to calculate the historical scheduling completion rate of historical task execution and scheduling records at each moment in the most recent historical period based on the historical scheduling effect value of historical task execution and scheduling records at each moment in the most recent historical period and the historical scheduling success rate of the two types of task queues at each moment in the most recent historical period.
[0013] Preferably, the method for the scheduling effect value determination subunit to obtain the supplier-side revenue value of each heterogeneous computing node at each moment in the most recent historical period and the demand-side revenue value of the two types of task queues at each moment in the most recent historical period includes: Based on historical task execution and scheduling records, the ratio of idle resources to total resources of each heterogeneous computing node at each moment in the most recent historical period is determined as the resource waste rate, and the ratio of current load to rated load of each heterogeneous computing node at each moment in the most recent historical period is determined as the overload risk. Based on the resource waste rate and overload risk, the supplier revenue value of each heterogeneous computing node at each moment in the most recent historical period is calculated. Based on the expected task completion rate and expected task delay rate of each type of task queue in the most recent historical period at each moment in the historical task execution and scheduling records, the demand-side revenue value of each type of task queue at each moment in the most recent historical period is calculated.
[0014] Preferably, the method for determining the scheduling effect value subunit to calculate the historical scheduling effect value of historical task execution and scheduling records at each moment in the most recent historical period based on the supplier-side revenue value of each heterogeneous computing node at each moment in the most recent historical period and the demand-side revenue value of the two types of task queues at each moment in the most recent historical period includes: Set the correlation coefficient between each heterogeneous computing node and each type of task queue; Based on the resource demand ratio of each type of task queue at each moment in the most recent historical period, the correlation coefficient between each heterogeneous computing node and each type of task queue, and the normalized value of the supply-side revenue of each heterogeneous computing node at each moment in the most recent historical period, the supply-side weight of each heterogeneous computing node and the demand-side weight of each type of task queue at each moment in the most recent historical period are calculated. Based on the normalized value of the supplier revenue of each heterogeneous computing node at each moment in the most recent historical period and the corresponding supplier weight, the supplier weighted revenue at each moment in the most recent historical period is calculated. Based on the normalized value of the demand-side revenue value of each type of task queue at each moment in the most recent historical period and the corresponding demand-side weight, the demand-side weighted revenue at each moment in the most recent historical period is calculated. Based on the scenario balance coefficient at each moment in the most recent historical period, the supply-side weighted revenue and demand-side weighted revenue at each moment in the most recent historical period are integrated to obtain the historical scheduling effect value of the historical task execution and scheduling records at each moment in the most recent historical period.
[0015] This invention provides an intelligent scheduling method for data centers with heterogeneous computing resources, comprising: Based on multi-dimensional operational data and historical task time data of heterogeneous data centers, the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a future preset period are obtained. Classify tasks in heterogeneous data centers and construct AI training task queues and general computing task queues; Based on the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset time period, the resource quotas of each heterogeneous computing node for the two types of task queues are dynamically adjusted. When the node load of a heterogeneous computing node exceeds a preset load threshold, the containers running on the corresponding heterogeneous computing node are migrated to a low-load heterogeneous computing node, and the data storage path of the migrated containers is adjusted based on the multi-path optimization strategy of the distributed storage system.
[0016] The beneficial effects of this invention compared to existing technologies are as follows: By analyzing multi-dimensional operational data and historical task latency data from heterogeneous data centers, the real-time resource load status and future resource gap prediction results of each heterogeneous computing node are obtained, providing accurate data support for subsequent scheduling and making scheduling decisions more forward-looking. Task classification constructs AI training task queues and general computing task queues, enabling refined task management and improving the processing efficiency of different types of tasks. Based on the real-time status of computing nodes and resource gap predictions, resource quotas for the two types of task queues are dynamically adjusted, flexibly adapting to task requirements and resource conditions, and improving the overall utilization rate of heterogeneous computing resources. When the load of heterogeneous computing nodes exceeds the threshold, containers are migrated to low-load nodes, and the data storage path is adjusted using a multi-path optimization strategy of the distributed storage system to ensure continuous task operation, enhance system stability and reliability, ensure that the data center can quickly recover normal operation when facing abnormal node loads, reduce the risk of task interruption, and comprehensively improve the operational efficiency and service quality of the data center.
[0017] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.
[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the functional modules contained in the intelligent scheduling system for data centers with heterogeneous computing resources according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the functional sub-modules contained in the resource perception module in an embodiment of the present invention; Figure 3 This is a schematic diagram of the functional modules contained in the queue construction module in an embodiment of the present invention. Detailed Implementation
[0020] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0021] like Figure 1 As shown, this invention provides an intelligent scheduling system for data centers with heterogeneous computing resources, comprising: The resource awareness module is used to obtain the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset future period based on multi-dimensional operational data and historical task time data of heterogeneous data centers. The queue building module is used to classify tasks in heterogeneous data centers and build AI training task queues and general computing task queues. The hierarchical scheduling module is used to dynamically adjust the resource quotas of each heterogeneous computing node for the two types of task queues based on the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset period of the future. The fault self-healing module is used to migrate containers running on the corresponding heterogeneous computing nodes to low-load heterogeneous computing nodes when the node load of the heterogeneous computing nodes exceeds a preset load threshold, and adjust the data storage path of the migrated containers based on the multi-path optimization strategy of the distributed storage system.
[0022] In this embodiment, a heterogeneous data center refers to a center that contains different types of computing devices, such as central processing units that are good at general computing and graphics processing units that are suitable for parallel computing, which can meet a variety of different computing needs, such as artificial intelligence training and regular data processing being carried out simultaneously.
[0023] In this embodiment, the multi-dimensional operating data includes various aspects of the hardware device's operating information, such as the CPU utilization rate, the GPU memory usage, and the device's real-time power consumption. For example, the CPU utilization rate is the proportion of the currently used cores to the total number of cores.
[0024] In this embodiment, historical task time data is a record of the time spent on previous task executions. For example, a previous data processing task took 2 hours to complete. This data can provide a reference for estimating the time of subsequent tasks.
[0025] In this embodiment, heterogeneous computing nodes are hardware devices composed of different computing architectures, such as central processing unit nodes based on complex instruction set architectures and graphics processing unit nodes based on stream processor architectures, each suitable for different types of computing tasks.
[0026] In this embodiment, the real-time resource load status reflects the current resource utilization of the computing node. For example, if the graphics processor is using 80% of its video memory at this moment, it indicates its real-time resource load status.
[0027] In this embodiment, the future preset time period is a predetermined range of future time, such as the next 30 minutes, which is used for subsequent prediction of resource conditions, etc.
[0028] In this embodiment, the resource gap prediction result for the future preset time period is based on current and historical data, which estimates the difference between the demand and supply of computing node resources in the future preset time period. For example, it is predicted that the central processing unit resource demand will be 20% more computing power than the existing resources in the next hour.
[0029] In this embodiment, the AI training task queue is a queue specifically for storing tasks related to artificial intelligence training, such as neural network training tasks, which typically have high demands on graphics processing unit resources.
[0030] In this embodiment, the general computing task queue is used to place general computing tasks, such as simple data statistical calculations, which mainly rely on the central processing unit for computation.
[0031] In this embodiment, the resource quota for each heterogeneous computing node to the two types of task queues refers to the proportion of resources allocated by different computing nodes to the artificial intelligence training task queue and the general computing task queue. For example, the graphics processing unit node allocates 70% of the resources to the artificial intelligence training task queue and 30% of the resources to the general computing task queue.
[0032] In this embodiment, the node load of a heterogeneous computing node refers to the current resource usage status of the node. Taking the central processing unit as an example, a high load may mean that most of the cores are busy with computation.
[0033] In this embodiment, the preset load threshold is a pre-set standard value for measuring the load level of a node. For example, if the central processing unit load reaches 85%, it is considered to exceed the threshold and corresponding measures should be taken.
[0034] In this embodiment, containers running on corresponding heterogeneous computing nodes are migrated to low-load heterogeneous computing nodes, and the data storage path of the migrated containers is adjusted based on the multi-path optimization strategy of the distributed storage system. That is, when a computing node is overloaded, the running containers are transferred to nodes with low load. At the same time, in the distributed storage system, by comparing the transmission speed, stability and other factors of different paths, the optimal path is selected to store the container data, such as selecting the fastest and most stable path to store the data.
[0035] like Figure 2 As shown, in order to provide accurate data support for scheduling, a resource awareness module is proposed, including: The current resource awareness submodule is used to collect multi-dimensional operational data from heterogeneous data centers in real time through hardware monitoring interfaces; The future resource awareness submodule is used to obtain the real-time resource load status of each heterogeneous computing node and the resource gap prediction results within a preset future period based on multi-dimensional operation data and historical task time data of heterogeneous data centers. The resource gap prediction results include resource gap values in multiple dimensions, where each resource gap value is the difference between the task request resource amount and the remaining available resource amount of the node in the corresponding dimension.
[0036] In this embodiment, the hardware monitoring interface is a channel that can acquire information about the operating status of hardware devices. Through it, data such as CPU utilization and GPU memory usage can be collected, which is like a window for the hardware to transmit its own information to the outside world, allowing the system to know the real-time status of various aspects of the hardware.
[0037] In this embodiment, the resource gap prediction result covers resource gap values from multiple aspects. Different dimensions refer to different types of resources such as CPU and GPU. Each resource gap value is calculated by subtracting the amount of that resource currently available on the node from the amount requested by a task in the corresponding dimension (such as the GPU dimension). For example, if a node has a total GPU memory of 8GB, 3GB has been used, and 5GB remains, and a task requests 6GB of GPU memory, then the resource gap value for the GPU dimension is 6GB - 5GB = 1GB.
[0038] like Figure 3 As shown, in order to achieve fine-grained task classification and management, a queue construction module is proposed, including: The demand determination submodule determines the computing power demand type based on the task characteristics of each task in the heterogeneous data center. The task classification submodule is used to build AI training task queues and general computing task queues based on the computing power requirements of all tasks in heterogeneous data centers. The AI training task queue is adapted to tasks with high GPU requirements, while the general computing task queue is adapted to tasks with high CPU requirements.
[0039] When a task contains instructions to call deep learning frameworks (such as TensorFlow and PyTorch) and the GPU computing power requirement accounts for ≥60%, it is classified into the AI training task queue; when a task contains only general computing instructions and the CPU computing power requirement accounts for ≥70%, it is classified into the general computing task queue.
[0040] In this embodiment, the task characteristics of each task refer to the relevant elements that can describe the characteristics of the task, such as the data type involved, algorithm complexity, and computational logic. These characteristics affect the task's computational power requirements. For example, the data for an image recognition task is image information, and the algorithm involves convolution operations, which is what distinguishes it from other tasks.
[0041] In this embodiment, determining the computing power requirement type based on the task characteristics of each task in the heterogeneous data center means judging which type of computing power the task mainly needs based on the specific data type, algorithm complexity, and other characteristics of the task. For example, deep learning tasks, due to their large number of matrix operations and parallel processing requirements, can be determined to have a computing power requirement type that leans towards the computing power of a graphics processing unit (GPU) based on these task characteristics.
[0042] In this embodiment, the computing power requirement type refers to the category of computing capabilities primarily relied upon during task execution. Common types include those relying on a central processing unit (CPU) for general-purpose computing and those relying on a graphics processing unit (GPU) for parallel computing. For example, for scientific computing tasks, if the task mainly involves complex logical operations, the computing power requirement type might be CPU-based; if it mainly involves large-scale data parallel processing, it might be GPU-based.
[0043] In this embodiment, high GPU-demand tasks refer to tasks that require significant computing power from the graphics processing unit (GPU) during execution. For example, neural network training tasks in deep learning require processing large amounts of data and performing numerous parallel calculations simultaneously. The parallel processing capabilities of the GPU can greatly improve its operating efficiency, so these types of tasks are considered high GPU-demand tasks.
[0044] In this embodiment, high CPU-demand tasks refer to tasks that require a high level of computing power from the central processing unit (CPU) during runtime. For example, traditional database queries and simple text processing tasks mainly perform sequential logical operations and control flow operations, relying more on the CPU's general-purpose computing power, and therefore belong to high CPU-demand tasks.
[0045] To improve resource scheduling efficiency, a hierarchical scheduling module is proposed, including: The model building submodule is used to build scheduling decision models based on reinforcement learning algorithms; The hierarchical scheduling submodule is used to dynamically adjust the resource quotas of heterogeneous computing nodes for two types of task queues based on the scheduling decision model with the optimization objectives of maximizing resource utilization and minimizing task latency, combined with the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset period of the future. The state space of the agent in the reinforcement learning algorithm includes the resource utilization rate of each heterogeneous computing node and the task queue length; the action space includes the resource quota adjustment ratio of each heterogeneous computing node; and the reward function is set as follows: , in: , .
[0046] In this embodiment, a scheduling decision model is constructed based on a reinforcement learning algorithm, which allows the model to learn through continuous interaction with the environment. In this example, the scheduling decision model continuously tries different scheduling strategies in a heterogeneous data center environment, adjusts the strategies based on the feedback received (such as resource utilization and task completion status), and gradually builds a model that can effectively schedule resources.
[0047] In this embodiment, based on a scheduling decision model with the optimization objectives of maximizing resource utilization and minimizing task latency, the resource quotas allocated to two types of task queues on heterogeneous computing nodes are dynamically adjusted by combining the real-time resource load status of each heterogeneous computing node with the predicted resource gap within a preset future time period. This means that the scheduling decision model constantly monitors the current resource utilization of each heterogeneous computing node, such as how much CPU, GPU, and FPGA resources are currently used, as well as the expected resource demand gap in the future. Then, with the goal of maximizing resource utilization while minimizing task execution time, the resource ratio allocated to the AI training task queue and the general computing task queue on different heterogeneous computing nodes is continuously changed. For example, if a large GPU resource gap is predicted in the future and the current GPU load is high, the GPU resources allocated to the high GPU demand task queue (AI training task queue) are reduced, and the resources are increased to other relatively resource-sufficient task queues.
[0048] In this embodiment, CPU, GPU, and FPGA usage refers to the actual amount of resources utilized by these three hardware devices at a specific point in time or within a specific time period. The total amount of CPU, GPU, and FPGA usage refers to the total amount of available resources possessed by these hardware devices.
[0049] To optimize the scheduling strategy, a hierarchical scheduling module is proposed, including: The history record acquisition submodule is used to acquire historical task execution and scheduling records of heterogeneous data centers within the most recent historical period; The historical record analysis submodule is used to determine the task priority ratio, business density ratio, and historical scheduling completion rate of the two types of task queues at each moment in the most recent historical period based on historical task execution and scheduling records. The linear model building submodule is used to construct a linear model between the priority ratio, business density ratio and scheduling completion rate of the two types of task queues based on the task priority ratio and business density ratio of the two types of task queues at each moment in the most recent historical period, as well as the historical scheduling completion rate of historical task execution and scheduling records at the corresponding moment. The resource quota adjustment submodule is used to dynamically adjust the resource quotas of each heterogeneous computing node for the two types of task queues based on the real-time resource load status of each heterogeneous computing node, the predicted resource gap within a preset time period, and the linear model between the priority ratio, business density ratio and scheduling completion rate of the two types of task queues.
[0050] In this embodiment, obtaining historical task execution and scheduling records of the heterogeneous data center within the most recent historical period involves collecting the specific details of task execution and task scheduling within the heterogeneous data center over a specific period in the past. This includes recording when each task started and ended, which computing nodes it was assigned to, and how resources were allocated during the scheduling process.
[0051] In this embodiment, a linear model is constructed based on the task priority ratio and business density ratio of the two types of task queues at each moment in the most recent historical period, as well as the historical scheduling completion rate of historical task execution and scheduling records at the corresponding moments. Here, the task priority ratio refers to the proportion of tasks with different priorities in the AI training task queue and the general computing task queue; the business density ratio refers to the proportional relationship between the number of tasks or task complexity of the two types of task queues at a specific moment. By analyzing these two ratios at each moment in the most recent historical period, and the actual historical scheduling completion rate at the corresponding moment—that is, the ratio of the number of tasks actually scheduled to the number of planned tasks—a linear regression method is used to find an approximate linear relationship between them, thus constructing a mathematical model. For example, calculations based on a large amount of data show that when the priority ratio increases by a certain value and the business density ratio increases by another value, the scheduling completion rate will change according to a certain linear law.
[0052] In this embodiment, based on the real-time resource load status of each heterogeneous computing node, the predicted resource gap within a preset future time period, and the linear model between the priority ratio, service density ratio, and scheduling completion rate of the two types of task queues, the resource quotas allocated to the two types of task queues by each heterogeneous computing node are dynamically adjusted. This means first understanding the current resource usage status of each heterogeneous computing node, such as the current occupancy level of resources like CPU, GPU, and FPGA, and predicting the difference between resource demand and supply in the future period, thereby determining the task priority ratio and service density ratio. Then, combined with the previously constructed linear model, with the goal of maximizing the scheduling completion rate (e.g., reaching 100%), the resource ratio allocated to the AI training task queue and the general computing task queue by each heterogeneous computing node is continuously changed in real time.
[0053] To provide a useful reference for resource scheduling, a historical record analysis submodule is proposed, including: The priority allocation determination unit is used to determine the task priority distribution set of each type of task queue at each moment in the most recent historical period based on historical task execution and scheduling records, and to calculate the task priority allocation ratio of the two types of task queues at each moment in the most recent historical period based on the task priority distribution set of the two types of task queues at each moment in the most recent historical period. The business density ratio determination unit is used to determine the business density of each type of task queue at each moment in the most recent historical period based on the task submission frequency and resource consumption of the two types of task queues at each moment in the most recent historical period in the historical task execution and scheduling records, and to calculate the business density ratio of the two types of task queues at each moment in the most recent historical period based on the business density of the two types of task queues at each moment in the most recent historical period. The scheduling completion rate analysis unit is used to determine the historical scheduling completion rate of each type of task queue at each moment in the most recent historical period based on historical task execution and scheduling records.
[0054] In this embodiment, based on historical task execution and scheduling records, the task priority distribution set for each type of task queue at each moment within the most recent historical period is determined. Then, based on the task priority distribution sets of the two types of task queues at each moment within the most recent historical period, the task priority ratio of the two types of task queues at each moment within the most recent historical period is calculated. This means first examining the records of task execution and scheduling within a specific historical time period, and for both the AI training task queue and the general computing task queue, determining the number of tasks of different priorities at each time point, forming a task priority distribution set. For example, at a certain moment, high-priority tasks account for 30%, medium-priority tasks account for 50%, and low-priority tasks account for 20% in the AI training task queue. Then, the priority distribution of the two types of task queues at the same moment is weighted and compared before calculation to obtain the task priority ratio.
[0055] In this embodiment, based on the task submission frequency and resource consumption of the two types of task queues at each moment in the most recent historical period from the historical task execution and scheduling records, the business density of each type of task queue at each moment in the most recent historical period is determined. Then, based on the business density of the two types of task queues at each moment in the most recent historical period, the business density ratio of the two types of task queues at each moment in the most recent historical period is calculated. Specifically, based on the task execution and scheduling records over a period of time, the number of times each task was submitted by the AI training task queue and the general computing task queue, as well as the amount of resources consumed to complete these tasks, are counted at each time point to determine the business density of each type of task queue at that moment. For example, at a certain moment, the AI training task queue submitted 10 tasks and consumed 50 units of resources; the business density of the AI training task queue at that moment is obtained by multiplying the two. Then, the business densities of the two types of task queues at the same moment are compared and calculated to obtain the business density ratio.
[0056] In this embodiment, the historical scheduling completion rate of each type of task queue is determined at each moment within the most recent historical period based on historical task execution and scheduling records. Specifically, based on task execution and scheduling records within a specific historical time period, for both the AI training task queue and the general computing task queue, the ratio of the number of tasks actually scheduled to the number of tasks planned to be scheduled is calculated at each time point. This ratio is the historical scheduling completion rate at that moment. For example, if the AI training task queue plans to schedule 10 tasks and actually schedules 8, then the historical scheduling completion rate of the AI training task queue at that moment is 8 ÷ 10 = 80%.
[0057] To provide quantitative indicators for adjusting scheduling strategies, a scheduling completion rate analysis unit is proposed, including: The scheduling success rate determination subunit is used to determine the historical scheduling success rate of each type of task queue at each moment in the most recent historical period based on historical task execution and scheduling records. The scheduling effect value determination subunit is used to calculate the historical scheduling effect value of historical task execution and scheduling records at each moment in the most recent historical period based on the supply-side revenue value of each heterogeneous computing node at each moment in the most recent historical period and the demand-side revenue value of the two types of task queues at each moment in the most recent historical period. The scheduling completion rate determination subunit is used to calculate the historical scheduling completion rate of historical task execution and scheduling records at each moment in the most recent historical period based on the historical scheduling effect value of historical task execution and scheduling records at each moment in the most recent historical period and the historical scheduling success rate of the two types of task queues at each moment in the most recent historical period.
[0058] In this embodiment, the historical scheduling success rate of each type of task queue at each moment in the most recent historical period is determined based on historical task execution and scheduling records. Specifically, based on detailed records of task execution and scheduling over a past period, the proportion of successfully scheduled tasks to the planned number of tasks is calculated for both the AI training task queue and the general computing task queue at each specific moment. For example, if the AI training task queue plans to schedule 20 tasks at a certain moment, and 15 are successfully scheduled, then the historical scheduling success rate of the AI training task queue at that moment is 15 ÷ 20 = 75%. "Successful scheduling" is a prerequisite for "complete scheduling." Only tasks that are successfully scheduled have a chance of being scheduled, but successful scheduling does not guarantee completion. Tasks may fail midway through execution due to various unforeseen circumstances, preventing completion. "Successful scheduling" refers to a task being successfully assigned to the target node, while "complete scheduling" refers to a task being successfully executed at the target node.
[0059] In this embodiment, the historical scheduling completion rate of the historical task execution and scheduling records at each moment within the most recent historical period is calculated based on the historical scheduling effect value of the historical task execution and scheduling records at each moment within the most recent historical period and the historical scheduling success rate of the two types of task queues at each moment within the most recent historical period. This involves a weighted calculation of the three factors; for example, the weights of the historical scheduling effect value of the historical task execution and scheduling records at each moment within the most recent historical period, and the historical scheduling success rate of the two types of task queues at each moment within the most recent historical period are 0.4, 0.3, and 0.3, respectively.
[0060] To provide a quantitative basis for evaluating scheduling effectiveness, a method is proposed for the scheduling effectiveness value determination subunit to obtain the supplier-side revenue value of each heterogeneous computing node at each moment in the most recent historical period and the demand-side revenue value of the two types of task queues at each moment in the most recent historical period, including: Based on historical task execution and scheduling records, the ratio of idle resources to total resources for each heterogeneous computing node at each moment within the most recent historical period is determined as the resource waste rate. The ratio of current load to rated load for each heterogeneous computing node at each moment within the most recent historical period is also determined as the overload risk. Based on the resource waste rate and overload risk, the supplier-side revenue value for each heterogeneous computing node at each moment within the most recent historical period is calculated. First, review the historical task execution and scheduling records. For each heterogeneous computing node, at each moment within the most recent historical period, calculate the ratio of its idle resources to its total resources; this ratio is the resource waste rate. For example, if a heterogeneous computing node has a total resource quantity of 100 units and idle resources of 20 units at a certain moment, then the resource waste rate at that moment is 20 ÷ 100 = 20%. Simultaneously, calculate the ratio of the node's current load to its rated load capacity at that moment; this is the overload risk. Suppose that at a certain moment, the current load of this node is 80 units, and the rated load is 100 units. The overload risk is 80 ÷ 100 = 80%. Then, through a certain calculation method, the resource waste rate and the overload risk are combined to calculate the supplier's revenue value. For example, a formula can be set: Supplier revenue value = 1 - Resource waste rate - Overload risk × Coefficient. If the coefficient is set to 0.5, according to the previous example, the supplier's revenue value = 1 - 20% - 80% × 0.5 = 1 - 0.2 - 0.4 = 0.4.
[0061] Based on the expected task completion rate and expected task latency rate of each task queue type in the historical task execution and scheduling records at each moment within the most recent historical period, the demand-side revenue value of each task queue type at each moment within the most recent historical period is calculated. Information is obtained from the historical task execution and scheduling records. For each task queue type (such as the AI training task queue and the general computing task queue), at each moment within the most recent historical period, the demand-side revenue value is calculated using the proportion of tasks expected to be completed (expected task completion rate) and the proportion of tasks expected to be delayed (expected task latency rate). For example, the demand-side revenue value can be set as: Demand-side revenue value = Expected task completion rate - Expected task latency rate × weight. Assuming that at a certain moment, the expected task completion rate of the AI training task queue is 90%, the expected task latency rate is 10%, and the weight is set to 0.8, then the demand-side revenue value of the AI training task queue at that moment = 90% - 10% × 0.8 = 90% - 8% = 82%.
[0062] To comprehensively and accurately evaluate scheduling effectiveness, a method is proposed for determining the historical scheduling effectiveness value of the historical task execution and scheduling records at each moment in the most recent historical period. This method is based on the supply-side revenue value of each heterogeneous computing node at each moment in the most recent historical period and the demand-side revenue value of the two types of task queues at each moment in the most recent historical period. The method includes: Assignment coefficients between heterogeneous computing nodes and each type of task queue. Assume the heterogeneous computing nodes include CPU nodes, GPU nodes, and FPGA nodes, and the task queues include AI training task queues and general computing task queues. The correlation coefficient between CPU nodes and the AI training task queue is set to 0.3. This is because AI training tasks primarily rely on GPUs for parallel computation, with the CPU assisting in some tasks; the correlation is relatively weak. The correlation coefficient between CPU nodes and the general computing task queue is set to 0.8. General computing tasks mostly involve sequential logic operations, with the CPU as the primary execution unit; the correlation is strong. The correlation coefficient between GPU nodes and the AI training task queue is set to 0.9. AI training tasks highly rely on the parallel computing capabilities of GPUs; the correlation is extremely strong. The correlation coefficient between GPU nodes and the general computing task queue is set to 0.2. General computing tasks have low GPU requirements, resulting in a low correlation. The correlation coefficient between FPGA nodes and the AI training task queue is set to 0.5. FPGAs can accelerate some AI training tasks through customized configurations, but the overall dependence is not as high as that of GPUs; the correlation is moderate. The correlation coefficient between FPGA nodes and the general computing task queue is set to 0.4. FPGAs are not as widely used as CPUs in general computing scenarios, and their correlation with CPUs is relatively low.
[0063] Based on the resource demand ratio of each task queue type at each moment in the most recent historical period, the correlation coefficient between each heterogeneous computing node and each task queue type, and the normalized value of the supply-side revenue of each heterogeneous computing node at each moment in the most recent historical period, the supply-side weight of each heterogeneous computing node and the demand-side weight of each task queue type at each moment in the most recent historical period are calculated. At each 5-minute timeframe, the sum of the resource demand ratios of the AI training task queue and the general computing task queue is 1. For example, at a certain moment, the resource demand ratio of the AI training task queue is 0.7, then the resource demand ratio of the general computing task queue is 1 minus 0.7, which is 0.3.
[0064] For each type of heterogeneous computing node, the minimum and maximum supply-side revenue for all moments within the 30-day period must first be calculated. Then, the minimum revenue is subtracted from the revenue value at each moment, and the result is divided by the difference between the maximum and minimum revenue. This maps the supply-side revenue value to the range of 0 to 1. For example, the minimum supply-side revenue for a GPU node is 0.5, the maximum is 0.9, and the revenue value at a certain moment is 0.7. After normalization, the result is (0.7-0.5) ÷ (0.9-0.5), which equals 0.5.
[0065] When calculating the supply-side weight, for each 5-minute timeframe, the correlation coefficient between the heterogeneous computing node and the two types of task queues is multiplied by the resource demand ratio of the corresponding task queue at that time, and then the two products are added together. The higher the correlation, the larger the demand ratio of the corresponding task queue, and the higher the weight of this heterogeneous computing node.
[0066] When calculating the demand-side weight, for each 5-minute timeframe, the correlation coefficient between the task queue and various heterogeneous computing nodes is multiplied by the proportion of the normalized supply-side revenue value of the corresponding heterogeneous computing node in the sum of the normalized supply-side revenue values of all heterogeneous computing nodes. These products are then summed. The higher the correlation, the higher the normalized revenue of the corresponding node, and the higher the weight of this task queue.
[0067] These steps allow us to calculate the supplier weight of each heterogeneous computing node at each 5-minute interval within the last 30 days, as well as the demand weight of each type of task queue at each 5-minute interval within the last 30 days. These weights are crucial for subsequent calculations of historical scheduling performance values and help heterogeneous data centers to rationally allocate and schedule resources.
[0068] Based on the normalized supply-side revenue values and corresponding supply-side weights of each heterogeneous computing node at each moment in the most recent historical period, the supply-side weighted revenue at each moment in the most recent historical period is calculated. This is done by multiplying the normalized supply-side revenue values of all heterogeneous computing nodes by their corresponding supply-side weights calculated previously, and then summing the results. For example, if a heterogeneous computing node has a normalized supply-side revenue value of 0.6 and a supply-side weight of 0.5, then the supply-side weighted revenue of that heterogeneous computing node at that moment is 0.6 × 0.5 = 0.3. The final result is obtained by summing the supply-side weighted revenues of all heterogeneous computing nodes at that moment.
[0069] Based on the normalized demand-side revenue value and corresponding demand-side weight for each task queue type at each moment in the most recent historical period, the demand-side weighted revenue for each moment in the most recent historical period is calculated. Similar to calculating the supply-side weighted revenue, the normalized demand-side revenue value for each task queue type is multiplied by its corresponding demand-side weight, and then summed. For example, if at a certain moment the normalized demand-side revenue value for the general computing task queue is 0.7 and its demand-side weight is 0.4, then the demand-side weighted revenue for the general computing task queue at that moment is 0.7 × 0.4 = 0.28. Adding the demand-side revenues of the two task queue types together gives the final result.
[0070] Based on the scenario balance coefficient at each moment in the most recent historical period, the supply-side weighted revenue and demand-side weighted revenue at each moment in the most recent historical period are integrated to obtain the historical scheduling effect value of the historical task execution and scheduling records at each moment in the most recent historical period.
[0071] Among them, the scene balance coefficient at each moment in the most recent historical period The ratio of business density at that moment is determined by the business density ratio: the higher the business density, The closer the value is to 0.6, the more priority is given to ensuring the supplier's profits; the lower the business density, the better. The closer it is to 0.4, the more priority is given to ensuring the benefits for the demand side.
[0072] in, ,in The total business density at that moment. (This represents the maximum business density within a historical period), ensuring that the weights are dynamically adjusted according to business scenarios.
[0073] The scenario balance coefficient is a value used to balance the importance of the supply and demand sides. It is calculated by multiplying the weighted revenue of the supply side by the scenario balance coefficient, multiplying the weighted revenue of the demand side by (1 - scenario balance coefficient), and then adding the two together to obtain the historical scheduling effect value. For example, if at a certain moment the weighted revenue of the supply side is 0.3, the weighted revenue of the demand side is 0.28, and the scenario balance coefficient is 0.6, then the historical scheduling effect value = 0.3 × 0.6 + 0.28 × (1 - 0.6) = 0.18 + 0.112 = 0.292.
[0074] This invention provides an implementation method for a data center intelligent scheduling method for heterogeneous computing resources, comprising: Based on multi-dimensional operational data and historical task time data of heterogeneous data centers, the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a future preset period are obtained. Classify tasks in heterogeneous data centers and construct AI training task queues and general computing task queues; Based on the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset time period, the resource quotas of each heterogeneous computing node for the two types of task queues are dynamically adjusted. When the node load of a heterogeneous computing node exceeds a preset load threshold, the containers running on the corresponding heterogeneous computing node are migrated to a low-load heterogeneous computing node, and the data storage path of the migrated containers is adjusted based on the multi-path optimization strategy of the distributed storage system.
[0075] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.
Claims
1. A data center intelligent scheduling system for heterogeneous computing resources, characterized in that, include: The resource awareness module is used to obtain the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset future period based on multi-dimensional operational data and historical task time data of heterogeneous data centers. The queue building module is used to classify tasks in heterogeneous data centers and build AI training task queues and general computing task queues. The hierarchical scheduling module is used to dynamically adjust the resource quotas of each heterogeneous computing node for the two types of task queues based on the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset period of the future. The fault self-healing module is used to migrate containers running on the corresponding heterogeneous computing nodes to low-load heterogeneous computing nodes when the node load of the heterogeneous computing nodes exceeds a preset load threshold, and adjust the data storage path of the migrated containers based on the multi-path optimization strategy of the distributed storage system.
2. The intelligent scheduling system for data centers oriented towards heterogeneous computing resources according to claim 1, characterized in that, The resource awareness module includes: The current resource awareness submodule is used to collect multi-dimensional operational data from heterogeneous data centers in real time through hardware monitoring interfaces; The future resource awareness submodule is used to obtain the real-time resource load status of each heterogeneous computing node and the resource gap prediction results within a preset future period based on multi-dimensional operation data and historical task time data of heterogeneous data centers. The resource gap prediction results include resource gap values in multiple dimensions, where each resource gap value is the difference between the task request resource amount and the remaining available resource amount of the node in the corresponding dimension.
3. The intelligent scheduling system for data centers oriented towards heterogeneous computing resources according to claim 1, characterized in that, The queue building module includes: The demand determination submodule determines the computing power demand type based on the task characteristics of each task in the heterogeneous data center. The task classification submodule is used to build AI training task queues and general computing task queues based on the computing power requirements of all tasks in heterogeneous data centers. The AI training task queue is adapted to tasks with high GPU requirements, while the general computing task queue is adapted to tasks with high CPU requirements.
4. The intelligent scheduling system for data centers oriented towards heterogeneous computing resources according to claim 1, characterized in that, The hierarchical scheduling module includes: The model building submodule is used to build the scheduling decision model. The hierarchical scheduling submodule is used to dynamically adjust the resource quotas of heterogeneous computing nodes for two types of task queues based on the scheduling decision model with the optimization objectives of maximizing resource utilization and minimizing task latency, combined with the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset period of the future.
5. The intelligent scheduling system for data centers oriented towards heterogeneous computing resources according to claim 1, characterized in that, The hierarchical scheduling module includes: The history record acquisition submodule is used to acquire historical task execution and scheduling records of heterogeneous data centers within the most recent historical period; The historical record analysis submodule is used to determine the task priority ratio, business density ratio, and historical scheduling completion rate of the two types of task queues at each moment in the most recent historical period based on historical task execution and scheduling records. The linear model building submodule is used to construct a linear model between the priority ratio, business density ratio and scheduling completion rate of the two types of task queues based on the task priority ratio and business density ratio of the two types of task queues at each moment in the most recent historical period, as well as the historical scheduling completion rate of historical task execution and scheduling records at the corresponding moment. The resource quota adjustment submodule is used to dynamically adjust the resource quotas of each heterogeneous computing node for the two types of task queues based on the real-time resource load status of each heterogeneous computing node, the predicted resource gap within a preset time period, and the linear model between the priority ratio, business density ratio and scheduling completion rate of the two types of task queues.
6. The intelligent scheduling system for data centers oriented towards heterogeneous computing resources according to claim 5, characterized in that, The historical record analysis submodule includes: The priority allocation determination unit is used to determine the task priority distribution set of each type of task queue at each moment in the most recent historical period based on historical task execution and scheduling records, and to calculate the task priority allocation ratio of the two types of task queues at each moment in the most recent historical period based on the task priority distribution set of the two types of task queues at each moment in the most recent historical period. The business density ratio determination unit is used to determine the business density of each type of task queue at each moment in the most recent historical period based on the task submission frequency and resource consumption of the two types of task queues at each moment in the most recent historical period in the historical task execution and scheduling records, and to calculate the business density ratio of the two types of task queues at each moment in the most recent historical period based on the business density of the two types of task queues at each moment in the most recent historical period. The scheduling completion rate analysis unit is used to determine the historical scheduling completion rate of each type of task queue at each moment in the most recent historical period based on historical task execution and scheduling records.
7. The intelligent scheduling system for data centers oriented towards heterogeneous computing resources according to claim 6, characterized in that, The scheduling completion rate analysis unit includes: The scheduling success rate determination subunit is used to determine the historical scheduling success rate of each type of task queue at each moment in the most recent historical period based on historical task execution and scheduling records. The scheduling effect value determination subunit is used to calculate the historical scheduling effect value of historical task execution and scheduling records at each moment in the most recent historical period based on the supply-side revenue value of each heterogeneous computing node at each moment in the most recent historical period and the demand-side revenue value of the two types of task queues at each moment in the most recent historical period. The scheduling completion rate determination subunit is used to calculate the historical scheduling completion rate of historical task execution and scheduling records at each moment in the most recent historical period based on the historical scheduling effect value of historical task execution and scheduling records at each moment in the most recent historical period and the historical scheduling success rate of the two types of task queues at each moment in the most recent historical period.
8. The intelligent scheduling system for data centers oriented towards heterogeneous computing resources according to claim 7, characterized in that, The method for determining the scheduling effect value subunit to obtain the supplier-side revenue value of each heterogeneous computing node at each moment in the most recent historical period and the demand-side revenue value of the two types of task queues at each moment in the most recent historical period includes: Based on historical task execution and scheduling records, the ratio of idle resources to total resources of each heterogeneous computing node at each moment in the most recent historical period is determined as the resource waste rate, and the ratio of current load to rated load of each heterogeneous computing node at each moment in the most recent historical period is determined as the overload risk. Based on the resource waste rate and overload risk, the supplier revenue value of each heterogeneous computing node at each moment in the most recent historical period is calculated. Based on the expected task completion rate and expected task delay rate of each type of task queue in the most recent historical period at each moment in the historical task execution and scheduling records, the demand-side revenue value of each type of task queue at each moment in the most recent historical period is calculated.
9. The intelligent scheduling system for data centers oriented towards heterogeneous computing resources according to claim 7, characterized in that, The scheduling effect value determination subunit calculates the historical scheduling effect value of historical task execution and scheduling records at each moment in the most recent historical period based on the supplier-side revenue value of each heterogeneous computing node at each moment in the most recent historical period and the demand-side revenue value of the two types of task queues at each moment in the most recent historical period. The method includes: Set the correlation coefficient between each heterogeneous computing node and each type of task queue; Based on the resource demand ratio of each type of task queue at each moment in the most recent historical period, the correlation coefficient between each heterogeneous computing node and each type of task queue, and the normalized value of the supply-side revenue of each heterogeneous computing node at each moment in the most recent historical period, the supply-side weight of each heterogeneous computing node and the demand-side weight of each type of task queue at each moment in the most recent historical period are calculated. Based on the normalized value of the supplier revenue of each heterogeneous computing node at each moment in the most recent historical period and the corresponding supplier weight, the supplier weighted revenue at each moment in the most recent historical period is calculated. Based on the normalized value of the demand-side revenue value of each type of task queue at each moment in the most recent historical period and the corresponding demand-side weight, the demand-side weighted revenue at each moment in the most recent historical period is calculated. Based on the scenario balance coefficient at each moment in the most recent historical period, the supply-side weighted revenue and demand-side weighted revenue at each moment in the most recent historical period are integrated to obtain the historical scheduling effect value of the historical task execution and scheduling records at each moment in the most recent historical period.
10. A data center intelligent scheduling method for heterogeneous computing resources, characterized in that, include: Based on multi-dimensional operational data and historical task time data of heterogeneous data centers, the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a future preset period are obtained. Classify tasks in heterogeneous data centers and construct AI training task queues and general computing task queues; Based on the real-time resource load status of each heterogeneous computing node and the predicted resource gap within a preset time period, the resource quotas of each heterogeneous computing node for the two types of task queues are dynamically adjusted. When the node load of a heterogeneous computing node exceeds a preset load threshold, the containers running on the corresponding heterogeneous computing node are migrated to a low-load heterogeneous computing node, and the data storage path of the migrated containers is adjusted based on the multi-path optimization strategy of the distributed storage system.