Heterogeneous computing resource optimization scheduling and management method
By combining the computing power monitoring module, resource management module, and scheduling module, and using machine learning algorithms for dynamic management and intelligent scheduling of heterogeneous computing resources, the problem of resource idleness and low management efficiency in traditional computing centers is solved, and efficient resource utilization and fault prevention are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional computing centers suffer from idle resources, inefficient management of heterogeneous hardware, and difficulty in rapidly monitoring and scheduling resources for ultra-large clusters.
By combining a computing power monitoring module, a resource management module, and a scheduling module, and using machine learning algorithms for dynamic management and intelligent scheduling of the resource pool, combined with FCFS+Backfilling, genetic algorithms, and ant colony algorithms, efficient resource utilization and load balancing are achieved.
It enables dynamic expansion and efficient resource utilization of computing clusters, prevents single points of failure, provides unified monitoring and intelligent alarms, and improves the management and scheduling efficiency of computing resources.
Smart Images

Figure CN121785751A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method for optimizing scheduling and management of heterogeneous computing resources. Background Technology
[0002] In recent years, computing technology has continuously evolved and upgraded, undergoing a development from low to high, from shallow to deep, and from quantitative to qualitative change. Computing power has been continuously improved, computing methods have been continuously optimized, and computing data has grown exponentially. Advanced computing is an evolution from single computing devices and technologies to diversified computing systems and applications, characterized by its advanced nature, ubiquity, and diversity. It is also the essence and core of the new generation of information technology industry.
[0003] Traditional computing centers commonly suffer from underutilized resources, making it difficult to fully utilize computing power. Against this backdrop, new computing platforms, primarily represented by large-scale data centers, AI computing centers, and supercomputing centers, are being deployed towards higher computing density and higher power density. This extends from spatial density based on hardware deployments such as servers and storage devices to computing power density that improves overall performance through server upgrades, optimized power allocation, and efficient heat dissipation; and from energy-efficient density based on energy consumption to power density that efficiently handles high concurrency and allocates resources on demand. The operational efficiency and efficiency of these new computing platforms are continuously evolving.
[0004] Currently, large-scale heterogeneous hardware suffers from low management efficiency, necessitating the design of efficient resource scheduling algorithms for different tasks to enable rapid collection, analysis, and utilization of monitoring data from ultra-large clusters. Ultimately, this will achieve the management and scheduling of ultra-large-scale computing and storage resources. Summary of the Invention
[0005] In view of this, the present invention provides a method for optimizing scheduling and management of heterogeneous computing resources, applied to an advanced computing platform integration system. The advanced computing platform integration system includes a computing resource management module, a computing scheduling module, and a computing monitoring module. The computing monitoring module has communication connections with the computing resource management module and the computing scheduling module, respectively. The computing resource management module includes a storage sub-cluster, a computing sub-cluster, and a management sub-cluster.
[0006] The computing power monitoring module provides hardware performance data and job operation data to the computing power scheduling module and computing power resource management module; the job operation data includes computing resource information and job logs for various types of jobs, and the hardware performance data includes real-time performance data;
[0007] The computing resource management module analyzes and processes the computing resource information of various types of jobs provided by the computing monitoring module to form a dynamic resource pool. The dynamic resource pool includes a high-performance computing resource pool, a general computing resource pool, a big data resource pool, and a high IO resource pool. The computing resource management module dynamically manages nodes, storage, and networks based on the dynamic resource pool. By adding, reducing, and migrating cluster nodes and adjusting the range of queue resources, the module achieves dynamic expansion of computing clusters and virtualization clusters, as well as efficient utilization and load balancing of resources within the cluster.
[0008] The computing power scheduling module uses machine learning algorithms to analyze and determine the computing power requirements of each service based on job logs and real-time performance data, thereby automatically triggering relevant actions to allocate different types and quantities of computing power resources according to different jobs.
[0009] Preferably, the storage sub-cluster includes storage capacity, storage utilization, storage node IOPS, and storage node read / write speed;
[0010] The computing sub-cluster includes node CPU, node memory, and network traffic;
[0011] The management sub-cluster includes the cluster service status, management node CPU, and management node memory.
[0012] The computing power scheduling module uses machine learning algorithms to analyze user behavior and job characteristics based on the detailed job history and real-time information uploaded by the computing power monitoring module, thereby obtaining the different resource requirements of different jobs, and then scheduling resources according to the FCFS+Backfilling algorithm, genetic algorithm or ant colony algorithm.
[0013] The computing power monitoring module includes a computing power monitoring server and a computing power monitoring client. The computing power monitoring server collects hardware performance data including CPU usage, DCU usage, memory usage, and IO usage. The computing power monitoring server collects job execution data including job queuing, job execution, job type, and job time.
[0014] The computing power monitoring client adopts a RESTful style interaction mode to reduce the performance overhead of data interaction; the computing power monitoring server adopts a distributed architecture, using server clusters to collect massive amounts of monitoring data, thereby improving data collection efficiency.
[0015] Compared with the prior art, the present invention has the following beneficial effects:
[0016] The resource management module dynamically manages nodes, storage, and networks based on resource pools. By adding, reducing, and migrating cluster nodes and adjusting queue resource ranges, it enables dynamic expansion of computing clusters and virtualization clusters, as well as efficient utilization and load balancing of resources within the cluster.
[0017] The computing power scheduling service is a decentralized architecture that uses a distributed processing mechanism to solve problems such as scale and reliability, effectively preventing the risk of a single point of failure causing the entire system to crash.
[0018] The computing power monitoring module manages and monitors computing, storage, network, basic business components, and application components through layered and partitioned management. It also processes massive amounts of heterogeneous data, providing data support and technical framework support for artificial intelligence and big data platforms. By combining AI and operations and maintenance (O&M), it achieves efficient and intelligent O&M and operations, quickly and accurately detecting anomalies and locating faults. It provides unified monitoring, intelligent alarms, fault analysis, multi-dimensional views, and statistical reports. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] In the picture:
[0021] Figure 1 This is a schematic diagram of the advanced computing platform integration system in an embodiment of the present invention;
[0022] Figure 2 This is a schematic diagram of the workflow of the computing power scheduling module in an embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0024] Figure 1 This is a schematic diagram of an advanced computing platform integration system according to an embodiment of the present invention. The heterogeneous computing resource optimization scheduling and management method of the present invention is applied to an advanced computing platform integration system. The advanced computing platform integration system includes a computing resource management module, a computing scheduling module, and a computing monitoring module. The computing monitoring module has communication connections with both the computing resource management module and the computing scheduling module. The computing resource management module includes a storage sub-cluster, a computing sub-cluster, and a management sub-cluster.
[0025] The computing power monitoring module provides hardware performance data and job operation data to the computing power scheduling module and computing power resource management module; the job operation data includes computing resource information and job logs for various types of jobs, and the hardware performance data includes real-time performance data.
[0026] The computing resource management module analyzes and processes the computing resource information of various types of jobs provided by the computing monitoring module to form a dynamic resource pool. The dynamic resource pool includes a high-performance computing resource pool, a general computing resource pool, a big data resource pool, and a high IO resource pool. The computing resource management module dynamically manages nodes, storage, and networks based on the dynamic resource pool. By adding, reducing, and migrating cluster nodes and adjusting the range of queue resources, the module achieves dynamic expansion of computing clusters and virtualization clusters, as well as efficient utilization and load balancing of resources within the cluster.
[0027] The computing power scheduling module uses machine learning algorithms to analyze and determine the computing power requirements of each service based on job logs and real-time performance data, thereby automatically triggering relevant actions to allocate different types and quantities of computing power resources according to different jobs.
[0028] In this way, the computing power monitoring module provides the computing power scheduling module and resource management module with real-time hardware performance data, job execution data, and historical logs. This assists the resource management module in effectively allocating and managing resources, while the computing power scheduling module analyzes the monitoring data to perform efficient computing power scheduling, meeting the resource requirements of jobs and ensuring their rapid and stable operation. All modules work collaboratively to ensure the efficient operation of the advanced computing power platform integrated system.
[0029] The working principle of the present invention will be further explained below.
[0030] Regarding the computing power resource management module: Current applications have vastly different computing power requirements, ranging from a few cores that can complete a task in a few hours to tens of thousands of cores working together for several months. Faced with such complex computing power demands and structures, and the need to further improve resource utilization, it is necessary to optimize computing power resource management. This invention addresses these diverse computing power needs by establishing a computing power management module based on the structure of heterogeneous supercomputing. This module enables real-time monitoring of application computing power requirements and allows for elastic scaling based on application needs.
[0031] By leveraging massive amounts of real-time performance data acquired from a large-scale cluster computing power monitoring platform, computing resources are divided into different dynamic computing power resource pools. This enables elastic scaling and dynamic expansion of resources, ensuring efficient resource allocation and enhancing the synergy between hardware and software. Furthermore, resource management strategies are optimized in reverse based on changes in monitoring data resulting from management policies, continuously iterating to enhance the computing power resource management module's ability to efficiently utilize computing, storage, and network resources.
[0032] Regarding the computing power scheduling module:
[0033] The massive data processing and computing power demands generated in the digital age have placed higher requirements on the performance of supercomputers in all aspects. As industries develop, their performance requirements for computing power are becoming increasingly complex and diverse. To meet these new demands, supercomputer structures are becoming more complex, and efficiently matching different computing power needs and allocating computing resources has become a key factor affecting computational efficiency.
[0034] like Figure 2 As shown, the computing power scheduling module fully utilizes the detailed job history and real-time information uploaded by the computing power monitoring module, such as job logs and real-time performance data. Then, machine learning algorithms are used to analyze user behavior and job characteristics to obtain the different resource requirements of different jobs, thereby enabling intelligent computing power scheduling. For example, the FCFS+Backfilling algorithm is used for intelligent resource scheduling. Building on this, heuristic algorithms based on animal behavior, such as genetic algorithms and ant colony algorithms, can also be used for the job scheduling process. These algorithms simulate the job scheduling process as genetic evolution and animal foraging behaviors, achieving better resource allocation efficiency.
[0035] Regarding the computing power monitoring module: Existing monitoring software has significant problems when applying ultra-large computing power clusters. On the one hand, these software programs cannot directly obtain underlying information, switch information, and switch chip information. On the other hand, when facing ultra-large computing power clusters, the scale and dimensions of the data collected by monitoring sensors are enormous, resulting in severe performance overhead, easily causing monitoring data loss, and even directly rendering the monitoring system unusable. This invention establishes a computing power monitoring module for heterogeneous supercomputing, combined with a computing power management and scheduling module, to achieve efficient management and utilization of computing power.
[0036] The computing power monitoring module includes a computing power monitoring server and a computing power monitoring client. The computing power monitoring server collects hardware performance data including CPU usage, DCU usage, memory usage, and IO usage. The computing power monitoring server collects job execution data including job queuing, job execution, job type, and job time.
[0037] The computing power monitoring client adopts a RESTful interaction model, meaning that the interaction between the client and the server does not generate additional state, thereby reducing the performance overhead of data interaction. The computing power monitoring server adopts a distributed architecture, utilizing a server cluster to collect massive amounts of monitoring data, thereby improving data collection efficiency.
[0038] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0039] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0040] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0041] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention. The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0042] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0043] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0044] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for optimizing scheduling and management of heterogeneous computing resources, applied to an advanced computing platform integration system, characterized in that, The advanced computing platform integrated system includes a computing resource management module, a computing scheduling module, and a computing monitoring module. The computing monitoring module is communicatively connected to the computing resource management module and the computing scheduling module, respectively. The computing resource management module includes a storage sub-cluster, a computing sub-cluster, and a management sub-cluster. The computing power monitoring module provides hardware performance data and job operation data to the computing power scheduling module and computing power resource management module; the job operation data includes computing resource information and job logs for various types of jobs, and the hardware performance data includes real-time performance data; The computing resource management module analyzes and processes the computing resource information of various types of jobs provided by the computing monitoring module to form a dynamic resource pool. The dynamic resource pool includes a high-performance computing resource pool, a general computing resource pool, a big data resource pool, and a high IO resource pool. The computing resource management module dynamically manages nodes, storage, and networks based on the dynamic resource pool. By adding, reducing, and migrating cluster nodes and queue resource ranges, the module achieves dynamic expansion of computing clusters and virtualization clusters, as well as efficient utilization and load balancing of resources within the cluster. The computing power scheduling module uses machine learning algorithms to analyze and determine the computing power requirements of each service based on job logs and real-time performance data, thereby automatically triggering relevant actions to allocate different types and quantities of computing power resources according to different jobs.
2. The method as described in claim 1, characterized in that, The storage sub-cluster includes storage capacity, storage utilization, storage node IOPS, and storage node read / write speed; The computing sub-cluster includes node CPU, node memory, and network traffic; The management sub-cluster includes the cluster service status, management node CPU, and management node memory.
3. The method as described in claim 2, characterized in that, The computing power scheduling module uses machine learning algorithms to analyze user behavior and job characteristics based on the detailed job history and real-time information uploaded by the computing power monitoring module, thereby obtaining the different resource requirements of different jobs, and then scheduling resources according to the FCFS+Backfilling algorithm, genetic algorithm or ant colony algorithm.
4. The method as described in claim 1, characterized in that, The computing power monitoring module includes a computing power monitoring server and a computing power monitoring client. The computing power monitoring server collects hardware performance data including CPU usage, DCU usage, memory usage, and IO usage. The computing power monitoring server collects job execution data including job queuing, job execution, job type, and job time. The computing power monitoring client adopts a RESTful style interaction mode to reduce the performance overhead of data interaction; the computing power monitoring server adopts a distributed architecture, using server clusters to collect massive amounts of monitoring data, thereby improving data collection efficiency.