Resource management method, system and equipment for high-performance computing cluster and medium
By collecting the number of queued jobs and resource usage data of the high-performance computing cluster, generating resource management information and sending reminder messages, the problem of unbalanced resource usage of clusters is solved and the efficiency of computing clusters is improved.
Patent Information
- Application Number
- CN202510159079.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-06
AI Technical Summary
Due to the large scale and numerous jobs in high-performance computing clusters, administrators cannot perceive resource usage overall, resulting in unbalanced resource usage, affecting the efficiency of the computing cluster usage and job efficiency.
By calling the scheduler device to collect and calculate the number of queued jobs in the queue, and combining the computing node resource usage data collected by the performance information collection device, the resource utilization rate of the cluster and queue is generated, and resource management information is generated, and a reminder message is sent to the administrator through the resource reminder device to adjust the resource configuration in time.
Timely management and optimization of high-performance computing cluster resources is achieved, and resource utilization and overall efficiency of computing clusters are improved.
Smart Images

Figure CN120104314A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computing technology, and in particular to a resource management method, system, device and medium for a high-performance computing cluster. Background Art
[0002] High Performance Computing (HPC) uses computer clusters integrated with various technologies to jointly process computing problems of massive data, greatly improving computing speed and accuracy, and is widely used in various fields. However, due to the large scale of HPC clusters and the large number of jobs, cluster administrators cannot perceive the overall resource usage within the HPC cluster, resulting in unbalanced resource usage.
[0003] Existing technical solutions usually collect resource usage data and display it to the administrator, who then determines the resource usage based on the resource usage data of each computing node. This results in poor management efficiency and makes it impossible to adjust cluster resources in a timely manner, which affects the cluster's usage efficiency and operating efficiency.
[0004] Therefore, an efficient computing resource management method is urgently needed to solve the above technical problems. Summary of the invention
[0005] Based on this, it is necessary to provide a resource management method, system, device and medium for a high-performance computing cluster to address the above technical issues.
[0006] In a first aspect, the present application provides a resource management method for a high-performance computing cluster, wherein the high-performance computing cluster includes multiple computing queues, each computing queue includes one or more computing nodes, and is applied to a resource early warning judgment device. In response to the arrival of a preset task cycle, the method includes:
[0007] Calling the queued job numbers of each computing queue collected by the scheduler device and summarizing the queued job numbers of the high-performance computing cluster to generate the cluster queued job number;
[0008] Calling the node resource usage data on each computing node collected by the performance information collection device, wherein the node resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device;
[0009] Determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data;
[0010] Generate resource management information based on the number of queued jobs, queue resource utilization, cluster resource utilization, and the number of cluster queued jobs;
[0011] The resource management information is sent to the resource reminder device to trigger the resource reminder device to generate a reminder message.
[0012] In some embodiments, resource management information is generated according to the number of queued jobs, the queue resource utilization rate, the cluster resource utilization rate, and the number of cluster queued jobs, including:
[0013] Determine the cluster resource level of the high-performance cluster according to the cluster resource utilization rate and the number of cluster queued jobs, where the cluster resource level includes a first cluster resource level, a second cluster resource level, and a third cluster resource level;
[0014] In response to detecting that the high-performance cluster is at the first cluster resource level, generating resource management information for adding a cluster computing node;
[0015] In response to detecting that the high-performance cluster is at a second cluster resource level, determining and generating resource management information based on queue resource levels of respective computing queues in the high-performance cluster and based on the queue levels;
[0016] In response to detecting that the high performance cluster is at the third cluster resource level, resource management information for increasing the cluster workload is generated.
[0017] In some embodiments, determining the cluster resource level of the high-performance cluster according to the cluster resource utilization rate and the number of cluster queued jobs includes:
[0018] In response to detecting that the cluster resource usage is greater than or equal to a first threshold and the number of cluster queued jobs is greater than or equal to a second threshold, determining that the cluster resource usage is a first cluster resource level;
[0019] In response to detecting that the cluster resource usage is less than a first threshold and the number of cluster queued jobs is less than a second threshold, determining that the cluster resource usage is a second cluster resource level;
[0020] In response to detecting that the cluster resource usage is less than a third threshold, it is determined that the cluster resource usage is a third cluster resource level.
[0021] In some embodiments, the queue resource level includes a first queue resource level, a second queue resource level, and a third queue resource level, and determining and based on the queue resource level of each computing queue in the high-performance cluster includes:
[0022] In response to detecting that the number of queued jobs of the computing queue is greater than or equal to a fourth threshold and the queue resource usage rate is greater than or equal to a fifth threshold, determining that the computing queue is at a first queue resource level;
[0023] In response to detecting that the number of queued jobs of the computing queue is greater than or equal to a fourth threshold and the queue resource usage is less than a fifth threshold, determining that the computing queue is at a second queue resource level;
[0024] In response to detecting that the number of queued jobs of the computing queue is less than a sixth threshold, the computing queue is determined to be at a third queue resource level.
[0025] In some embodiments, generating resource management information according to the queue level includes:
[0026] In response to detecting that there is a first computing queue with a first queue resource level, generating resource management information for adding a computing node in the first computing queue;
[0027] In response to detecting the presence of a second computing queue at a second queue resource level, generating resource management information for managing job objects of the second computing queue;
[0028] In response to detecting the existence of a third queue whose resource level is a third computing queue, resource management information for increasing the workload of the third computing queue is generated.
[0029] In some embodiments, resource usage data is collected by a performance information reporting device deployed on each computing node and reported to a performance information collecting device, including:
[0030] The performance information reporting device encrypts the collected resource usage data corresponding to the computing node and the first node configuration data based on an asymmetric encryption algorithm to generate first encrypted data;
[0031] The performance information reporting device calls the first interface address to report the first encrypted data to the performance information collecting device at every preset time period;
[0032] The performance information collection device decrypts the first encrypted data based on an asymmetric encryption algorithm to obtain resource usage data of the computing node and first node configuration data.
[0033] In some embodiments, before determining the queue resource usage rate corresponding to each computing queue and the cluster resource usage rate of the high performance computing cluster according to the node resource usage data, the method further includes:
[0034] calling second node configuration data of the computing node stored in the scheduler device;
[0035] Calling the first node configuration data on each computing node collected by the performance information collection device;
[0036] In response to detecting that the first node configuration information and the second node configuration data of the same computing node do not match, a node configuration update instruction is generated and sent to the scheduler device to trigger the scheduler device to update the second node configuration data according to the first node configuration data.
[0037] In a second aspect, the present application provides a resource management system for a high-performance computing cluster, the system comprising a resource early warning judgment device, a scheduler device, a performance information reporting device, a performance information collection device, and a resource reminder device;
[0038] The resource early warning judgment device is used to call the queued job numbers of each computing queue collected by the scheduler device and summarize the queued job numbers of the high-performance computing cluster to generate the cluster queued job number;
[0039] A resource early warning judgment device, used to call the node resource usage data on each computing node collected by the performance information collection device, wherein the resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device;
[0040] A resource early warning judgment device is used to determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data;
[0041] A resource early warning judgment device is used to generate resource management information according to the number of queued jobs, the queue resource utilization rate, the cluster resource utilization rate and the number of cluster queued jobs;
[0042] The resource early warning judgment device is used to send resource management information to the resource reminder device to trigger the resource reminder device to generate a reminder message.
[0043] In a third aspect, the present application provides a computer program product, which implements the steps of the following method when the computer program is executed by a processor:
[0044] Calling the queued job numbers of each computing queue collected by the scheduler device and summarizing the queued job numbers of the high-performance computing cluster to generate the cluster queued job number;
[0045] Calling the node resource usage data on each computing node collected by the performance information collection device, wherein the node resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device;
[0046] Determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data;
[0047] Generate resource management information based on the number of queued jobs, queue resource utilization, cluster resource utilization, and the number of cluster queued jobs;
[0048] The resource management information is sent to the resource reminder device to trigger the resource reminder device to generate a reminder message.
[0049] In a fourth aspect, the present application provides an electronic device, the electronic device comprising: one or more processors;
[0050] and a memory associated with one or more processors, the memory being used to store program instructions, which, when read and executed by the one or more processors, perform the following operations:
[0051] Calling the queued job numbers of each computing queue collected by the scheduler device and summarizing the queued job numbers of the high-performance computing cluster to generate the cluster queued job number;
[0052] Calling the node resource usage data on each computing node collected by the performance information collection device, wherein the node resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device;
[0053] Determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data;
[0054] Generate resource management information based on the number of queued jobs, queue resource utilization, cluster resource utilization, and the number of cluster queued jobs;
[0055] The resource management information is sent to the resource reminder device to trigger the resource reminder device to generate a reminder message.
[0056] In a fifth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and the computer program enables a computer to perform the following operations:
[0057] Calling the queued job numbers of each computing queue collected by the scheduler device and summarizing the queued job numbers of the high-performance computing cluster to generate the cluster queued job number;
[0058] Calling the node resource usage data on each computing node collected by the performance information collection device, wherein the node resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device;
[0059] Determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data;
[0060] Generate resource management information based on the number of queued jobs, queue resource utilization, cluster resource utilization, and the number of cluster queued jobs;
[0061] The resource management information is sent to the resource reminder device to trigger the resource reminder device to generate a reminder message.
[0062] The beneficial effects achieved by this application are:
[0063] The present application provides a resource management method for a high-performance computing cluster, which is applied to a resource early warning judgment device, including calling a scheduler device to collect the queue queue job numbers of each computing queue and aggregating the queue queue job numbers of the high-performance computing cluster to generate a cluster queue job number; calling a performance information collection device to collect node resource usage data on each computing node, wherein the node resource usage data is collected by a performance information reporting device deployed on each computing node and reported to the performance information collection device; determining the queue resource usage rate corresponding to each computing queue and the cluster resource usage rate of the high-performance computing cluster according to the node resource usage data; generating resource management information according to the queue queue job number, the queue resource usage rate, the cluster resource usage rate and the cluster queue job number; and sending the resource management information to a resource reminder device to trigger the resource reminder device to generate a reminder message. By setting up a performance data reporting device on the computing node, it is ensured that the node resource usage data can be reported in a timely manner, and the timeliness of the reporting is ensured by using a multi-threaded method; the resource early warning judgment device summarizes and analyzes the node resource usage data reported by the performance data reporting device, judges the resource level of the current computing queue and computing cluster, and further generates matching resource management information, which is displayed to the management personnel in a timely manner through the resource reminder device so that the resources within the cluster can be managed and optimized. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work, among which:
[0065] Figure 1 It is a schematic diagram of a resource management method for a high-performance computing cluster provided in an embodiment of the present application;
[0066] Figure 2 This is a resource management system architecture diagram of a high-performance computing cluster provided in an embodiment of the present application;
[0067] Figure 3 This is another resource management system architecture diagram of a high performance computing cluster provided in an embodiment of the present application;
[0068] Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0070] It should be understood that in the description of the present application, unless the context clearly requires otherwise, words such as "include", "comprises", and the like in the entire specification and claims should be interpreted as inclusive rather than exclusive or exhaustive; that is, the meaning of "including but not limited to".
[0071] It should also be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.
[0072] It should be noted that the terms "S1", "S2", etc. are only used for the purpose of describing the steps, and do not specifically refer to the order or sequence, nor are they used to limit the present application. They are only for the convenience of describing the method of the present application, and cannot be understood as indicating the order of the steps. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in this field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0073] As mentioned in the background technology, high performance computing has been widely used in many fields such as simulation modeling and big data analysis; high performance computing clusters are large-scale computing clusters composed of hundreds or even thousands of servers; and by grouping the cluster machines into multiple queues for management, the jobs are assigned to the servers of a single queue to run. Due to the large size of the cluster and the large number of jobs, the cluster administrator cannot perceive the overall cluster resource usage and the resource usage of each queue, and there will be problems with unbalanced resource usage. For example, there is a situation where cluster resource usage is too low, and some queues use less resources, resulting in idle resources for some machines; while some queues have more job tasks and are seriously short of resources. It is necessary to remind the resource administrator to shut down the nodes with low resource usage, remove the queue, or migrate them to the queues with tight resources, so as to reduce server energy consumption and improve resource utilization.
[0074] However, resource usage is constantly changing, and administrators cannot keep track of cluster resource usage. Currently implemented solutions mostly collect resource usage data through resource monitoring and display it to administrators, who then judge resource usage based on each node's resource usage data and their own experience. However, this solution requires administrators to have considerable cluster usage and management experience, and to keep track of resource data, making it impossible to adjust cluster resources in a timely manner, which affects the usage efficiency of high-performance clusters and the efficiency of job operations.
[0075] Embodiment 1
[0076] The embodiment of the present application provides a resource management method for a high-performance computing cluster. Specifically, in a resource early warning judgment device, when a preset task cycle arrives, such as Figure 1 As shown, the method disclosed in the embodiment of the present application is used to realize reasonable management of resources of the computing queue in the computing cluster, including the following contents:
[0077] S1. Call the queue job numbers of each computing queue collected by the scheduler device and summarize the queue job numbers of the high-performance computing cluster to generate the cluster queue job number.
[0078] The above-mentioned task cycle is set by technicians in this field according to the duration of job operation and the actual scenario, and is preferably set to 10 minutes, which is not limited in this application. It is understandable that the task cycle is regularly maintained in the configuration file of the above-mentioned resource early warning judgment device, and the above-mentioned resource early warning judgment device is deployed in the cluster management node. The above-mentioned resource early warning judgment device calls the scheduler device in an internal call mode to obtain the number of queued jobs of each computing queue according to the set task cycle when the task cycle arrives.
[0079] Among them, the scheduler device is deployed in the cluster management node, such as the SLURM (Simple Linux Utility for Resource Management) scheduler or interface. The specific deployment method can be achieved by downloading the SLURM scheduler installation package, executing the compilation and installation commands; this application does not limit the specific scheduler type. The scheduler device is responsible for providing the queue queue job number and cluster queue job number interface of the queue; the above-mentioned resource early warning judgment device can directly call the scheduler device to obtain the queue queue job number of each computing queue in the current high-performance computing cluster; it can be understood that the scheduler device also maintains the second node configuration data of each computing node initially input manually, and also maintains the relationship data between the queue and the node, that is, which computing nodes a computing queue includes. In order to facilitate searching, the scheduler device can also save the above-mentioned maintained data in the database.
[0080] S2. Calling the performance information collection device to collect the node resource usage data on each computing node.
[0081] Similar to the call of the above-mentioned scheduler device, when the task cycle arrives, the above-mentioned resource warning judgment device calls the node resource usage data of each computing node collected by the performance information collection device in an internal call manner. Specifically, the above-mentioned node resource usage data includes but is not limited to the total number of node CPU cores, the number of GPU cards, node memory, the number of CPU cores used, the number of GPU cards used, and memory usage.
[0082] The resource usage data is collected by the performance information reporting device deployed at each computing node and reported to the performance information collection device. The resource usage data is collected by the performance information reporting device deployed at each computing node and reported to the performance information collection device. Specifically, the performance information reporting device encrypts the collected resource usage data corresponding to the computing node and the first node configuration data based on an asymmetric encryption algorithm to generate first encrypted data, such as the RAS algorithm; the performance information reporting device calls the first interface address to report the first encrypted data to the performance information collection device at every preset time period; the performance information collection device decrypts the first encrypted data based on the asymmetric encryption algorithm to obtain the resource usage data of the computing node and the first node configuration data. The above-mentioned preset time period is less than the above-mentioned task cycle, and is preferably set to 1 minute. It is set by those skilled in the art according to actual needs, and this application does not limit this. In some implementation scenarios, the performance information collection device updates the resource usage information and first node configuration information of the node recorded in the cache based on the collected resource usage data and first node configuration data, and saves the information to the database, updates the first node configuration information matching the computing node in the database, and facilitates the subsequent formation of node performance reports; saves the latest node performance data to the cache to increase the speed of reading the latest node performance data. It is understandable that because the calling frequency is fixed and not very frequent, the performance information reporting device and the performance information collection device preferably transmit data through a short connection such as the http protocol; the data is transmitted in the form of an RSA encrypted json string.
[0083] It is understandable that if the performance information reporting device cannot collect information within the specified time, it means that the node is abnormal, including network abnormality, state abnormality (abnormal shutdown), device abnormality, etc., and needs to be checked according to the situation. Network abnormality can be determined by ping, ssh, etc., state abnormality can be checked according to the power status, and device abnormality can be checked according to the device operation log. At this time, the performance information reporting device generates node abnormality information and sends it to the resource reminder device. The above node abnormality information includes at least abnormal contacts and abnormal types; then the resource reminder device can remind the management personnel that the computing node is abnormal through station letters, emails, etc. The above-mentioned specified time is determined by technical personnel in this field based on the actual cluster configuration data, and this application does not limit this.
[0084] S3. Determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high performance computing cluster according to the node resource utilization data.
[0085] Specifically, for a computing queue, the node resource usage data of all computing nodes in the computing queue is summarized, and the number of CPU cores used, the number of GPU cards used, and the memory usage of the computing nodes are summarized according to different data types to determine the number of CPU cores used in the queue, the number of GPU cards used in the queue, and the memory usage of the queue corresponding to the computing queue; then the total number of node CPU cores, the number of GPU cards, and the node memory of all computing nodes in the computing queue are summarized to determine the total number of CPU cores, the number of GPU cards, and the queue memory available in the computing queue. The first usage rate is further determined by the ratio of the number of CPU cores used in the queue to the total number of CPU cores in the queue, the second usage rate is determined according to the ratio of the number of GPU cards used in the queue to the number of GPU cards in the queue, and the third usage rate is determined according to the ratio of the queue memory usage to the queue memory; the queue resource usage rate of the computing queue is further generated by weighted summing of the first usage rate, the second usage rate, and the third usage rate, wherein the specific weight selection is configured by those skilled in the art according to the actual scenario. Furthermore, the method for determining the cluster resource utilization rate includes: calculating the product of the resource proportion of the computing queue in the cluster and the calculated queue resource utilization rate to generate a fourth utilization rate, and summing the fourth utilization rates matched by each computing queue included in the cluster to generate the cluster resource utilization rate. Through the above method, when the resource early warning judgment device responds to the arrival of the task cycle, the accurate cluster resource utilization rate of the high-performance computing cluster in the current state and the queue resource utilization rate of each computing queue under the cluster are timely obtained.
[0086] In some implementation scenarios, the above-mentioned performance information collection device can also determine the CPU usage rate according to the ratio of the number of CPU cores used by the uploaded computing node to the total number of CPU cores of the node, the ratio of the number of GPU cards used to the number of GPU cards to determine the GPU usage rate, the ratio of memory usage to node memory and node memory usage rate and other node performance data, and then save the node performance data, CPU cores used, GPU cards used, memory usage; then save the node performance data to the node performance history record table and cache (memory, reids, etc.), the main fields in the table include the total number of CPU cores of the node, the number of GPU cards, memory size, CPU cores used, GPU cards used, memory usage, CPU usage rate, GPU, usage rate, memory usage rate, in order to reduce the database storage pressure, the partition and table scheme is used to process the data, that is, a database table partition is generated every month, and last year's table is modified to a historical table every year, and a new current database table is generated. It can be understood that only the latest node performance information of the node is saved in the cache. By forming resource and job report data for nodes, queues, and clusters, that is, providing node performance history curve reports, it helps administrators to promptly discover resource allocation imbalances, resource waste, resource shortages, etc.; it is convenient for administrators to track performance history and provide data support for administrators to adjust cluster resources.
[0087] In some implementation scenarios, before the resource early warning judgment device determines the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data, it is also necessary to correct the scheduler device: call the second node configuration data of the computing node stored in the scheduler device; call the first node configuration data on each computing node collected by the performance information collection device; in response to detecting that the first node configuration information and the second node configuration data of the same computing node do not match, generate a node configuration update instruction and send it to the scheduler device to trigger the scheduler device to update the second node configuration data according to the first node configuration data. That is, compare the total number of node cpu cores, the number of gpu cards, and the memory size in the node configuration information collected by the performance information collection device with the total number of node cpu cores, the number of gpu cards, and the memory size stored in the scheduler device. If the two do not match, modify the relevant data in the scheduler device. By correcting the scheduler device, the accuracy of the queue resource utilization rate and the cluster resource utilization rate determined subsequently is ensured, and the accuracy of the early warning is further improved.
[0088] S4. Generate resource management information according to the number of queued jobs, the queue resource utilization rate, the cluster resource utilization rate and the number of cluster queued jobs.
[0089] Specifically, the resource management information is generated according to the number of queued jobs, the queue resource utilization rate, the cluster resource utilization rate and the number of cluster queued jobs, including:
[0090] According to the cluster resource utilization rate and the number of cluster queued jobs, the cluster resource level of the high-performance cluster is determined, and the cluster resource level includes the first cluster resource level, the second cluster resource level and the third cluster resource level; in response to detecting that the high-performance cluster is the first cluster resource level, resource management information for adding cluster computing nodes is generated; in response to detecting that the high-performance cluster is the second cluster resource level, the queue resource level of each computing queue in the high-performance cluster is determined and resource management information is generated according to the queue level; in response to detecting that the high-performance cluster is the third cluster resource level, resource management information for increasing the cluster job volume is generated. Wherein, according to the cluster resource utilization rate and the number of cluster queued jobs, the cluster resource level of the high-performance cluster is determined, including: in response to detecting that the cluster resource utilization rate is greater than or equal to the first threshold and the number of cluster queued jobs is greater than or equal to the second threshold, the cluster resource utilization rate is determined to be the first cluster resource level; in response to detecting that the cluster resource utilization rate is less than the first threshold and the number of cluster queued jobs is less than the second threshold, the cluster resource utilization rate is determined to be the second cluster resource level; in response to detecting that the cluster resource utilization rate is less than the third threshold, the cluster resource utilization rate is determined to be the third cluster resource level. The first threshold is the maximum utilization rate of cluster resources that the high-performance cluster can provide normal computing power, which is determined according to the actual high-performance cluster configuration; the second threshold is the maximum number of cluster queued jobs allowed by the high-performance cluster, which is determined according to the actual high-performance cluster configuration; the third threshold is the idle resource utilization rate pre-set in the high-performance cluster, that is, if it is less than the idle resource utilization rate, it means that the high-performance cluster is an idle cluster, and the idle resource utilization rate is determined according to the actual high-performance cluster configuration. The present application pre-analyzes the resource usage of the entire high-performance computing cluster, and directly triggers the generation of resource management information when the resource usage is too much or too little, accelerating the management process of cluster resource anomalies. In addition, when the computing cluster can operate normally, it further analyzes the resource usage of each computing queue, speeds up the entire resource analysis management process, and avoids unnecessary waste of analysis resources.
[0091] Further, the above-mentioned queue resource level includes a first queue resource level, a second queue resource level and a third queue resource level. The above-mentioned determination and based on the queue resource level of each computing queue in the high-performance cluster specifically includes: in response to detecting that the number of queued jobs of the computing queue is greater than or equal to the fourth threshold and the queue resource utilization rate is greater than or equal to the fifth preset threshold, determining that the computing queue is at the first queue resource level; in response to detecting that the number of queued jobs of the computing queue is greater than or equal to the fourth threshold and the queue resource utilization rate is less than the fifth preset threshold, determining that the computing queue is at the second queue resource level; in response to detecting that the number of queued jobs of the computing queue is less than the sixth preset threshold, determining that the computing queue is at the third queue resource level. The above-mentioned fourth threshold is the maximum utilization rate of cluster resources that the computing queue can provide normal computing power, which is determined according to the actual configuration of the computing queue; the above-mentioned fifth threshold is the maximum number of cluster queued jobs allowed by the computing queue, which is determined according to the actual configuration of the computing queue; the above-mentioned sixth threshold is the idle resource utilization rate pre-set in the computing queue, that is, if it is less than the idle resource utilization rate, it means that the computing queue is an idle cluster, and the above-mentioned idle resource utilization rate is determined according to the actual configuration of the computing queue. It is understandable that in specific scenarios, the fourth threshold, fifth threshold, and sixth threshold matched by different computing queues may be different. In specific implementation scenarios, the present application also provides a web-based interface, where managers can directly set the first to sixth thresholds, making it easy to adjust the thresholds in real time according to actual usage scenarios.
[0092] In a specific implementation scenario, the above-mentioned generation of resource management information according to the queue level specifically includes: in response to detecting the existence of a first computing queue with a first queue resource level, generating resource management information for adding computing nodes in the first computing queue, that is, the queue may have insufficient resources, and it is necessary to call a resource reminder device to remind the administrator that there is a risk of insufficient resources related to the computing queue, and it is necessary to migrate computing nodes from a third computing queue or a computing queue outside the cluster to the first computing queue; in response to detecting the existence of a second computing queue with a second queue resource level, generating resource management information for managing the job objects of the second computing queue, that is, the resource configuration of the computing queue is unreasonable, and the job application resources submitted to the queue are too large, and it is necessary to call a resource reminder device to remind the administrator to manage the relationship between the job objects and queues of the second computing queue, reduce the job objects of the second computing queue, and bind the reduced job objects to other queues with sufficient resources; in response to detecting the existence of a third queue with a third computing queue resource level, generating resource management information for increasing the job volume of the third computing queue, that is, there are idle resources in the third computing queue, and it is necessary to call a resource reminder device to remind the administrator to execute instructions to remove some computing nodes from the third computing queue or shut down some scattered nodes according to the resource relationship information. By using the method for generating management information provided by the resource early warning judgment device, the management information of the computing queue can be generated in a timely manner and sent to the resource reminder device, which will then display it to the administrator to solve the problem of unbalanced resource allocation in the cluster queues, thereby improving the overall operating efficiency of the computing cluster.
[0093] S5. Send the resource management information to the resource reminder device to trigger the resource reminder device to generate a reminder message.
[0094] Specifically, the resource warning judgment device sends the generated multiple resource management information to the resource reminder device through internal calling, and the resource reminder device displays the resource management information to the user according to a preset reminder method, such as email, so that the user can execute the corresponding management instructions.
[0095] The resource management method provided in the present application ensures that the node resource usage data can be reported in a timely manner by setting up a performance data reporting device on the computing node, and uses a multi-threaded approach to ensure the timeliness of the reporting; the resource early warning judgment device summarizes and analyzes the node resource usage data reported by the performance data reporting device, judges the resource level of the current computing queue and computing cluster, and further generates matching resource management information, which is displayed to the management personnel in a timely manner through the resource reminder device so that the resources within the cluster can be managed and optimized.
[0096] Embodiment 2
[0097] Corresponding to the above-mentioned embodiment 1, the present application also provides a resource management system for a high-performance computing cluster, such as Figure 2 The architecture diagram shown specifically includes: a resource early warning judgment device, a scheduler device, a performance information reporting device, a performance information collection device, and a resource reminder device;
[0098] The resource early warning judgment device is used to call the queued job numbers of each computing queue collected by the scheduler device and summarize the queued job numbers of the high-performance computing cluster to generate the cluster queued job number;
[0099] A resource early warning judgment device, used to call the node resource usage data on each computing node collected by the performance information collection device, wherein the resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device;
[0100] A resource early warning judgment device is used to determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data;
[0101] A resource early warning judgment device is used to generate resource management information according to the number of queued jobs, the queue resource utilization rate, the cluster resource utilization rate and the number of cluster queued jobs;
[0102] The resource early warning judgment device is used to send resource management information to the resource reminder device to trigger the resource reminder device to generate a reminder message.
[0103] In some implementation scenarios, the resource warning judgment device is also used to determine the cluster resource level of the high-performance cluster based on the cluster resource utilization rate and the number of cluster queued jobs, the cluster resource levels including the first cluster resource level, the second cluster resource level and the third cluster resource level; the resource warning judgment device is also used to generate resource management information for adding cluster computing nodes in response to detecting that the high-performance cluster is at the first cluster resource level; the resource warning judgment device is also used to determine and generate resource management information based on the queue resource level of each computing queue within the high-performance cluster in response to detecting that the high-performance cluster is at the second cluster resource level; the resource warning judgment device is also used to generate resource management information for increasing the cluster job volume in response to detecting that the high-performance cluster is at the third cluster resource level.
[0104] In some implementation scenarios, the resource warning judgment device is also used to determine that the cluster resource utilization rate is a first cluster resource level in response to detecting that the cluster resource utilization rate is greater than or equal to a first threshold and the number of cluster queued jobs is greater than or equal to a second threshold; the resource warning judgment device is also used to determine that the cluster resource utilization rate is a second cluster resource level in response to detecting that the cluster resource utilization rate is less than the first threshold and the number of cluster queued jobs is less than the second threshold; the resource warning judgment device is also used to determine that the cluster resource utilization rate is a third cluster resource level in response to detecting that the cluster resource utilization rate is less than a third threshold.
[0105] In some implementation scenarios, the resource warning judgment device is also used to determine that the computing queue is at the first queue resource level in response to detecting that the number of queued jobs of the computing queue is greater than or equal to a fourth threshold and the queue resource utilization rate is greater than or equal to a fifth threshold; the resource warning judgment device is also used to determine that the computing queue is at the second queue resource level in response to detecting that the number of queued jobs of the computing queue is greater than or equal to a fourth threshold and the queue resource utilization rate is less than a fifth threshold; the resource warning judgment device is also used to determine that the computing queue is at the third queue resource level in response to detecting that the number of queued jobs of the computing queue is less than a sixth threshold.
[0106] In some implementation scenarios, the resource warning judgment device is also used to generate resource management information for increasing computing nodes within the first computing queue in response to detecting the existence of a first computing queue with a first queue resource level; the resource warning judgment device is also used to generate resource management information for managing job objects of the second computing queue in response to detecting the existence of a second computing queue with a second queue resource level; the resource warning judgment device is also used to generate resource management information for increasing the workload of the third computing queue in response to detecting the existence of a third queue with a third queue resource level.
[0107] In some implementation scenarios, the performance information reporting device encrypts the collected resource usage data corresponding to the computing node and the first node configuration data based on an asymmetric encryption algorithm to generate first encrypted data; the performance information reporting device calls the first interface address to report the first encrypted data to the performance information collection device at every preset time period; the performance information collection device decrypts the first encrypted data based on the asymmetric encryption algorithm to obtain the resource usage data of the computing node and the first node configuration data.
[0108] In some implementation scenarios, the resource early warning judgment device is also used to call the second node configuration data of the computing node stored in the scheduler device; the resource early warning judgment device is also used to call the first node configuration data on each computing node collected by the performance information collection device; the resource early warning judgment device is also used to generate a node configuration update instruction in response to detecting that the first node configuration information and the second node configuration data of the same computing node do not match, and send it to the scheduler device to trigger the scheduler device to update the second node configuration data according to the first node configuration data.
[0109] In some implementation scenarios, such as Figure 3As shown, the above-mentioned resource management system also includes a database and a cache. The performance information collection device updates the resource usage information and the first node configuration information of the node recorded in the cache according to the collected resource usage data and the first node configuration data, and saves the information to the database, and updates the first node configuration information matching the computing node in the database to facilitate the subsequent formation of a node performance report; the latest node performance data is saved to the cache to improve the speed of reading the latest node performance data.
[0110] Embodiment 3
[0111] The present application provides a computer program product, which implements the following steps when the computer program is executed by a processor:
[0112] Calling the queued job numbers of each computing queue collected by the scheduler device and summarizing the queued job numbers of the high-performance computing cluster to generate the cluster queued job number;
[0113] Calling the node resource usage data on each computing node collected by the performance information collection device, wherein the node resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device;
[0114] Determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data;
[0115] Generate resource management information based on the number of queued jobs, queue resource utilization, cluster resource utilization, and the number of cluster queued jobs;
[0116] The resource management information is sent to the resource reminder device to trigger the resource reminder device to generate a reminder message.
[0117] Embodiment 4
[0118] Corresponding to all the above embodiments, an embodiment of the present application provides an electronic device, including: one or more processors; and a memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the following operations are performed:
[0119] Calling the queued job numbers of each computing queue collected by the scheduler device and summarizing the queued job numbers of the high-performance computing cluster to generate the cluster queued job number;
[0120] Calling the node resource usage data on each computing node collected by the performance information collection device, wherein the node resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device;
[0121] Determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data;
[0122] Generate resource management information based on the number of queued jobs, queue resource utilization, cluster resource utilization, and the number of cluster queued jobs;
[0123] The resource management information is sent to the resource reminder device to trigger the resource reminder device to generate a reminder message.
[0124] in, Figure 4 The architecture of the electronic device is exemplarily shown, which may include a processor 410, a video display adapter 411, a disk drive 412, an input / output interface 413, a network interface 414, and a memory 420. The processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, and the memory 420 may be communicatively connected via a bus 430.
[0125] Among them, the processor 410 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solution provided in this application.
[0126] The memory 420 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 420 can store an operating system 421 for controlling the execution of the electronic device 400, and a basic input and output system (BIOS) 422 for controlling the low-level operations of the electronic device 400. In addition, a web browser 423, a data storage management system 424, and an icon font processing system 425, etc. can also be stored. The above-mentioned icon font processing system 425 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided in the present application is implemented by software or firmware, the relevant program code is stored in the memory 420 and is called and executed by the processor 410.
[0127] The input / output interface 413 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0128] The network interface 414 is used to connect to a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0129] The bus 430 comprises a pathway for transmitting information between the various components of the device (eg, the processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, and the memory 420).
[0130] In addition, the electronic device 400 can also obtain information on specific collection conditions from the virtual resource object collection condition information database for use in condition judgment, etc.
[0131] It should be noted that, although the above device only shows a processor 410, a video display adapter 411, a disk drive 412, an input / output interface 413, a network interface 414, a memory 420, a bus 430, etc., in the specific implementation process, the device may also include other components necessary for normal execution. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.
[0132] It can be seen from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application can essentially or in other words, the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a disk, an optical disk, etc., including several instructions for a computer device (which can be a personal computer, a cloud service end, or a network device, etc.) to execute the methods of various embodiments of the present application or certain parts of the embodiments.
[0133] Embodiment 5
[0134] Corresponding to all the above embodiments, the embodiments of the present application further provide a computer-readable storage medium storing a computer program, which enables a computer to perform the following operations:
[0135] Calling the queued job numbers of each computing queue collected by the scheduler device and summarizing the queued job numbers of the high-performance computing cluster to generate the cluster queued job number;
[0136] Calling the node resource usage data on each computing node collected by the performance information collection device, wherein the node resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device;
[0137] Determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data;
[0138] Generate resource management information based on the number of queued jobs, queue resource utilization, cluster resource utilization, and the number of cluster queued jobs;
[0139] The resource management information is sent to the resource reminder device to trigger the resource reminder device to generate a reminder message.
[0140] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, in which the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0141] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A resource management method for a high performance computing cluster, characterized in that: The high-performance computing cluster includes a plurality of computing queues, each of which includes one or more computing nodes, and is applied to a resource early warning judgment device. In response to the arrival of a preset task cycle, the method includes: Calling the queued job numbers of each computing queue collected by the scheduler device and summarizing the queued job numbers of the high-performance computing cluster to generate the cluster queued job number; Calling the node resource usage data on each computing node collected by the performance information collection device, wherein the node resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device; Determine the queue resource usage rate corresponding to each computing queue and the cluster resource usage rate of the high performance computing cluster according to the node resource usage data; Generate resource management information according to the number of queued jobs in the queue, the resource utilization rate of the queue, the resource utilization rate of the cluster, and the number of queued jobs in the cluster; The resource management information is sent to the resource reminder device to trigger the resource reminder device to generate a reminder message.
2. The method according to claim 1, characterized in that The generating resource management information according to the number of queued jobs in the queue, the queue resource utilization rate, the cluster resource utilization rate and the number of queued jobs in the cluster includes: Determining a cluster resource level of the high-performance cluster according to the cluster resource usage rate and the number of cluster queued jobs, wherein the cluster resource level includes a first cluster resource level, a second cluster resource level, and a third cluster resource level; In response to detecting that the high-performance cluster is at the first cluster resource level, generating resource management information for adding a cluster computing node; In response to detecting that the high-performance cluster is at a second cluster resource level, determining and generating resource management information according to queue resource levels of respective computing queues within the high-performance cluster and according to the queue levels; In response to detecting that the high-performance cluster is at the third cluster resource level, resource management information for increasing the cluster workload is generated.
3. The method according to claim 2, characterized in that The determining the cluster resource level of the high-performance cluster according to the cluster resource utilization rate and the number of cluster queued jobs includes: In response to detecting that the cluster resource usage is greater than or equal to a first threshold and the number of cluster queued jobs is greater than or equal to a second threshold, determining that the cluster resource usage is the first cluster resource level; In response to detecting that the cluster resource usage is less than the first threshold and the number of cluster queued jobs is less than the second threshold, determining that the cluster resource usage is the second cluster resource level; In response to detecting that the cluster resource usage is less than a third threshold, determining that the cluster resource usage is the third cluster resource level.
4. The method according to claim 2, characterized in that: The queue resource level includes a first queue resource level, a second queue resource level, and a third queue resource level, and the determining and determining the queue resource level of each computing queue in the high-performance cluster includes: In response to detecting that the number of queued jobs of the computing queue is greater than or equal to a fourth threshold and the queue resource usage rate is greater than or equal to a fifth threshold, determining that the computing queue is at a first queue resource level; In response to detecting that the number of queued jobs of the computing queue is greater than or equal to a fourth threshold and the queue resource usage rate is less than the fifth threshold, determining that the computing queue is at a second queue resource level; In response to detecting that the number of queued jobs of the computing queue is less than a sixth threshold, the computing queue is determined to be at a third queue resource level.
5. The method according to claim 4, characterized in that The generating resource management information according to the queue level includes: In response to detecting that there is a first computing queue with a first queue resource level, generating resource management information for adding computing nodes in the first computing queue; In response to detecting the presence of a second computing queue of a second queue resource level, generating resource management information for managing job objects of the second computing queue; In response to detecting the existence of a third queue resource level being a third computing queue, resource management information for increasing the workload of the third computing queue is generated.
6. The method according to claim 1, characterized in that The resource usage data is collected by a performance information reporting device deployed on each computing node and reported to the performance information collecting device, including: The performance information reporting device encrypts the collected resource usage data corresponding to the computing node and the first node configuration data based on an asymmetric encryption algorithm to generate first encrypted data; The performance information reporting device calls the first interface address to report the first encrypted data to the performance information collecting device at intervals of a preset time period; The performance information collection device decrypts the first encrypted data based on the asymmetric encryption algorithm to obtain the resource usage data of the computing node and the first node configuration data.
7. The method according to claim 6 is applied to a resource early warning judgment device, characterized in that: Before determining the queue resource usage rate corresponding to each computing queue and the cluster resource usage rate of the high performance computing cluster according to the node resource usage data, the method further includes: calling second node configuration data of the computing node stored in the scheduler device; Calling the first node configuration data on each computing node collected by the performance information collection device; In response to detecting that the first node configuration information and the second node configuration data of the same computing node do not match, a node configuration update instruction is generated and sent to the scheduler device to trigger the scheduler device to update the second node configuration data according to the first node configuration data.
8. A resource management system for a high performance computing cluster, characterized in that: The system includes a resource early warning judgment device, a scheduler device, a performance information reporting device, a performance information collection device and a resource reminder device; The resource early warning judgment device is used to call the queued job numbers of each computing queue collected by the scheduler device and summarize the queued job numbers of the high-performance computing cluster to generate the cluster queued job number; A resource early warning judgment device, used to call the node resource usage data on each computing node collected by the performance information collection device, wherein the resource usage data is collected by the performance information reporting device deployed on each computing node and reported to the performance information collection device; A resource early warning judgment device, used to determine the queue resource utilization rate corresponding to each computing queue and the cluster resource utilization rate of the high-performance computing cluster according to the node resource utilization data; A resource early warning judgment device, used for generating resource management information according to the number of queued jobs in the queue, the resource utilization rate of the queue, the resource utilization rate of the cluster and the number of queued jobs in the cluster; The resource early warning judgment device is used to send the resource management information to the resource reminder device to trigger the resource reminder device to generate a reminder message.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; And a memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the method according to any one of claims 1-7 is executed.
10. A computer-readable storage medium, characterized in that: The computer stores a computer program, wherein the computer program enables a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Multi-service queue elastic consumption regulation and control method based on database load awareness
CN120295795A