Low-energy-consumption high-efficiency cluster management method and system for heterogeneous computing units

By strengthening the hyper-heuristic prediction method and dynamic resource management online, the problems of unreasonable resource allocation and excessive energy consumption in heterogeneous computing clusters are solved, and efficient and low-energy computing effects are achieved.

CN119988022AInactive Publication Date: 2025-05-13BORNSALES SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510094923.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

There are problems in the management of existing heterogeneous computing clusters that are unreasonable resource allocation, lack of dynamic adjustment mechanisms and excessive energy consumption.

Method used

The online enhanced hyper-heuristic prediction method is used to predict the number requirements of heterogeneous computing units, dynamically configure and manage resource buffers according to requirements, and pre-scheduling decisions of heterogeneous computing units are generated to optimize resource allocation.

Benefits of technology

Through reasonable task allocation and dynamic adjustment mechanisms, computing efficiency and system performance are significantly improved, energy consumption is reduced, and system flexibility is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988022A_ABST
    Figure CN119988022A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of heterogeneous computing unit cluster management, in particular to a low-energy-consumption efficient cluster management method and system for heterogeneous computing units. The method comprises the following steps: creating a reserved resource buffer area in a cluster based on a quantity demand rt + 1 of heterogeneous computing units in a (t + 1) th period, and dynamically configuring and managing the resource buffer area according to a task resource demand change; a pre-scheduling decision of the heterogeneous computing unit is generated according to a task demand submitted by a user, and the heterogeneous computing unit is scheduled based on the pre-scheduling decision. Through innovative design of task analysis and classification, a dynamic allocation strategy and a dynamic adjustment mechanism, the energy efficiency can be remarkably improved, the system flexibility is enhanced, the resource utilization rate is optimized, and a new solution is provided for management of heterogeneous computing clusters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of heterogeneous computing unit cluster management, and in particular to a low-energy consumption and high-efficiency cluster management method and system for heterogeneous computing units. Background Art

[0002] Currently, in the management of heterogeneous computing clusters, although there are some methods and systems for resource scheduling and task allocation, there are still some significant problems and shortcomings.

[0003] Unreasonable resource allocation: Existing resource management systems often do not fully consider the characteristics of tasks and the advantages of computing units, resulting in unreasonable resource allocation and low energy efficiency. For example, assigning tasks that require high parallel computing capabilities to CPU processing will greatly reduce computing efficiency and energy efficiency.

[0004] Lack of dynamic adjustment mechanism: With the dynamic arrival of tasks and changes in cluster load, existing resource management systems often lack dynamic adjustment mechanisms and are unable to optimize resource allocation in real time based on actual conditions.

[0005] Excessive energy consumption: Due to unreasonable resource allocation and lack of dynamic adjustment mechanism, existing heterogeneous computing clusters often have the problem of excessive energy consumption, which not only increases operating costs but also violates the concept of green computing.

[0006] In order to solve the above problems, a low-energy and high-efficiency cluster management method and system for heterogeneous computing units came into being. Summary of the invention

[0007] The purpose of the present invention is to provide a low-energy and high-efficiency cluster management method and system for heterogeneous computing units: to solve the technical problems of unreasonable resource allocation, lack of dynamic adjustment mechanism and excessive energy consumption in the management of heterogeneous computing clusters in existing solutions.

[0008] The purpose of the present invention can be achieved through the following technical solutions:

[0009] In one aspect, a low-energy and high-efficiency cluster management method for heterogeneous computing units is provided, the method comprising:

[0010] Get the actual number of heterogeneous computing units required in the tth period r t , based on the online enhanced hyper-inspiration prediction method, predict the number of heterogeneous computing units required in the t+1 period r t+1 ;

[0011] Based on the number of heterogeneous computing units required in the t+1th cycle r t+1 Does it exceed the preset number of heterogeneous computing units in the cluster? If so, based on the heterogeneous computing unit number requirement r in the t+1th cycle t+1Create a reserved resource buffer in the cluster, dynamically configure and manage the resource buffer according to changes in task resource requirements; if not, generate pre-scheduling decisions for heterogeneous computing units based on the task requirements submitted by users, and schedule heterogeneous computing units based on the pre-scheduling decisions.

[0012] Furthermore, based on the online enhanced hyper-inspiration prediction method, the number of heterogeneous computing units required in the t+1 period is predicted. t+1 The specific process includes the following:

[0013] Receive the actual number of heterogeneous computing units required in the tth period r t ;

[0014] Based on r t Calculate the loss Loss(k, t) of the kth underlying prediction method in the tth period:

[0015]

[0016] Where Pre(k, t) is the prediction result of the kth underlying prediction method in the tth period, δ is the impact factor that the predicted demand for the number of heterogeneous computing units is greater than the actual demand, ε is the impact factor that the predicted demand for the number of heterogeneous computing units is less than the actual demand, and the prediction result is the demand for the number of heterogeneous computing units;

[0017] Sort in descending order based on Loss(k, t) to obtain the descending rank rank(k, t) corresponding to Loss(k, t);

[0018] Calculate the overall performance P of the kth underlying prediction method in the first t cycles based on descending rank rank(k, t) t (k):

[0019] P t (k) = γ × P t-1 (k)+(1-γ)×rank(k,t);

[0020] Among them, γ is the learning factor, 0≤γ≤1. The larger the value of γ, the more the performance of the prediction method is affected by its historical performance. t-1 (k) is the overall performance of the kth underlying prediction method in the first t-1 cycles;

[0021] At the beginning of the t+1th cycle, the upper heuristic method obtains the number of heterogeneous computing units required r according to the following formula t+1 :

[0022]

[0023] Among them, Z tIt is a set of sub-prediction methods selected by the upper-level heuristic method based on the overall performance of each underlying prediction method. is the sum of all presidential performances in the sub-prediction method set, where the kth underlying prediction method belongs to Z t , Pre(z, t) is the prediction result of the z-th sub-prediction method in the t-th cycle, where the prediction result is the required number of heterogeneous computing units.

[0024] Furthermore, based on the number of heterogeneous computing units required in the t+1th cycle r t+1 Creating a reserved resource buffer in a cluster includes the following steps:

[0025]

[0026] Among them, C hc Indicates the capacity of the resource buffer, R MAX represents the peak resource demand of delay-sensitive tasks, R P represents the resource requirements of delay-sensitive tasks predicted by the Prophet model, n 1 Indicates the preset number of heterogeneous computing units.

[0027] Furthermore, dynamically configuring and managing the resource buffer according to changes in task resource requirements specifically includes the following processes:

[0028] Calculate the amount of resources required to fill the resource buffer according to the amount of resource buffer resources and the capacity of the resource buffer;

[0029] Get the resource buffer resources and the working nodes where the allocated reserved resources are located and count the idle resources of the working nodes. If the amount of idle resources on the working node is less than or equal to the amount of resources required to fill the buffer, put these idle resources into the buffer and update the amount of resources required to fill the buffer. If the amount of idle resources on the working node is greater than the amount of resources required to fill the buffer, sort these working nodes in descending order according to the amount of idle resources of the nodes, and put the idle resources of the nodes into the buffer in turn until the resource buffer is filled.

[0030] Furthermore, generating a pre-scheduling decision for a heterogeneous computing unit according to the task requirements submitted by the user specifically includes the following process:

[0031] Obtain the delay-sensitive tasks in the queue in the cluster according to the task requirements submitted by the user, and determine whether there is any correlation between the delay-sensitive tasks. If so, group the delay-sensitive tasks with correlation into one group;

[0032] Determine whether the idle resources in the cluster can meet the resource requirements of all tasks to be scheduled in the task group. If so, set the idle resources of the node as the available resources of the working node; if not, set the sum of the idle resources and redundant resources on the node as the available resources of the working node:

[0033]

[0034] Among them, r avlb Represents the available resources on the working node, r idle represents the amount of idle resources on the working node, R idle Represents the idle resources in the entire cluster, R rqst Represents the sum of resource requirements of all tasks to be scheduled;

[0035] Determine whether there are any associated tasks of the task group to be scheduled among the tasks in the running state of the cluster. If so, use the working node where the associated running task is located as the candidate working node for scheduling tasks in the group. Otherwise, select a candidate working node from the entire cluster.

[0036] When pre-allocating work nodes for tasks, determine whether there are multiple work nodes among the candidate work nodes whose available resources meet the total resource requirements of the unscheduled tasks in the current group. If so, select the work node with available resources close to the total resource requirements as the scheduling target node for these tasks; if not, select the node with the most available resources from the candidate work nodes as the target node for task scheduling, and place the most tasks in the task group on this work node, and then repeat the operation in this way until all tasks in the group are scheduled.

[0037] Furthermore, determining whether there is a correlation between delay-sensitive tasks specifically includes the following process:

[0038] Field names and field types for latency-sensitive tasks;

[0039] The association between fields is determined based on the field names and field types according to association rule mining.

[0040] Furthermore, selecting a candidate working node from the entire cluster specifically includes the following process:

[0041]

[0042] Among them, r avlb-i It represents the available resources of the ith working node after all working nodes in the cluster are arranged in descending order of available resources. rqst It represents the total resource demand of unscheduled tasks in the current group, and n represents the minimum number of working nodes that can meet the resource demand of the tasks to be scheduled.

[0043] Furthermore, the heterogeneous computing unit includes a CPU and a GPU.

[0044] On the other hand, a low-energy and high-efficiency cluster management system for heterogeneous computing units includes:

[0045] Heterogeneous computing unit quantity prediction unit, used to obtain the actual quantity demand r of heterogeneous computing units in the tth period t , based on the online enhanced hyper-inspiration prediction method, predict the number of heterogeneous computing units required in the t+1 period r t+1 ;

[0046] The judgment unit is used to determine the number of heterogeneous computing units required in the t+1th cycle r t+1 Whether the number of heterogeneous computing units in the cluster exceeds the preset number;

[0047] Cluster management unit, used to calculate the number of heterogeneous computing units required in the t+1th period r t+1 Create a reserved resource buffer in the cluster, dynamically configure and manage the resource buffer according to changes in task resource requirements; generate pre-scheduling decisions for heterogeneous computing units based on task requirements submitted by users, and schedule heterogeneous computing units based on the pre-scheduling decisions.

[0048] Compared with the existing solutions, the present invention achieves the following beneficial effects:

[0049] The present invention is based on the number of heterogeneous computing units required in the t+1th cycle r t+1 A reserved resource buffer is created in the cluster, and the resource buffer is dynamically configured and managed according to changes in task resource requirements; a pre-scheduling decision for heterogeneous computing units is generated according to the task requirements submitted by the user, and the heterogeneous computing units are scheduled based on the pre-scheduling decision. The present invention achieves more efficient computing effects and system performance through reasonable task allocation and dynamic adjustment mechanisms, thereby significantly reducing energy consumption.

[0050] Furthermore, the flexibility of the system is enhanced: the present invention has a high degree of flexibility and can dynamically allocate computing resources according to the specific requirements of the task to adapt to various scenarios and requirements.

[0051] In summary, through the innovative design of task analysis and classification, dynamic allocation strategy and dynamic adjustment mechanism, the present invention can significantly improve energy efficiency, enhance system flexibility and optimize resource utilization, and provide a new solution for the management of heterogeneous computing clusters. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0053] Figure 1 It is a workflow diagram of a low-energy-consumption and high-efficiency cluster management method for heterogeneous computing units according to an embodiment of the present invention;

[0054] Figure 2 It is a workflow diagram of another low-energy-consumption and high-efficiency cluster management method of heterogeneous computing units according to an embodiment of the present invention;

[0055] Figure 3 It is a workflow diagram of another low-energy-consumption and high-efficiency cluster management method of heterogeneous computing units according to an embodiment of the present invention;

[0056] Figure 4 It is a system block diagram of a low-energy-consumption and high-efficiency cluster management system of a heterogeneous computing unit according to an embodiment of the present invention. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0058] In addition, the described features, structures or characteristics may be combined in one or more example embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the example embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, steps, etc. may be adopted. In other cases, well-known structures, methods, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0059] This embodiment provides a low-energy and high-efficiency cluster management method for heterogeneous computing units. Figure 1 is a workflow diagram of a low-energy-consumption and high-efficiency cluster management method for heterogeneous computing units according to an embodiment of the present invention, such as Figure 1 As shown, the method comprises the following steps:

[0060] Step S101: Obtain the actual number of heterogeneous computing units required in the tth period, and predict the number of heterogeneous computing units required in the t+1th period based on the online enhanced hyper-heuristic prediction method, where r t The actual number of heterogeneous computing units required in the tth period is represented by r t+1 Indicates the number of heterogeneous computing units required in the t+1th cycle;

[0061] Step S102: Based on the number of heterogeneous computing units required in the t+1th cycle r t+1 Whether the number of heterogeneous computing units preset in the cluster exceeds the number, if yes, proceed to step S103, if no, proceed to step S104;

[0062] Step S103: creating a reserved resource buffer in the cluster based on the number of heterogeneous computing units required in the t+1th cycle, and dynamically configuring and managing the resource buffer according to changes in task resource requirements;

[0063] Step S104: Generate a pre-scheduling decision for the heterogeneous computing unit according to the task requirements submitted by the user, and schedule the heterogeneous computing unit based on the pre-scheduling decision.

[0064] In summary, the present invention is based on the number of heterogeneous computing units required in the t+1th cycle r t+1 A reserved resource buffer is created in the cluster, and the resource buffer is dynamically configured and managed according to changes in task resource requirements; a pre-scheduling decision for heterogeneous computing units is generated according to the task requirements submitted by the user, and the heterogeneous computing units are scheduled based on the pre-scheduling decision. Through the innovative design of task analysis and classification, dynamic allocation strategy and dynamic adjustment mechanism, the present invention can significantly improve energy efficiency, enhance system flexibility and optimize resource utilization, and provide a new solution for the management of heterogeneous computing clusters.

[0065] In some embodiments, the number of heterogeneous computing units required in the t+1th period is predicted based on an online enhanced hyper-heuristic prediction method. t+1 The specific process includes the following:

[0066] Receive the actual number of heterogeneous computing units required in the tth period r t ;

[0067] Based on r t Calculate the loss Loss(k, t) of the kth underlying prediction method in the tth period:

[0068]

[0069] Where Pre(k, t) is the prediction result of the kth underlying prediction method in the tth period, δ is the impact factor that the predicted demand for the number of heterogeneous computing units is greater than the actual demand, ε is the impact factor that the predicted demand for the number of heterogeneous computing units is less than the actual demand, and the prediction result is the demand for the number of heterogeneous computing units;

[0070] Sort in descending order based on Loss(k, t) to obtain the descending rank rank(k, t) corresponding to Loss(k, t);

[0071] Calculate the overall performance P of the kth underlying prediction method in the first t cycles based on descending rank rank(k, t) t (k):

[0072] P t (k) = γ × P t-1 (k)+(1-γ)×rank(k,t);

[0073] Among them, γ is the learning factor, 0≤γ≤1. The larger the value of γ, the more the performance of the prediction method is affected by its historical performance. t-1 (k) is the overall performance of the kth underlying prediction method in the first t-1 cycles;

[0074] At the beginning of the t+1th cycle, the upper heuristic method obtains the number of heterogeneous computing units required r according to the following formula t+1 :

[0075]

[0076] Among them, Z t It is a set of sub-prediction methods selected by the upper-level heuristic method based on the overall performance of each underlying prediction method. is the sum of all presidential performances in the sub-prediction method set, where the kth underlying prediction method belongs to Z t , Pre(z, t) is the prediction result of the z-th sub-prediction method in the t-th cycle, where the prediction result is the required number of heterogeneous computing units.

[0077] In some embodiments, Figure 2 is a workflow diagram of another low-energy-consumption and high-efficiency cluster management method for heterogeneous computing units according to an embodiment of the present invention, such as Figure 2 As shown, dynamically configuring and managing the resource buffer according to the changes in task resource requirements specifically includes the following steps:

[0078] Step S201: Calculate the amount of resources required to fill the resource buffer according to the amount of resource buffer resources and the capacity of the resource buffer; wherein the difference between the capacity of the resource buffer and the amount of resource buffer resources is recorded as the amount of resources required to fill the resource buffer;

[0079] Step S202: Obtain resource buffer resources and the working nodes where the allocated reserved resources are located and count the idle resources of the working nodes;

[0080] Step S203: If the amount of free resources on the working node is less than or equal to the amount of resources required to fill the buffer, put these free resources into the buffer and update the amount of resources required to fill the buffer. If the amount of free resources on the working node is greater than the amount of resources required to fill the buffer, arrange these working nodes in descending order according to the amount of free resources on the node, and put the free resources of the node into the buffer in turn until the resource buffer is filled.

[0081] In some embodiments, generating a pre-scheduling decision for a heterogeneous computing unit according to a task requirement submitted by a user specifically includes the following process:

[0082] Obtain the delay-sensitive tasks in the queue in the cluster according to the task requirements submitted by the user, and determine whether there is any correlation between the delay-sensitive tasks. If so, group the delay-sensitive tasks with correlation into one group;

[0083] Determine whether the idle resources in the cluster can meet the resource requirements of all tasks to be scheduled in the task group. If so, set the idle resources of the node as the available resources of the working node; if not, set the sum of the idle resources and redundant resources on the node as the available resources of the working node:

[0084]

[0085] Among them, r avlb Represents the available resources on the working node, r idle represents the amount of idle resources on the working node, R idle Represents the idle resources in the entire cluster, R rqst Represents the sum of resource requirements of all tasks to be scheduled;

[0086] Determine whether there are any associated tasks of the task group to be scheduled among the tasks in the running state of the cluster. If so, use the working node where the associated running task is located as the candidate working node for scheduling tasks in the group. Otherwise, select a candidate working node from the entire cluster.

[0087] When pre-allocating work nodes for tasks, determine whether there are multiple work nodes among the candidate work nodes whose available resources meet the total resource requirements of the unscheduled tasks in the current group. If so, select the work node with available resources close to the total resource requirements as the scheduling target node for these tasks; if not, select the node with the most available resources from the candidate work nodes as the target node for task scheduling, and place the most tasks in the task group on this work node, and then repeat the operation in this way until all tasks in the group are scheduled.

[0088] In some embodiments, Figure 3 is a workflow diagram of another low-energy-consumption and high-efficiency cluster management method for heterogeneous computing units according to an embodiment of the present invention, such as Figure 3 As shown, determining whether there is a correlation between delay-sensitive tasks specifically includes the following steps:

[0089] Step S301: extracting the field name and field type of the delay-sensitive task;

[0090] Specifically, all relevant data fields of latency-sensitive tasks are identified based on NLP (Natural Language Processing): field name, field Chinese name, field type. Data fields are columns or attributes used to store data in a database. According to the type of data stored and its purpose, data fields can be divided into different types, including primary keys, foreign keys, text fields, and numeric fields.

[0091] Step S302: Determine the association between fields based on the field names and field types according to association rule mining.

[0092] Specifically, data preparation: select the data set to be analyzed and ensure data quality; determine the fields to be analyzed, which can be numerical, categorical, or textual. Perform necessary preprocessing on the data, such as data cleaning, conversion, and discretization (for continuous data).

[0093] Define support and confidence:

[0094] Support: The frequency with which an item set appears in all transactions. In field correlation analysis, it can be understood as the frequency with which two or more field values ​​appear at the same time.

[0095] Confidence: The conditional probability that Y is also included in the transaction containing X, expressed as Support(X,Y) / Support(X). In field correlation analysis, it can be understood as the probability that field Y takes a certain value when field X takes a certain value.

[0096] Apply the association rule mining algorithm:

[0097] Association rule mining algorithms such as Apriori and FP-Growth are used to identify frequent item sets and generate association rules. The association between fields is determined based on field names and field types based on association rule mining.

[0098] In some embodiments, selecting a candidate working node from the entire cluster specifically includes the following process:

[0099]

[0100] Among them, r avlb-i It represents the available resources of the ith working node after all working nodes in the cluster are arranged in descending order of available resources. rqst It represents the total resource demand of unscheduled tasks in the current group, and n represents the minimum number of working nodes that can meet the resource demand of the tasks to be scheduled.

[0101] In some embodiments, the heterogeneous computing unit includes a CPU and a GPU.

[0102] In some embodiments, Figure 4 is a system block diagram of a low-energy-consumption and high-efficiency cluster management system of a heterogeneous computing unit according to an embodiment of the present invention. Figure 4 As shown, the system includes:

[0103] Heterogeneous computing unit quantity prediction unit, used to obtain the actual quantity demand r of heterogeneous computing units in the tth period t , based on the online enhanced hyper-inspiration prediction method, predict the number of heterogeneous computing units required in the t+1 period r t+1 ;

[0104] The judgment unit is used to determine the number of heterogeneous computing units required in the t+1th cycle r t+1 Whether the number of heterogeneous computing units in the cluster exceeds the preset number;

[0105] Cluster management unit, used to calculate the number of heterogeneous computing units required in the t+1th period r t+1 Create a reserved resource buffer in the cluster, dynamically configure and manage the resource buffer according to changes in task resource requirements; generate pre-scheduling decisions for heterogeneous computing units based on task requirements submitted by users, and schedule heterogeneous computing units based on the pre-scheduling decisions.

[0106] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0107] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0108] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0109] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only some logical function divisions. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0110] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0111] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A low-energy and high-efficiency cluster management method for heterogeneous computing units, characterized in that: Methods include: Get the actual number of heterogeneous computing units required in the tth period r t , based on the online enhanced hyper-inspiration prediction method, predict the number of heterogeneous computing units required in the t+1 period r t+1 ; Based on the number of heterogeneous computing units required in the t+1th cycle r t+1 Does it exceed the preset number of heterogeneous computing units in the cluster? If so, based on the heterogeneous computing unit number requirement r in the t+1th cycle t+1 Create a reserved resource buffer in the cluster, dynamically configure and manage the resource buffer according to changes in task resource requirements; if not, generate pre-scheduling decisions for heterogeneous computing units based on the task requirements submitted by users, and schedule heterogeneous computing units based on the pre-scheduling decisions.

2. A low-energy and high-efficiency cluster management method for heterogeneous computing units according to claim 1, characterized in that: Predict the number of heterogeneous computing units required in the t+1th period based on the online enhanced hyper-heuristic prediction method t+1 The specific process includes the following: Receive the actual number of heterogeneous computing units required in the tth period r t ; Based on r t Calculate the loss Loss(k, t) of the kth underlying prediction method in the tth period: Where Pre(k, t) is the prediction result of the kth underlying prediction method in the tth period, δ is the impact factor that the predicted demand for the number of heterogeneous computing units is greater than the actual demand, ε is the impact factor that the predicted demand for the number of heterogeneous computing units is less than the actual demand, and the prediction result is the demand for the number of heterogeneous computing units; Sort in descending order based on Loss(k, t) to obtain the descending rank rank(k, t) corresponding to Loss(k, t); Calculate the overall performance P of the kth underlying prediction method in the first t cycles based on descending rank rank(k, t) t (k): P t (k)=γ×P t-1 (k)+(1-γ)×rank(k,t); Among them, γ is the learning factor, 0≤γ≤1. The larger the value of γ, the more the performance of the prediction method is affected by its historical performance. t-1 (k) is the overall performance of the kth underlying prediction method in the first t-1 cycles; At the beginning of the t+1th cycle, the upper heuristic method obtains the number of heterogeneous computing units required r according to the following formula t+1 : Among them, Z t is the set of sub-prediction methods selected by the upper-level heuristic method according to the overall performance of each underlying prediction method, ∑ k∈Zt P t (z) is the sum of all presidential performances in the sub-prediction method set, where the kth underlying prediction method belongs to Z t , Pre(z, t) is the prediction result of the z-th sub-prediction method in the t-th cycle, where the prediction result is the required number of heterogeneous computing units.

3. The low-energy and high-efficiency cluster management method of heterogeneous computing units according to claim 1 is characterized in that: Based on the number of heterogeneous computing units required in the t+1th cycle r t+1 Creating a reserved resource buffer in a cluster includes the following steps: Among them, C hc Represents the capacity of the resource buffer, R MAX represents the peak resource demand of delay-sensitive tasks, R P It represents the resource demand of delay-sensitive tasks predicted by the Prophet model, and n1 represents the number of preset heterogeneous computing units.

4. The low-energy and high-efficiency cluster management method of heterogeneous computing units according to claim 2, characterized in that: Dynamically configuring and managing resource buffers based on changes in task resource requirements includes the following processes: Calculate the amount of resources required to fill the resource buffer according to the amount of resource buffer resources and the capacity of the resource buffer; Get the resource buffer resources and the working nodes where the allocated reserved resources are located and count the idle resources of the working nodes. If the amount of idle resources on the working node is less than or equal to the amount of resources required to fill the buffer, put these idle resources into the buffer and update the amount of resources required to fill the buffer. If the amount of idle resources on the working node is greater than the amount of resources required to fill the buffer, sort these working nodes in descending order according to the amount of idle resources of the nodes, and put the idle resources of the nodes into the buffer in turn until the resource buffer is filled.

5. The low-energy and high-efficiency cluster management method of heterogeneous computing units according to claim 1, characterized in that: The pre-scheduling decision of heterogeneous computing units generated according to the task requirements submitted by the user specifically includes the following processes: Obtain the delay-sensitive tasks in the queue in the cluster according to the task requirements submitted by the user, and determine whether there is any correlation between the delay-sensitive tasks. If so, group the delay-sensitive tasks with correlation into one group; Determine whether the idle resources in the cluster can meet the resource requirements of all tasks to be scheduled in the task group. If so, set the idle resources of the node as the available resources of the working node; if not, set the sum of the idle resources and redundant resources on the node as the available resources of the working node: Among them, r avlb Represents the available resources on the working node, r idle represents the amount of idle resources on the working node, R idle Represents the idle resources in the entire cluster, R rqst Represents the sum of resource requirements of all tasks to be scheduled; Determine whether there are any associated tasks of the task group to be scheduled among the tasks in the running state of the cluster. If so, use the working node where the associated running task is located as the candidate working node for scheduling tasks in the group. Otherwise, select a candidate working node from the entire cluster. When pre-allocating work nodes for tasks, determine whether there are multiple work nodes among the candidate work nodes whose available resources meet the total resource requirements of the unscheduled tasks in the current group. If so, select the work node with available resources close to the total resource requirements as the scheduling target node for these tasks; if not, select the node with the most available resources from the candidate work nodes as the target node for task scheduling, and place the most tasks in the task group on this work node, and then repeat the operation in this way until all tasks in the group are scheduled.

6. A low-energy-consumption and high-efficiency cluster management method for heterogeneous computing units according to claim 5, characterized in that: Determine whether there is a correlation between delay-sensitive tasks The process includes: Field names and field types for latency-sensitive tasks; The association between fields is determined based on the field names and field types according to association rule mining.

7. A low-energy consumption and high-efficiency cluster management method for heterogeneous computing units according to claim 6, characterized in that: The process of selecting a candidate working node from the entire cluster specifically includes the following: Among them, r avlb-i It represents the available resources of the ith working node after all working nodes in the cluster are arranged in descending order of available resources. rqst It represents the total resource demand of the unscheduled tasks in the current group, and n represents the minimum number of working nodes that can meet the resource demand of the tasks to be scheduled.

8. The low-energy and high-efficiency cluster management method of heterogeneous computing units according to claim 1, characterized in that: Heterogeneous computing units include CPU and GPU.

9. A low-energy and high-efficiency cluster management system for heterogeneous computing units, characterized in that: A low-energy-consumption and high-efficiency cluster management method for heterogeneous computing units applicable to any one of claims 1 to 8, the system comprising: Heterogeneous computing unit quantity prediction unit, used to obtain the actual quantity demand r of heterogeneous computing units in the tth period t , based on the online enhanced hyper-inspiration prediction method, predict the number of heterogeneous computing units required in the t+1 period r t+1 ; The judgment unit is used to determine the number of heterogeneous computing units required in the t+1th cycle r t+1 Whether the number of heterogeneous computing units in the cluster exceeds the preset number; Cluster management unit, used to calculate the number of heterogeneous computing units required in the t+1th period r t+1 Create a reserved resource buffer in the cluster, dynamically configure and manage the resource buffer according to changes in task resource requirements; generate pre-scheduling decisions for heterogeneous computing units based on task requirements submitted by users, and schedule heterogeneous computing units based on the pre-scheduling decisions.

Citation Information

Cited By

  • Heterogeneous computing cluster deployment method and collaborative scheduling system

    CN120429130A