HPC cluster benchmark energy efficiency adjusting and optimizing method based on software and hardware joint optimization
By adopting the joint optimization method of software and hardware in the HPC cluster, the energy efficiency sensitive parameters are screened and the parameter optimization process is optimized with the energy efficiency prediction model, and combined with the dynamic frequency regulation strategy, the problems of insufficient universality and large overhead in the existing technology are solved, and the benchmark energy efficiency of HPC clusters is improved.
Patent Information
- Application Number
- CN202411875555.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The existing HPC cluster benchmark energy efficiency optimization technology has problems such as insufficient generality, high overhead for model training and acquisition, and lack of a global optimization perspective.
Using a method based on joint software and hardware optimization, the energy efficiency sensitive parameters are screened through software optimization, and combined with the sampling algorithm in the optimization process of the energy efficiency prediction model optimization parameter, reducing the overhead of optimization and model training. In terms of hardware, dynamic frequency regulation is achieved to improve energy efficiency by extracting feature of CPU core group utilization and training.
It improves the overall energy efficiency level of HPC cluster benchmark operation, reduces the overhead of the parameter optimization process and model training time, and realizes low-cost offline parameter optimization.
Smart Images

Figure CN120010990A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of HPC cluster benchmark energy efficiency tuning in a benchmark test scenario, and in particular to an HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization. Background Art
[0002] With the rapid development of computer technology, high-performance computing has become an indispensable tool in many scientific, engineering and commercial fields. However, with the continuous increase in computing tasks and the expansion of data scale, the energy consumption problem of HPC clusters has gradually become prominent. How to control energy consumption within an acceptable range while improving performance is also one of the main challenges facing the future development of E-class supercomputers. This has triggered an urgent need for energy efficiency tuning of HPC clusters. Among the numerous HPC applications, HPC benchmarks are usually typical representatives of high-performance applications and represent typical HPC workloads (such as matrix calculations, FFT, molecular dynamics simulations, etc.). Optimizing them can provide valuable experience and reference for energy efficiency optimization of practical applications.
[0003] In terms of software tuning, there have been many studies in the past that have optimized computing methods and task load division for specific benchmarks such as HPL. There have also been many studies on specific optimization and adaptation of HPC underlying operator libraries such as BLAS libraries. Although these studies have achieved many good optimization results, their optimization goals have certain specificity and may not be universal for other benchmarks. The behavior of HPC applications depends on highly configurable software environments. Finding their optimal parameterization is a complex task because the size of their parameter space and the nonlinear behavior of HPC systems make manual tuning, theoretical modeling, or exhaustive sampling unsuitable in most cases. Automatic tuning methods based on black-box optimization eliminate the complex task of tuning the system manually or through theoretical models to find the best combination of algorithm, application, and hardware parameters. There have been some studies in the HPC field that apply black-box parameter optimization methods to optimize application-related parameters. Robert S et al. proposed an online auto-tuner that relies on black-box optimization to find the best parameters of an IO accelerator for a given HPC application within a limited number of iterations without making any assumptions about the behavior of the tuning system. Mittal S et al. applied Bayesian optimization (BO) to jointly tune the HPL benchmark parameters and hardware parameters of a multi-GPU cluster, and based on this tuning method, configured the supercomputer system kukai and achieved the second place in the 2017 Green500 list. Lin W et al. proposed a multi-benchmark driven parameter tuning server energy efficiency tuning method for mixed loads in data centers. It collects runtime data of multiple benchmarks, establishes a multi-label load classification model to achieve the mapping between the actual mixed load and multiple benchmarks, and trains energy efficiency optimization models of multiple different benchmarks to optimize the energy efficiency performance of the actual running mixed load. An important problem with these methods is that although they get rid of the inefficiency of manual optimization, they need to run multiple times on the actual machine to obtain the optimal parameters or train multiple different benchmark energy efficiency optimization models. This process may bring a large model training overhead, so it is necessary to optimize the optimization algorithm and reduce the search space to improve the optimization convergence speed and reduce the model training overhead.
[0004] In terms of hardware optimization, in the optimization scheme of voltage and frequency adjustment of benchmark hardware in the HPC field, Jinpyo K et al. determined the GPU operating frequency with the best energy efficiency for HPL operation through multiple sets of experiments, fixed the GPU at this frequency to improve energy efficiency during HPL operation, and proposed that the optimal energy efficiency of HPL of the cluster is not achieved at the maximum performance of the CPU / GPU, so it is necessary to optimize the GPU / CPU appropriately. Schoonhoven R et al. compared the effects of power capping and fixed frequency energy-saving strategies under the calculation of the general matrix multiplication kernel (GEMM), pointed out that fixed frequency has more precise control over power consumption and supports a larger power consumption adjustment range, and gave the most energy-efficient clock frequency of the GPU by establishing a GPU power consumption model, which greatly reduced the large adjustment search space. Like many current methods, these two methods fix the optimal GPU frequency before the benchmark is run, without considering the dynamic changes of the load globally and formulating the corresponding optimal frequency adjustment strategy, which is more suitable for tasks with little change in load level during operation.
[0005] Specifically, for the current benchmark hardware and software energy efficiency optimization under HPC clusters, software optimization is usually targeted at specific benchmarks (such as HPL) or operator libraries (such as BLAS), lacks universality, and has a complex parameter space. Although black-box optimization methods avoid the complexity of manual tuning and theoretical modeling, they require a large number of real machine runs to collect data, resulting in high overhead and slow convergence in the optimization process. In terms of hardware optimization, existing methods usually adopt fixed frequency or power capping strategies, fail to fully consider the dynamic changes of the load, and have difficulty adjusting hardware parameters in real time to adapt to different task requirements. They often lack a global perspective, resulting in the inability to effectively optimize the energy efficiency of the entire system. In addition, the collaborative optimization of software and hardware is insufficient, the scalability of existing optimization methods in large-scale systems is poor, and the optimization tools between different benchmarks lack standardization, which limits their universal applicability and effectiveness in practical applications. Summary of the invention
[0006] The main purpose of the present invention is to overcome the problems of insufficient versatility of existing HPC benchmark energy efficiency optimization technologies, large model training and collection overhead, and lack of global optimization perspective, and to provide an HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization. The present invention screens and obtains benchmark energy efficiency sensitive parameters in software, reducing the parameter search space. The method can adapt to different benchmark input parameters and system kernel parameters, and combines the sampling algorithm in the parameter optimization process of the energy efficiency prediction model to reduce the optimization and model training overhead, and realizes offline low-cost parameter optimization. In terms of hardware, based on the optimal parameters given in software, the present invention extracts features from the utilization of the CPU core group during the benchmark operation, and the trained clustering model can guide the benchmark to perform dynamic frequency modulation from a global perspective. The software and hardware joint optimization measures combined with the two can effectively improve the overall energy efficiency level of the benchmark operation under the cluster.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] The present invention provides a HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization, comprising the following steps: including software optimization and hardware optimization parts, specifically:
[0009] S1. Software Optimization:
[0010] S11. Collect benchmark operation energy efficiency values under multiple groups of different optimization parameters on the actual machine, and screen out energy efficiency sensitive parameters in the benchmark operation energy efficiency values through parameter importance analysis;
[0011] S12. Use the Bayesian parameter optimization method to perform an iterative search on the screened energy efficiency sensitive parameters, and record the energy efficiency value under each parameter combination;
[0012] S13. Use a data-driven ensemble learning method to establish a parameter prediction model for energy efficiency, and compare the actual machine collection value with the model accuracy index to evaluate the model accuracy;
[0013] S14. Combine the energy efficiency prediction model to select and optimize the sampling algorithm of the Bayesian parameter optimization process, and provide feedback to optimize the actual machine Bayesian parameter process;
[0014] S15. The optimized Bayesian parameter optimization model is used to guide the tuning of the benchmark energy efficiency sensitive parameters under the cluster, and enters the offline tuning stage without the need for actual machine operation;
[0015] S2. Hardware optimization:
[0016] S21, running and collecting CPU core utilization changes of different nodes in the cluster at fixed time intervals, and extracting features of the cores in the group according to the CPU cores that share voltage and frequency;
[0017] S22, drawing a distribution scatter plot of CPU core utilization within the same group during the benchmark operation, and determining a main load mode according to the utilization distribution of the CPU core group;
[0018] S23, taking all the features of the CPU core group after feature extraction as multi-dimensional features, and using a clustering algorithm to obtain the load pattern existing during the benchmark operation;
[0019] S24, combining the load patterns obtained by analysis and clustering, introducing a filtering layer to optimize the clustering model performance, and using the clustering model to give load patterns of different CPU core groups;
[0020] S25. Select a load mode with energy-saving space and formulate a frequency modulation action to guide feedback energy saving.
[0021] As a preferred technical solution, in step S11:
[0022] A random parameter search is used to generate a discrete or continuous parameter combination of optimization parameters. The energy efficiency value of the tuning benchmark under this parameter combination is collected on a real machine. The relationship between parameters and energy efficiency is learned using a random forest model, and the characteristic importance index of each parameter's impact on energy efficiency is output. Energy efficiency sensitive parameters are screened out to narrow the parameter search space; the energy efficiency sensitive parameters are parameters with higher importance.
[0023] As a preferred technical solution, in step S12,
[0024] A Bayesian parameter tuning method based on Gaussian process is used to search the new parameter space. During this process, the pruning method is used to speed up the convergence process of Bayesian parameter optimization, and the benchmark energy efficiency value and the historical best energy efficiency value are recorded as the number of search iterations increases. When the historical best energy efficiency value no longer increases or tends to be stable, the parameter optimization is considered to have converged.
[0025] As a preferred technical solution, in step S13,
[0026] When establishing an energy efficiency prediction model, multiple sets of results obtained by the Bayesian parameter optimization process are collected on the actual machine to establish a data-driven energy efficiency prediction model. The model accuracy evaluation index and the accuracy difference between the prediction parameter combination and the actual machine model are used to evaluate the effects of multiple integrated learners, and the one with the best comprehensive effect is selected as the energy efficiency prediction model. The model accuracy evaluation index includes MSE, MAE and R 2 .
[0027] As a preferred technical solution, in step S14,
[0028] After the energy efficiency prediction model is established, the energy efficiency value output by the energy efficiency prediction model according to the given parameter combination is used as the objective function of Bayesian parameter optimization to perform low-overhead parameter optimization based on the prediction model. Under this condition, the convergence speed and convergence accuracy of the sampling algorithms used in the Bayesian parameter optimization process are compared, and different sampling algorithms suitable for the offline stage and the real machine collection stage are selected and fed back to the real machine parameter optimization process.
[0029] As a preferred technical solution, in step S21,
[0030] The collection method is to group the CPU cores that can share the voltage frequency during frequency adjustment, collect the utilization of the cores sharing the voltage frequency in the same group when the benchmark needs to be adjusted, and process the data as the same group. The data is collected once every time interval T, and the CPU core utilization of each group is extracted according to the mean, range, standard deviation, maximum value, and median to obtain a set of eigenvalues χ i = {U mean ,U range ,U std ,U max ,U med}.
[0031] As a preferred technical solution, in step S22,
[0032] According to the collection time interval T and the core grouping situation, draw a scatter plot of the core utilization distribution within each CPU core group during the benchmark operation, identify the main load mode, and obtain frequency modulation feedback through experiments to confirm the adjustable mode and non-adjustable mode. The key to distinguishing whether it is an adjustable mode is to adjust the core group in this mode by reducing the frequency so that the benchmark operation energy efficiency is improved and the performance degradation is within the set threshold dg max the following.
[0033] As a preferred technical solution, in step S23,
[0034] Taking multi-dimensional features as input, comparing the combined effects of multiple feature values, using clustering performance indicators and WCSS to select the feature dimension input with the best clustering effect, and finding the optimal number of clusters based on the elbow principle of the clustering algorithm.
[0035] As a preferred technical solution, in step S24,
[0036] Combined with the adjustable mode, a filtering layer is added before the clustering model output to enhance the discrimination of the adjustable mode, so that when an input feature is given, the frequency modulation action required for the output point can be quickly determined according to the clustering model.
[0037] As a preferred technical solution, in step S24,
[0038] There are several types of CPU load modes, of which the last two are adjustable modes:
[0039] Ⅰ. The utilization of most cores in the group is higher than the set threshold U min , this state can be maintained at the highest frequency;
[0040] Ⅱ. Most CPUs in the group are in a low utilization state, but a small number of them are extremely high in utilization. In this state, frequency reduction has a greater impact on performance.
[0041] Ⅲ. Most CPUs are in low utilization / idle state, and a small number of CPUs are in medium to high utilization. In this state, the frequency needs to be reduced, and the frequency reduction gear can be based on the utilization of a small number of CPUs;
[0042] IV. The utilization of the combined cores is low, and the frequency is directly reduced to the lowest frequency.
[0043] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0044] (1) The present invention proposes an input parameter and system kernel parameter optimization method based on Bayesian optimization. The method is not designed for a specific benchmark, but can be applied to the automatic optimization of parameters during the operation of various HPC benchmarks. It can provide the optimized parameter combination of benchmark energy efficiency sensitive parameters without human intervention, thereby improving the efficiency of parameter optimization.
[0045] (2) The present invention proposes a strategy that combines data-driven training of an energy efficiency prediction model in the parameter optimization process. This strategy outputs a benchmark energy efficiency under a specific parameter combination by combining an energy efficiency prediction model established with historical sample values. This strategy can be used for offline parameter optimization and sampling algorithms in the optimization process of clusters of the same size, thereby helping to alleviate the power and time cost issues of the black-box parameter optimization method that requires real-machine tuning on large-scale HPC clusters, and improving the optimization effect of the tuning algorithm.
[0046] (2) Based on the dynamic operating characteristics of the benchmark during operation, the present invention proposes a dynamic frequency modulation energy efficiency optimization strategy that changes according to the CPU core group utilization characteristics. By identifying the adjustable load mode of the CPU voltage and frequency group and combining the clustering model to guide the dynamic frequency modulation action, the benchmark's operating energy efficiency can be effectively improved from a global perspective. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0048] Figure 1 Flowchart of the HPC cluster benchmark energy efficiency tuning method based on joint optimization of software and hardware.
[0049] Figure 2 Schematic diagram of the Bayesian parameter optimization process during real machine acquisition.
[0050] Figure 3 Schematic diagram of the optimization process of the Bayesian parameter optimization method combined with the energy efficiency prediction model.
[0051] Figure 4 This is an example of a scatter plot showing the distribution of core utilization over time for a CPU core group during the benchmark run. DETAILED DESCRIPTION
[0052] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0053] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0054] The server configuration of this embodiment includes the configuration of hardware information, operating system and dependent libraries, as shown in Table 1.
[0055] Table 1 Server configuration of the embodiment
[0056]
[0057] like Figure 1As shown, this embodiment is based on the HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization. The method is divided into two aspects: software tuning and hardware tuning. In terms of software, after obtaining the energy efficiency sensitive parameters under the benchmark real machine operation through parameter importance analysis, the method combines the sampling algorithm in the parameter optimization process with the energy efficiency prediction model to reduce the optimization and model training overhead and improve the model convergence speed, thereby realizing offline low-cost parameter optimization; in terms of hardware, based on the optimal operation parameters of the cluster, the method extracts features from the utilization of the CPU core group during the benchmark operation, and trains a clustering model for identifying the CPU load mode. It can judge the resource requirements of the load from a global perspective, thereby performing corresponding dynamic frequency adjustment to improve the energy efficiency operation performance of the cluster. The method is specifically as follows:
[0058] S1. Software Optimization:
[0059] First, for the benchmark to be optimized, its adjustable parameters are determined, including input parameters specific to the benchmark and general system kernel parameters. A random search algorithm is used to generate adjustable parameter combinations collected on the actual machine. After obtaining the operating energy efficiency under multiple sets of different adjustable parameter combinations, the random forest algorithm is used to output the importance analysis of the impact of each adjustable parameter on energy efficiency, so as to select energy-sensitive parameters as the parameter space for subsequent optimization. This step narrows the search space of the parameter optimization algorithm, thereby improving the convergence speed of subsequent optimization. In the example, the energy-sensitive parameters under the HPL-AI (HPL-derived benchmark adapted to mixed precision scenarios) benchmark are searched using this algorithm as shown in Table 2.
[0060] Table 2 Energy efficiency sensitive parameters obtained after parameter importance analysis and screening
[0061]
[0062] Then, the Bayesian parameter optimization algorithm is used to optimize the energy efficiency sensitive parameters. The Bayesian parameter optimization process needs to define the parameter search space, step size, sampling algorithm, objective function, etc. The parameter search space is given by the energy efficiency sensitive parameter combination determined in the previous step. The sampling algorithm can first use the default sampling algorithm, and the objective function is Figure 2 The energy efficiency collection module shown in the figure is used for collection. In the optimization process, the recorded historical sample points are used as training samples for the energy efficiency prediction model. The energy efficiency prediction model is trained using a data-driven ensemble learning algorithm. The model accuracy obtained by various ensemble learners (xgboost, lightgbm, catboost, randomforest, gradientboost) is compared. The model accuracy evaluation indicators (MSE, MAE, R 2) and the accuracy difference between the predicted parameter combination energy efficiency and the actual machine model to select the best ensemble learner to establish the energy efficiency prediction model.
[0063] As the Bayesian parameter optimization process progresses, the accuracy of the energy efficiency prediction model continues to improve. In order to minimize the number of real machine samples, this method evaluates the impact of the number of training samples on the model prediction accuracy. As shown in Table 3, this evaluation method finds that when the number of samples is about 60 groups, R 2 An energy efficiency prediction model with a prediction accuracy of 0.9327 and 93.4% is developed. As the number of optimization samples continues to increase, its prediction accuracy still has room for improvement, indicating that the method used can achieve higher prediction accuracy with a smaller number of samples, thereby alleviating the problem that the traditional black-box parameter optimization process requires a large number of real machine runs to collect data, resulting in a large overhead in the optimization process.
[0064] Table 3 Energy efficiency prediction model performance changes with the increase of record sample interval
[0065] Sampling sample interval (0,20) (0,40) (0,60) (0,80) (0,100) <![CDATA[Prediction accuracy R 2 > -0.5302 0.8748 0.9327 0.9495 0.9512 Predict the best energy efficiency (Gflops / W) 1.344 1.639 1.8108 1.8584 1.8685 Comparison with actual value error (%) 30.5 15.3 6.4 3.9 3.4
[0066] After obtaining an energy efficiency prediction model with high prediction accuracy, it can be combined with the Bayesian parameter optimization process. At this time, offline Bayesian parameter optimization can be performed. The objective function in the Bayesian parameter optimization process is given by the energy efficiency prediction model. The Bayesian parameter optimization process combined with the energy efficiency prediction model is as follows: Figure 3 As shown in Figure 4, the model can also be used to quickly iterate and compare the sampling algorithms of the Bayesian parameter optimization process with extremely low overhead, so as to select a sampling algorithm that is more suitable for the current optimization situation. The convergence results of different sampling algorithms (TPE, CMAES, GPEI, GPPI, GPLCB) obtained by using this model are shown in Table 4. For these five different algorithms, the overhead of TPE and CMAES is much smaller than the latter three algorithms based on the GP model. The GP model algorithm improves faster in the early stage (especially the GPPI model) and can quickly converge to a better value, but its solution time has a cumulative effect as time increases. With the increase in sampling rounds, the GP-based model needs to use more historical data for prediction. When sampling, it involves calculating the inverse of the covariance matrix. The increase in the amount of data will lead to a significant increase in computational complexity, and its time complexity is usually O(n 3), so it is not suitable as a sampler when there are more rounds. Therefore, GPPI is suitable as a sampling algorithm when real machine operation is required in the data collection stage. Once a high-precision energy efficiency model is established, the TPE sampling algorithm with lower overhead should be selected for multiple rounds of sampling. TPE is not as fast as the GP-based model in the early stage, but its convergence speed is second only to the GPPI sampler, and its time overhead is much less than the sampling algorithm based on the GP model. After obtaining the comparison results of the sampling algorithms, the most suitable sampling algorithm can be selected in the actual machine operation stage and offline operation stage of the benchmark in the future to improve the optimization efficiency.
[0067] Table 4 Comparison of convergence rounds and convergence time of different sampling algorithms
[0068] Sampling algorithm type CMA-ES TPE GPEI GPPI GPLCB Convergence rounds 111 73 77 55 75 Convergence time (s) 15.88 7.05 28.89 9.7 15.47
[0069] S2. Hardware optimization:
[0070] After the software tuning obtains the best operating parameters of the benchmark, the best operating parameters are used as the benchmark operating parameters for hardware tuning. In the hardware tuning stage, it is necessary to run and collect the CPU utilization changes of different nodes in the cluster at fixed time intervals. The data is collected once every time interval T, and the CPU cores that share the voltage frequency during frequency modulation are grouped. The CPU core utilization of each group is extracted according to the mean, range, standard deviation, maximum value, and median to obtain a set of utilization feature values χ i = {U mean ,U range ,U std ,U max ,U med}, the feature value is used to train the k-means clustering model, taking all the collected feature points as input and using clustering performance indicators (Silhouette Score, WCSS) to select the feature dimension input with the best clustering effect, and find the optimal number of clusters according to the elbow principle of the clustering algorithm.
[0071] In the step of identifying the load pattern, the collected data is used to draw a scatter plot of the distribution of core utilization within the CPU core group over time during the benchmark run (such as Figure 4 ), identify the main load modes (shown in dashed boxes of different colors), and obtain frequency modulation feedback through experiments to confirm the adjustable mode and non-adjustable mode. The adjustable mode adjusts the core group in this mode by reducing the frequency so that the benchmark operation energy efficiency is improved and the performance degradation is within the set threshold dg max For the HPL-AI benchmark in the embodiment, the CPU load modes during its operation are as follows, of which the last two are adjustable modes:
[0072] Ⅰ. The utilization of most cores in the group is higher than the set threshold U min , this state can be maintained at the highest frequency;
[0073] Ⅱ. Most CPUs in the group are in a low utilization state, but a small number (1-2) have particularly high utilization. In this state, frequency reduction has a greater impact on performance;
[0074] Ⅲ. Most CPUs are in low utilization / idle state, and a small number of CPUs are in medium to high utilization (10%-70%). In this state, the frequency needs to be reduced, and the frequency reduction gear can be based on the utilization of a small number of CPUs;
[0075] IV. The utilization of the combined cores is low and can be directly reduced to the minimum frequency.
[0076] After obtaining the load adjustable mode, a feature filtering layer can be written to filter the CPU core group utilization characteristics during the benchmark operation according to the characteristics of the adjustable mode, and then a clustering model can be used for clustering to obtain load classification, so as to formulate frequency modulation actions in combination with the load mode and classification. The energy efficiency optimization results of the frequency modulation strategy combined with the load mode grouping in the embodiment are shown in Table 5. The energy efficiency after combining software parameter optimization and dynamic frequency modulation is improved by 9.84% compared with the server's default energy-saving mode.
[0077] Table 5 Energy efficiency improvement of the mode-specific dynamic frequency modulation method combined with the parameter optimization method
[0078]
[0079] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0080] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0081] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization, characterized by: The following steps are included: including software optimization and hardware optimization parts, specifically: S1. Software Optimization: S11. Collect benchmark operation energy efficiency values under multiple groups of different optimization parameters on the actual machine, and screen out energy efficiency sensitive parameters in the benchmark operation energy efficiency values through parameter importance analysis; S12. Use the Bayesian parameter optimization method to perform an iterative search on the screened energy efficiency sensitive parameters, and record the energy efficiency value under each parameter combination; S13. Use a data-driven ensemble learning method to establish a parameter prediction model for energy efficiency, and compare the actual machine collection value with the model accuracy index to evaluate the model accuracy; S14. Combine the energy efficiency prediction model to select and optimize the sampling algorithm of the Bayesian parameter optimization process, and provide feedback to optimize the actual machine Bayesian parameter process; S15. The optimized Bayesian parameter optimization model is used to guide the tuning of the benchmark energy efficiency sensitive parameters under the cluster, and enters the offline tuning stage without the need for actual machine operation; S2. Hardware optimization: S21, running and collecting CPU core utilization changes of different nodes in the cluster at fixed time intervals, and extracting features of the cores in the group according to the CPU cores that share voltage and frequency; S22, drawing a distribution scatter plot of CPU core utilization within the same group during the benchmark operation, and determining a major load mode according to the utilization distribution of the CPU core group; S23, taking all the features of the CPU core group after feature extraction as multi-dimensional features, and using a clustering algorithm to obtain the load pattern existing during the benchmark operation; S24, combining the load patterns obtained by analysis and clustering, introducing a filtering layer to optimize the clustering model performance, and using the clustering model to give load patterns of different CPU core groups; S25. Select a load mode with energy-saving space and formulate a frequency modulation action to guide feedback energy saving.
2. The HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization according to claim 1 is characterized in that: In step S11: A random parameter search is used to generate a discrete or continuous parameter combination of optimization parameters. The energy efficiency value of the tuning benchmark under this parameter combination is collected on a real machine. The relationship between parameters and energy efficiency is learned using a random forest model, and the characteristic importance index of each parameter's impact on energy efficiency is output. Energy efficiency sensitive parameters are screened out to narrow the parameter search space; the energy efficiency sensitive parameters are parameters with higher importance.
3. The HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization according to claim 1 is characterized in that: In step S12, A Bayesian parameter tuning method based on Gaussian process is used to search the new parameter space. During this process, the pruning method is used to speed up the convergence process of Bayesian parameter optimization, and the benchmark energy efficiency value and the historical best energy efficiency value are recorded as the number of search iterations increases. When the historical best energy efficiency value no longer increases or tends to be stable, the parameter optimization is considered to have converged.
4. The HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization according to claim 1 is characterized in that: In step S13, When establishing an energy efficiency prediction model, multiple sets of results obtained by the Bayesian parameter optimization process are collected on the actual machine to establish a data-driven energy efficiency prediction model. The model accuracy evaluation index and the accuracy difference between the prediction parameter combination and the actual machine model are used to evaluate the effects of multiple integrated learners, and the one with the best comprehensive effect is selected as the energy efficiency prediction model. The model accuracy evaluation index includes MSE, MAE and R 2 .
5. The HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization according to claim 1 is characterized in that: In step S14, After the energy efficiency prediction model is established, the energy efficiency value output by the energy efficiency prediction model according to the given parameter combination is used as the objective function of Bayesian parameter optimization to perform low-overhead parameter optimization based on the prediction model. Under this condition, the convergence speed and convergence accuracy of the sampling algorithms used in the Bayesian parameter optimization process are compared, and different sampling algorithms suitable for the offline stage and the real machine collection stage are selected and fed back to the real machine parameter optimization process.
6. The HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization according to claim 1 is characterized in that: In step S21, The collection method is to group the CPU cores that can share the voltage frequency during frequency adjustment, collect the utilization of the cores sharing the voltage frequency in the same group when the benchmark needs to be adjusted, and process the data as the same group. The data is collected once every time interval T, and the CPU core utilization of each group is extracted according to the mean, range, standard deviation, maximum value, and median to obtain a set of eigenvalues χ i = {U mean ,U range ,U std ,U max ,U med }.
7. The HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization according to claim 1 is characterized in that: In step S22, According to the collection time interval T and the core grouping situation, draw a scatter plot of the core utilization distribution within each CPU core group during the benchmark operation, identify the main load mode, and obtain frequency modulation feedback through experiments to confirm the adjustable mode and non-adjustable mode. The key to distinguishing whether it is an adjustable mode is to adjust the core group in this mode by reducing the frequency so that the benchmark operation energy efficiency is improved and the performance degradation is within the set threshold dg max the following.
8. The HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization according to claim 1 is characterized in that: In step S23, Taking multi-dimensional features as input, comparing the combined effects of multiple feature values, using clustering performance indicators and WCSS to select the feature dimension input with the best clustering effect, and finding the optimal number of clusters based on the elbow principle of the clustering algorithm.
9. The HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization according to claim 1 is characterized in that: In step S24, Combined with the adjustable mode, a filtering layer is added before the clustering model output to enhance the discrimination of the adjustable mode, so that when an input feature is given, the frequency modulation action required for the output point can be quickly determined according to the clustering model.
10. The HPC cluster benchmark energy efficiency tuning method based on software and hardware joint optimization according to claim 1, characterized in that: In step S24, There are several types of CPU load modes, of which the last two are adjustable modes: Ⅰ. The utilization of most cores in the group is higher than the set threshold U min , this state can be maintained at the highest frequency; Ⅱ. Most CPUs in the group are in a low utilization state, but a small number of them are extremely high in utilization. In this state, frequency reduction has a greater impact on performance. Ⅲ. Most CPUs are in low utilization / idle state, and a small number of CPUs are in medium to high utilization. In this state, the frequency needs to be reduced, and the frequency reduction gear can be based on the utilization of a small number of CPUs; IV. The utilization of the combined cores is low, and the frequency is directly reduced to the lowest frequency.
Citation Information
Patent Citations
Automatic hyper-parameter tuning method based on genetic algorithm
CN115952417A
Hybrid load-oriented multi-reference drive parameter adjustment server energy efficiency optimization method and device
CN116737360A
Performance prediction using machine learning
US20230368067A1
Cited By
Terminal software comprehensive analysis method and system based on artificial intelligence
CN120448247A