Region-aware power and energy regulation
By adopting regional perceived power and energy adjustment technology in HPC systems, dynamically adjusting the CPU frequency and power upper limit, the problem of high power consumption in HPC systems is solved and higher performance and electrical efficiency are achieved.
Patent Information
- Application Number
- CN202410975780.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-10
- Filing Date
- 2024-07-19
- Publication Date
- 2025-05-13
AI Technical Summary
High-performance computing (HPC) systems face challenges in improving electrical power efficiency and energy efficiency, especially due to high power consumption caused by high performance requirements, which in turn increase upfront and operational costs and may cause environmental and regulatory issues.
The region-aware power and energy regulation technology is used to determine the computational restrictive parameters for each area of the application, and dynamically adjust the CPU frequency and power upper limit to achieve optimal performance and electrical efficiency of each area.
It achieves higher performance and electrical efficiency than traditional application-aware optimization methods, reduces power consumption, while avoiding performance degradation, and improving the overall efficiency of the system.
Smart Images

Figure CN119987519A_ABST
Abstract
Description
Background Art
[0001] Power consumption and efficiency are becoming increasingly important issues for computing devices. Power consumption can be quantified in terms of electrical power consumption and electrical energy consumption. Electrical power consumption refers to the instantaneous power consumption (voltage times current) measured in watts (W) or similar units. Electrical energy consumption refers to the integral of electrical power consumption over time, measured, for example, in joules (J), watt-hours (W·h), or similar units.
[0002] In general, it is desirable to improve the electrical power efficiency and energy efficiency of computing systems. This is particularly true for high performance computing (HPC) systems. HPC systems aggregate computing power to provide much higher performance than what can be achieved by a typical single computer or workstation. Typically, HPC systems network multiple computing devices (also known as nodes) together to create a high-performance architecture. Applications are executed simultaneously on the networked computing devices, so that their performance is improved relative to that which can be achieved by a single device. Due to the high performance of such HPC systems, they tend to have very high electrical power / energy requirements, and these requirements are expected to increase as system performance becomes better and HPC jobs become larger and / or more complex.
[0003] The power requirements of an HPC system (or other computing system) may have some negative effects. For example, systems with higher power consumption requirements may incur higher upfront costs and higher ongoing operating costs. The upfront costs of high-power systems may be higher because they may need to be equipped with better-performing power supply units and cooling solutions to accommodate the power requirements. Operating costs may also be higher due to the cost of electricity itself (especially in locations where electricity costs are higher), and may also be higher due to the need to supply more coolant to the system and / or cool the coolant to a lower temperature during operation. In addition, the large amount of electricity consumed by HPC systems may raise regulatory or environmental issues. Therefore, there is a need to improve electrical efficiency in computing systems, and particularly in HPC systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Alone or with Figure 1 The present disclosure can be understood from the following detailed description. These drawings are included to provide a further understanding of the present disclosure, and are incorporated into and constitute a part of this specification. The drawings illustrate one or more examples of the present teachings and together with the description explain certain principles and operations. In the drawings:
[0005] Figure 1 is a block diagram schematically illustrating an example information processing system configured to perform a zone-aware power / energy regulator.
[0006] Figure 2is a block diagram schematically illustrating an example computing node configured to perform a zone-aware power / energy regulator.
[0007] Figure 3 is a block diagram schematically illustrating an example HPC system including an HPC system control node configured to execute a zone-aware power / energy regulator.
[0008] Figure 4 is a process flow diagram illustrating an example method of zone-aware power / energy regulation.
[0009] Figure 5 is a process flow diagram illustrating another example method of zone-aware power / energy regulation.
[0010] Figure 6 is a process flow diagram illustrating an example method for adjusting non-core frequency as part of locality aware power / energy regulation.
[0011] Figure 7 is a process flow diagram illustrating an example method for determining whether a region needs to be characterized as part of region-aware power / energy regulation.
[0012] Figure 8 is a block diagram schematically illustrating an example non-transitory computer-readable medium including zone-aware power / energy regulation instructions. DETAILED DESCRIPTION
[0013] One of the largest consumers of electrical power / energy in a computing device is the processor (e.g., CPU). Therefore, one way to reduce the overall power consumption of a computing device or a system having multiple computing devices (e.g., an HPC system) is to adjust the operating parameters of the CPU(s) to reduce power consumption. In particular, one CPU parameter that can be adjusted to save power is the CPU core clock frequency. Specifically, some power saving methods reduce the CPU frequency to a value lower than the normal operating value, which in many cases will reduce the amount of power consumed (more or less, as the case may be, as will be described below).
[0014] Although reducing the CPU frequency can reduce power consumption, it can also degrade system performance. In other words, there is usually a trade-off between saving power and system performance. In some cases, saving power at the expense of performance degradation may be considered acceptable, provided that the magnitude of performance degradation is small relative to the power savings. However, if the performance degradation is severe and / or if the amount of energy savings is small, this trade-off may be considered unacceptable. Therefore, many power saving methods that rely on reducing the CPU frequency attempt to determine a CPU frequency that will achieve a desired balance between system performance and power consumption. This may be referred to as "optimizing" the CPU frequency.
[0015] In practice, it may be difficult to optimize the CPU frequency to produce the desired balance between system performance and power consumption. This is because it is usually impossible to know in advance how much performance will be degraded or how much power will be saved in response to a certain reduction in CPU frequency. Reducing the CPU frequency by a given amount may produce different results in different situations, depending on what the CPU is doing at the time, that is, depending on the current workload. For some workloads, reducing the CPU frequency by a given amount may cause only a slight degradation in performance - for example, if the CPU is currently waiting for data to be transferred from the memory, then the reduction in the CPU core frequency will not have a significant impact on performance. For other workloads, reducing the CPU frequency by exactly the same amount may greatly degrade performance - for example, if the CPU is currently executing a series of instructions that do not require a lot of waiting for data, then the magnitude of the performance degradation may be proportional to the magnitude of any frequency reduction. In addition, for some workloads, reducing the CPU frequency may reduce both power and energy usage, while in other cases, the same reduction in the CPU frequency may reduce power but have little effect on energy usage (or even increase energy usage). Therefore, there may not be any CPU frequency that is optimal in all cases.
[0016] In particular, the system's response to a CPU frequency reduction may vary from application to application. For example, some primarily memory-bound applications (such as lbm) may suffer very little performance degradation and may achieve significant power / energy savings when the CPU frequency is reduced - for example, in one test system, a CPU frequency reduction of approximately 58% saved 37% power and 36% energy, with only a 1.6% performance loss. On the other hand, primarily compute-bound applications (such as imagick) may suffer greater performance degradation and may not achieve significant energy savings when the CPU frequency is reduced - for example, in one test system, a CPU frequency reduction of 52% saved 51% power, but at the expense of a 7.7% increase in energy and a 53% decrease in performance. Therefore, some methods for optimizing CPU frequency may attempt to account for differences between applications by characterizing the applications based on their sensitivity to CPU frequency, and then controlling the CPU frequency based on the type of application being executed. In other words, the "optimal" frequency is determined on a per-application basis. For example, if the application being executed is characterized as being compute-bound, the CPU frequency may be reduced less (or not at all) to avoid performance degradation, whereas if the application is characterized as being memory-bound, the CPU frequency may be reduced more to save power and energy with little performance degradation. This may be referred to as application-aware power / energy optimization.
[0017] However, application-aware power / energy optimization may not always produce the best results. In particular, the system's response to CPU frequency can vary not only from application to application, but also within different areas (e.g., functions, routines, loops, or other areas) of the same application. Many (perhaps most) applications contain some mix of memory-bound and compute-bound areas - few applications are uniformly memory-bound or uniformly compute-bound. Even in applications that are primarily compute-bound or memory-bound as a whole, there will often be at least one area (or sometimes multiple areas) of the application that does not follow the trend of the application as a whole. Accordingly, if a single CPU frequency is set for the entire application, it is very likely that this frequency will be suboptimal for at least some areas of the application.
[0018] For example, if an application is characterized as memory-constrained under the application-aware optimization method and a reduced frequency is set for the application, this may produce the desired power savings when executing the memory-constrained regions. However, whenever one of the compute-constrained regions in the application is executed, the performance of the system will be affected due to the lower frequency. Therefore, the overall performance of the system (i.e., the time required to complete the job) will be degraded. Conversely, if an application is characterized as compute-constrained under the application-aware optimization method and a higher frequency is set, this may achieve good performance during the execution of the compute-constrained regions. However, whenever one of the memory-constrained regions is executed, the higher frequency will result in unnecessary power consumption. Therefore, the electrical efficiency of the system will be reduced. Accordingly, application-aware optimization may not achieve all the power savings and / or system performance that could theoretically be achieved.
[0019] To address these and other problems, this document discloses a regional-aware power and energy optimization technique that can be implemented in a regional-aware power and energy regulator. In the regional-aware power and energy optimization technique, a computational constraint parameter (%CB) is determined for each region of an application (e.g., function, routine, etc.), and then the optimal CPU parameters (e.g., CPU frequency, power cap, etc.) are determined for each region based on the computational constraint of each region. Regions with higher computational constraints can be assigned higher frequencies (or higher power caps, etc.) to slow down performance degradation, while regions with lower computational constraints can be assigned lower frequencies (or lower power caps, etc.) to save power. Throughout the execution of an application, the parameters of the CPU (e.g., frequency) can be repeatedly changed to different values according to the region currently being executed, so that at any given time, the current CPU parameters (e.g., frequency) are equal to the optimal parameters (e.g., optimal frequency) of the current region being executed. This allows for higher performance and electrical efficiency than that possible under the application-aware optimization method.
[0020] In some examples, the compute-limitedness parameter %CB and the associated optimal frequency for a region can be determined during runtime when executing an application. Thus, no pre-characterization of the platform or application is required. Additionally, the region-aware power and energy optimization technique does not require a user to attempt to analyze or characterize the application or its regions, thereby allowing the approach to be more user-friendly (e.g., approaches that rely on users to define or characterize aspects of an application are rarely, if ever, put into practice because users often do not want to do this extra work when submitting a job).
[0021] In some examples, a compute-limitedness parameter (%CB) for a region quantifies the compute-limitedness of the region. For example, %CB can be a percentage value between 0 and 100, where a value of 100% means that the region is purely compute-limited and 0% means that the region is not compute-limited at all (i.e., purely memory-limited). In some examples, to determine the compute-limitedness parameter %CB for a given region, a number of instructions per second (IPS) executed by a processor during a plurality of sampling periods associated with different CPU frequencies during execution of the given region is determined, and the compute-limitedness parameter %CB for the given region is determined based on the IPS measurements. For example, when the CPU frequency is at a predetermined high value Freq high When the first IPS measurement value IPS high Sampling is performed, and then, when the CPU frequency is at a predetermined low value Freq low When the second IPS measurement value IPS low sampling (both are sampled during execution of the same given region), and can be based on IPS high and IPS low To determine the computational constraint parameter %CB for a given area - for example, the IPS high ,IPS low 、Freq high and Freq low An equation (eg, Equation 1 described below) is associated with the calculated limitability parameter %CB as an output variable to determine %CB.
[0022] IPS is a metric that is readily available on almost every (if not every) platform, and therefore, in contrast to many machine learning approaches that rely on hardware counters that are specific to one type of system (e.g., specific to Intel processors or specific to AMD processors), locality-aware optimization techniques can be used in almost any type of computing system. Among other things, this allows locality-aware power and energy optimization techniques to be vendor- and platform-agnostic.
[0023] In some examples, the optimal frequency for a given region is determined not only based on the computational constraints parameter %CB for the region but also based on a performance degradation parameter (PD). The performance degradation parameter PD represents an acceptable level of performance degradation relative to the default performance that would be achievable at the default CPU frequency (without any adjustment to save power). For example, a PD of 5% would indicate that a performance degradation of 5% is acceptable—that is, the performance is 95% of the default performance level. In some examples, the optimal frequency for a given region can be determined by evaluating an equation (e.g., Equation 2 described below) that relates %CB and PD as input variables to the optimal frequency as an output variable. In some examples, the performance degradation parameter PD can be specified by a user, such as when a user submits a job to be executed. In this way, region-aware optimization is easily customizable to achieve a desired balance between power savings and performance.
[0024] In some examples, a simple algebraic equation (see Equation 1 described below) can be used to determine %CB based on the IPS value, and another simple algebraic equation (see Equation 2 described below) can be used to determine the optimal frequency based on %CB and PD, and thus, the optimal frequency can be determined very quickly and easily (e.g., on the order of 100 ms). Therefore, the method is relatively light in terms of computational overhead and flexible enough to react to rapidly changing areas of the application. In contrast, machine learning-based methods (such as application-aware methods) tend to be more computationally intensive and take longer to converge to a solution (e.g., on the order of seconds, minutes, or longer, depending on the complexity of the model), which makes this method less useful in some cases.
[0025] In addition to setting the CPU core frequency based on %CB as described above, some examples may also include other region-based optimizations. For example, in some embodiments, the non-core frequency may also be adjusted based on %CB for some regions. The non-core frequency refers to the frequency of the parts of the CPU other than the core, which may include L3 cache, memory controller, etc. In some examples, the non-core frequency may be reduced in regions with high %CB.
[0026] In some examples, the region-aware power / energy regulator is also aware of message passing interface (MPI) regions, and is configured to optimize CPU frequencies specifically for these regions in a manner that may be slightly different from other regions.MPI is a message passing standard for parallel computing architectures (such as HPC systems).Applications configured to utilize MPI (such as many HPC applications) may sometimes be referred to as MPI applications.During the execution of such applications, a process may occasionally need to communicate with another process using an MPI function.Due to the special properties of these functions, if they are characterized using the method described above without any special considerations, erroneous results may sometimes be received.For example, some MPI wait functions may frequently poll other processes while waiting for data, which increases the measured IPS (because each polling cycle includes the execution of instructions), and therefore, it can be determined that the MPI wait function has a high computational constraint %CB, which will result in assigning high frequencies during MPI calls.However, despite having high IPS, no useful work is performed during the MPI wait period, and therefore, high frequencies will be wasted. Accordingly, examples disclosed herein may override standard optimization behavior when such an MPI wait function is encountered and may instead set the CPU frequency to a predetermined low value during the function. Furthermore, in some examples, such a change in the CPU frequency of an MPI wait function to a low value may be performed only when certain criteria are met, such as when the MPI wait function is expected to have a duration longer than a specified minimum value.
[0027] Turning now to the drawings, various apparatus, systems, and methods in accordance with aspects of the present disclosure will be described.
[0028] Figure 1 1 is a block diagram schematically illustrating the information processing system 100. It should be understood that Figure 1 Certain shapes, sizes, or other structural details are not intended to be illustrated accurately or to scale, and implementations of information handling system 100 may have a different number and arrangement of the components illustrated and may also include other parts not illustrated.
[0029] like Figure 1 As shown in , the information processing system 100 includes a processor 110 and a storage medium 120 communicatively connected to the processor 110. The processor 110 may include a central processing unit (CPU), a system on a chip (SoC), or any other hardware processing resource. The storage medium 120 includes a non-transitory computer-readable medium, such as a hard disk drive (HDD), a solid-state drive (SSD), a flash memory, a random access memory (RAM), or any other non-transitory computer-readable medium.
[0030] The storage medium 120 stores zone-aware power and energy regulation instructions 130 executable by the processor 110. When the processor 110 executes these instructions 130, a zone-aware power / energy regulator 140 is instantiated. The zone-aware power / energy regulator 140 performs operations described herein related to zone-aware power / energy optimization and regulation. Such regulation includes, among other things, characterizing various zones of an application (e.g., an HPC application) being executed by the CPU and determining an optimal CPU frequency for each zone to be used by the CPU during execution of the zone.
[0031] The CPU that is executing the application program may be the processor 110 itself, or the CPU may be part of some other device that is controlled and monitored by the system 100. For example, Figure 2 An example implementation of the system 100 is illustrated in the form of a computing node 200, where the same processor 210 instantiates both the application 250 and the regulator 240. As another example, Figure 3 Another example implementation of system 100 is illustrated in the form of HPC system 300 , where the processor on which regulator 340 is instantiated is different from the processor on which application 350 is executed.
[0032] return Figure 1 , instructions 130 include region identification instructions 131. Region identification instructions 131, when executed, cause regulator 140 to identify the region of the application that is currently being executed. As used herein, a "region" of an application is defined as: (A) any function (also referred to as a routine or subroutine) of an application or any part of an application that is equivalent to a function; or (B) the smallest discrete part of an application that can be identified and characterized. Examples of regions include functions (routines, subroutines), loops, subroutines, etc. For example, region information may be monitored (periodically received or requested) by regulator 140, wherein the region information indicates the current region being executed. In some examples, the region information may include the current address of the instruction pointer of the process, which may be cross-referenced with the known address of the region of the application to determine the region that is currently being executed.
[0033] For example, the regulator 140 may start by identifying which regions the application contains and saving the addresses associated with these regions for quick lookup. For example, if the application contains identifiable functions, the regulator 140 may determine and store the addresses of each of these functions. In some examples, the regulator 140 may also identify other types of regions of the application, such as portions of contiguous memory of a predetermined size (e.g., the minimum size that can be characterized), and store the addresses associated with these regions. The regulator 140 may then begin to continuously monitor the application to identify the current region by obtaining the current instruction pointer at predetermined intervals. By default, the regulator 140 identifies the current region as: (a) if the address of the instruction pointer is associated with a function, then the function containing the address of the instruction pointer, or (b) if the address of the instruction pointer is not associated with a function, then the portion of the contiguous memory address containing the address of the instruction pointer. By default, the size of the contiguous portion of the memory address is a predetermined size (e.g., the minimum size that can be characterized). The interval for monitoring the instruction pointer may be as short as possible to account for entering / leaving different regions, such as every 50 μs in some examples.
[0034] In some examples, obtaining the regional address involves parsing the binary symbol table, which is more complicated when there is a shared library. Because the shared library is loaded at the different addresses of each process (and each subsequent operation), the regulator 140 first uses the address space of the process such as ptrace. The link table of the shared library ELF file information can be found in the link map of the loaded application program, and the address of the link map is located in the global offset table (GOT). Then, the dynamic symbol table of each ELF file can be parsed to obtain the function name and address.
[0035] If multiple regions that include adjacent portions of consecutive addresses have been characterized and these regions have similar %CB, then these adjacent portions can be merged together into a single region. In this context, similarity can be defined as their %CB differing by less than a threshold amount. The threshold amount can be predetermined or user definable. In some examples, the threshold amount is 5%. In some examples, the threshold amount is 10%.
[0036] The region identification instruction 131 may also be configured to cause the regulator 140 to determine whether the identified region needs to be characterized. (In this context, being characterized refers to determining the computational constraints parameters of the region). Some regions may have been previously characterized, in which case no further characterization is required. Other regions may not have been characterized yet, but may be omitted from the characterization due to one or more exceptions. For example, in some embodiments, only regions that are considered important are characterized. In some examples, a region is considered important if the amount of time spent in the region exceeds a predetermined threshold. In some examples, the regulator 140 uses Linux perf_event to determine when a region is encountered and how much time is spent in the region. By using the PERF_COUNT_SW_CPU_CLOCK counter, the regulator 140 can reliably obtain the current instruction pointer at regular time intervals (e.g., 50μs) and measure the approximate time spent in each region. This also allows the regulator 140 to know when to periodically resample the region, which may be important for long-running applications.
[0037] The instructions 130 further include region computational constraint determination instructions 132. These instructions 132 may be executed when it is determined that a given region needs to be characterized. The region computational constraint determination instructions 132 include instructions for determining a computational constraint parameter %CB for the current region being characterized based on the IPS measurement. In some examples, the computational constraint parameter %CB for the region quantifies the computational constraint of the region. For example, %CB may be a percentage value between 0 and 100, where a value of 100% means that the region is purely computationally constrained and 0% means that the region is not computationally constrained at all (i.e., purely memory constrained).
[0038] In some examples, when it is determined that a given region needs to be characterized, the regulator 140 will perform a sampling procedure in which the CPU frequency is set to a predetermined value for the duration of a sampling period (during which the region continues to be executed) and the IPS is measured after the sampling period is completed. For example, when the CPU frequency is at a predetermined high value Freq high When the first IPS measurement value IPS high Sampling is performed, and then, when the CPU frequency is at a predetermined low value Freq low When the second IPS measurement value IPS low IPS-based high and IPS low The calculation limit parameter %CB for a given area is determined, for example, by evaluating the following equation:
[0039]
[0040] In Equation 1, %CB n is the computational constraint parameter of the nth region of the currently executing application (in this context, “n” is an arbitrary index used herein to identify a given region), IPS high_n is the high IPS measurement value obtained for the nth area, IPS low_n is the low IPS measurement value obtained for the nth region, Freq high_n It is for IPS high_n The high frequency at which the sampling is performed, and Freq low_n It is for IPS low_n The low frequency at which sampling is performed. %CB n is limited to a value between 0 and 100%. In some examples, a memory-limited parameter %MB may also be calculated, where %MB=100%-%CB. The IPS ratio in Equation 1 informs how frequency changes affect performance, which is then scaled by the frequency ratio. To intuitively understand the formula, if the frequency ratio is 1.5 (i.e., 3GHz / 2GHz), but the IPS ratio is 1.25 (i.e., performance at 3GHz is only 25% faster than performance at 2GHz), then regulator 140 considers this to be 50% compute-limited. In addition, the metric can dynamically adapt to platform capabilities, i.e., if a different processor has less memory bandwidth per core, the application may become more memory-limited and the IPS ratio will be smaller.
[0041] In some examples, Freq high is the turbine frequency, and Freq low is any lower frequency (eg, highest non-turbo frequency).
[0042] In some examples, the IPS measurement value may be received from a processor executing an application (or more specifically, from an operating system operating on the processor executing the application). In some examples, the processor executing the application may be the same as processor 110, or different from processor 110, as described above. In some examples, the regulator 140 uses the Linux perf_event interface to access the IPS measured variable and as a tool to obtain the instruction pointer and timestamp for each sample. For example, in some embodiments, to measure IPS, the regulator 140 may define a sampling period according to a defined number of instructions, and then, how long it takes to execute the defined number of instructions. This will produce a number representing the number of seconds per instruction, which is the inverse of IPS. IPS can therefore be calculated as 1 divided by this number. For example, a sampling period may include 100,000 instructions, and if a given sampling period takes X seconds to complete, the IPS for that sampling period is 100,000 / X. The regulator 140 may use the PERF_COUNT_HW_INSTRUCTIONS performance counter to determine when a predetermined number of instructions have been executed. Perf also allows tracking of specific cores or processes. Using the Linux ps command, the regulator 140 can automatically identify both the core and pid of any process running the target application, and then supply it to perf_event.
[0043] In some examples, the IPS are sequentially executed during a single given execution of the region being characterized. high and IPS low In other examples, the IPS high and IPS low The sampling of IPsec can be spread across multiple different instances of the execution region to avoid too frequent sampling. For example, in some embodiments, the first time a region is executed, IPsec is sampled. high and IPS low One of them is sampled and the next time the region is executed, the IPS high and IPS low In addition, in some examples, multiple IPS can be collected high Samples and / or multiple IPS low samples, and the value used in Equation 1 may be a statistical aggregation of the values (eg, an average).
[0044] As described above, in some embodiments, not every region is considered important in order to determine whether it should be characterized. In some embodiments, such a restriction can be imposed to avoid performing sampling too frequently. Sampling too frequently can be disadvantageous. If too much high sampling is performed, this may consume power / energy savings if the region is memory-constrained. Similarly, if too much low sampling is performed, this may harm performance if the region is compute-constrained. In some embodiments, the importance threshold can be those regions that occupy at least 5% of the current running time. In addition, important regions will sometimes lack enough samples of only one type (high or low), so if there are not enough high samples, 'downtime' can be further reduced by performing only high sampling, and vice versa.
[0045] The instructions 130 further include frequency setting instructions 133. The frequency setting instructions 133 include instructions for determining the optimal CPU frequency for the currently executing region based on the computational constraint %CB of the currently executing region and instructions for commanding the system 100 to set the CPU frequency to the optimal frequency. Highly computationally constrained regions may be assigned higher frequencies to mitigate performance degradation, whereas less computationally constrained regions may be assigned lower frequencies to save power. Throughout the execution of the application, the frequency of the CPU may be repeatedly changed to different values depending on the currently executing region, such that at any given time, the current CPU frequency is equal to the optimal frequency for the current region being executed.
[0046] In some cases, the optimal frequency for a region will be known, for example, because the region has been characterized, in which case the regulator 140 can generate a frequency setting command to set the frequency to the known optimal frequency. In other cases, the optimal frequency for a region is not yet known, in which case the optimal frequency can be calculated based on %CB.
[0047] In some examples, the optimal frequency for a given region is determined not only based on the computational constraints parameter %CB for the region but also based on a performance degradation parameter (PD). The performance degradation parameter PD represents an acceptable level of performance degradation relative to the default performance that would be achievable at the default CPU frequency (without any adjustments to save power). For example, a PD of 5% would indicate that a 5% performance degradation is acceptable—that is, the performance is 95% of the default performance level. Therefore, in some examples, the optimal frequency for a given region is determined by evaluating an equation that relates %CB and PD as input variables to the optimal frequency as an output variable. For example, in some embodiments, the optimal frequency is given by the following equation:
[0048]
[0049] In Equation 2, Freq n represents the optimal or ideal frequency of the nth region, Freq high_n represents the highest (turbine) frequency in the nth region, PD is the performance degradation parameter, and %CB n is the computational constraint parameter of the nth region. In some examples, the performance degradation parameter PD can be specified by a user, for example, when the user submits a job to be executed. In this way, region-aware optimization is easily customizable to achieve a desired balance between power savings and performance.
[0050] Equation 2 can be derived based on the insight that compute-boundedness %CB reflects the sensitivity of a region's performance to frequency changes. That is, the time to complete a given amount of work may vary in proportion to the CPU frequency, where %CB serves as a proportionality constant. Intuitively, if the region is completely memory-bound (%CB=0), any changes in CPU frequency will have no effect on performance (time to completion). Conversely, if the region is completely CPU-bound, changes in CPU frequency will affect performance in a 1-to-1 manner. Based on these intuitions, the following functional relationship can be inferred:
[0051]
[0052] In Equation 3, Time low It refers to the low frequency Freq low The time to complete a given amount of work, and Time high It refers to the high frequency Freq high In addition, the performance degradation parameter PD can be compared with Time high and Time low Related, as follows:
[0053]
[0054] Combine Equation 3 and Equation 4, rearrange the result, and use Freq n Replace Freq low , we get equation 2.
[0055] In some embodiments, instructions 133 may cause regulator 140 to optimize another CPU parameter for each region based on the corresponding %CB of those regions, wherein changes in the other CPU parameter affect performance and power / energy consumption depending (at least in part) on the computational constraints of the region. Examples of such other CPU parameters include power cap parameters, thermal limit parameters (e.g., the CPU temperature at which throttling begins and / or acceleration / turbo behavior is limited), or other parameters. Typically, these other CPU parameters may affect performance and power consumption in part by indirectly limiting the CPU frequency, which may be useful in systems where direct control of the CPU frequency is difficult or undesirable. In some embodiments, optimization of another CPU parameter may be performed in addition to optimizing frequency. In some embodiments, optimization of another CPU parameter may be performed instead of optimizing frequency. Equation 2 may be used to optimize another CPU parameter, except that Freq is replaced by the value of another CPU parameter. high and Freq low (For example, replace Freq with a high power limit high , and replace Freq with a low power cap low ).
[0056] Once the optimal frequency Freq is determined for the nth region n , then the frequency setting instruction 133 can instruct the CPU to use the optimal frequency Freq whenever the nth region is executed. n . In some embodiments, dynamic voltage frequency scaling (DVFS) is used to adjust the frequency of the processor. In some embodiments, the regulator 140 uses the CPUFreq interface to modify the CPU frequency. Specifically, the cpupower frequency-set command allows the root user to set the maximum frequency of all cores at once, rather than modifying individual system files for each core. The current frequency can also be checked with cpupower or by reading the corresponding system file. Importantly, changing the frequency does not occur immediately and the regulator 140 must take the switching delay into account. For example, in the Intel Sapphire Rapids test system, the CPUFreq specification has a switching delay of 10μs, and in the AMDGenoa test system, the CPUFreq specification has a switching delay of 8μs.
[0057] In addition, in some cases, the calculated ideal frequency may not be an available frequency option in the CPUFreq interface. In examples where the available frequency settings do not fully match the calculated optimal frequency, the next closest available frequency setting can be used as the optimal frequency setting. However, in some cases, if the available frequency settings are too coarse (e.g., in the Genoa test system, the available frequency options are in the coarse granularity of 3.7 GHz (turbo), 2.4 GHz, 1.9 GHz, and 1.5 GHz), instead of using the next closest available frequency setting, the optimal frequency can be simulated by alternating between the next lowest frequency setting and the next highest frequency setting with a timing ratio that results in a weighted average frequency equal to the optimal frequency.
[0058] In those examples where another CPU parameter is optimized for a region based on the %CB for that region, instructions 133 may include instructions for commanding a system running an application to set the other CPU parameter to the determined optimal value for that parameter.
[0059] In addition to setting the CPU core frequency based on %CB as described above, in some examples, the regulator 140 may also include instructions for determining the optimal non-core frequency for each region and setting the non-core frequency based thereon. Non-core frequency refers to the frequency of the parts of the CPU other than the core, which may include L3 cache, memory controller, etc. In some examples, the non-core frequency may be reduced in regions with high %CB. For example, in some embodiments, the non-core frequency may be set based on a computationally limited parameter %CB similar to the CPU frequency, except that the relationship between the non-core frequency and %CB may be reversed compared to the relationship between the CPU frequency and %CB. In other words, if the region has a low %CB (i.e., is highly memory-constrained), the non-core frequency is set high to avoid performance loss, but if the region has a high %CB (i.e., is highly computationally constrained), the non-core frequency may be set low to save power without significant performance loss. In some examples, the non-core frequency is set using the following equation:
[0060]
[0061] In some examples, the zone-aware power / energy regulator 140 is also aware of message passing interface (MPI) functions and is configured to optimize CPU frequency specifically for these functions in a manner that may be slightly different from other functions. In some examples, the regulator 140 can override the standard optimization behavior when an MPI wait function is encountered and can instead set the CPU frequency to a predetermined low value during the function. In addition, in some examples, this change of the CPU frequency to a low value for an MPI wait function can only be performed when certain criteria are met, such as when the MPI wait function is expected to have a duration longer than a specified minimum value.
[0062] In addition, MPI applications can, for example, use sub-communicators to assign different processes (ranks) to perform different tasks. Therefore, there may be a situation where rank A is in a computationally limited area and rank B is in a wait function, and if regulator 140 is monitoring rank B and reducing the CPU frequency of all cores, rank A will be adversely affected. A solution to this point used in some embodiments is to use per-core DVFS, where the CPU frequency is changed only for a specific core rather than the entire socket. For example, since the integration of per-core voltage regulators in Haswell, per-core DVFS has been available on most Intel platforms. In order to utilize per-core DVFS, system 100 can deploy multiple regulator 140 processes, where each regulator 140 process monitors a single MPI rank. In this way, each regulator 140 process will still properly characterize function / phase, but any frequency changes will not affect other ranks, and the lightweight nature of regulator 140 will minimize overhead. In other embodiments, per-core DVFS may not be available, in which case the regulator 140 may run the CPU at the highest CPU frequency when a sub-communicator call is detected.
[0063] The regulator 140 may allow substantial power and energy savings to be achieved while still allowing desired performance levels to be maintained. For example, in two test systems (an Intel Saphire Rapids system and an AMD Genoa system), the regulator 140 achieved promising results, summarized below.
[0064] For the Intel test system, the regulator 140 achieved 10% to 20% energy and power savings for memory-constrained applications with a performance loss of approximately 5%. In some applications (e.g., lbm), performance was lost by up to 7%, accompanied by 33% energy savings and 38% power savings. Importantly, other methods did not detect these available savings and obtained almost the same results as the baseline. For more compute-constrained applications, the regulator 140 was able to save 5% to 10% of energy by improving performance in compute-constrained applications. For more mixed applications, the results seen by the regulator 140 were mostly on par with the baseline, with approximately <3% performance loss, but some applications saw significant improvements in energy savings, such as 5% and 9% in xz and xalancbmk.
[0065] For the AMD test systems, most compute-bound applications achieve within 3% of the desired performance level, with power savings proportional to the allowed performance loss and relative energy consumption about the same as the baseline. More memory-bound applications tend to save about 30% of power and energy and achieve 95%+ relative performance. More mixed applications have greater variability in achieving appropriate performance degradation, but overall power / energy savings.
[0066] Now turn to Figure 2 , an example computing node 200 will be described. The computing node 200 is Figure 1 1 is an example implementation of information processing system 100 of FIG. 1 . Some components of computing node 200 correspond to (e.g., are similar to or are configurations of) components of system 100, and these components are given similar reference numerals with the same last two digits, such as 110 and 210. Unless otherwise indicated or logically contradictory, descriptions of components of system 100 apply to similar components of node 200, and duplicate descriptions are omitted. Although computing node 200 is an example of system 100, system 100 is not limited to computing node 200.
[0067] Compute node 200 represents an example implementation of system 100 in which the same processor running the application for which optimization is sought is also the processor on which the locality-aware power / energy regulator is instantiated. Figure 2, processor 210 instantiates both application 250 and regulator 240. In some examples, compute node 200 may be a separate compute node of a larger HPC system, where jobs are processed simultaneously by the entire HPC system, where each compute node (including compute node 200) executes applications (including application 250) to perform multiple portions of the job simultaneously. In some examples, compute node 200 may be a separate computing device (e.g., a server) that is not necessarily part of any larger HPC system.
[0068] Application 250 includes multiple regions, including regions 251-1 to 251-N, where N is any integer equal to or greater than 2. In some examples, region 251 can be a function. In other examples, region 251 can be a portion of a continuous memory address. In other examples, some regions 251 can be functions and other portions of a continuous memory address. Regulator 240 is configured to determine the optimal frequency of regions 251-1 to 251-N, as described above with respect to regulator 140. In this example, region identification information (e.g., instruction address pointer) and IPS measurement value information are provided to regulator 240 by operating system interface 260, and a CPU frequency setting command is sent from regulator 240 to operating system interface 260. Operating system interface 260 can be part of an operating system, BIOS, firmware, or other system management system of node 200 and can be instantiated by processor 210. Operating system interface 260 can include perf_event, CPUFreq, and other interfaces and tools mentioned above with respect to regulator 140.
[0069] Note that node 200 also includes a storage medium (similar to storage medium 120) having instructions (similar to instructions 130) for instantiating regulator 240, but Figure 2 The views in FIG. 1 and 2 omit these components.
[0070] Now turn to Figure 3 , an example HPC system 300 will be described. The HPC system 300 is Figure 1 300 is an example implementation of an information processing system 100 of the present invention. Some components of HPC system 300 correspond to (e.g., are similar to or are configurations of) components of system 100, and these components are given similar reference numerals with the same last two digits, such as 140 and 340. Unless otherwise indicated or logically contradictory, the description of the components of system 100 is applicable to similar components of HPC system 300, and repeated descriptions are omitted. Although HPC system 300 is an example of system 100, system 100 is not limited to HPC system 300.
[0071] HPC system 300 represents an example implementation of system 100 in which the application seeking optimization and the locality-aware power / energy regulator are instantiated by different processors, specifically by different processors of different nodes of the HPC system.
[0072] Specifically, the HPC system 300 includes a plurality of computing nodes 380-1 to 380-P (where P is an integer equal to or greater than 2) that perform computing tasks of jobs submitted to the HPC system 300, and an HPC system control node 370 that controls the operation of the entire system (including orchestrating jobs). In some examples, the HPC system control node 370 is also a computing node that is delegated to be responsible for the system control area, but in other examples, the system control node 370 is a node dedicated only to the system control area. Each computing node 380 includes a processor 381 configured to execute an HPC application 350 (e.g., node 380-1 includes a processor 381-1 that executes application 350-1, etc.). Each HPC application 350 includes multiple areas.
[0073] The HPC system control node includes a processor 371 configured to instantiate a region-aware power / energy regulator 340. The regulator 340 may be similar to the regulator 140 described above. In this example, the regulator 340 receives region identification information and IPS measurements from an external source (i.e., from nodes 380-1 to 380-P). For example, the operating system interface of these nodes may provide this information to the regulator 340. Node 380-1 may provide region identification information region-1 indicating its current execution region and IPS measurement value IPS-1 measured based on its processor 381-1, and the regulator 340 may determine the optimal frequency of the region and send the frequency setting instruction Frequency-1 to the node 380-1. Similarly, node 380-P may provide region indication information region-P indicating its current execution region and IPS measurement value IPS-P and corresponding IPS measurement result, and the regulator 340 may return the frequency setting instruction Frequency-P. In this way, each node 380 may be a frequency that is individually optimized based on its current execution region. In some examples, the same region can be executed on multiple nodes 380 (concurrently, or at different times), and in some examples, when this occurs, the optimal frequency determined for one node 380 is applied to another node without having to characterize the region again for the other node 380 - in other words, in some examples, the regulator 340 can reuse information learned about one node 380 when regulating another node 380. In some examples, a single instance of the regulator 340 can be responsible for optimizing each node 380 (receiving input data from the node 380, characterizing the region of the node 380, and sending frequency setting commands to the node 380). In other examples, multiple instances of the regulator 340 can be instantiated, where each instance of the regulator 340 regulates a corresponding node in the nodes 380.
[0074] In some examples, the HPC system control node 370 also includes a job scheduler 372. The job scheduler 372 receives job requests from users, which may include an indication of an application desired to be run and a data set to be used for the application. The job scheduler 372 may then schedule the job on the node 380. In some examples, the job scheduler 372 may be configured to allow a user to specify a performance degradation parameter PD when entering a job, and may transmit this information to the regulator 340 so that the regulator 340 can use this information when calculating the optimal frequency of the region of the application.
[0075] Now turn to Figure 4, an example method 400 will be described. The method 400 may be performed by a zone-aware power / energy regulator (such as any of the regulators 140, 240, and 340 described above). In some examples, the zone-aware power / energy regulation instructions 130 include instructions for performing the operations of the method 400.
[0076] The method begins at block 401. In block 401, the regulator identifies the current execution region of the application, which is denoted herein as Reg n (where "n" is an arbitrary index used to identify the region). This identification of the current execution region may include, for example, looking up a current instruction address pointer and cross-referencing the current instruction address pointer with known addresses of the region of the application. In some examples, block 401 may be performed periodically to update the identification of the current execution region. The method then proceeds to block 402.
[0077] In block 402, the identified region Reg n During this time, the regulator measures the instructions per second (IPS) values, and these IPS values may be denoted herein as IPS n IPS n Values can include Reg n Multiple IPS measurements, including those obtained at different CPU frequencies, such as Figure 5 The method then proceeds to block 403.
[0078] In block 403, the regulator determines the region Reg n The calculation limitation parameter is expressed as %CB in this paper. n IPS-based n Measure the value to determine the parameter %CB n %CB n The compute-boundedness of the region is quantified. Specifically, compute-boundedness refers to the sensitivity of the region to changes in CPU frequency, where regions whose performance (e.g., completion time) is highly sensitive to CPU frequency are highly compute-bounded, and regions whose performance is not sensitive to CPU frequency are not compute-bounded (i.e., memory-bound). This sensitivity can be characterized as a percentage, where 100% means a 1-to-1 decrease in performance in response to a decrease in CPU frequency (e.g., a 50% decrease in frequency results in a 50% decrease in performance), and 0% means no change in performance in response to a decrease in CPU frequency. In some examples, %CB is determined in block 403. n This may include assessments such as those discussed above Figure 1 Described in Equation 1.
[0079] The method then proceeds to block 404. In block 404, the regulator determines the value of the %CBn To determine Reg n The optimal CPU frequency is expressed as Freq in this article. n For example, in block 404, Reg n The optimal CPU frequency may include evaluating as above about Figure 1 The method then proceeds to block 405 .
[0080] In block 405, the regulator instructs the system executing the application to execute Reg n The CPU frequency is set to Freq n .
[0081] Now turn to Figure 5 , another example method 500 will be described. The method 500 is an example implementation of the method 400.
[0082] Method 500 begins at block 501. In block 501, similar to block 401 described above, a current execution region Reg n The method then proceeds to block 506 .
[0083] In block 506, the regulator determines the region Reg n Whether it needs to be characterized. For example, if Reg n has been previously characterized, then the Reg n does not need to be characterized again. As another example, if the exception applies, then Reg n does not need to be characterized. An exception could be, for example, that the region does not exceed a threshold of importance (e.g., is executed for more than a threshold amount of time). n needs to be characterized, the process continues along the "yes" path to blocks 507 to 510. n Does not need to be characterized, the process continues along the “no” path to block 511 .
[0084] Blocks 507 to 510 correspond to an implementation example of block 402 in method 400. In block 507, the regulator sets the CPU frequency to a predetermined high value Freq n-high For example, this may be the maximum (ie, turbo) frequency. In block 508, while the CPU frequency is still at Freq n-high And the region Reg n In the case of still being performed, the IPS is measured, where the result is denoted as IPS in this paper. n-high In block 509, the regulator sets the CPU frequency to a predetermined low value Freq n-low For example, this can be any value below the high frequency. In block 510, while the CPU frequency is still at Freqn-low And the region Reg n In the case of still being performed, the IPS is measured, where the result is denoted as IPS in this paper. n-low The operations of blocks 507 to 510 may be performed in a different order than shown, for example, the low frequency IPS n-low In high frequency IPS n-high After completing blocks 507 to 510 (in whatever order they happen to be executed), the process then continues to block 503.
[0085] In block 503, the regulator is based on the IPS n-high and IPS n-low To determine Reg n The calculation limit parameter %CB n For example, Equation 1 can be used to determine %CB n The process then continues to block 504 .
[0086] In block 504, the regulator calculates the value based on %CB n To determine Reg n The optimal CPU core frequency is expressed as Freq in this article. n For example, Equation 2 can be used to determine Freq n .
[0087] In block 505, the regulator instructs the system to execute Reg n Set the CPU core frequency to Freq n .
[0088] In block 511 (reached from the "No" path following block 506), if the region Reg n The optimal CPU frequency Freq n (i.e., previously characterized Reg n ), and if no other exceptions apply, the CPU core frequency may be set to the previously determined optimal value Freq n In some embodiments, it may be possible to prevent the use of a previously determined Freq n An example of an anomaly is the region Reg n As an MPI wait region and the estimated duration of the region is less than the threshold, in this case, Freq n Can be covered and a predetermined high frequency can be used.
[0089] Now turn to Figure 6, another method 600 will be described. The method 600 can be used in conjunction with any of the methods 400 and 500 to supplement its area. In particular, the method 600 can be performed after executing the block 403 or 503 in the method 400 or 500.
[0090] In block 612, for region Reg n Determine the optimal CPU uncore frequency Freq n-uncore The optimal CPU uncore frequency Freq can be determined based on %CB n-uncore More specifically, based on the memory constraint %MB equal to 1-%CB. For example, Equation 5 described above can be used to determine Freq n-uncore .
[0091] In block 613, the regulator instructs the system to set the CPU uncore frequency to Freq n-uncore .
[0092] Now turn to Figure 7 , method 700 is described. Method 700 compresses an example implementation of block 506 of method 500 .
[0093] In block 716, it is determined whether any exceptions apply to the current execution region Reg n In some examples, the anomaly may include a region Reg n As a short MPI call (e.g., an MPI wait region predicted to have a duration less than a specified threshold). If the exception applies, the process proceeds along the "yes" path to block 718, where the region Reg n Does not need to be characterized. If no exception applies, the process proceeds along the “no” path to block 714 .
[0094] In block 714, the region Reg is determined n Has it been previously characterized? In this context, previous characterization includes determining whether Reg n The calculation limit parameter %CB n And confirm Reg n The optimal frequency Freq n In some examples, partially completed characterization (e.g., some IPS data has been sampled, but not enough to calculate %CB n ) will not count as the region being previously characterized, but only the completed characterization will suffice. n , the process continues along the "yes" path to box 718, in which the region Reg is determined n does not need to be characterized. If Reg has not been previously characterized n , the process proceeds along the “no” path to box 715.
[0095] In block 715, the region Reg is determined n In some examples, if the time taken to execute a region exceeds a threshold, Reg n is determined to be significant. In some examples, this threshold is 5% of the total execution time of the application. n is not significant, the process continues along the “NO” path to block 718, where the region Reg n does not need to be characterized. If Reg n is significant, the process proceeds along the “yes” path to block 717, where the region Reg n Does not need to be characterized.
[0096] Now turn to Figure 8 , describes a non-transitory computer readable medium 820. The non-transitory computer readable medium 820 includes region-aware power and energy adjustment instructions 830. The instructions 830 are similar to the instructions 130 described above and may include instructions for performing any of the methods 400, 500, 600, and 700 described above. In particular, the instructions 830 include region indication instructions 831 that may be similar to the instructions 131, region computational limitability determination instructions 832 that may be similar to the instructions 132, and frequency setting instructions 833 that may be similar to the instructions 133.
[0097] In the above description, various types of electronic circuits are described. As used herein, "electronic" is intended to be broadly understood to include all types of circuits that utilize electricity, including digital and analog circuits, direct current (DC) and alternating current (AC) circuits, and circuits for converting electricity into another form of energy, as well as circuits for performing other areas using electricity. In other words, as used herein, there is no distinction between "electronic" circuits and "electrical" circuits.
[0098] It should be understood that both the general description and the detailed description provide examples that are illustrative in nature and are intended to provide an understanding of the present disclosure without limiting the scope of the present disclosure. Various mechanical, compositional, structural, electronic and operational changes may be made without departing from the scope of the present specification and claims. In some instances, well-known circuits, structures and techniques are not shown or described in detail to avoid blurring these examples. In two or more drawings, the same number represents the same or similar element.
[0099] In addition, unless the context indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms as well. In addition, the terms "comprises", "comprising", "includes", etc. specify the presence of stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups. Unless otherwise specifically stated, components described as connected may be directly electrically or mechanically connected, or may be indirectly connected via one or more intermediate components. Unless the context of the specification indicates otherwise, mathematical and geometric terms do not necessarily have to be used in accordance with their strict definitions, as one of ordinary skill in the art will understand that, for example, substantially similar elements that are divided into regions in a substantially similar manner can easily fall within the scope of a descriptive term even if the term also has a strict definition.
[0100] And / or: Occasionally, the phrase "and / or" is used herein in conjunction with a list of items. This phrase is meant to include any combination of the items in the list - from a single item to all items, and any permutations in between. Thus, for example, "A, B, and / or C" means one of "{A}, {B}, {C}, {A, B}, {A, C}, {C, B}, and {A, C, B}".
[0101] Elements and related aspects described in detail with reference to one example may be included in other examples where they are not specifically shown or described when feasible. For example, if an element is described in detail with reference to one example and is not described with reference to a second example, the element may still be claimed to be included in the second example.
[0102] In addition, unless otherwise stated herein or the context implies otherwise, when approximate terms such as "substantially," "approximately," "about," "around," "roughly," and the like are used, this should be understood to mean that mathematical precision is not required, but rather refers to a range of variation that includes but is not strictly limited to the stated value, characteristic, or relationship. In particular, in addition to any ranges (if any) explicitly recited herein, the range of variation implied by the use of such approximate terms includes at least any insignificant variations and those variations that are typical for items of the type in question due to manufacturing tolerances or other tolerances in the relevant art. In any case, unless otherwise indicated, the range of variation may at least include values within ±1% of the recited value, characteristic, or relationship.
[0103] In view of the disclosure herein, further modifications and changes will be clear to those of ordinary skill in the art. For example, devices and methods may include additional components or steps omitted from the figures and descriptions for operational clarity. Accordingly, this description is only interpreted as illustrative and is to teach those skilled in the art the general manner of performing this teaching. It should be understood that the various examples shown and described herein will be considered exemplary. Elements and materials, and the arrangement of these elements and materials may be used to replace those shown and described herein, components and processes may be reversed, and certain features of this teaching may be utilized independently, all of which will be clear to those skilled in the art after benefiting from the description herein. Without departing from the scope of this teaching and the appended claims, the elements described herein may be changed.
[0104] It is to be understood that the specific examples set forth herein are non-limiting and that modifications in structure, dimensions, materials, and methods may be made without departing from the scope of the present teachings.
[0105] Other examples from this disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered exemplary only, and that the following claims be given their fullest breadth, including equivalents under applicable law.
Claims
1. An information processing system, comprising: processor; A non-transitory storage medium comprising instructions executable by the processor to instantiate a zone-aware power / energy regulator, the zone-aware power / energy regulator being configured to perform the following operations: periodically identifying, from a plurality of regions of an application, a region of the application currently being executed by a given processor; measuring instructions per second (IPS) of the given processor during execution of the identified region; determining a computational limitation parameter of the identified region based on the measured IPS; determining an optimal frequency for the identified region based on a computational constraint parameter for the identified region; as well as The given processor is instructing to set a frequency of the given processor to the optimal frequency during execution of the identified region.
2. The information processing system according to claim 1, in, Measuring the IPS of the given processor includes setting the frequency of the given processor to a high frequency and measuring the IPS to determine a high frequency IPS, and setting the frequency of the given processor to a low frequency and measuring the IPS to determine a low frequency IPS, and The computational limitation parameter is determined based on the high-frequency IPS and the low-frequency IPS.
3. The information processing system according to claim 2, in, Determining a computational limitation parameter for the identified region includes evaluating an equation relating the high frequency, the low frequency, the high frequency IPS, and the low frequency IPS as input variables to the computational limitation parameter as an output variable.
4. The information processing system according to claim 3, in, Determining the optimal frequency for the identified region includes evaluating an equation relating the computational limitations parameter and the performance degradation parameter as independent variables to the optimal frequency as a dependent variable, and The performance degradation parameter indicates an acceptable performance degradation level relative to a default performance.
5. The information processing system according to claim 4, in, The regulator is configured to receive user input specifying the performance degradation parameter.
6. The information processing system according to claim 1, in, The regulator is configured to: determine an optimal non-core frequency for the identified region based on a computational limitability parameter of the identified region; as well as During execution of the identified region, the given processor is instructed to set an uncore frequency of the given processor to the optimal frequency.
7. The information processing system according to claim 1, in, The regulator is configured to determine whether the identified region is an MPI call, and in response to determining that the region is an MPI call, determine an expected duration of the MPI call.
8. The information processing system according to claim 7, in, The regulator is configured to set the frequency of the given processor to a predetermined frequency if the identified region is an MPI call that is not expected to last longer than a specified threshold, regardless of the determined optimal frequency for the region.
9. The information processing system according to claim 8, in, The regulator is configured to set the frequency of the given processor according to the determined optimal frequency for the region if the identified region is an MPI call that is expected to last longer than the specified threshold.
10. The information processing system according to claim 1, in, The given processor executing the identified region is the processor.
11. The information processing system according to claim 10, in, The information processing system includes a computing node of a high performance computing (HPC) system.
12. The information processing system according to claim 1, in, The given processor executing the identified region is different from the processor.
13. The information processing system according to claim 12, in, The information handling system comprises a high performance computing (HPC) system, the processor is part of a system controller node of the HPC system, and the given processor is part of a compute node of the HPC system.
14. A region-aware power / energy regulation method, comprising: periodically identifying, from a plurality of regions of an application, a region of the application currently being executed by a given processor; measuring instructions per second (IPS) of the given processor during execution of the identified region; determining a computational limitation parameter of the identified region based on the measured IPS; determining an optimal frequency for the identified region based on a computational constraint parameter for the identified region; as well as During execution of the identified region, the given processor is instructed to set a frequency of the given processor to the optimal frequency.
15. The method of claim 14, in, Measuring the IPS of the given processor includes setting the frequency of the given processor to a high frequency and measuring the IPS to determine a high frequency IPS, and setting the frequency of the given processor to a low frequency and measuring the IPS to determine a low frequency IPS, and The computational limitation parameter is determined based on the high-frequency IPS and the low-frequency IPS.
16. The method of claim 14, further comprising: determining an optimal non-core frequency for the identified region based on a computational constraint parameter for the identified region; as well as The given processor is instructed to set an uncore frequency of the given processor to the optimal frequency during execution of the identified region.
17. The method of claim 14, further comprising: A determination is made as to whether the identified region is an MPI call and in response to determining that the region is an MPI call, an expected duration of the MPI call is determined.
18. The method of claim 17, in, In response to the identified region being an MPI call that is not expected to last longer than a specified threshold, setting the frequency of the given processor to a predetermined frequency regardless of the determined optimal frequency for the region.
19. The method of claim 17, in, The regulator is configured to, in response to the identified region being an MPI call that is expected to last longer than a specified threshold, set the frequency of the given processor according to the determined optimal frequency for the region.
20. A non-transitory storage medium comprising instructions executable by a processor to instantiate a locality-aware power / energy regulator, the locality-aware power / energy regulator being configured to: periodically identifying, from a plurality of regions of an application, a region of the application currently being executed by a given processor; measuring instructions per second (IPS) of the given processor during execution of the identified region; determining a computational limitation parameter of the identified region based on the measured IPS; determining an optimal frequency for the identified region based on a computational constraint parameter for the identified region; as well as The given processor is instructing to set a frequency of the given processor to the optimal frequency during execution of the identified region.