Dynamic guard band with timing protection and performance protection
The dynamic adjustment of loop sensor parameters in processors addresses voltage droop-related errors, improving processor efficiency and yield by reducing recoverable errors and optimizing timing and performance.
Patent Information
- Application Number
- JP2025500277
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-19
- Filing Date
- 2023-07-17
- Publication Date
- 2025-07-25
AI Technical Summary
Existing processor technologies face challenges in maintaining optimal timing and performance due to voltage droops, leading to recoverable and irrecoverable errors, which are addressed by implementing a dynamic guard band that adjusts loop sensor parameters to manage voltage and timing margins dynamically.
A method that monitors core recovery events and adjusts loop sensor parameters to maintain optimal timing and performance by increasing or decreasing the voltage setpoint and loop sensor delay to prevent irrecoverable errors, thereby optimizing processor efficiency and yield.
This approach reduces the occurrence of recoverable errors, improves processor yield, and enhances performance by dynamically managing voltage and timing margins, allowing more cores to operate simultaneously while minimizing power consumption.
Smart Images

Figure 2025523791000001_ABST
Abstract
Description
Background Art
[0001] The present invention generally relates to a computer-implemented method, a computer system, and a computer program product configured and arranged to provide a dynamic guard band with timing protection and / or performance protection to a computer system, and more specifically.
[0002] In a distributed computing environment, there can be numerous jobs or queries that arrive as workloads to be processed on a processor in the computing environment. A processor core is a processing unit that reads instructions and performs specific actions. The instructions are linked together, and as a result, when the instructions operate in real time on the processor, the processor executes the desired workload configured by the instructions. A multi-core processor is a computer processor on a single integrated circuit that has two or more separate processing units, each of which reads and executes program instructions. The instructions are normal instructions (e.g., addition, data movement, branching, etc.), but a single processor can operate instructions simultaneously on separate cores, increasing the overall speed of a program that supports multi-threading or other parallel computing techniques.
[0003] Problems can occur in the operation of the processor, and the cores of the processor operate to avoid these problems. Techniques are needed to improve the timing and / or performance of the processor.
Summary of the Invention
[0004] Embodiments of the present invention are directed to a computer-implemented method for a dynamic guard band with timing protection and / or performance protection. A non-limiting computer-implemented method includes, in response to monitoring a processor during operation, detecting, by a computer, a first number of core recovery events in the processor, and determining, by the computer, that the first number of core recovery events satisfies a first condition for a first core recovery event threshold. The method includes modifying, by the computer, a value of at least one loop sensor calibration or adjustment parameter of the processor by a first amount, the at least one loop sensor calibration or adjustment parameter affecting sensitivity to a voltage droop. The method includes, in response to modifying the value of the at least one loop sensor calibration or adjustment parameter by the first amount, detecting, by the computer, a second number of core recovery events in the processor, and determining, by the computer, that the second number of core recovery events satisfies a second condition for a second core recovery event threshold. The method includes modifying, by the computer, the value of the at least one loop sensor calibration or adjustment parameter of the processor by a second amount.
[0005] This can provide an improvement over known methods for using a static guardband by reducing the timing margin until a first number of core recovery events occur, thereby reducing power and allowing for a higher yield (i.e., a greater number of processor cores operating simultaneously in the processor), which improves the flow instructions for the workload. Next, to handle a first number of core recovery events that coincide with the processing of a large workload, the timing margin and / or voltage margin is increased. After a second number of core recovery events (which may be no core recovery events), the timing margin is decreased, thereby saving power and improving yield (e.g., increasing the number of processor cores operating simultaneously on the processor). With regard to the adjustment of the voltage margin, in a memory array, the circuit may be impaired by the voltage margin, and one or more embodiments may similarly adjust the voltage margin to avoid failures.
[0006] In addition to, or alternatively to, one or more of the features described above or below, a further embodiment of the invention discloses that the step of modifying the value of the at least one loop sensor delay parameter of the processor by the first amount comprises increasing the value of the at least one loop sensor calibration or adjustment parameter of the processor by the first amount. Thereby, advantageously, when a first condition for a first core recovery event threshold is met to prevent a processor core from experiencing an irrecoverable error that results in a service interruption, the execution of workload instructions is reduced / delayed. When there is an "irrecoverable error" that is a failure of a server and mainframe system, a checkpoint may be included or may occur. Further, increasing the digital loop sensor delay prevents recovery events and / or at least reduces their rate. Too many recovery events can affect performance, and if a recovery event occurs during a recovery action, this can result in an irrecoverable error. In some cases, a checkpoint may be used synonymously with a (more recoverable or) irrecoverable error.
[0007] In addition to, or as an alternative to, one or more of the features described above or below, a further embodiment of the present invention discloses that the step of modifying the value of the calibration or adjustment parameter of the at least one droop sensor of the processor by the second amount has the step of decreasing the value of the calibration or adjustment parameter of the at least one droop sensor of the processor by the second amount. Advantageously, this means that since a second condition for the second core recovery event threshold is satisfied, the number of instructions of the processor being executed increases / the speed increases, which means that more instructions are processed without the concern of irrecoverable errors.
[0008] In addition to, or as an alternative to, one or more of the features described above or below, a further embodiment of the present invention discloses that satisfying the first condition for the first core recovery event threshold includes that the first number of core recovery events is greater than the first core recovery event threshold. Advantageously, this means that when the first condition for the first core recovery event threshold is satisfied to prevent the processor core from having irrecoverable errors, the execution of the workload instructions is reduced / delayed.
[0009] In addition to, or as an alternative to, one or more of the features described above or below, a further embodiment of the present invention discloses that satisfying the second condition for the second core recovery event threshold includes that the second number of core recovery events is less than the second core recovery event threshold. Advantageously, this means that since the second condition for the second core recovery event threshold is satisfied, the number of instructions of the processor being executed increases / the speed increases, which means that more instructions are processed without the concern of having irrecoverable errors.
[0010] In addition to, or alternatively to, one or more of the features described above or below, a further embodiment of the invention discloses that the step of modifying the value of the calibration or adjustment parameter of the at least one droop sensor of the processor by the second amount has the step of returning to a baseline value for the calibration or adjustment parameter of the at least one droop sensor. Advantageously, thereby, since a second condition for a second core recovery event threshold is satisfied, the number of instructions of the processor executed increases / the speed increases, which means that more instructions are processed without the concern of having an irrecoverable error. The system can return to the baseline if no recoverable errors are seen in the system over a predetermined period, and as a result, the system can then reduce the margin threshold. As described herein, since the baseline is the starting point at which the system was first started, the baseline is "safe". For example, if the power, temperature, or current of the chip or system begins to approach safety limits, this can further cause the system to return to the baseline. Therefore, the core / processor returns to the baseline margin threshold and stays within the range of these limits.
[0011] In addition to, or alternatively to, one or more of the features described above or below, a further embodiment of the invention discloses that the second amount is greater than the first amount, equal to the first amount, or less than the first amount.
[0012] Other embodiments of the invention implement the features of the above method in a computer system and a computer program product.
[0013] Further technical features and benefits are realized by the technology of the invention. Embodiments and aspects of the invention are described in detail herein and are considered part of the claimed subject matter. Refer to the detailed description and drawings for a better understanding.
Brief Description of the Drawings
[0014] The details of the exclusive rights described in this specification are particularly pointed out and are clearly claimed in the claims at the conclusion of the specification. The above and other features and advantages of the embodiments of the present invention will be apparent from the following detailed description taken in conjunction with the accompanying drawings.
[0015]
Figure 1
[0016]
Figure 2
[0017]
Figure 3
[0018]
Figure 4
[0019]
Figure 5
[0020]
Figure 6
[0021]
Figure 7A
[0022]
Figure 7B
[0023]
Figure 8A
[0024]
Figure 8B
[0025]
Figure 9
[0026]
Figure 10
[0027]
Figure 11
[0028]
Figure 12
[0029]
Figure 13
[0030] One or more embodiments of the present invention are described for a computer-implemented method, a computer system, and a computer program product configured and arranged to provide a dynamic guard band with timing protection for a processor core on a processor and / or a dynamic guard band with performance protection. The terms processor core, core, or core unit may be used interchangeably. The terms processor, chip, processor chip, and integrated circuit may be used interchangeably. Timing guard band protection sometimes allows the timing guard band to be reduced, thereby providing an improvement in core yield, an improvement in power dissipation, and / or an improvement in timing margin. Although timing margin may be discussed herein, it should be noted that one or more embodiments are equally applicable to timing and / or voltage margins. The digital loop sensor can be digital or analog and can be a timing sensor or a voltage sensor. In one or more embodiments, for example, a dynamic guard band with timing protection allows the processor to operate at its nominal voltage with little or no voltage droop under normal conditions such as a steady-state workload, nominal temperature, etc. In particular, since the power supply noise is minimal and there is no large voltage droop, the processor operates with a reduced guard band (i.e., with a reduced voltage margin and / or a reduced timing margin). When its workload switches directly from an idle state to a highly active workload, this change in activity appears as the worst-case voltage droop (as further discussed herein).
[0031] Typically, in a core, there are some critical circuit paths that may malfunction and cause errors when the voltage of the circuit drops to a critical value. In a core designed for high robustness, there is an error-checking circuit that detects these errors. When an error is detected, the system can be returned to the checkpoint state before the error for retry. As a result, most errors are recoverable. Even though most circuit errors are recoverable, a high rate of error recovery can affect performance. Furthermore, if there is an error during the error recovery process itself, this can result in an unrecoverable error, which can cause a more serious system interruption.
[0032] Without the technical benefits of one or more embodiments of the present invention, a way to address it is to operate (constantly) at a fixed voltage regulator setpoint voltage that is high enough to address its worst-case voltage droop. The voltage setpoint required for each desired operating frequency is set during the manufacturing test of each chip and stored in a non-volatile memory called VPD (vital product data). Even if the voltage setpoint is kept constant, it is impossible to keep the voltage at the circuit level constant due to the limitations of the voltage regulator and the resistance and inductance of the parasitic board, socket, and package. When the VDD current increases rapidly over time (also known as di / dt), the power supply droops at the on-chip circuit location. The voltage difference (i.e., delta) from the voltage setpoint that causes an error with respect to the actual voltage setpoint is called the voltage margin as illustrated in FIGS. 7A and 7B discussed herein.
[0033] Generally, the critical path accelerates with the increase in voltage, and thus the timing margin increases with the increase in voltage. In manufacturing tests, it has been found that the Vmin workload operates without any circuit errors when the VDD set point is at or above the minimum value Vmin. To add margin during manufacturing tests to account for different workloads, temperature variations, device degradation over life, etc., the VDD set point is increased by N% up to Vmin + N * 0.01V, resulting in an N% timing margin (in units of V%) for the Vmin workload in the manufacturing environment for the Vmin workload. Next, all loop sensors are calibrated when operating this Vmin workload such that the minimum output value of each loop sensor is calibrated to the desired value corresponding to the N% margin. Next, the worst-case loop workload (with events having a large di / dt that cause the worst-case loop) is operated, and the loop sensors are adjusted to mitigate these worst-case loops such that the minimum loop sensor output never drops below the desired calibrated value, and thus an N% margin is maintained at all loop sensor positions during the loop workload. This adjustment of loop sensing is adjusted to provide the desired N% margin over the range of VDD set point values. At lower VDD set points, when a higher current workload with events having a larger di / dt is operated, the loop sensors will detect loops more frequently, and loop mitigation will be required more frequently to maintain the desired N% timing margin at all sensor positions. Each loop mitigation event or period of constant loop mitigation generally results in a clock frequency reduction or instruction slot throttling, which can reduce performance. Therefore, after loop sensor calibration and adjustment, the next step is to select a VDD set point high enough such that loop mitigation events are rare, resulting in an acceptable performance loss not occurring for the selected performance workload.The calibration and adjustment parameters of each droop sensor, as well as the VDD set point, which result in acceptable performance, are stored in the VPD. When the Vmin workload, droop workload, and performance workload accurately reflect the customer's environment, and when the temperature environment in the power delivery network (PDN) and manufacturing test environment exactly matches the customer's environment, there will be an N% timing margin in the customer's application. Unfortunately, it is difficult to predict and address the characteristics of all customer workloads, PDNs, and temperature environments during manufacturing tests. Therefore, to ensure the reliability of the system, typically N% of the timing margin is increased. This results in higher voltage and higher power. Voltage is typically limited due to concerns about the long-term reliability of devices and insulators, and power is also limited by cooling capacity constraints. Therefore, increasing N% to increase the timing margin generally reduces chip yield and achievable system performance. One or more embodiments are configured to provide a reduction in the timing margin that is beneficial for static and dynamic power consumption and chip yield. In other words, reducing the timing margin allows the processor to operate at a higher clock frequency, or alternatively, more cores or processors can be configured while staying within voltage and power constraints, thereby improving the functionality of the computer system. In particular, one or more embodiments use dynamic guardbands to handle idle state workloads (or small workloads) that require little or no processing, as well as extreme workloads that require intensive processing. The dynamic guardband uses the droop sensor trip point to handle any variations in the devices on the chip, the chip itself, the power delivery network, the chip temperature environment, or the workload operating on the chip, such that the dynamic guardband can be automatically changed (continuously) on-the-fly in the computing environment.
[0034] The droop sensor can be digital or analog, can directly sense voltage, or they can sense a circuit delay that is sensitive to voltage as well as cycle time. In one or more embodiments, a digital droop sensor that is sensitive to both voltage and cycle time is used. The digital delay sensor is calibrated by adding delay elements to or subtracting delay elements from a delay line. Any type of analog or digital droop sensor can be used to calibrate and adjust the droop sensor. In one embodiment, the droop sensor is calibrated using a Vmin workload, and then a separate adjustment process is used to select a threshold for triggering a droop reduction response in the case of a large and fast droop in the worst case. In other embodiments, different methods and sequences can be used to calibrate and / or adjust the droop sensor to provide the desired droop reduction behavior. Further, the droop sensor can be any kind of margin sensor, such as an analog timing sensor, a digital voltage sensor, or an analog voltage sensor. Generally, any type of margin sensor can be utilized in one or more embodiments. Since the digital droop sensor is a timing margin sensor, the system can adjust the threshold of this timing margin sensor by adjusting / calibrating the delay of the digital droop sensor. It should be understood that other margin sensors with various types of thresholds can be used in one or more embodiments. In one or more embodiments, it can be found that the timing or voltage margin is sensitive to temperature. Therefore, the system can adjust the margin threshold as a function of temperature to avoid recoverable errors before they occur at high temperatures.
[0035] According to one or more embodiments, a dynamic guard band with performance protection prevents and / or reduces core performance degradation. A chip (e.g., a processor) has a given digital loop sensor trip point. When the voltage level crosses the digital loop sensor trip point due to the workload on the core, a given performance degradation will occur. According to one or more embodiments, by increasing the voltage of the core and assuming that the digital loop sensor trip point remains at that voltage level (exactly), as illustrated in FIGS. 8A and 8B, the distance between the new voltage set point for the chip and the digital loop sensor trip point increases. This means that in the case of a workload that previously crossed the digital loop sensor trip point and thus had an early performance hit, at the new voltage set point, the same workload will no longer cross the digital loop sensor trip point, and thus, the performance degradation (e.g., performance hit) will no longer be seen. Therefore, a dynamic guard band with performance protection can dynamically change the voltage set point (also referred to as VDD voltage, drain voltage, positive supply voltage) to provide dynamic performance protection.
[0036] In one or more embodiments, a dynamic guard band with timing protection that dynamically increases / decreases the calibration and adjustment of the digital loop sensor within each core can be integrated with a dynamic guard band with performance protection that dynamically increases / decreases the voltage set point (e.g., VDD voltage) of the core. For purposes of explanation and to facilitate understanding, the dynamic guard band with timing protection and the dynamic guard band with performance protection can be discussed separately, but both functionalities are intended to be integrated for use in improving a computer system.
[0037] One or more embodiments of the present invention provide improvements in processors, and in particular, improvements in cores on the processor. By optimizing the timing guard band, the probability or rate of margin-related circuit errors is reduced while minimizing voltage and power for higher efficiency, while the performance guard band prevents performance degradation on the cores of the processor due to excessive amounts of droop reduction actions. This operation of the core results in an improvement in the computer system itself by finely tuning the operation of the core on the processor during runtime when the processor core is executing instructions. Also, the dynamic guard band with timing protection and / or the dynamic guard band with performance protection are configured to operate the processor strongly at a high level while reducing the probability or rate of recoverable or non-recoverable circuit errors.
[0038] Referring now to FIG. 1, a computer system 100 according to one or more embodiments of the invention is generally shown. The computer system 100 can be an electronic computer framework that includes and / or uses computing devices and networks that utilize various communication technologies in any number and combination, as described herein. The computer system 100 can be easily scalable, extensible, and modular due to its ability to change for different services or reconfigure some features independently of others. The computer system 100 can be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smartphone. In some examples, the computer system 100 can be a cloud computing node. The computer system 100 can be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, and data structures, etc., that perform specific tasks or implement specific abstract data types. The computer system 100 can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be arranged on both local computer system storage media and remote computer system storage media, including memory storage devices.
[0039] As shown in FIG. 1, computer system 100 has one or more central processing units (CPUs) 101a, 101b, 101c, etc. (collectively or generically referred to as processor 101). Processor 101 can be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. Processor 101, also referred to as a processing circuit, is coupled to system memory 103 and various other components via system bus 102. System memory 103 can include read-only memory (ROM) 104 and random access memory (RAM) 105. ROM 104 is coupled to system bus 102 and can include a basic input / output system (BIOS) that controls certain basic functions of computer system 100, or its successor such as the Unified Extensible Firmware Interface (UEFI). RAM is a read-write memory coupled to system bus 102 for use by processor 101. System memory 103 provides a temporary memory space for the operation of the above instructions during operation. System memory 103 can include random access memory (RAM), read-only memory, flash memory, or any other suitable memory system.
[0040] Computer system 100 includes an input / output (I / O) adapter 106 and a communication adapter 107 coupled to system bus 102. I / O adapter 106 can be a small computer system interface (SCSI) adapter that communicates with hard disk 108 and / or any other similar component. I / O adapter 106 and hard disk 108 are collectively referred to herein as mass storage 110.
[0041] Software 111 for execution on computer system 100 may be stored within mass storage 110. Mass storage 110 is an example of a tangible storage medium readable by processor 101, where software 111 is stored as instructions for execution by processor 101 to cause computer system 100 to operate as described hereinafter in this specification with respect to the various figures. Examples of computer program products and execution of such instructions are discussed in more detail herein. Communication adapter 107 interconnects system bus 102 with network 112, which may be an external network, enabling computer system 100 to communicate with other such systems. In one embodiment, system memory 103 and a portion of mass storage 110 collectively store an operating system, which may be any suitable operating system for coordinating the functions of the various components shown in FIG. 1.
[0042] Additional input / output devices are shown as being connected to system bus 102 via display adapter 115 and interface adapter 116. In one embodiment, adapters 106, 107, 115, and 116 may be connected to one or more I / O buses that are connected to system bus 102 via an intermediate bus bridge (not shown). A display 119 (e.g., a screen or display monitor), which may include a graphics controller and a video controller to improve the performance of graphics-intensive applications, is connected to system bus 102 by display adapter 115. A keyboard 121, a mouse 122, a speaker 123, etc. may be interconnected with system bus 102 via interface adapter 116, which may include, for example, a Super I / O chip that integrates multiple device adapters into a single integrated circuit. Appropriate I / O buses for connecting peripheral devices such as hard disk controllers, network adapters, and graphics adapters typically include common protocols such as Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe). Thus, as configured in FIG. 1, computer system 100 includes a processing function in the form of processor 101, a storage function including system memory 103 and mass storage 110, input means such as keyboard 121 and mouse 122, and output functions including speaker 123 and display 119.
[0043] In some embodiments, the communication adapter 107 can transmit data using any suitable interface or protocol, such as, among other things, an Internet Small Computer System Interface. The network 112 can be, among other things, a cellular network, a wireless network, a wide area network (WAN), a local area network (LAN), or the Internet. An external computing device can be connected to the computer system 100 through the network 112. In some examples, the external computing device can be an external web server or a cloud computing node.
[0044] It should be understood that the block diagram of FIG. 1 is not intended to show that the computer system 100 includes all of the components shown in FIG. 1. Rather, the computer system 100 can include any suitable fewer or additional components not shown in FIG. 1 (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.). Further, the embodiments described herein in connection with the computer system 100 can be implemented with any suitable logic. The logic referred to herein can include, in various embodiments, any suitable hardware (e.g., among other things, a processor, an embedded controller, or an application specific integrated circuit), software (e.g., among other things, an application), firmware, or any suitable combination of hardware, software, and firmware.
[0045] FIG. 2 illustrates a block diagram of an exemplary computer system 202 configured to provide a dynamic guard band with timing protection and / or a dynamic guard band with performance protection for processor cores on a processor, according to one or more embodiments of the present invention. A drawer with a plurality of processors 204 connected thereto may be configured, and a plurality of drawers may be connected to each other. A computer system 202 having a plurality of processors 204 may be regarded as a drawer. Many features of the computer system 100 including hardware and software may be integrated into the computer system 202. The computer system 202 includes a processor 204, and details of an example of the processor are shown. The processor 204 has a plurality of cores 220A to 220N, where N represents the last number of the foregoing elements. The processor cores 220A to 220N may generally be referred to as processor cores 220. Each of the processor cores 220A to 220N has its own digital droop sensor (DDS) 222A to 222N, throttle meter 224A to 224N, and firmware (FW) 226A to 226N. The digital droop sensors 222A to 222N may generally be referred to as digital droop sensors 222. The throttle meters 224A to 224N may generally be referred to as throttle meters 224. Similarly, the firmware 226A to 226N may generally be referred to as firmware 226.
[0046] One or more power circuits 230 are controlled by a controller 232 to provide power to respective cores 220A - 220N on a processor 204. Firmware 236 may be utilized to control one or more operations of the power circuit 230 and / or the controller 232. In one or more embodiments, any one or more of firmware 226A - 226N and firmware 236 may be configured as (separate) state machines. For example, firmware 226A - 226N and firmware 236 may be implemented as on - chip state machines. A circuit that operates according to a particular sequence of events is called a state machine or a sequence circuit. A state machine requires memory to store information about past actions and uses that memory to determine the next action to take.
[0047] The computer system 202 may include and / or represent various software applications such as software 111 that can be executed as instructions on one or more processors 101 to be implemented according to one or more embodiments of the present invention. Although not shown, the processor 204 includes all hardware and software elements for functioning as understood by those skilled in the art, which includes logical units, caches, registers, fetch circuits, decode circuits, execution circuits, clocks, buses, etc. The computer system 202 may represent one or more portions of the cloud computing environment 50 illustrated in FIG. 12. A dynamic guard band with timing protection and / or a dynamic guard band with performance protection may be incorporated and / or integrated into the hardware and software layer 60 illustrated in FIG. 13. FIGS. 12 and 13 are further discussed herein.
[0048] For purposes of illustration rather than limitation, in the exemplary scenario, Core 220A is used to show the use of a dynamic guard band with timing protection and / or performance protection. By analogy, it should be understood that a dynamic guard band with timing protection and / or performance protection may be implemented simultaneously in Cores 220A - 220N according to one or more embodiments. Similarly, a dynamic guard band with timing protection and / or performance protection may be implemented simultaneously in Processor 204.
[0049] FIG. 3 is a flowchart of a process 300 that uses a dynamic guard band with timing protection for Processor 204 according to one or more embodiments. Any of the figures discussed herein may be referenced.
[0050] In block 302 of process 300, the firmware 226A of core 220A is configured to monitor the operation of core 220. The firmware 226A can communicate with the digital loop sensor 222A to monitor the voltage droop (or decrease) in core 220A while core 220A processes workload instructions. The firmware 226A can communicate with the throttle meter 224A to monitor the performance degradation of core 220A while the core processes the workload. FIG. 7A illustrates a graph of the voltage of core 220 executing a workload over time. As can be seen in FIG. 7A, core 220 has a default digital loop trip point at a set voltage level (e.g., VPD). In one or more embodiments, each of cores 220A - 220N can have the same digital loop trip point at their respective digital loop sensors 222A - 220N. In one or more embodiments, one or more of cores 220A - 220N can be set to have different digital loop trip points, and / or the digital loop trip points of one or more cores can change over time according to their operation. When core 220 operates above the digital loop trip point, there is a performance hit and no action is taken. In other words, the voltage level of the core has not dropped to the digital loop trip point. However, when the processor (or core 220) reaches or falls below the digital loop trip point, a stall execution event is inserted into some of cores 220. When the execution event is stalled for core 220, the effect is that those cores 220 are no longer executed in each clock cycle and instead pause for a given cycle only. This reduces the voltage droop but also affects performance. In other words, the execution of instructions is delayed for a given clock cycle only, thereby being identified as a stall execution event.
[0051] In FIG. 7A, the core 220 may be allowed to have a voltage droop (decrease) to the core recovery (zone) a predetermined number of times, but the core 220 is not allowed to droop (decrease) to a checkstop or an irrecoverable error (zone) at all, and / or is prevented from doing so. Core recovery is a process on the processor 202 that resets the core to the last known good architectural state (checkpoint state). Core recovery may include clearing the cache (i.e., array built-in self-test (ABIST)), resetting the state machine, and restoring any shadow copies of the architectural registers to the last known good state. In other words, the processor 204 is configured to store the architectural state of a given register. Once that is done, the processor 204 is configured to reset a given core 220. To that end, the processor 204 is configured to clear the cache (which is done by ABIST), reset a given state machine to IDLE, reload the architectural state into a given register, and then continue execution therefrom. All of these occur while the core 220 remains in the operating state.
[0052] An irrecoverable error of the core occurs when core recovery fails. The irrecoverable error causes a checkstop, which in turn stops the clock of that core 220, and the core 220 can no longer proceed. In the case of some processors 204, there is a process of further evacuating the workload that was operating on that core to a spare core, or adding this workload to an already operating core 220 if there is no available spare core.
[0053] Referring back to FIG. 3, in block 304, the firmware 226A of core 220A is configured to detect a core recovery event. A core recovery event occurs when the instructions of a workload are being processed, and the intensity or requirements of the workload cause the voltage of core 220 to droop to and / or below a core recovery voltage threshold for core recovery. Under normal conditions, such as a steady-state workload, nominal temperature, etc., core 220 operates at its nominal voltage (e.g., at or near the supply voltage) with little to no voltage droop. In this case, when the workload switches from an idle state to a fully executing state (immediately), this change in activity appears as the worst-case voltage droop, as illustrated by a large voltage droop across the core recovery voltage threshold in FIG. 7A. Since the DDS detection time and the droop mitigation reaction time allow the voltage to continue to drop after the voltage has passed the DDS trip point, the DDS trip point must be set higher than the checkpoint stop threshold.
[0054] In response to detecting a core recovery event, the firmware 226A of core 220A is configured to activate core recovery. The processor 204 is configured to activate core recovery as discussed herein. Further, the processor 204 may execute any core recovery process as would be understood by one of ordinary skill in the art.
[0055] In block 306, if the number of core recovery events in a set period meets or exceeds a core recovery event threshold, the flow proceeds to block 308. If the number of core recovery events in the set period is less than or does not exceed the core recovery event threshold, the flow proceeds to block 302.
[0056] In block 308, the firmware 226A of core 220A is configured to check whether the value of one or more loop sensor calibration or adjustment parameters in the digital loop sensor 222A has reached its maximum value. In other words, the firmware 226A can check whether the loop sensor calibration or adjustment parameter is already at the maximum delay, which means that no further delay is desired. In one or more embodiments, block 308 may not be triggered until a predetermined number of core recovery events occur for the core. For example, two core recovery events may be used to trigger block 308, because the first recovery event may have been an aberration, while the second recovery event is an indication that further action is guaranteed for the core. In one or more embodiments, the predetermined number of core recovery events required to trigger block 308 may range from about 1 to 5 core recovery events.
[0057] There can be a range of delay units for the loop sensor delay parameter. For example, the range of delay units can be in the range of 0 to 255, where each delay unit can be a delay of 5 picoseconds (ps), and / or each delay unit can correspond to the addition of a delay element. For illustrative purposes, the nominal value of the loop sensor delay parameter can be 100 delay units or 500 ps, which, when the core has a cycle time of 4 GHz, corresponds to two clock cycles of 250 ps each. Although exemplary values have been discussed, larger or smaller values can be utilized for the delay of the delay unit and the cycle time of the core. For illustrative purposes, FIG. 11 illustrates an exemplary digital loop sensor according to one or more embodiments. The programmable delay can be increased (or decreased) as discussed herein, and an increase in the programmable delay delays the execution of the workload's instructions, while a decrease in the programmable delay does not delay (or reduces the delay of) the execution of the workload's instructions. The programmable delay has a maximum value. The maximum delay can also be a delay set to limit the increase in voltage, current, and power that can result when the loop sensor delay increases. When both timing protection and voltage protection are active, increasing the timing protection with an increase in the loop sensor delay can result in more performance loss, so then the VDD can be increased to a dynamic margin with performance protection. Without a limit on the DDS delay, this can result in excessive voltage, current, or power limits and contribute to thermal runaway. Also, the larger the programmable delay setting, the less the propagation of the signal through the edge detector circuit illustrated in FIG. 11 (and the lower the detected edge value). Since the lower edge detection values are correlated with the voltage droop event, they trigger the throttling of instruction execution. Therefore, the programmable delay can be used to indirectly set the threshold of the voltage level at which throttling will be triggered.
[0058] In block 310, if the maximum value of the (no) parameter has not been reached, the firmware 226A of the core 220A is configured to increase the delay value of the loop sensor calibration or adjustment parameter in the digital loop sensor 220A. The digital loop sensor 222 (circuit) has a delay element. To increase the digital loop trip point to a higher voltage set point as illustrated in FIG. 7B, the firmware 226A increments the loop sensor delay and / or trigger threshold in the digital loop sensor 222. By increasing the DDS delay, the voltage level of the trip point or the threshold at which instruction execution slot throttling is triggered is indirectly increased. An example of the nominal delay value is 100. Next, to increase the sensitivity of the digital loop sensor trip point, the firmware 226A increases the value (of the digital loop sensor) from 100 to 101, 102, 103, 104, etc. as needed to control the rate of circuit errors and recovery events. When the firmware 226A increases the digital loop sensor trip point, this causes the execution slot throttling to be triggered at a higher voltage for the loop reduction reaction. This throttling often stops and increases the voltage from drooping, thereby giving the core 220 more timing margin during voltage droop or during high current operation of a large workload. Further, the firmware 226A is configured to dynamically reduce the timing margin as discussed below in FIG. 4 if the rate of the recovery event drops below the threshold. This reduces the performance caused by the loop reduction instruction throttling. Since the performance loss is reduced, this also results in a dynamic margin by the performance protection control loop that reduces the voltage, thereby reducing the current and power for efficiency improvement.
[0059] As further seen in FIG. 7B, a graph of the voltage of core 220 executing a workload over time is illustrated, and the voltage set point as the digital loop sensor trip point is increased to a new digital loop sensor trip point. The new voltage set point corresponds to an increase in the amount of delay for core 220. As seen in FIG. 7B, there is a maximum voltage at which the digital loop sensor trip point can be moved from the original DDS sensor input corresponding to the original DDS delay stored in the VPD. In one or more embodiments, the maximum voltage change of the voltage set point can be 10 millivolts (mV), 15 mV, 20 mV, 25 mV, 30 mV, etc. The firmware 226A can have a preset step / increment for moving the DDS delay and the resulting DDS sensor trip point before stopping at the maximum DDS delay. Note that these figures represent operation at a single frequency because the relationship between the DDS delay and the trip point voltage changes when the clock frequency changes. For example, for each recovery event (or a preset number of recovery events), the firmware 226A can move the DDS trip point up to a voltage level of "J" mV until reaching the maximum amount of voltage change for the digital loop sensor trip point, where J is the step / increment and J can be 5 mV, 10 mV, etc.
[0060] FIG. 4 is a flowchart of a process 400 that uses a dynamic guard band with timing protection for a processor 204 according to one or more embodiments. FIG. 4 can continue the process discussed in FIG. 3. In block 402 of process 400, the firmware 226A of core 220A is configured to monitor the operation of core 220. The firmware 226A can communicate with the digital loop sensor 222A to monitor the voltage droop (or decrease) in core 220A while core 220A processes a workload.
[0061] In block 404, the firmware 226A of core 220A is configured to determine / check the number of core recovery events within a predetermined time after an increase to a new / updated voltage level of the voltage setpoint. The predetermined time after the last increase in the voltage setpoint of the digital loop sensor trip point can range from about 3 minutes to about 50 minutes. As described herein, a core recovery event can occur when a workload instruction is being processed, and depending on the intensity or requirements of the workload, the voltage of core 220 can be drooped to and / or below a core recovery voltage threshold for core recovery.
[0062] In block 406, the firmware 226A of core 220A is configured to check whether the number of core recovery events at a predetermined time meets and / or falls below a decrease delay core recovery event threshold. In one or more embodiments, the decrease delay recovery event threshold can be 0 recovery events at a predetermined time. In one or more embodiments, the decrease delay recovery event threshold can be a number of core recovery events less than the number of core recovery events that caused the increase in the voltage setpoint of the digital loop sensor trip point. In one or more embodiments, the decrease delay recovery event threshold can be less than a predetermined number of core recovery events required to trigger block 308 in FIG. 3. If the number of core recovery events at a predetermined time does not meet and / or does not fall below the decrease delay core recovery event threshold, the flow returns to block 402 and the monitoring of processor 204 continues.
[0063] In block 408, if the number of core recovery events at a given time (yes) decreases and satisfies and / or falls below the reduced delay core recovery event threshold, the firmware 226A of core 220A is configured to check whether the values of one or more digital loop sensor delay parameters have reached their minimum values. The range of delay units can be in the range of 0 to 255 such that the minimum value (or lowest value) of the digital loop sensor delay parameter is 0 while the maximum value is 255. It should be understood that different ranges can be utilized. (Yes) If the minimum value of the digital loop sensor delay parameter is reached, the flow returns to block 402.
[0064] In block 410, if (no) the value of the digital loop sensor delay parameter has not reached its minimum value, the firmware 226A of core 220A is configured to decrease the value of one or more digital loop sensor delay parameters by a predetermined amount, thereby decreasing the voltage set point of the digital loop sensor trip point. In one or more embodiments, the sensor trip point of the digital loop sensor trip point can be decremented by the same step / unit by which the DDS sensor point can be incremented. The firmware 226A can have a preset step / decrement for moving the digital loop sensor trip point before stopping at the minimum digital loop sensor trip point. For example, until the amount of maximum voltage change for the digital loop sensor trip point is reached, the firmware 226A can move the digital loop sensor trip point to a voltage level of "J" mV, where J can be 5 mV, 10 mV, etc. as a step. In one or more embodiments, the firmware 226A can decrement the voltage set point of the digital loop sensor trip point and the entire amount of the maximum voltage change at once. In one or more embodiments, the maximum voltage change can be 10 millivolts (mV), 15 mV, 20 mV, 25 mV, 30 mV, etc.
[0065] In one or more embodiments, the firmware 226 of the core 220 is configured to operate a digital loop sensor delay test on the digital loop sensor 222, either as a manufacturing test or in the field, identify one or more digital loop sensor delay parameter sets, evaluate one or more digital loop sensor delay parameter sets, and load a preferred digital loop sensor delay parameter set. Evaluating one or more digital loop sensor delay parameter sets further includes the firmware 226 of the core 220 comparing one or more digital loop sensor delay parameter sets, identifying a particular digital loop sensor delay parameter, and selecting the digital loop sensor delay parameter set having the lowest particular digital loop sensor delay parameter as the preferred digital loop sensor delay parameter set. The firmware 226 of the core 220 is configured to detect a first number of core recovery events and a second number of core recovery events using timing checks. The digital loop sensor delay parameters of the processor are selected from the group including yield and power.
[0066] FIG. 5 is a flowchart of a process 500 using a dynamic guard band with performance protection for the processor 204, according to one or more embodiments. Reference may be made to any of the figures discussed herein. In FIGS. 5 and 6, an exemplary scenario using the core 220A may be continued for ease of understanding and consistency. Here also, it should be understood that all of the cores 220A - 220N of the processor 204 may simultaneously execute a dynamic guard band with performance protection and a dynamic guard band with timing protection as discussed herein. Similarly, a dynamic guard band with performance protection and a dynamic guard band with timing protection may be simultaneously implemented for the processor 204 in the computer system 202.
[0067] In block 502 of process 500, the firmware 226A of core 220A is configured to monitor the operation of core 220. The firmware 226A can communicate with the digital loop sensor 222A to monitor the voltage droop (or decrease) in core 220A while core 220A processes the workload. The firmware 226A can communicate with the throttle meter 224A to monitor the performance degradation of core 220A while the core processes the workload. In one or more embodiments, droop reduction can be achieved using frequency reduction instead of instruction slot throttling. In this case, the performance loss is caused by core frequency reduction instead of instruction slot throttling. In this embodiment, the droop event results in a frequency reduction instead of instruction slot throttling. In this embodiment, instead of monitoring instruction slot throttling to determine performance loss, the firmware monitors frequency reduction to determine performance loss. In one or more embodiments, both instruction slot throttling and frequency reduction can be used simultaneously to reduce droop. Therefore, the firmware monitors both instruction slot throttling and frequency reduction to determine performance loss.
[0068] In block 504, the firmware 226A of core 220A is configured to detect a first amount of slot throttling within a predetermined time to measure slot throttling. The throttle meter 224A measures the number of cycles in which slot throttling measurement is active at a predetermined time. In this case, the first slot throttling threshold can range from 1 to millions. In one or more embodiments, the slot throttling meter measures the number of slot throttling amounts at a predetermined time.
[0069] The throttle meter 224 is a circuit (which may include and / or be coupled to a counter) that provides an indication of how many suspend execution cycles have been asserted for each core within the processor. The reading of each throttle meter corresponds to a (single) suspend execution cycle. This number of suspend execution cycles increases or decreases in response to a performance degradation. Consequently, the level of performance degradation of core 220 directly corresponds to a predetermined number of suspend execution cycles that core 220 is receiving.
[0070] In one or more embodiments, the predetermined time for checking the throttling amount can range from a few microseconds to several minutes or hours. In one or more embodiments, the predetermined time for checking the throttling amount can range from about 1 minute to about 1 hour. In one or more embodiments, since a throttling amount that is less or zero is detected from a previous check, the predetermined time for checking the throttling amount can shift from a smaller number such as 1 minute to a larger number such as 5 minutes.
[0071] In block 506, the firmware 226A of core 220A is configured to check whether the first amount of throttling within a predetermined time for checking throttling has an associated performance degradation that is greater than a first performance degradation threshold. The first performance degradation threshold can be set to a 1% performance degradation of core 220A. In one or more embodiments, the first performance degradation threshold can range from about 0.1% to about 3% performance degradation of core 220A. The performance degradation in the case of the first amount of throttling can correspond to a specific number of suspend execution cycles, and the firmware 226A can convert the number of suspend execution cycles of core 220A to a performance degradation rate using, for example, a table in the firmware or elsewhere. The firmware 226A is configured to check whether the performance degradation (rate) or the number of suspend execution cycles in the case of the first amount of throttling is greater than the first performance degradation threshold, for example, greater than a 1% performance degradation.
[0072] In one or more embodiments, the first performance degradation threshold may correspond to the hibernation execution cycle threshold. In one or more embodiments, the firmware 226A may check whether a first amount of throttling having several hibernation execution cycles (in one example there may be no hibernation execution cycles) is greater than a first performance degradation threshold which is the number of hibernation execution cycles as a threshold.
[0073] In block 514, if the first amount of throttling within a predetermined time for checking the throttling amount has an associated performance degradation that is not greater than the first performance degradation threshold, the firmware 226A of the core 220A is configured to keep the supply voltage at its current voltage level (i.e., its current voltage setting). For illustrative purposes, FIG. 8A illustrates a graph of the voltage of the core 220 executing a workload over time. As can be seen, the workload has its apex or peak at the original / current supply voltage before the voltage droop to the core recovery (zone). Assume that no change is made to the supply voltage supplied to the core 220A based on blocks 506, 514.
[0074] In block 508, if the first amount of throttling within a predetermined time for checking the throttling amount has an associated performance degradation that is greater than the first performance degradation threshold, the firmware 226A of the core 220A is configured to check whether the power usage conditions are met to increase the supply voltage.
[0075] Core 220A is on processor (chip) 204. There may be multiple processor chips 204 within a drawer (or computer system 202). The drawer is interconnected with other drawers using known methods as understood by those skilled in the art. The power supply usage (PSU) condition is that the PSU of the drawer containing Core 220A is less than the power supply usage threshold (i.e., the PSU of the drawer < PSU threshold). In one or more embodiments, the PSU threshold can be in the range of about 3000 watts (W) to about 3900 W. If the power supply usage condition is not met (i.e., the PSU of the drawer > PSU threshold), the flow proceeds to block 514 without increasing the supply voltage of Core 220A.
[0076] As an additional check for the power supply usage condition that can be optionally added to block 508, the power supply usage may also include verifying that the PSU of the drawer (computer system 202) containing Core 220A is not greater than the maximum power supply usage of the drawer. If not, this part of the condition is satisfied or fulfilled to increase the supply voltage of Core 220A. On the other hand, if the PSU of the drawer (computer system 202) containing Core 220A is greater than the maximum power supply usage of that drawer, the firmware 226A of Core 220A is configured to cause the firmware 236 of the power supply circuit 230 to revert to the default settings for the supply power to all cores 220A - 220N in the processor 204.
[0077] Furthermore, it should be noted that the firmware 236 can be integrated with the controller 232 to control the supply voltage supplied to each of the cores 220A - 220N on the processor 204. In one or more embodiments, the firmware 226 operably communicates with the firmware 236 of the power supply circuit 240 to provide and change the supply voltage provided to the cores 220A - 220N.
[0078] In block 510, when the power supply usage conditions for increasing the supply voltage are met (i.e., the PSU of the drawer < PSU threshold) (and optionally including that the PSU of the drawer including core 220A is not greater than the maximum power supply usage of that drawer), the firmware 226A of core 220A is configured to check whether the supply voltage supplied to core 220A has increased to the maximum supply voltage of core 220A. As described herein, as shown in FIGS. 8A and 8B, there is an acceptable maximum supply voltage change. The firmware 226A of core 220A checks whether the supply voltage has already increased to the permitted maximum supply voltage. (Yes) If the supply voltage supplied to core 220A has increased to the maximum supply voltage of core 220A, the flow proceeds to block 514.
[0079] In block 512, if the supply voltage supplied to core 220A has not increased to the maximum supply voltage of core 220A, the firmware 226A of core 220A is configured to increase the supply voltage to core 220A by a predetermined amount / step. FIG. 8B illustrates a graph of the voltage over time of core 220 executing a workload, where the supply voltage is increasing. In FIG. 8B, when droop reduction is not required, by increasing the supply voltage vertically, the entire graph of the workload shifts up by the amount of increase in the increased voltage set point. In FIG. 8B, the dashed curve illustrates the old position of the graph of the workload, while the solid curve is the new position of the graph of the workload, indicating that the peak is now at the new supply voltage (new VDD). In FIG. 8B, the digital droop sensor trip point remains at the voltage set point. Therefore, the voltage in the DDS sensor crosses the DDS trip point at a later time, and droop reduction stops the droop at approximately the same minimum voltage. As described herein, by increasing the supply voltage of that core 220 and assuming that the digital droop sensor trip point just remains at a fixed voltage level, as illustrated in FIG. 8B, the distance between the new supply voltage for that core 220 and the digital droop sensor trip point is increasing. This means that a workload that previously crossed the digital droop sensor trip point (and thus had a performance hit) will now, with the new supply voltage setting, no longer cross the digital droop sensor trip point, and thus no longer exhibit a performance hit / drop. Generally, the performance loss will decrease with an increase in the VDD set point.
[0080] In one or more embodiments, the maximum change to increase the supply voltage to core 220 can be in the range of about 10 mV to 30 mV. In one or more embodiments, there is a predetermined number of steps to increase the supply voltage of core 220. In one or more embodiments, for a maximum change of 20 mV, the predetermined amount / step to increase the supply voltage to core 220 can be 5 mV each for 4 steps. In one or more embodiments, processor 204 can collectively have a total maximum change of about 20 mV (for all cores 220). In one or more embodiments, a drawer (or computer system 202) having a plurality of processors 204 can have a maximum number of predetermined amounts / steps. For example, if there are 20 predetermined steps of 5 mV each for the drawer, the drawer can increase by a total of 100 mV.
[0081] FIG. 6 is a flowchart of a process 600 using a dynamic guard band with performance protection for a processor 204, according to one or more embodiments. FIG. 6 can continue the process discussed in FIG. 5. Any of the figures discussed herein may be referenced. In block 602 of process 600, the firmware 226A of core 220A is configured to monitor the operation of core 220A. The firmware 226A can communicate with the digital loop sensor 222A to monitor the voltage droop (or decrease) in core 220A while core 220A processes a workload. The firmware 226A can communicate with the throttle meter 224A to monitor the performance degradation of core 220A while the core processes a workload.
[0082] In block 604, the firmware 226A of core 220A is configured to detect a second amount of throttling within a predetermined time for measuring throttling. The second amount of throttling is the same as or less than the first amount of throttling measured by throttle meter 224A. The second amount of throttling may be 1, 2, 3, 4, 5, 6 to 10 times less than the first amount of throttling. As described herein, the predetermined time for checking the amount of throttling may be about every 2 minutes. In one or more embodiments, the predetermined time for checking the amount of throttling may range from about (several) microseconds to 24 hours. In one or more embodiments, since a lesser or zero amount of throttling is detected from a previous check, the predetermined time for checking the amount of throttling may shift from a lesser number such as 2 minutes to a greater number such as 5 minutes.
[0083] In block 606, the firmware 226A of core 220A is configured to check whether the second amount of throttling within a predetermined time for checking the amount of throttling has an associated performance degradation that is less than a second performance degradation threshold. The second performance degradation threshold may be set to a 0.1% performance degradation of core 220A. In one or more embodiments, the second performance degradation threshold may range from about 0.00% to about 1.0% performance degradation of core 220A. Similar to the above discussion, the performance degradation in the case of the second amount of throttling may identify a specific number of halted execution cycles, and the firmware 226A may convert the number of halted execution cycles of core 220A to a performance degradation rate using, for example, a table in the firmware or elsewhere. The firmware 226A is configured to check whether the performance degradation (rate) or the number of halted execution cycles in the case of the second amount of throttling is less than the second performance degradation threshold, for example, less than 0.1% performance degradation.
[0084] In one or more embodiments, the second performance degradation threshold may correspond to the hibernation execution cycle threshold. In one or more embodiments, the firmware 226A may check that a second amount of throttling having several hibernation execution cycles (in one example there may be no hibernation execution cycles) is less than a second performance degradation threshold which is the number of hibernation execution cycles (as a threshold).
[0085] In block 608, if the second amount of throttling within a predetermined time for checking the throttling amount has an associated performance degradation that is less than the second performance degradation threshold, the firmware 226A of the core 220A is configured to reduce the supply voltage. Similar to the maximum change for increasing the supply voltage to the core 220 and the predetermined amount / step for each increase discussed in FIG. 5, the same may apply to the decrease of the supply voltage to the core. In one or more embodiments, the maximum change for reducing the supply voltage to the core 220 may be in the range of about 10 mV to 30 mV. In one or more embodiments, there is a predetermined number of predetermined amounts / steps for reducing the supply voltage of the core 220. In one or more embodiments, for a maximum change of 20 mV for reducing the supply voltage to the core 220, the predetermined amount / step may be 5 mV each for 4 steps. In one or more embodiments, the processor 204 may collectively have a total maximum change of about 20 mV (for all cores 220).
[0086] In one or more embodiments, the supply voltage may be reduced to the default supply voltage setting. In one example, the default supply voltage setting may be the original supply voltage setting illustrated in FIG. 8B.
[0087] In block 610, if the second amount of throttling within a predetermined time for checking the throttling amount has an associated performance degradation that is greater than the second performance degradation threshold, the firmware 226A of the core 220A is configured to maintain the supply voltage at the current setting.
[0088] Figure 9 is a flowchart of a computer-implemented method 900 of a dynamic guard band with timing protection for a processor core 220 of a processor 204 according to one or more embodiments. Any of the figures discussed herein may be referenced. In block 902, firmware 226 of core 220 (of computer system 202) is configured to detect a first number of core recovery events in processor 204 in response to monitoring processor 204 during operation. The first number of core recovery events is that of core 220, such as core 220A. In one or more embodiments, the first number (and the following second number) may collectively be that of cores 220A - 220N.
[0089] In block 904, firmware 226 of core 220 (of computer system 202) is configured to determine that the first number of core recovery events satisfies a first condition for a first core recovery event threshold. For example, as discussed in block 308 of FIG. 3, a predetermined number of core recovery events occur, thereby satisfying the first condition for the first core recovery event threshold.
[0090] In block 906, the firmware 226 of the core 220 (of the computer system 202) is configured to modify the value of at least one digital loop sensor delay parameter of the processor 204 (of the digital loop sensor 222) by a first amount, and the at least one digital loop sensor delay parameter affects the execution of one or more instructions on the processor 204. For example, reference may be made to the discussion of block 310 in FIG. 3. As a technical solution / benefit, the firmware 226 is configured to adjust the loop threshold (e.g., by increasing the delay in the digital loop sensor 222), thereby correspondingly increasing the voltage droop and the sensitivity to low voltage, so that the probability or rate of a recovery event is reduced. The digital loop sensor 222 has delay adjustment to adjust the loop threshold, but one or more embodiments may utilize other types of loop sensors having a voltage adjustment knob (or parameter), or any other method of adjusting the loop threshold. For example, there may be an analog loop sensor that triggers a loop response for loop reduction. Therefore, the digital loop sensor parameter may be implemented as a voltage adjustment (increase / decrease) that increases and decreases in much the same way as the delay in the digital loop sensor 222 according to one or more embodiments.
[0091] In block 908, the firmware 226 of the core 220 (of the computer system 202) is configured to detect a second number of core recovery events in the processor 204 in response to modifying the value of at least one digital loop sensor delay parameter of the digital loop sensor 222 by a first amount.
[0092] In block 910, the firmware 226 of the core 220 (of the computer system 202) is configured to determine that the second number of core recovery events satisfies a second condition for a second core recovery event threshold. For example, as discussed in block 406 of FIG. 4, the second condition for the second core recovery event threshold (e.g., the decreased latency core recovery event threshold) is satisfied.
[0093] In block 912, the firmware 226 of the core 220 (of the computer system 202) is configured to modify the value of at least one digital loop sensor delay parameter of the processor by a second amount. For example, reference may be made to block 410 of FIG. 4.
[0094] Modifying the value of at least one digital loop sensor delay parameter of the processor 204 by a first amount includes increasing the value of at least one digital loop sensor delay parameter of the processor by the first amount. For example, the firmware 226A may instruct the digital loop sensor 222A to increase the value of at least one digital loop sensor delay parameter. Modifying the value of at least one digital loop sensor delay parameter of the processor 204 by a second amount includes decreasing the value of at least one digital loop sensor delay parameter of the processor by the second amount. For example, the firmware 226A may instruct the digital loop sensor 222A to decrease the value of at least one digital loop sensor delay parameter.
[0095] The k that satisfies the first condition for the first core recovery event threshold includes that the first number of core recovery events is greater than the first core recovery event threshold. For example, the firmware 226 may determine that a predetermined number of core recovery events is greater than the first core recovery event threshold as discussed in block 308. Satisfying the second condition for the second core recovery event threshold includes that the second number of core recovery events is less than the second core recovery event threshold. For example, the firmware 226 may determine that the number of core recovery events at a predetermined time has decreased below the reduced delay core recovery event threshold.
[0096] Modifying the value of at least one digital loop sensor delay parameter of the processor 204 by a second amount includes returning to a baseline value for the at least one digital loop sensor delay parameter. The firmware 226 may instruct the digital loop sensor 222 to return to the baseline value for the at least one digital loop sensor delay parameter(s). The baseline value may be a delay of 0. The baseline value may be 100 delay elements. The second amount by which the value of the at least one digital loop sensor delay parameter decreases may be greater than the first amount, equal to the first amount, or less than the first amount. The control loop discussed herein may continue indefinitely during operation of the processor. Also, although the recovery event may be for a core, it should be understood that the recovery event may be triggered by a recovery event in a circuit smaller or larger than the "core".
[0097] FIG. 10 is a flowchart of a computer-implemented method 1000 of a dynamic guard band with performance protection for a processor core 220 of a processor 204 according to one or more embodiments. Any of the figures discussed herein may be referenced.
[0098] In block 1002, the firmware 226 of the core 220 (of the computer system 202) is configured to detect a first amount of throttling in the processor 204 in response to monitoring the processor 204 during operation. The first amount of throttling is a predetermined number of throttle meter readings of the throttle meter 224. An example is discussed in block 504 of FIG. 5.
[0099] In block 1004, the firmware 226 of the core 220 (of the computer system 202) is configured to determine that the first amount of throttling satisfies a first condition regarding a throttling amount threshold. The first amount of throttling has an associated performance degradation. For example, as discussed in block 506 of FIG. 5, the first condition is satisfied because the associated performance degradation for the first amount of throttling is greater than a first performance degradation threshold. The first performance degradation threshold can be set to a performance degradation of 1%, 2%, 3%, etc. of the core 220. The performance degradation of the first amount of throttling can correspond to a specific number of halt execution cycles.
[0100] In block 1006, the firmware 226 of the core 220 (of the computer system 202) is configured to modify the voltage level of the processor 204 by the first amount. The firmware 226 can command and / or communicate with the firmware 236 of the power supply circuit 230 to modify the voltage level. An example of the modification of the voltage level is illustrated in block 512 of FIG. 5.
[0101] In block 1008, the firmware 226 of the core 220 (of the computer system 202) is configured to detect a second amount of throttling in the processor 204 in response to modifying the voltage level of the processor 204 by the first amount. The second amount of throttling is a predetermined threshold of the throttle meter 224 following the modification (e.g., increase) of the voltage level.
[0102] In block 1010, the firmware 226 of the core 220 (of computer system 202) is configured to determine that a second amount of throttling satisfies a second condition regarding a throttling amount threshold.
[0103] The second amount of throttling has an associated performance degradation. For example, as discussed in block 606 of FIG. 6, the second condition is satisfied because the associated performance degradation for the second amount of throttling is less than a second performance degradation threshold. As an example, the second performance degradation threshold can be set to a 0.1% performance degradation of core 220, or the second performance degradation threshold can be any number in the range of from about 0.0% to about 1.0% performance degradation of core 220.
[0104] In block 1012, the firmware 226 of the core 220 (of computer system 202) is configured to modify the voltage level of the processor 204 by a second amount.
[0105] The firmware 226 may instruct and / or communicate with the firmware 236 of the power supply circuit 230 to modify the voltage level. An example of modifying the voltage level is illustrated in block 608 of FIG. 6.
[0106] Modifying the voltage level of the processor 204 by a first amount includes, for example, increasing the voltage level of the processor by the first amount, as illustrated in block 512 of FIG. 5. Modifying the voltage level of the processor 204 by a second amount includes, for example, decreasing the voltage level of the processor by the second amount, as illustrated in block 608 of FIG. 6.
[0107] The firmware 226 of the core 220 is configured to confirm that the power consumption is less than the power consumption threshold before modifying the voltage level of the processor by a first amount. The firmware 226 of the core 220 checks whether the power supply usage (PSU) is greater than the power supply usage threshold (e.g., block 508), rejects the request to modify the voltage level by a first amount (e.g., by the firmware 236 and / or the firmware 226), and is configured to remain at the current voltage level (e.g., block 514) in response to the power supply usage being greater than the power supply usage threshold. The firmware 226 of the core 220 is configured to provide conditions for returning the voltage level to the default voltage level in response to determining that the power supply usage is greater than the maximum power supply usage threshold. The first amount for modifying the voltage level of the processor 204 ranges from about 5 millivolts (mV) to 10 mV.
[0108] The present disclosure includes a detailed description regarding cloud computing, but it should be understood that the implementations of the teachings recited herein are not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment, whether currently known or developed in the future.
[0109] Cloud computing is a service - delivery model that enables convenient on - demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0110] The characteristics are as follows.
[0111] On-demand self-service: Cloud consumers can automatically and on an as-needed basis provision computing capabilities such as server time and network storage unilaterally, without the need for human interaction with a service provider.
[0112] Broad network access: Capabilities are available over a network and accessed through standard mechanisms that promote use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs (trademark)).
[0113] Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with various physical and virtual resources dynamically assigned and reassigned according to demand. Consumers generally have no control or knowledge over the exact location of the provided resources, but may be able to specify location at a higher level of abstraction (e.g., country, state, or data center), which provides a certain degree of location independence.
[0114] Rapid elasticity: Capabilities can be provisioned quickly and elastically, and in some cases automatically, scaling out rapidly and also being released and scaled in quickly. In many cases, to the consumer, the capabilities available for provisioning appear to be unlimited and can be purchased in any quantity at any time.
[0115] Measured service: The cloud system automatically controls and optimizes resource use by leveraging measurement capabilities appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts) at some level of abstraction. Resource utilization is monitored, controlled, and reported, providing transparency to both the provider and consumer of the utilized services.
[0116] The service model is as follows.
[0117] Software as a Service (SaaS): The ability provided to the consumer is to use the provider's applications running on the cloud infrastructure. The applications are accessible from various client devices via a client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, which includes the network, servers, operating systems, storage, or even the individual application capabilities, except for limited user-specific application configuration settings in some cases.
[0118] Platform as a Service (PaaS): The ability provided to the consumer is to deploy the applications created or acquired by the consumer, which are created using the programming languages and tools supported by the provider, onto the cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, which includes the network, servers, operating systems, or storage, but controls the deployed applications and, in some cases, the configuration of the application hosting environment.
[0119] Infrastructure as a Service (IaaS): The ability provided to the consumer is to provision processing, storage, network, and other basic computing resources, and the consumer can deploy and run any software that may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but controls the operating systems, storage, deployed applications, and, in some cases, limited control over selected networking components (e.g., host firewall).
[0120] The deployment model is as follows.
[0121] Private cloud: The cloud infrastructure is operated solely for an organization. It may be managed by the organization or a third party and may exist on-premises or off-premises.
[0122] Community cloud: The cloud infrastructure is shared by multiple organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies, and regulatory compliance considerations). The community cloud may be managed by those organizations or a third party and may exist on-premises or off-premises.
[0123] Public cloud: The cloud infrastructure is made available to the general public or a large industry group and is owned by an organization that sells cloud services.
[0124] Hybrid cloud: This cloud infrastructure is a composite of two or more clouds (private, community, or public), which remains a unique entity but is joined together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load distribution between clouds).
[0125] The cloud computing environment is service-oriented, emphasizing statelessness, loose coupling, modularity, and semantic interoperability. At the core of cloud computing, there is an infrastructure that includes a network of interconnected nodes.
[0126] Referring now to FIG. 12, an exemplary cloud computing environment 50 is illustrated. As shown, cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers (e.g., mobile information terminal (PDA) or cellular phone 54A, desktop computer 54B, laptop computer 54C and / or in-vehicle computer system 54N, etc.) can communicate. Nodes 10 may communicate with each other. They may be physically or virtually grouped in one or more networks (such as private cloud, community cloud, public cloud and / or hybrid cloud as described above, or combinations thereof, etc.) (not shown). Thereby, cloud computing environment 50 can provide infrastructure, platform and / or software as services for which cloud consumers do not need to maintain resources on local computing devices. The types of computing devices 54A - N shown in FIG. 12 are merely intended to be exemplary, and it is understood that cloud computing nodes 10 and cloud computing environment 50 can communicate with any type of computerized device via any type of network and / or network addressable connection (e.g., using a web browser).
[0127] Referring now to FIG. 13, a set of functional abstraction layers provided by cloud computing environment 50 (FIG. 12) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 13 are merely intended to be exemplary, and embodiments of the invention are not limited thereto. As shown, the following layers and corresponding functions are provided.
[0128] The hardware and software layer 60 comprises hardware and software components. Examples of hardware components include mainframe 61; RISC (Reduced Instruction Set Computer) architecture-based server 62; server 63; blade server 64; storage device 65; and network and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0129] The virtualization layer 70 provides an abstraction layer that can provide the following examples of virtual entities: virtual server 71; virtual storage 72; virtual network 73 including a virtual private network; virtual applications and operating systems 74; and virtual client 75.
[0130] In one example, the management layer 80 may provide the functions described below. Resource provisioning 81 provides for the dynamic procurement of computing resources and other resources utilized to perform tasks within a cloud computing environment. Metering and pricing 82 provides for cost tracking when resources are utilized within a cloud computing environment and accounting or billing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, and protection of data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides for cloud computing resource allocation and management such that the required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides for the upfront commitment and procurement of cloud computing resources where future requirements are anticipated to conform to the SLA.
[0131] The workload layer 90 provides examples of functions that can be utilized in a cloud computing environment. Examples of workloads and functions that can be provided from this layer include mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analysis processing 94; transaction processing 95; and workloads and functions 96.
[0132] Various embodiments of the present invention are described herein with reference to the associated drawings. Alternative embodiments can be devised without departing from the scope of the present invention. In the following description and drawings, various connection and positional relationships between elements (e.g., above, below, adjacent, etc.) are described. However, those skilled in the art will understand that many of the positional relationships described herein are independent of orientation, and the functionality described will be maintained even if the orientation changes. These connections and / or positional relationships can be direct or indirect, unless otherwise specified, and the present invention is not intended to be limited in this regard. Thus, the coupling between entities can refer to a direct or indirect coupling, and the positional relationship between entities can be a direct or indirect positional relationship. As an example of an indirect positional relationship, when it is said in this description that layer "A" is formed on layer "B", one or more intermediate layers (e.g., layer "C") are included in the situation existing between layer "A" and layer "B", provided that the relevant characteristics and functionality of layer "A" and layer "B" are not substantially changed by that intermediate layer.
[0133] For the sake of brevity, the prior art related to creating and using aspects of the present invention may or may not be described in detail herein. In particular, various aspects of computing systems and specific computer programs for implementing the various technical features described herein are well known. Thus, for the purpose of brevity, many details of conventional implementations are only briefly mentioned herein or are completely omitted without providing details of well-known systems and / or processes.
[0134] In some embodiments, various functions or operations may be performed at a given location and / or in connection with the operation of one or more devices or systems. In some embodiments, a portion of a given function or operation may be performed at a first device or location, and the remaining function or operation may be performed at one or more additional devices or locations.
[0135] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises" and / or "comprising", when used herein, specify the presence of the stated features, integers, steps, operations, and / or element components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, element components, and / or groups thereof.
[0136] Corresponding structures, materials, acts, and equivalents of any means-plus-function elements or step-plus-function elements in the following claims are intended to include any structure, material, or act for performing the recited function in combination with other claimed elements as specifically claimed. Although the present disclosure has been presented for purposes of illustration and description, it is not intended to be exhaustive or limited to the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the present disclosure. Embodiments were chosen and described in order to best explain the principles of the present disclosure and the practical application, and to enable others of ordinary skill in the art to understand the present disclosure for various embodiments with various modifications as are suited to the particular use contemplated.
[0137] The figures illustrated in this specification are exemplary. Without departing from the spirit of the present disclosure, many changes can be made to the figures or the steps (or operations) described therein. For example, actions can be performed in a different order, or actions can be added, deleted, or modified. Also, the term "coupled" as used herein represents that there is a signal path between two elements and does not imply a direct connection between elements without intervening elements / connections between them. All of these changes are considered to be part of the present disclosure.
[0138] The following definitions and abbreviations are used in the interpretation of the claims and the specification. As used herein, the terms "comprises", "comprising", "includes", "including", "has", "having", "contains", or "containing", or any other variation thereof, are intended to cover non-exclusive inclusion. For example, a composition, mixture, process, method, article, or apparatus that includes a list of elements is not necessarily limited to only those elements, but may include other elements not expressly listed, or other elements inherent to such composition, mixture, process, method, article, or apparatus.
[0139] Furthermore, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration". Any embodiment or design described as "exemplary" herein is not necessarily to be construed as more preferred or advantageous than other embodiments or designs. The terms "at least one" and "one or more" are understood to include any integer greater than or equal to 1, i.e., 1, 2, 3, 4, etc. The term "plurality" is understood to include any integer greater than or equal to 2, i.e., 2, 3, 4, 5, etc. The term "connected" can include both indirect "connection" and direct "connection".
[0140] The terms "about," "substantially," "approximately," and variations thereof are intended to include a degree of error associated with the measurement of a particular quantity based on the equipment available at the time of filing. For example, "about" can include a range of ±8%, or 5%, or 2% of a given value.
[0141] The present invention can be a system, method, and / or computer program product integrated at any possible technical detail level. The computer program product can include one (or more) computer-readable storage media having computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0142] The computer-readable storage media can be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage media can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage media includes the following: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disks (DVDs), memory sticks, floppy disks, punch cards, mechanically encoded devices such as raised structures within grooves in which instructions are recorded, and any suitable combination of the foregoing. The computer-readable storage media should not be construed herein as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through an electrical wire.
[0143] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or an external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium in each respective computing / processing device.
[0144] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including, for example, assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or any combination of the foregoing, such as the programming languages Smalltalk®, C++, or similar object-oriented programming languages, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuit in order to implement aspects of the present invention.
[0145] Aspects of the invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0146] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium containing the instructions comprises a manufacture including instructions for implementing the aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0147] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to create a computer-implemented process, and thus the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0148] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions that include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may be performed in an order different from that noted in the drawings. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order depending on the related functionality. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0149] The description of the various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein were chosen in order to best explain the principles of the embodiments, the practical application, or technical improvements made to the technology found in the marketplace, or to enable other ordinary skill in the art to understand the embodiments described herein.
Claims
1. In response to monitoring the processor during operation, detecting, by a computer, a first number of core recovery events of the processor; Determining, by the computer, that the first number of core recovery events satisfies a first condition for a first core recovery event threshold; Modifying, by the computer, a value of at least one loop sensor parameter of the processor by a first amount, the at least one loop sensor parameter affecting sensitivity to a voltage droop; In response to modifying the value of the at least one loop sensor parameter by the first amount, detecting, by the computer, a second number of core recovery events in the processor; Determining, by the computer, that the second number of core recovery events satisfies a second condition for a second core recovery event threshold; and Modifying, by the computer, the value of the at least one loop sensor parameter of the processor by a second amount A computer-implemented method comprising.
2. The step of modifying the value of the at least one loop sensor parameter of the processor by the first amount includes increasing the value of the at least one loop sensor parameter of the processor by the first amount, the computer-implemented method according to claim 1.
3. The step of modifying the value of the at least one loop sensor parameter of the processor by the second amount includes decreasing the value of the at least one loop sensor parameter of the processor by the second amount, the computer-implemented method according to any of the preceding claims.
4. Satisfying the first condition for the first core recovery event threshold includes that the first number of core recovery events is greater than the first core recovery event threshold, the computer-implemented method according to any of the preceding claims.
5. Satisfying the second condition for the second core recovery event threshold includes that the second number of core recovery events is less than the second core recovery event threshold, the computer-implemented method according to any of the preceding claims.
6. The step of modifying the value of the at least one loop sensor parameter of the processor by the second amount has a step of returning to a baseline value for the at least one loop sensor parameter, the computer-implemented method according to any of the preceding claims.
7. The computer-implemented method according to any of the preceding claims, wherein the second amount is greater than, equal to, or less than the first amount.
8. A memory having computer-readable instructions; and A computer for executing the computer-readable instructions, the computer-readable instructions being: In response to monitoring the processor during operation, detecting, by the computer, a first number of core recovery events of the processor; Determining, by the computer, that the first number of core recovery events satisfies a first condition for a first core recovery event threshold; Modifying, by the computer, a value of at least one loop sensor parameter of the processor by a first amount, the at least one loop sensor parameter affecting sensitivity to voltage droop; In response to modifying the value of the at least one loop sensor parameter by the first amount, detecting, by the computer, a second number of core recovery events in the processor; Determining, by the computer, that the second number of core recovery events satisfies a second condition for a second core recovery event threshold; and Modifying, by the computer, the value of the at least one loop sensor parameter of the processor by a second amount Controlling the computer to perform an operation including A system comprising.
9. The system according to claim 8, wherein the step of modifying the value of the at least one loop sensor parameter of the processor by the first amount has a step of increasing the value of the at least one loop sensor parameter of the processor by the first amount.
10. The system according to any of the preceding claims 8 or 9, wherein the step of modifying the value of the at least one loop sensor parameter of the processor by the second amount has a step of decreasing the value of the at least one loop sensor parameter of the processor by the second amount.
11. The system according to any one of the preceding claims 8 to 10, wherein satisfying the first condition for the first core recovery event threshold includes that the first number of core recovery events is greater than the first core recovery event threshold.
12. The system according to any one of the preceding claims 8 to 11, wherein satisfying the second condition for the second core recovery event threshold includes that the second number of core recovery events is less than the second core recovery event threshold.
13. The system according to any one of the preceding claims 8 to 12, wherein the step of modifying the value of the at least one loop sensor parameter of the processor by the second amount includes returning to a baseline value for the at least one loop sensor parameter.
14. The system according to any one of the preceding claims 8 to 14, wherein the second amount is greater than, equal to, or less than the first amount.
15. A computer program product comprising a computer-readable storage medium having program instructions embodied thereon, the program instructions: detecting, by the computer, a first number of core recovery events of the processor in response to monitoring the processor during operation; determining, by the computer, that the first number of core recovery events satisfies a first condition for a first core recovery event threshold; modifying, by the computer, the value of at least one loop sensor parameter of the processor by a first amount, the at least one loop sensor parameter affecting sensitivity to voltage droop; detecting, by the computer, a second number of core recovery events in the processor in response to modifying the value of the at least one loop sensor parameter by the first amount; determining, by the computer, that the second number of core recovery events satisfies a second condition for a second core recovery event threshold; and modifying, by the computer, the value of the at least one loop sensor parameter of the processor by a second amount such that the operations are executable by the computer to cause the computer to perform the operations Computer program product.
16. The step of modifying the value of the at least one loop sensor parameter of the processor by the first amount includes the step of increasing the value of the at least one loop sensor parameter of the processor by the first amount, the computer program product according to claim 15.
17. The step of modifying the value of the at least one loop sensor parameter of the processor by the second amount includes the step of decreasing the value of the at least one loop sensor parameter of the processor by the second amount, the computer program product according to any one of the preceding claims 15 to 16.
18. Satisfying the first condition for the first core recovery event threshold includes that the first number of core recovery events is greater than the first core recovery event threshold, the computer program product according to any one of the preceding claims 15 to 17.
19. Satisfying the second condition for the second core recovery event threshold includes that the second number of core recovery events is less than the second core recovery event threshold, the computer program product according to any one of the preceding claims 15 to 18.
20. The step of modifying the value of the at least one loop sensor parameter of the processor by the second amount includes the step of returning to the baseline value for the at least one loop sensor parameter, the computer program product according to any one of the preceding claims 15 to 19.