Enhancing Central Processing Unit Service Quality Assurance in Serving Accelerator Requests

By monitoring the CPU overhead of accelerator system service requests on the CPU and dynamically adjusting the processing delay, the problem of interference between accelerator system service requests on the performance and energy efficiency of the CPU application is solved, and more efficient system performance and resource utilization is achieved.

CN112041822BActive Publication Date: 2025-05-30ADVANCED MICRO DEVICES INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201980029389.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-04-16
Filing Date
2019-02-14
Publication Date
2025-05-30
Estimated Expiration
2039-02-14

AI Technical Summary

Technical Problem

In modern systems on chip, accelerator system service requests significantly interfere with the performance and energy efficiency of CPU applications, resulting in performance degradation.

Method used

By monitoring and tracking the CPU overhead of accelerator system service requests on the CPU, the processing delays of kernel worker threads to reduce interference to CPU applications.

Benefits of technology

It effectively reduces the performance and energy efficiency degradation of the accelerator system service request on CPU applications, and improves the overall performance and resource utilization efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112041822B_ABST
    Figure CN112041822B_ABST
Patent Text Reader

Abstract

Systems, devices, and methods are disclosed for enhancing processor quality of service assurance when servicing system service requests (SSRs). A system includes a first processor that executes an operating system and a second processor that executes an application that generates an SSR for the first processor to service. The first processor monitors the number of cycles spent servicing an SSR during a previous time interval, and if the number of cycles is greater than a threshold, the first processor begins to delay servicing subsequent SSRs. In one implementation, if the previous delay was non-zero, the first processor increases the delay used in servicing subsequent SSRs. If the number of cycles is less than or equal to the threshold, the first processor services the SSR without delay. As the delay increases, the second processor begins to stall and its SSR generation rate decreases, thereby reducing the load on the first processor.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This invention was made with government support under the PathForward Project with Lawrence Livermore National Security awarded by the U.S. Department of Energy (prime contract No. DE-AC52-07NA27344, subcontract No. B620717).

[0002] Description of Related Art

[0003] Modern System-on-Chip (SoC) typically integrates a large number of different types of components on a single chip or on a multi-chip module. For example, a typical SoC includes a main processor (e.g., a central processing unit (CPU)) and accelerators (such as an integrated graphics processing unit (GPU) and a media engine). As these accelerators become more and more powerful, it is expected that they will directly invoke complex operating system (OS) services, such as page faults, file system access, or network access. However, the OS does not run on the accelerators. Therefore, these accelerator system service requests (SSR) need to be serviced by the OS running on the CPU. These accelerator SSRs seriously interfere with the concurrent CPU applications. Generally, due to the disruptive interference caused by servicing the SSRs from the accelerators, significant performance and energy efficiency degradation will occur in the concurrent CPU applications. Brief Description of the Drawings

[0004] The advantages of the methods and mechanisms described herein can be better understood by reference to the following description in conjunction with the accompanying drawings, in which:

[0005] Figure 1 is a block diagram of one implementation of a computing system.

[0006] Figure 2 is a block diagram of one implementation of a computing system with a CPU and an accelerator.

[0007] Figure 3 is a timing diagram showing the overhead associated with processing accelerator SSRs according to one implementation.

[0008] Figure 4 is a block diagram of another implementation of a scheme for processing accelerator SSRs.

[0009] Figure 5 is a general flowchart showing one implementation of a method for enhancing quality of service assurance on a CPU when processing requests from an accelerator.

[0010] Figure 6 is a general flowchart showing one implementation of a method for dynamically adjusting the latency added to the service of a request.

[0011] Figure 7 It is a general flowchart showing an implementation of a method for processing system service requests (SSRs). Detailed implementation

[0012] In the following description, numerous specific details are set forth to provide a thorough understanding of the methods and mechanisms presented herein. However, one of ordinary skill in the art should recognize that various implementations may be practiced without these specific details. In some instances, well-known structures, components, signals, computer program instructions, and techniques have not been shown in detail to avoid obscuring the methods described herein. It should be understood that, for simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements.

[0013] Disclosed herein are various systems, devices, methods, and computer-readable media for enhancing central processing unit (CPU) quality of service (QoS) assurance in the face of accelerator system service requests. In one implementation, the system includes at least a CPU and an accelerator. The accelerator is a graphics processing unit (GPU) or other type of processing unit. In some implementations, the system includes multiple accelerators. In one implementation, the CPU executes an operating system (OS), and the accelerator executes an application program. When the accelerator application needs the help of the OS, the accelerator application sends a system service request (SSR) to the CPU for service. For certain applications, the accelerator sends a large number of SSRs to the CPU for service. The CPU is configured to monitor the amount of time the OS spends servicing SSRs from one or more accelerators. In various implementations, the amount of time is measured in cycles. In other implementations, different time metrics are used. For ease of discussion, cycle-based tracking of time will be used herein. In one implementation, the OS routines involved in servicing SSRs track their CPU usage cycles.

[0014] In one implementation, the kernel background thread wakes up periodically to determine whether the number of CPU cycles spent servicing SSRs in a previous time interval is greater than a specified limit. In one implementation, the limit is specified by an administrator. In another implementation, the software application dynamically adjusts the value of the limit based on operating conditions. In one implementation, the kernel worker thread adds an adjustable amount of delay to the processing of newly received SSRs, where the delay is calculated based on the CPU overhead (e.g., percentage of CPU time) spent servicing SSRs in a previous time interval. For example, in one implementation, when starting to process an SSR, the kernel worker thread checks whether the percentage of CPU time spent processing the SSR is higher than a specified threshold. The kernel worker thread uses the information collected by the kernel background thread to perform this check. If the percentage of CPU time spent processing the SSR is below the specified threshold, the kernel worker thread sets the desired delay to zero and immediately continues processing the SSR. Otherwise, if the percentage of CPU time spent processing the SSR is greater than the specified threshold, the kernel worker thread sets the delay for processing the SSR based on exponential backoff. For example, in one implementation, if the desired delay was previously greater than zero, the kernel worker thread increases the delay to a value greater than the previous delay value. For example, in various implementations, the new delay value can be a multiple (e.g., 2x, 3x, etc.) of the previous value. In other implementations, the new delay value can be a larger value that is not a multiple of the previous delay value. Otherwise, if the desired delay was previously zero, the kernel worker thread sets the new delay to an initial nominal value (e.g., 10 μs). The processing of the SSR is then delayed by this amount.

[0015] As the delay increases, one or more accelerators will start to stall and the SSR rate will eventually decline. When the CPU overhead drops below the set limit, SSRs will be serviced again without any artificial delay. Additionally, in one implementation, the servicing of accelerator SSRs is automatically throttled only when it interferes with one or more CPU applications. In this implementation, if the CPU is otherwise idle, the SSRs are serviced as quickly as possible even if the CPU overhead is above the limit. In one implementation, this check is performed by querying the OS scheduler for other processes waiting in the run list.

[0016] Now refer to Figure 1, a block diagram showing one implementation of computing system 100. In one implementation, computing system 100 includes at least processors 105A - 105N, input / output (I / O) interface 120, bus 125, one or more memory controllers 130, network interface 135, and one or more memory devices 140. In other implementations, computing system 100 includes other components and / or computing system 100 is arranged differently. Processors 105A - 105N represent any number of processors included in system 100.

[0017] In one implementation, processor 105A is a general - purpose processor, such as a central processing unit (CPU). In this implementation, processor 105N is an accelerator engine. For example, in one implementation, processor 105N is a data - parallel processor with a highly parallel architecture. Data - parallel processors include graphics processing units (GPUs), digital signal processors (DSPs), field - programmable gate arrays (FPGAs), application - specific integrated circuits (ASICs), etc. In some implementations, processors 105A - 105N include multiple accelerator engines. The multiple accelerator engines are configured to send system service requests (SSRs) to processor 105A for processing. Processor 105A is configured to monitor the overhead associated with processing these SSRs. Depending on the implementation, processor 105A monitors the overhead in terms of CPU cycles, percentage of total CPU cycles, amount of time, and / or based on other metrics. If the overhead in a previous time interval exceeds a threshold, processor 105A delays the processing of SSRs and / or otherwise reduces the amount of resources used to process SSRs.

[0018] One or more memory controllers 130 represent any number and type of memory controllers that can be accessed by processors 105A - 105N and I / O devices (not shown) coupled to I / O interface 120. One or more memory controllers 130 are coupled to any number and type of one or more memory devices 140. One or more memory devices 140 represent any number and type of memory devices. For example, the types of memory in one or more memory devices 140 include dynamic random - access memory (DRAM), static random - access memory (SRAM), NAND flash memory, NOR flash memory, ferroelectric random - access memory (FeRAM), etc.

[0019] The I / O interface 120 represents any number and type of I / O interfaces (e.g., Peripheral Component Interconnect (PCI) bus, PCI Extended (PCI-X), PCI Express (PCIE) bus, Gigabit Ethernet (GBE) bus, Universal Serial Bus (USB)). Various types of peripheral devices are coupled to the I / O interface 120. Such peripheral devices include (but are not limited to) displays, keyboards, mice, printers, scanners, joysticks or other types of game controllers, media recording devices, external storage devices, network interface cards, etc. The network interface 135 is used to receive and send network messages across the network.

[0020] In various implementations, the computing system 100 is any one of a computer, a laptop computer, a mobile device, a game console, a server, a streaming device, a wearable device, or various other types of computing systems or devices. It should be noted that the number of components of the computing system 100 varies according to the implementation. For example, in other implementations, there are more or fewer of each component compared to the number shown. It should also be noted that in other implementations, the computing system 100 includes Figure 1 other components not shown in. Additionally, in other implementations, the computing system 100 is structured in a different way from Figure 1 that shown. Figure 1 shown.

[0021] Now turning to Figure 2 , a block diagram of one implementation of a system 200 having a CPU and an accelerator is shown. The CPU 205 is coupled to the accelerator 215, and both the CPU 205 and the accelerator 215 are coupled to the system memory 220. The CPU 205 includes cores 210A - 210C representing any number of cores. The cores 210A - 210C are also referred to herein as "execution units". Figure 2 An instance of the process by which the CPU 205 processes system service requests (SSRs) generated by the accelerator 215 is shown in. The accelerator 215 begins by setting parameters in a queue 230 in the system memory 220. Then, the accelerator 215 sends the SSR to core 210A. In one implementation, core 210A executes an upper half interrupt handler that schedules a lower half interrupt handler on core 210B. The lower half interrupt handler sets work queues in a software work queue 225 and the queue 230, and queues a kernel work thread on core 210C. The kernel work thread processes the SSR by accessing the software work queue 225. Then, the kernel work thread processes the SSR and generates a system service response and passes it to the accelerator 215.

[0022] It should be understood that Figure 2An example of a processing accelerator SSR according to one implementation is shown. In other implementations, other schemes for processing accelerator SSRs that include other steps and / or other step orders are used. Figure 2 One of the disadvantages of the method shown is that the CPU 205 lacks the ability to throttle requests generated by the accelerator 215 in cases where the number of requests begins to affect the performance of the CPU 205. For example, in another implementation, the CPU 205 is coupled to multiple accelerators, and each accelerator generates a large number of SSRs in a short period of time.

[0023] Now referring to Figure 3 , an implementation of a timing diagram 300 is shown, which shows the overhead associated with the processing accelerator SSR. Figure 3 The timing diagram 300 of Figure 2 corresponds to the steps for the CPU 205 to process the SSRs generated by the accelerator 215 shown. The first row of the timing diagram 300 shows the timing of events for the accelerator 215. In one implementation, the accelerator 215 generates an interrupt 305 that is processed by the core 210A. Region 310 represents the indirect CPU overhead that the CPU 205 spends transitioning between user mode and kernel mode. Region 315 represents the time spent scheduling the bottom half interrupt handler. After scheduling the bottom half interrupt handler, region 320 represents another indirect CPU overhead for transitioning between kernel mode and user mode. Region 325 represents the time spent running at a lower instructions per cycle (IPC) rate in user mode due to the kernel using various resources of the processor (such as cache and translation lookaside buffer (TLB) space) in the processing of the SSR. This ultimately reduces the resources available for other CPU tasks. In one implementation, when a worker thread calculates the CPU overhead spent servicing the SSR, the worker thread includes the indirect CPU overhead associated with transitioning between user mode and kernel mode and the time spent running at a lower IPC rate in user mode when calculating the overhead.

[0024] For the row shown for core 210B, region 330 represents the indirect CPU overhead spent transitioning from user mode to kernel mode. Region 335 represents the time spent executing the bottom half interrupt handler. Region 340 represents another indirect CPU overhead for transitioning from kernel mode to user mode. Region 345 represents the indirect CPU overhead spent running at a lower IPC rate in user mode due to the kernel using various resources of the processor in the processing of the SSR.

[0025] The lower half interrupt handler initiates a kernel worker thread on core 210C. Region 350 represents the indirect CPU overhead incurred in transitioning from user mode to kernel mode. Region 355 represents the time taken by the kernel worker thread to service the SSR of the accelerator. After servicing the SSR of the accelerator, the kernel worker thread incurs an indirect overhead (represented by region 360) in transitioning from kernel mode to user mode and the time taken to run at a lower IPC rate in user mode (represented by region 365).

[0026] In addition to the time taken by CPU 205 in processing the SSR, accelerator 215 experiences stalls due to the latency of CPU 205 in processing the SSR, as shown by duration 370. Although accelerator 215 has latency hiding capabilities, the latency of CPU 205 in processing the SSR may be longer than what accelerator 215 can hide. As seen from the events represented by timing diagram 300, the accelerator SSR directly and indirectly affects CPU performance, and the CPU's handling of the SSR also affects the performance of the accelerator.

[0027] Now turning to Figure 4 , a block diagram showing another implementation of a scheme for handling accelerator SSRs is shown. Similar to the scheme shown in Figure 2 , CPU 405 is coupled to accelerator 415, and both CPU 405 and accelerator 415 are coupled to system memory 420. CPU 405 includes cores 410A - 410C representing any number of cores. Accelerator 415 starts by setting parameters in queue 430 in system memory 420. Then, accelerator 415 sends the SSR to core 410A. In one implementation, core 410A executes an upper half interrupt handler that schedules a lower half interrupt handler on core 410B. The lower half interrupt handler sets up work queues in software work queue 425 and queue 430 and queues the kernel worker thread on core 410C. However, in contrast to the scheme shown in Figure 2 , regulator 440 determines how long to delay the kernel worker thread before it services the SSR. In one implementation, regulator 440 uses counter 442 to implement the delay. Counter 442 can be implemented using the count of clock cycles or any other suitable measure of time. After this delay, the kernel worker thread services the SSR and generates a system service response and passes it to accelerator 415. Depending on the implementation, regulator 440 is implemented as an OS thread, part of a driver, or any suitable combination of hardware and / or software.

[0028] In one implementation, since the processing of SSRs that have already arrived is delayed when the amount of CPU time spent on processing SSRs is higher than the desired rate, the SSR rate is slowed down. This delay will eventually backpressure accelerator 415 to stop generating new SSR requests. By adding a delay to serving SSRs instead of directly rejecting SSRs from the accelerator, this scenario can be achieved without any modification to how the accelerator generates SSRs. In one implementation, governor 440 decides whether to delay the processing of SSRs based on the amount of CPU time spent on processing SSRs. In one implementation, governor 440 is implemented as a kernel worker thread.

[0029] In one implementation, all OS routines involved in serving SSRs keep track of their CPU cycles. This information is then used by a kernel background thread that wakes up periodically (e.g., every 10 μs) to calculate whether the number of CPU cycles spent on serving SSRs during a period exceeds a specified limit. In one implementation, the limit is specified by an administrator. In another implementation, the limit is dynamically set by the OS based on the characteristics of the applications running at any given time.

[0030] Additionally, a kernel worker thread processes SSRs, as Figure 4 shown. When starting to process an SSR, the worker thread checks whether the CPU cycles spent on processing the SSR are higher than a specified threshold. The worker thread uses the information collected by the background thread to determine whether the CPU cycles spent on processing the SSR are higher than the specified threshold. If the number of CPU cycles spent on processing the SSR is less than or equal to the specified threshold, the worker thread sets the desired delay to zero and immediately continues processing the SSR. Otherwise, if the CPU time spent on processing the SSR is higher than the specified threshold, the worker thread determines the amount of delay to add to the processing of the SSR. The worker thread then waits for this amount of delay before processing subsequent SSRs. By delaying the service of SSRs, governor 440 causes accelerator 415 to throttle its SSR generation rate. For example, accelerator 415 typically has limited space to store the state associated with each SSR. Therefore, delaying SSRs causes accelerator 415 to reduce its SSR generation rate.

[0031] In one implementation, the worker thread uses an exponential backoff scheme to set the amount of delay for processing SSRs. The following is in regard to Figure 5An example of a worker thread using an exponential backoff scheme is described in more detail in the discussion of method 500. For example, in one implementation, if the desired latency was previously greater than zero, the worker thread increases the latency. Otherwise, if the desired latency was previously zero, the worker thread sets the latency to an initial value (e.g., 5 μs). Then, the processing latency of the SSR is determined by the amount determined by the worker thread. As the latency increases, accelerator 415 begins to stall, and the SSR generation rate eventually drops. When the overhead drops below a threshold, the SSR will be serviced again without any artificial latency.

[0032] In other implementations, governor 440 uses other techniques to implement the QoS guarantee mechanism. For example, in another implementation, governor 440 maintains a lookup table to determine how much latency to add to the servicing of SSRs. In this implementation, when the number of CPU cycles spent servicing an SSR is calculated, this number is used as an input to the lookup table to retrieve the corresponding latency value to add to the servicing of subsequent SSRs. In other implementations, governor 440 implements other suitable types of QoS guarantee mechanisms.

[0033] Now refer to Figure 5 , which shows one implementation of method 500 for enforcing QoS guarantees on the CPU when processing requests from the accelerator. For the purposes of discussion, the steps in this implementation and Figures 6 to 7 's steps are shown in sequential order. However, it should be noted that in various implementations of the described method, one or more of the described elements are executed simultaneously, executed in a different order than shown, or omitted entirely. Other additional elements are also executed as needed. Any of the various systems or devices described herein are configured to implement method 500.

[0034] The CPU determines whether the number of CPU cycles spent servicing a system service request (SSR) from the accelerator is greater than a threshold (conditional box 505). Note that the CPU is referred to herein as the first processor, and the accelerator is referred to herein as the second processor. In another implementation, the CPU tracks in conditional box 505 whether the number of CPU cycles spent servicing SSRs from multiple accelerators is. If the number of CPU cycles spent servicing the SSR is less than or equal to the threshold (conditional box 505, "no" branch), the CPU sets the latency equal to zero (box 510). Otherwise, if the number of CPU cycles spent servicing the SSR is greater than the threshold (conditional box 505, "yes" branch), the thread determines whether the latency is currently greater than zero (conditional box 515).

[0035] If the current delay is equal to zero (conditional box 515, "No" branch), the thread sets the delay to an initial value (e.g., 10 μs) (box 520). The initial value varies according to the implementation. Otherwise, if the current delay is greater than zero (conditional box 515, "Yes" branch), the thread increases the value of the delay (box 525). Next, after box 520 or 525, the thread receives a new SSR from the accelerator (box 530). Before servicing the new SSR, the thread sleeps for a duration equal to the current delay value (box 535). Note that the term "sleeps" used in box 535 refers to waiting for an amount of time equal to the current delay value before starting to service the new SSR. After sleeping for a duration equal to the "delay", the thread services the new SSR and returns the result to the accelerator (box 540). After box 540, method 500 ends.

[0036] Now turn to Figure 6 and shows an implementation of method 600 for dynamically adjusting the delay added to the servicing of a request. A first processor monitors the number of cycles that a thread of the first processor spent servicing requests generated by a second processor during a previous time interval (box 605). In another implementation, the first processor not only monitors the number of cycles but also the overhead involved in servicing the requests of the second processor, where the overhead includes multiple components. For example, the overhead includes the time spent actually servicing the request, the indirect CPU overhead spent transitioning between user mode and kernel mode, and the time spent running at a lower IPC rate in user mode due to the kernel using various resources of the processor. In one implementation, the first processor is a CPU and the second processor is an accelerator (e.g., GPU). In other implementations, the first processor and the second processor are other types of processors. The duration of the time interval during which the number of cycles is counted varies according to the implementation.

[0037] If the number of cycles that the first processor thread spent servicing requests generated by the second processor during the previous time interval is greater than a threshold (conditional box 610, "Yes" branch), the first processor adds a first delay amount to the servicing of subsequent requests from the second processor (box 615). Otherwise, if the number of cycles that the first processor thread spent servicing requests generated by the second processor during the previous time interval is less than or equal to the threshold (conditional box 610, "No" branch), the first processor adds a second delay amount to the servicing of subsequent requests from the second processor, where the second delay amount is less than the first delay amount (box 620). In some cases, the second delay amount is zero, such that the first processor immediately services the subsequent requests. After boxes 615 and 620, method 600 ends.

[0038] Now refer to Figure 7, showing an implementation of a method 700 for processing system service requests (SSRs). A first processor receives an SSR from a second processor (block 705). In one implementation, the first processor is a CPU and the second processor is an accelerator (e.g., a GPU). In other implementations, the first and second processors are other types of processors. In response to receiving the SSR from the second processor, the first processor determines whether a first condition has been detected (conditional block 710). In one implementation, the first condition is that the overhead on the first processor for servicing SSRs from the second processor (and optionally from one or more other processors) during a previous time interval is greater than a threshold. The overhead includes the cycles spent actually servicing the SSR, the cycles spent transitioning between user mode and kernel mode before and after servicing the SSR, the cycles spent running at a lower IPC rate in user mode due to degradation of the microarchitecture state (e.g., consuming processor resources such that fewer resources are available), etc. In other implementations, the first condition is any one or a combination of various other types of conditions.

[0039] If the first condition has been detected (conditional block 710, "yes" branch), the first processor waits for a first amount of time before initiating service of the SSR (block 715). Alternatively, when the first condition has been detected, the first processor assigns a first priority to the service of the SSR implementation. If the first condition has not been detected (conditional block 710, "no" branch), the first processor waits for a second amount of time before initiating service of the SSR, where the second amount of time is less than the first amount of time. (block 720). In some cases, the second amount of time is zero, such that the first processor immediately services the SSR. Alternatively, in another implementation, when the first condition has not been detected, the first processor assigns a second priority to the service of the SSR, where the second priority is higher than the first priority. After blocks 715 and 720, method 700 ends.

[0040] In various implementations, program instructions of a software application are used to implement the methods and / or mechanisms described herein. For example, program instructions executable by a general-purpose processor or a special-purpose processor are contemplated. In various implementations, such program instructions are represented in a high-level programming language. In other implementations, the program instructions are compiled from a high-level programming language into a binary form, an intermediate form, or other forms. Alternatively, program instructions are written that describe the behavior or design of hardware. Such program instructions are represented in a high-level programming language such as C. Alternatively, a hardware design language (HDL) such as Verilog is used. In various implementations, the program instructions are stored on any of a variety of non-transitory computer-readable storage media. The storage media may be accessed by a computing system during use to provide the program instructions to the computing system for program execution. Generally speaking, such a computing system includes at least one or more memories and one or more processors configured to execute the program instructions.

[0041] It should be emphasized that the above implementations are merely non-limiting examples of implementations. Once the above disclosure is fully understood, many variations and modifications will become obvious to those skilled in the art. The following claims are intended to be interpreted as covering all such variations and modifications.

Claims

1. A system, which comprises: A first processor, the first processor including circuitry configured to execute multiple threads of an operating system kernel; And A second processor, the second processor coupled to the first processor, wherein the second processor includes circuitry configured to execute an application program and send system service requests to the first processor for service; Wherein the first processor is configured to: Monitor the number of cycles spent by a thread executing on the first processor in serving system service requests during a previous time interval; and Calculate the overhead associated with serving system service requests during the previous time interval, the overhead including the number of cycles and the cycles spent in transitioning between user mode and kernel mode; And Dynamically adjust the amount of delay added to the service of a given system service request based on the overhead.

2. The system according to claim 1, wherein dynamically adjusting the amount of delay added to the service of the given system service request comprises: Adding a first amount of delay in response to determining that the number of cycles is greater than a threshold; And Adding a second amount of delay in response to determining that the number of cycles is less than or equal to the threshold, wherein the second amount of delay is less than the first amount of delay.

3. The system according to claim 2, wherein the circuitry of the first processor is configured to set the first amount of delay to a value greater than the previous amount of delay in response to the number of cycles being greater than the threshold and the previous amount of delay being greater than zero.

4. The system according to claim 1, wherein dynamically adjusting the amount of delay added to the service of the given system service request comprises: Waiting for a first duration before initiating the service of the given system service request in response to determining that the number of cycles is greater than a threshold; And Waiting for a second duration before initiating the service of the given system service request in response to determining that the number of cycles is less than or equal to the threshold, wherein the second duration is less than the first duration.

5. The system according to claim 4, wherein the threshold is dynamically adjusted by the operating system kernel.

6. The system according to claim 1, wherein the overhead further includes the cycles spent in running at a lower instructions per cycle (IPC) rate in user mode.

7. A method, which comprises: Monitoring, by a first processor, the number of cycles spent by a thread executing on the first processor in serving system service requests generated by a second processor during a previous time interval; And Calculating the overhead associated with serving system service requests during the previous time interval, the overhead including the number of cycles and the cycles spent in transitioning between user mode and kernel mode; And Dynamically adjusting the amount of delay added to the service of a given system service request based on the overhead.

8. The method according to claim 7, wherein dynamically adjusting the amount of delay added to the service of the given system service request comprises: Adding a first amount of delay in response to determining that the number of cycles is greater than a threshold; And Adding a second delay amount in response to determining that the number of cycles is less than or equal to the threshold, where the second delay amount is less than the first delay amount.

9. The method of claim 8, further comprising: Setting the first delay amount to a value greater than the previous delay amount in response to the number of cycles being greater than the threshold and the previous delay amount being greater than zero.

10. The method of claim 7, wherein dynamically adjusting the delay amount added to the service of the given system service request comprises: Waiting for a first duration before initiating the service of the given system service request in response to determining that the number of cycles is greater than the threshold; and Waiting for a second duration before initiating the service of the given system service request in response to determining that the number of cycles is less than or equal to the threshold, where the second duration is less than the first duration.

11. The method of claim 10, wherein the threshold is set by an operating system kernel.

12. The method of claim 7, wherein the overhead further includes cycles spent running at a lower instructions per cycle (IPC) rate in user mode.

13. An apparatus, which comprises: A memory; One or more execution units coupled to the memory; wherein the apparatus is configured to: Monitor the number of cycles spent by a thread executing on the one or more execution units in servicing system service requests during a previous time interval; and Calculate the overhead associated with servicing system service requests during the previous time interval, the overhead including the number of cycles and the cycles spent in transitioning between user mode and kernel mode; and Dynamically adjust the delay amount added to the service of a given system service request based on the overhead.

14. The apparatus of claim 13, wherein dynamically adjusting the delay amount added to the service of the given system service request comprises: Adding a first delay amount in response to determining that the number of cycles is greater than the threshold; and Adding a second delay amount in response to determining that the number of cycles is less than or equal to the threshold, where the second delay amount is less than the first delay amount.

15. The apparatus of claim 14, wherein the apparatus is further configured to set the first delay amount to a value greater than the previous delay amount in response to the number of cycles being greater than the threshold and the previous delay amount being greater than zero.

16. The apparatus of claim 13, wherein dynamically adjusting the delay amount added to the service of the given system service request comprises: Waiting for a first duration before initiating the service of the given system service request in response to determining that the number of cycles is greater than the threshold; and Waiting for a second duration before initiating the service of the given system service request in response to determining that the number of cycles is less than or equal to the threshold, where the second duration is less than the first duration.

17. The apparatus of claim 16, wherein the threshold is set by an operating system.

Citation Information

Patent Citations

  • glow wire candle

    DE620717C

  • Method, system, and computer program for managing a queuing system

    US20060161920A1

  • Dynamically adjusting wait periods according to system performance

    US20150234677A1

  • Call stack sampling for threads having latencies exceeding a threshold

    US8286139B2