Input / Output Stator Wake Alignment

By synchronizing request processing and managing shared resource access with urgency levels, the power management system addresses inefficiencies in multi-client computing systems, reducing power consumption and cooling needs.

JP2025522499APending Publication Date: 2025-07-15ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024574610
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-29
Filing Date
2023-05-04
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The increasing power consumption of modern integrated circuits due to frequent transitions between active and idle states in multi-client computing systems leads to higher cooling costs and inefficiencies, particularly in devices with high-performance microprocessors and heterogeneous integration.

Method used

Implementing a power management system where clients synchronize request processing by inserting urgency levels and using indicators to manage shared resource access, reducing unnecessary transitions and optimizing power consumption.

Benefits of technology

This approach reduces power consumption by minimizing active-to-idle state transitions, thereby lowering cooling system requirements and overall energy usage in multi-client computing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522499000001_ABST
    Figure 2025522499000001_ABST
Patent Text Reader

Abstract

An apparatus and method for efficiently performing power management of a multi-client computing system are provided. In various embodiments, the computing system includes a plurality of clients that process tasks corresponding to applications. The clients store generated requests of a particular type during the processing of the tasks. The client receives an indicator that designates that another client services requests of a particular type. In response to receiving this indicator, the client inserts a first urgency level into one or more stored requests of a particular type before sending the requests for service. When the client determines that a particular time interval has elapsed, the client sends an indicator to other clients that designates that requests of a particular type are being serviced. The client inserts a second urgency level, different from the first urgency level, into one or more stored requests of a particular type before sending the requests for service.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] (Description of Related Art) The power consumption of modern integrated circuits (ICs) has become a design challenge that increases with each generation of semiconductor chips. As power consumption increases, more expensive cooling systems such as larger fans and heat sinks must be used to remove excess heat and prevent IC failure. However, the cooling system increases the system cost. Suppressing the power loss of ICs is a challenge not only for portable computers and mobile communication devices, but also for desktop computers and servers that use high-performance microprocessors. These microprocessors include multiple processor cores or cores, and multiple pipelines within the cores.

[0002] Various computing devices, such as various servers, utilize heterogeneous integration that integrates multiple types of ICs to provide system functions. Each of these multiple types of ICs is called a "client". Multiple functions provided by multiple clients include audio / video (A / V) data processing, other advanced data parallel applications for the medical and business fields, instruction processing of general-purpose instruction set architecture (ISA), digital, analog, hybrid signal, and radio-frequency (RF) functions, and the like. There are various choices for system packaging to integrate multiple types of ICs. In some computing devices, a system-on-a-chip (SOC) is used, while in other computing devices, smaller and higher-yield chips are packaged as large chips within a multi-chip module (MCM). Multiple clients of a computing device provide more functions, but these multiple clients are also multiple sources of service requests targeting shared resources. A significant amount of time is spent and a significant amount of power is consumed to service these requests.

[0003] From the above perspective, an efficient method and system for performing efficient power management for a multi-client computing system are desired.

Brief Description of the Drawings

[0004]

Figure 1

Figure 2

Figure 3

Figure 4

Best Mode for Carrying Out the Invention

[0005] Although the present invention has room for various modifications and alternative forms, specific embodiments are shown in the drawings by way of example and are described in detail herein. However, it should be understood that the drawings and their detailed description are not intended to limit the present invention to the specific forms disclosed, but on the contrary, the present invention encompasses all modifications, equivalents, and alternatives included within the scope of the present invention as defined by the appended claims.

[0006] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, those skilled in the art should recognize that the present invention may be practiced without these specific details. In some instances, well-known circuits, structures, and techniques are not shown in detail to avoid obscuring the present invention. Further, for the sake of brevity and clarity of the description, it should be understood that the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements are exaggerated relative to other elements.

[0007] An apparatus and method for efficiently performing efficient power management for a multi-client computing system are contemplated. In various embodiments, a computing system includes a memory storing one or more applications of a workload, and a plurality of clients processing tasks corresponding to the one or more applications. As used herein, a "client" refers to an integrated circuit having a data processing circuit and local memory having a task assigned by a scheduler such as an operating system scheduler. Examples of clients include a general-purpose central processing unit (CPU), a parallel data processing engine having a relatively wide single-instruction-multiple-data (SIMD) microarchitecture, a multimedia engine, any of various types of application specific integrated circuits (ASICs), a digital signal processor (DSP), a display controller, a field programmable gate array (FPGA), and the like.

[0008] During processing of a workload task, a client stores generated requests of a particular type in a data storage area. Examples of data storage areas include a table or queue, a first-in-first-out (FIFO) buffer, a set of flip-flop circuits, and the like. An example of a generated request of a particular type is a system memory request serviced by a system memory controller communicating with the system memory of the computing system. Another example of a generated request of a particular type is an interrupt service request serviced by a particular processor core capable of executing the kernel of the operating system and executing a plurality of interrupt service routines (ISRs). In some embodiments, the particular processor core is a general-purpose core of a general-purpose central processing unit (CPU).

[0009] Of the plurality of clients, a particular client receives an indicator that specifies that at least another client of the plurality of clients causes a shared resource to service a particular type of request. Also, this indicator can specify that other clients should cause a shared resource to service a particular type of request, but the shared resource has not yet started servicing these requests. For example, other clients send requests to the shared resource within a relatively short period of time and send an indicator to the particular client. In some embodiments, the indicator is a sideband signal transmitted between two clients either directly or in a transfer manner with other clients used as communication hops. In one embodiment, other clients send requests to the shared resource before sending the sideband signal to the particular client, but the sideband signal arrives at the particular client before the shared resource receives the request or before the shared resource starts servicing the request. Thus, this indicator specifies that other clients should service a particular type of request. The shared resource may already have serviced these requests of other clients, but that is not necessary for the meaning of the indicator. In other embodiments, the indicator is a message within a packet transmitted via a communication fabric. In response to receiving this indicator, the particular client inserts a first urgency level into one or more stored requests of a particular type. The particular client inserts this first urgency level before sending one or more stored requests of a particular type for servicing.

[0010] The urgency level provides an indication of the expected amount of time to service the corresponding request of a particular type. In some embodiments, the urgency level can be one or more bits inserted into the request that give a value that is a combination of one or more of a priority level, a quality of service (QoS) level, an indication of an application type such as a real-time application, etc.

[0011] Also, one or more of the plurality of clients maintains the duration between scheduled services for a particular type of request. In some embodiments, a value corresponding to a time threshold or time interval is stored in a configuration register. A particular client counts clock cycles or otherwise measures the time since the last scheduled service for a particular type of request. The particular client compares the measured duration to the time interval, and if the measured duration exceeds the time interval, the particular client sends an indicator to one or more other clients of the plurality of clients specifying that a particular type of request is being serviced by a shared resource. Also, the particular client inserts a second urgency level different from the first urgency level into one or more stored requests of a particular type. The particular client inserts this second urgency level before sending one or more stored requests of a particular type to be serviced. In one embodiment, the second urgency level indicates a higher urgency than the first urgency level. Further details for efficiently performing efficient power management for a multi-client computing system are provided in the following description.

[0012] Referring to FIG. 1, a general block diagram of a computing system 100 that performs power management for a plurality of clients is shown. The computing system 100 includes a semiconductor chip 110 and a system memory 130. The semiconductor chip 110 (or chip 110) includes a plurality of types of integrated circuits. For example, the chip 110 includes at least a plurality of processor cores (or cores), such as cores in the processing unit 150 and cores 112 and 116 of the processing units (units) 115 and 119. There are various options for placing the circuits of the chip 110 in system packaging to integrate a plurality of types of integrated circuits. Some examples are system-on-a-chip (SOC), multi-chip module (MCM), and system-in-package (SiP). A clock source, such as a phase lock loop (PLL), an interrupt controller, a power controller, an interface for input / output (I / O) devices, etc., are not shown in FIG. 1 for simplicity of explanation.

[0013] In various embodiments, the units 115, 119, 150 are a plurality of clients within the chip 110. As used herein, "client" refers to an integrated circuit having a data processing circuit and local memory that has tasks assigned by a scheduler, such as an operating system scheduler. Examples of clients are general-purpose central processing units (CPUs), parallel data processing engines having a relatively wide single-instruction-multiple-data (SIMD) microarchitecture, multimedia engines, any of various types of application-specific integrated circuits (ASICs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), etc.

[0014] Chip 110 includes a system memory controller 120 for communicating with system memory 130. Chip 110 also includes interface logic 140, processing units (units) 115 and 119, a communication fabric 160, a shared cache memory subsystem 170, and a processing unit 150. Processing unit 115 includes a core 112 and a corresponding cache memory subsystem 114. Processing unit 119 includes a core 116 and a corresponding local memory 118. Although a single core is shown for each of units 115 and 119, in other embodiments, a different number of cores may be used. System memory 130 is shown as including operating system code 132. Note that various portions of operating system code 132 may be stored and resident in system memory 130, in cache 114, and on a non-volatile storage device such as a hard disk (not shown). In one embodiment, the exemplary functions of chip 110 are incorporated on a single integrated circuit.

[0015] Three clients (units 115, 119, 150) are shown, but chip 110 can and is contemplated to include a different number of clients. In various embodiments, cores 112 and 116 can execute one or more threads and share at least a shared cache memory subsystem 170, a processing unit 150, and a combined input / output (I / O) device connected to interface logic 140 (or interface 140). In some embodiments, units 115, 119, 150 use different microarchitectures.

[0016] The hardware such as the circuits of a specific core among the multiple cores within chip 110 executes the instructions of the operating system. In one embodiment, core 112 is a general-purpose core and unit 115 is a general-purpose CPU. In various embodiments, system memory controller 120 services system memory requests generated by multiple clients (units 115, 119, 150), and core 112 services interrupt service requests generated by multiple clients (units 119 and 150). System memory controller 120 and core 112 are examples of shared resources that service specific types of requests from multiple clients.

[0017] The client (cores within units 115, 119, 150) stores generated requests of a specific type in a data storage area during the processing of workload tasks. Examples of data storage areas are tables or queues, first-in first-out (FIFO) buffers, sets of flip-flop circuits, etc. Examples of generated requests of a specific type sent to shared resources are system memory requests and interrupt service requests. The shared resource is the corresponding component (e.g., any of system memory controller 120, core 112, etc.) that services requests of a specific type sent to the shared resource from multiple clients. A specific client, such as core 116 within unit 119, receives an indicator specifying that at least another client (cores within units 115, 119, 150) should service requests of a specific type. For example, one or more cores of units 115, 119, 150 have already sent or are currently sending requests of a specific type to the shared resource for servicing. In some embodiments, the indicator is a sideband signal transmitted between two clients either directly or in a transfer manner with other clients used as communication hops. In other embodiments, the indicator is a message within a packet transmitted via communication fabric 160. The indicator functions as an input / output (I / O) stutter hint that communicates to multiple clients that there is an opportunity to synchronize the processing of requests of a specific type. In response to receiving this indicator, a specific client (core 116) inserts a first urgency level into one or more stored requests of a specific type. Core 116 inserts this first urgency level before sending one or more stored requests of a given type to the shared resource for servicing.

[0018] The urgency level provides an indication of the amount of time expected to service a corresponding request of a particular type. In some embodiments, the urgency level can be one or more bits inserted into a request that give a value that is a combination of one or more of a priority level, a quality of service (QoS) level, an indication of an application type such as a real-time application, etc. Also, one or more of a plurality of clients (cores within units 115, 119, 150) maintain a duration between scheduled services of requests of a particular type. In some embodiments, a value corresponding to a time threshold or time interval is stored in a configuration register.

[0019] A particular client (core 116) counts clock cycles or otherwise measures the time since the last scheduled service of a request of a particular type. Core 116 compares the measured duration to a time interval and, if the measured duration exceeds the time interval, core 116 transmits to one or more other clients (cores within units 115, 119, 150) an indication specifying that a request of a particular type is being serviced by a shared resource corresponding thereto. In some embodiments, the indication is a sideband signal transmitted between two clients either directly or in a transfer scheme with other clients used as communication hops. In other embodiments, the indication is a message within a packet transmitted via communication fabric 160. Also, core 116 inserts a second urgency level different from a first urgency level into one or more stored requests of a particular type. Core 116 inserts this second urgency level prior to transmitting one or more stored requests of a given type for service. In one embodiment, the second urgency level indicates a higher urgency than the first urgency level.

[0020] In some embodiments, the core 116 further adds a second condition to the above first condition of determining that the measured duration exceeds a time interval for a particular type of request to be eligible to be sent to the shared resource and sending the metric to other clients. In one embodiment, this second condition is to determine that the number of requests of a particular type in hold exceeds a threshold number. If the number of requests of a particular type in hold does not exceed the threshold number, the core 116 does not send the requests of the particular type in hold to the shared resource and does not send the metric to other clients. The shared resource, which may be in an idle state, remains in the idle state. Thus, the power consumption of the computing system is reduced. Before proceeding with further details of efficiently scheduling tasks in a dynamic manner for a plurality of cores supporting heterogeneous computing architectures, a further description of the components of the computing system 100 is provided.

[0021] Interface 140 generally provides an interface for various types of input / output (I / O) devices outside the chip 110 to the shared cache memory subsystem 170 and the processing unit 115. Generally, interface logic 140 includes buffers for receiving packets from corresponding links and buffering packets to be transmitted on corresponding links. Any suitable flow control mechanism may be used to transmit and receive packets to and from the chip 110. System memory 130 may be used as system memory for the chip 110 and may include any suitable memory device such as one or more RAMBUS dynamic random access memories (DRAMs), synchronous DRAMs (SDRAMs), DRAMs, static RAMs, etc.

[0022] The address space of chip 110 is divided among multiple memories corresponding to multiple cores. In one embodiment, the address coherence point is system memory controller 120 that communicates with the memory storing the byte corresponding to the address. System memory controller 120 includes control circuitry for interfacing with the memory and a request queue for queuing memory requests. Generally speaking, communication fabric 160 responds to control packets received on the links of interface 140, generates control packets in response to cores 112 and 116 and / or cache memory subsystems 114 and 118, generates probe commands and response packets in response to a transaction selected by system memory controller 120 for servicing, and routes packets to other nodes via interface logic 140. The communication fabric supports various packet transmission protocols and includes one or more of a system bus, packet processing circuitry and packet selection arbitration logic, and queues for storing requests, responses and messages.

[0023] Cache memory subsystem 114 includes a relatively high-speed cache memory that stores data blocks. Cache memory subsystem 114 can be integrated within each high-performance core 112. Alternatively, cache memory subsystem 114 can be connected to high-performance core 112 in a backside cache configuration or an in-line configuration as needed. Cache memory subsystem 114 can be implemented as a cache hierarchy. In one embodiment, cache memory subsystem 114 represents an L2 cache structure and shared cache subsystem 170 represents an L3 cache structure. Local memory 118 can be implemented in any of the ways described above for cache memory subsystem 114, such as a local data store.

[0024] Referring to FIG. 2, a general block diagram of a timing diagram 200 is shown that depicts a period during which multiple clients within a computing system send a particular type of generated request to a shared resource. Timing diagram 200 includes a period 220 during which clients send a particular type of request to the shared resource without synchronizing the processing of the particular type of request. Further, timing diagram 200 includes a period 230 during which clients send a particular type of request to the shared resource while synchronizing the processing of the particular type of request. Clients 210, 212, 214 represent any type of client such as the examples provided previously or other types of clients. As described previously, an example of a particular type of generated request is a system memory request serviced by a system memory controller that communicates with the system memory of a computing system. Another example of a particular type of generated request is an interrupt service request serviced by a particular processor core, such as a CPU core, that executes the kernel of an operating system and can execute multiple interrupt service routines (ISRs). Thus, the system memory controller and the particular processor core are examples of shared resources.

[0025] The period during which at least one client is sending requests of a specific type to the shared resource is shown over time as block 204. The period during which none of the clients is sending requests of a specific type to the shared resource is shown over time as block 202. During period 220, clients 210, 212, and 214 do not attempt to synchronize when sending requests of a specific type to a shared resource such as a system memory controller. The shared resource is the corresponding component (e.g., any one of a system memory controller, a CPU core, etc.) that services requests of a specific type sent to the shared resource from multiple clients. These requests of a specific type are sent to the component during the period shown as time t1 (or time t1) and during the periods from time t2 to t11. When receiving a service request, the component (a shared resource such as a system memory controller) can be in an operating mode corresponding to an idle power performance state (P state). The component transitions to an operating mode corresponding to an active P state and consumes a significant amount of time to execute the step of servicing the received request.

[0026] Multiple clients provide more functions, but these multiple clients are also multiple sources of service requests. Sometimes, the request for a service is from a single client among multiple clients, for example, during the period from time t1 to t9. After the component (shared resource) services the request, the component returns to an operating mode corresponding to the idle P state. However, shortly thereafter, the component receives more service requests, which could be from a single client among multiple clients, for example, during the period from time t2 to t9. The component once again transitions to an operating mode corresponding to the active P state and executes the step of servicing the received requests. Considerable time is spent repeatedly waking up the component to service requests from multiple clients. Also, the corresponding component frequently transitions between an active state (awake state) and an idle state (sleep state). Therefore, the computing system increases power consumption by repeatedly waking up the component to service requests from multiple clients.

[0027] During period 230, clients 210, 212, 214 attempt to synchronize when sending requests of a particular type to the shared resource for servicing. For example, a particular client among a plurality of clients receives an indicator specifying that at least another client among the plurality of clients should have a request of a particular type serviced by the shared resource. As explained above, this indicator specifies that other clients should service requests of a particular type. The shared resource may already be servicing these requests of other clients, but that is not necessary in the sense of the indicator. In some embodiments, the indicator is a sideband signal transmitted between two clients either directly or via another client used as a communication hop. In other embodiments, the indicator is a message within a packet transmitted via the communication fabric. Also, the plurality of clients perform the above-described steps with respect to the clients of the computing system 100 (of FIG. 1) and the following method 400 (of FIG. 4). Accordingly, corresponding components (e.g., system memory controller, CPU core, etc.) transition between an active state (awake state) and an idle state (sleep state) less frequently. As shown, these requests of a particular type are transmitted to the component between times t20 and t25. Accordingly, the computing system reduces power consumption.

[0028] Referring to FIG. 3, a general block diagram of a table 300 that stores information used to perform power management of a plurality of clients and shared resources is shown. Table 300 includes a plurality of table entries (or entries), each storing information in a plurality of fields such as at least fields 302 to 306. Table 300 is implemented by any of a flip-flop circuit, a random access memory (RAM), a content addressable memory (CAM), a first-in first-out (FIFO) buffer, and the like. Although it is shown that specific information is stored in fields 302 to 306 in a specific consecutive order, in other embodiments, different orders are used and different numbers and types of information are stored. Here, two specific types of requests are managed, such as a system memory request shown as a direct memory access (DMA) request and an interrupt service request shown as a CPU request.

[0029] As shown, fields 302 and 304 store bits corresponding to a binary truth table indicating which requests of a particular type are permitted to be sent to a shared resource to be serviced by multiple clients. Field 306 stores steps supported based on the information stored in fields 302 - 304. These steps include when to use an indicator (such as a sideband signal), which functions as an input / output (I / O) stutter hint that enables multiple clients to operate in an I / O stutter wake alignment state. The indicator communicates an opportunity to synchronize the processing of requests of a particular type among multiple clients. In some embodiments, the term "urgent" refers to an urgency level inserted into a service request when a client determines that a measured duration exceeds a time interval, and the term "non-urgent" refers to an urgency level inserted into a service request when a client receives an indicator specifying that at least another one of multiple clients should service a request of a particular type to a shared resource. As explained above, this indicator specifies that other clients should service requests of a particular type. The shared resource may already be servicing these requests of other clients, but that is not the necessary meaning of the indicator. In another embodiment, the terms "urgent" and "non-urgent" have the reverse definitions. In yet another embodiment, the terms "urgent" and "non-urgent" have other definitions based on design requirements.

[0030] In one embodiment, each of a plurality of clients stores a copy of table 300 and further determines when to send a particular type of request for service to a shared resource using the information in fields 302-306. The information stored in table 300 is used by the client in combination with other conditions, such as receiving an indicator (e.g., sideband signal) from another client, determining that a measured duration exceeds a time interval, determining that the number of outstanding requests of a particular type exceeds a threshold number, etc. By using this combination of conditions, the shared resource can synchronize the service of requests of a particular type from a plurality of clients. As shown earlier in the timing diagram 200 (of FIG. 2), this synchronization reduces the transition of the shared resource between an active state (awake state) and an idle state (sleep state). Thus, the computing system reduces power consumption. In some embodiments, table 300 is programmable and a power manager updates the values (contents) stored in table 300. The power manager can notify the client of when the update is to be performed.

[0031] Next, referring to FIG. 4, a general block diagram of a method 400 for efficiently performing power management of a multi-client computing system is shown. For the sake of explanation, the steps in this embodiment are shown in order. However, in other embodiments, some steps occur in a different order than shown, some steps are executed simultaneously, some steps are combined with other steps, and some steps do not exist.

[0032] In various embodiments, a computing system includes (block 402) a memory that stores one or more applications of a workload, and a plurality of clients that process tasks corresponding to the one or more applications. During processing of the tasks of the workload, the clients store (block 404) generated requests of a particular type in a data storage area. Store the generated system memory requests (block 404). Examples of data storage areas include tables or queues, first-in first-out (FIFO) buffers, sets of flip-flop circuits, and the like.

[0033] An example of a generated request of a particular type is a system memory request serviced by a system memory controller that communicates with the system memory of the computing system. Another example of a generated request of a particular type is an interrupt service request serviced by a particular processor core that executes the kernel of an operating system and can execute a plurality of interrupt service routines (ISRs). In some embodiments, the particular processor core is a general-purpose core of a general-purpose central processing unit (CPU).

[0034] In various embodiments, another client sends a particular type of request to a particular shared resource to service a particular type of request. In one embodiment, another client sends a system memory request to a system memory controller before sending an indicator (such as a sideband signal) to the client. However, it is possible for the indicator to arrive at the client before the system memory controller receives the system memory request or before the system memory controller starts servicing the system memory request. Thus, this indicator designates that another client should service a particular type of request. The system memory controller may already be servicing a system memory request of another client, but that is not necessary in the context of the indicator. When a client receives an indicator that another client should service a particular type of request (conditional block 406: "Yes"), the client inserts a first urgency level into one or more stored requests of the particular type (block 408).

[0035] In one embodiment, the first urgency level indicates a lower urgency than the urgency level of a particular type of request from another client that has started servicing a particular type of request. In some embodiments, the urgency level can be one or more bits inserted into the request that give a value that is a combination of one or more of a priority level, a quality of service (QoS) level, an indicator of an application type such as a real-time application, etc. The client sends one or more stored requests of the particular type to the corresponding component to be serviced (block 410). As described above, the system memory controller services system memory requests generated by multiple clients, and the CPU core that executes the operating system services interrupt service requests generated by multiple clients.

[0036] If the client does not receive an indication that another client should service a particular type of request (condition block 406: "No"), the control flow of method 400 moves to condition block 412. If the client determines that a particular time interval has elapsed (condition block 412: "Yes"), the client sends an indication to one or more other clients specifying that a particular type of request has been sent for servicing (block 414). In some embodiments, the indication is a sideband signal sent between two clients, either directly or via a transfer method with other clients used as communication hops. In other embodiments, the indication is a message within a packet sent via a communication fabric. In one embodiment, the client uses a further prerequisite to determine whether to execute the steps of blocks 414-418 of method 400. An example of a further prerequisite is to determine that the number of pending requests of a particular type exceeds a threshold number. For the steps executed in blocks 406-410 of method 400, the number of pending requests of a particular type may have been reduced previously. If the client determines that the further prerequisite is not satisfied, the control flow of method 400 skips the steps of blocks 414-418 and returns to block 402.

[0037] The client inserts a second urgency level, different from the first urgency level, into one or more stored requests of a particular type (block 416). The client transmits one or more stored requests of a particular type to the corresponding component for servicing (block 418). In various embodiments, the client resets a counter used to measure the duration after the service scheduled at the end of the requests of a particular type. Thereafter, the control flow of method 400 returns to block 402. Similarly, if the client determines that a particular time interval has not yet elapsed (conditional block 412: "No"), the control flow of method 400 returns to block 402. As described above, in some embodiments, the client further adds a second condition to the above first condition of determining that the measured duration exceeds a time interval for obtaining eligibility to transmit requests of a particular type to the shared resource and transmitting an indicator to other clients. In one embodiment, this second condition is to determine that the number of pending requests of a particular type exceeds a threshold number. If the number of pending requests of a particular type does not exceed the threshold number, the client does not transmit the pending requests of a particular type to the shared resource and does not transmit an indicator to other clients. The shared resource, which may be in an idle state, remains in the idle state. Thus, the power consumption of the computing system is reduced.

[0038] In one embodiment, each of a plurality of clients stores a copy of a table, such as table 300 (of FIG. 3), and uses the information in the table to further determine when to send a particular type of request for service to a shared resource. The information stored in the table is used by the client in combination with other conditions, such as receiving an indicator (sideband signal, etc.) from another client, determining that a measured duration exceeds a time interval, determining that the number of requests of a particular type pending exceeds a threshold number, etc. By using this combination of conditions, the shared resource can synchronize the service of requests of a particular type from a plurality of clients. This synchronization reduces the transition of the shared resource between an active state (awake state) and an idle state (sleep state), as previously shown in timing diagram 200 (of FIG. 2). Thus, the computing system reduces power consumption.

[0039] It should be noted that one or more of the above-described embodiments include software. In such embodiments, the program instructions implementing the method and / or mechanism are carried or stored on a computer-readable medium. A number of types of media configured to store program instructions are available, including hard disks, floppy (registered trademark) disks, CD-ROMs, DVDs, flash memories, programmable ROMs (Programmable ROM, PROM), random access memories (random access memory, RAM), and various other forms of volatile or non-volatile storage devices. Generally speaking, a computer-accessible storage medium includes any storage medium that can be accessed by a computer during use to provide instructions and / or data to the computer. For example, computer-accessible storage media include magnetic or optical media such as disks (fixed or removable), tapes, CD-ROMs, DVD-ROMs, CD-Rs, CD-RWs, DVD-Rs, DVD-RWs, or storage media such as Blu-Ray (registered trademark). Storage media further include volatile or non-volatile memory media such as RAM (e.g., synchronous dynamic RAM (synchronous dynamic RAM, SDRAM), double data rate (DDR, DDR2, DDR3, etc.) SDRAM, low power DDR (LPDDR2, etc.) SDRAM, Rambus DRAM (Rambus DRAM, RDRAM), static RAM (SRAM), etc.), ROM, flash memory, non-volatile memory (e.g., flash memory) accessible via a peripheral interface such as a Universal Serial Bus (Universal Serial Bus, USB) interface, etc. Storage media include microelectromechanical systems (microelectromechanical system, MEMS), as well as storage media accessible via communication media such as networks and / or wireless links.

[0040] Additionally, in various embodiments, the program instructions include an operational level description of the hardware functionality or a register-transfer level (RTL) description in a high-level programming language such as C, a design language (HDL) such as Verilog or VHDL, or a database format such as the GDSII stream format (GDS II). In some cases, the description is read by a synthesis tool that synthesizes the description to produce a netlist that includes a list of gates from a synthesis library. The netlist includes a set of gates that also represent the functionality of the hardware including the system. The netlist can then be placed and routed to produce a dataset that describes the geometric shapes to be applied to a mask. The mask can then be used in various semiconductor manufacturing steps to produce a semiconductor circuit or circuits corresponding to the system. Alternatively, the instructions on the computer-accessible storage medium are, optionally, a netlist (with or without a synthesis library) or a dataset. Additionally, the instructions are utilized for emulation by a hardware-based type of emulator from a vendor such as Cadence®, EVE®, and Mentor Graphics®.

[0041] While the above embodiments have been described in considerable detail, many modifications and variations will become apparent to those skilled in the art upon a full understanding of the above disclosure. It is intended that the following claims be interpreted to embrace all such modifications and variations.

Claims

1. An apparatus, comprising a circuit, wherein the circuit is configured to: during processing of a task of a workload, store generated requests of a specific type in a data storage area; and in response to receiving a first indicator specifying that an external client services the requests of the specific type, transmit one or more stored requests of the specific type for servicing; and is configured to perform the above. The apparatus.

2. Before transmitting one or more stored requests of the specific type to a shared resource for servicing, the circuit is configured to insert an urgency level indicating an urgency lower than an urgency level of the requests of the specific type from the external client into the one or more stored requests of the specific type. The apparatus according to Claim 1.

3. In response to determining that a predetermined time interval has elapsed, the circuit is configured to transmit a second indicator specifying that the apparatus services the requests of the specific type to one or more external clients. The apparatus according to Claim 1.

4. Before transmitting one or more stored requests of the specific type to a shared resource for servicing, in response to receiving the first indicator, the circuit is configured to insert an urgency level indicating an urgency higher than an urgency level of the requests of the specific type transmitted by the apparatus into the one or more stored requests of the specific type. The apparatus according to Claim 3.

5. In response to determining that the number of stored requests of the specific type exceeds a threshold, the circuit is configured to transmit the second indicator to the one or more external clients. The apparatus according to Claim 3.

6. The requests of the specific type are system memory requests. The apparatus according to Claim 1.

7. The requests of the specific type are interrupt service requests. The apparatus according to Claim 1.

8. A method, comprising: a plurality of clients processing a task of a workload; and a first client among the plurality of clients storing generated requests of a specific type in a data storage area during processing of the task of the workload; In response to the first client receiving a first indicator specifying that a second client among the plurality of clients services requests of the specific type, transmitting one or more stored requests of the specific type for servicing. Method. **Claim 9** Before the first client transmits one or more stored requests of the specific type to a shared resource for servicing, inserting an urgency level indicating an urgency lower than the urgency level of requests of the specific type from the second client into the one or more stored requests of the specific type. The method of claim 8. **Claim 10** In response to the first client determining that a predetermined time interval has elapsed, transmitting a second indicator specifying that the first client services requests of the specific type to one or more of the plurality of clients. The method of claim 8. **Claim 11** Before the first client transmits one or more stored requests of the specific type to a shared resource for servicing, in response to receiving the first indicator, inserting an urgency level indicating an urgency higher than the urgency level of requests of the specific type transmitted by the first client into the one or more stored requests of the specific type. The method of claim 10. **Claim 12** The first client is configured to transmit the second indicator to the one or more external clients in response to determining that the number of stored requests of the specific type has exceeded a threshold. The method of claim 10. **Claim 13** The requests of the specific type are system memory requests. The method of claim 8. **Claim 14** The requests of the specific type are interrupt service requests. The method of claim 8. **Claim 15** A computing system comprising: A memory configured to store one or more applications of a workload; A plurality of clients each configured to process tasks of the workload. Among the plurality of clients, a first client: During processing of tasks of the workload, storing generated requests of a specific type in a data storage area. In response to receiving a first indicator that specifies that a second client among the plurality of clients services the specific type of request, transmitting one or more stored requests of the specific type for servicing; configured to perform; a computing system.

16. Before the first client transmits one or more stored requests of the specific type to a shared resource for servicing, the first client is configured to insert an urgency level indicating an urgency lower than the urgency level of the request of the specific type from the second client into the one or more stored requests of the specific type. The computing system of claim 15.

17. In response to determining that a predetermined time interval has elapsed, the first client is configured to transmit a second indicator that specifies that the first client services the specific type of request to one or more of the plurality of clients. The computing system of claim 15.

18. Before the first client transmits one or more stored requests of the specific type to a shared resource for servicing, in response to receiving the first indicator, the first client is configured to insert an urgency level indicating an urgency higher than the urgency level of the request of the specific type transmitted by the first client into the one or more stored requests of the specific type. The computing system of claim 17.

19. In response to determining that the number of stored requests of the specific type has exceeded a threshold, the first client is configured to transmit the second indicator to the one or more external clients. The computing system of claim 17.

20. The request of the specific type is a system memory request. The computing system of claim 15.