Dynamic configuration of processor subassemblies
By embedding control circuits in the processor hardware, real-time observation and dynamic management of the utilization of functional units, the problems of processor power consumption and heat generation are solved, and efficient power management and performance maintenance are achieved.
Patent Information
- Application Number
- CN202380083281.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-19
- Filing Date
- 2023-12-19
- Publication Date
- 2025-07-11
AI Technical Summary
The power consumption and heat generation of the processor adversely affect performance, and existing software-based throttling technologies respond slowly and have limited granularity.
By embedding control circuits in the processor hardware, the utilization of functional units and subcomponents is observed in real time, and throttling is dynamically performed and dethrotating is improved to improve power efficiency.
Efficient power management of processor components at low latency is achieved, reducing power consumption and heat generation while maintaining performance.
Smart Images

Figure CN120303640A_ABST
Abstract
Description
Background Art
[0001] The performance of a computing device while running a program may be limited in part by the performance of its processor. Various advancements have improved processor performance to increase the number of instructions per cycle that the processor can execute. For example, the addition of various circuits and components may allow the processor to perform actions in a more parallel manner, or otherwise reduce the requirement to perform actions in a specific order. Multicore processors may also allow for parallel computing and multithreading. A multicore processor may include multiple cores or processing units, each core or processing unit being capable of independently running a thread. To further utilize these processor features, software programs may be optimized for parallel computing. Brief Description of the Drawings
[0002] The drawings illustrate multiple exemplary embodiments and are part of the specification. Together with the following description, these drawings demonstrate and explain the various principles of the present disclosure.
[0003] Figure 1 is a block diagram of an exemplary processor that includes a control circuit for dynamically configuring subcomponents.
[0004] Figure 2 is a block diagram of an exemplary processor that includes a control circuit for dynamically configuring subcomponents.
[0005] Figure 3 is a block diagram of an exemplary multicore processor that includes a control circuit for dynamically configuring subcomponents.
[0006] Figure 4 is a block diagram of a dual pipeline for a micro-op cache.
[0007] Figure 5 is a flowchart of an exemplary method for dynamic configuration of processor subcomponents.
[0008] In all of the drawings, the same reference numerals and descriptions indicate like but not necessarily identical elements. While the exemplary embodiments described herein are susceptible to various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and will be described in detail herein. However, the exemplary embodiments described herein are not intended to be limited to the particular forms disclosed. Rather, the present disclosure covers all modifications, equivalents, and alternatives falling within the scope of the appended claims. Detailed Description
[0009] The present disclosure generally relates to dynamic configuration of processor sub-components for power efficiency. Physical considerations may limit the performance of a processor. When additional processor components operate, these components may cause power consumption of the processor. Power consumption and the associated heat generation may have an adverse effect on the performance of the processor. Various techniques may address power consumption via software-based conditions that may trigger throttling of the processor or its components. Such techniques are typically abstracted from the processor hardware such that throttling may take hundreds to thousands of cycles to respond to the trigger. Such techniques may also be limited in terms of the granularity at which the processor is observed and analyzed. As will be explained in more detail below, specific implementations of the present disclosure may implement a control circuit coupled to and physically proximate to a functional unit of the processor and / or a sub-component of the functional unit. The control circuit may observe whether the utilization of the functional unit and / or the sub-component is outside of a desired utilization range and throttle portions of the processor as needed. By using a control circuit in such an arrangement, the systems and methods described herein may advantageously improve power efficiency and heat generation at a granularity level with relatively low latency.
[0010] As will be described in more detail below, the present disclosure describes various systems and methods for dynamically configuring processor sub-components to improve power efficiency. The systems and methods described herein observe how functional units are utilized and throttle the functional units if the utilization of the functional units is outside of a desired utilization range.
[0011] In one example, an apparatus for dynamic configuration of a processor sub-component includes: a functional unit of a processor, the functional unit including at least one sub-component, the at least one sub-component including a target sub-component; and a control circuit coupled to the target sub-component. The control circuit is configured to observe the utilization of the target sub-component and detect that the utilization is outside of a desired utilization range. The controller is further configured to throttle the at least one sub-component of the functional unit in response to the detection to reduce the power consumption of the functional unit.
[0012] In some examples, the control circuit is further configured to observe the workload input to the target sub-component. In some examples, the at least one sub-component of the functional unit is selected for throttling based on reducing the workload input to the target sub-component. In some examples, the at least one sub-component of the functional unit is selected for throttling based on the type of workload of the workload. In some examples, the workload corresponds to a minimum workload, and the at least one sub-component corresponds to the target sub-component.
[0013] In some examples, throttling the at least one sub-component includes placing the at least one sub-component in a low power state. In some examples, throttling the at least one sub-component includes throttling the functional unit.
[0014] In some examples, the control circuit is further configured to detect that the utilization is within the desired utilization range, and in response to detecting that the utilization is within the desired utilization range, un-throttle the at least one sub-component.
[0015] In some examples, the target sub-component corresponds to a micro-operation queue for storing pre-decoded instructions, and the functional unit further includes at least a first sub-component and a second sub-component, and the first sub-component and the second sub-component each provide pre-decoded instructions to the micro-operation queue. Additionally, the at least one sub-component corresponds to at least one of the first sub-component and the second sub-component.
[0016] In some examples, the functional unit corresponds to at least one of an arithmetic logic unit (ALU), a floating point unit (FPU), or a load-store unit (LSU). In some examples, the processor corresponds to a multi-core processor, and the functional unit corresponds to a core of the multi-core processor.
[0017] In some examples, the device further includes a second functional unit of the processor, the second functional unit including a second target sub-component; a second control circuit coupled to the second target sub-component; and a higher-level control circuit coupled to the control circuit and the second control circuit. The higher-level control circuit is configured to detect throttling of the at least one sub-component of the functional unit, and to coordinate throttling of sub-components by the control circuit and the second control circuit.
[0018] In one example, a method for dynamic configuration of a processor sub-component includes: observing, using a control circuit coupled to a target sub-component of a functional unit of a processor, a utilization of the target sub-component based on a workload input to the target sub-component, and detecting that the utilization is outside a desired utilization range of the workload. The method further includes: throttling at least one sub-component of the functional unit to reduce power consumption of the functional unit.
[0019] In some examples, the at least one sub-component of the functional unit is selected for throttling based on reducing the workload input to the target sub-component. In some examples, the workload corresponds to a minimum workload, and the at least one sub-component corresponds to the target sub-component.
[0020] In some examples, the target sub-component corresponds to a micro-operation queue for storing pre-decoded instructions. The functional unit further includes at least a first sub-component and a second sub-component, and the first sub-component and the second sub-component each provide pre-decoded instructions to the micro-operation queue. The at least one sub-component corresponds to at least one of the first sub-component and the second sub-component.
[0021] In some examples, the method further includes: using a higher-level control circuit coupled to the control circuit to detect throttling of the at least one sub-component of the functional unit. The higher-level control circuit is further coupled to a second control circuit, and the second control circuit is coupled to a second target sub-component of a second functional unit of the processor. The method further includes: using the higher-level control circuit to coordinate throttling of sub-components by the control circuit and the second control circuit.
[0022] In one specific implementation, a system for dynamic configuration of a processor sub-component includes a physical memory and at least one physical processor, and the at least one physical processor includes a functional unit and a control circuit. The functional unit includes at least one sub-component having a target sub-component. The control circuit is coupled to the target sub-component and is configured to observe the workload and utilization rate of the target sub-component and detect that the utilization rate is outside a desired utilization rate range of the workload. The control circuit is further configured to throttle at least one sub-component of the functional unit to reduce power consumption of the physical processor.
[0023] In some examples, the at least one sub-component of the functional unit is selected for throttling based on reducing the workload input to the target sub-component. In some examples, the target sub-component corresponds to a micro-operation queue for storing pre-decoded instructions, the functional unit further includes at least a first sub-component and a second sub-component, the first sub-component and the second sub-component each provide pre-decoded instructions to the micro-operation queue, and the at least one sub-component corresponds to at least one of the first sub-component and the second sub-component.
[0024] According to the general principles described herein, the features of any of the specific implementations described herein can be used in combination with each other. These and other specific implementations, features, and advantages will be more fully understood after reading the following detailed description in conjunction with the accompanying drawings and the claims.
[0025] The following will refer to Figures 1 to 5 Provide a detailed description of a system and method for dynamic configuration of a processor sub-component. It will be combined with Figures 1 to 3 Provide a detailed description of various example processors having a control circuit for dynamic configuration of sub-components. It will be combined with Figure 4 Provide a detailed description of an example dual pipeline for a micro-operation cache. In addition, it will also be combined with Figure 5 Provide a detailed description of the corresponding method.
[0026] Figure 1 is a block diagram of an example system 100 for dynamic configuration of a processor sub-component for power efficiency. System 100 corresponds to a computing device, such as a desktop computer, laptop computer, server, tablet device, mobile device, smartphone, wearable device, augmented reality device, virtual reality device, network device, and / or electronic device. As Figure 1 illustrated, system 100 includes one or more memory devices, such as memory 105. Memory 105 generally represents any type or form of volatile or non-volatile storage device or medium capable of storing data and / or computer-readable instructions. Examples of memory 105 include, but are not limited to, random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), optical disc drive, cache, variations or combinations of one or more of the foregoing, and / or any other suitable memory.
[0027] As Figure 1 illustrated, example system 100 includes one or more physical processors, such as processor 110. Processor 110 generally represents any type or form of hardware-implemented processing unit capable of interpreting and / or executing computer-readable instructions. In some examples, processor 110 accesses and / or modifies data and / or instructions stored in memory 105. Examples of processor 110 include, but are not limited to, microprocessors, microcontrollers, central processing units (CPUs), graphics processing units (GPUs), field-programmable gate arrays (FPGAs) implementing soft-core processors, application-specific integrated circuits (ASICs), systems-on-a-chip (SoCs), portions of one or more of the foregoing, variations or combinations of one or more of the foregoing, and / or any other suitable physical processor.
[0028] In some embodiments, the term "instruction" refers to computer code that can be read and executed by a processor. Examples of instructions include, but are not limited to, macro-instructions (e.g., program code that requires the processor to decode into processor instructions that the processor can execute directly) and micro-operations (e.g., low-level processor instructions that are decoded from and form part of a macro-instruction).
[0029] As Figure 1As further illustrated, the processor 110 includes a control circuit 120 and a functional unit 130. The control circuit 120 includes circuitry and / or instructions for the dynamic configuration of processor sub-components such as the functional unit 130. The functional unit 130 corresponds to any component and / or sub-component of the processor 110 as further described herein. The functional unit 130 corresponds to execution units such as an arithmetic logic unit (ALU), a floating point unit (FPU), a load-store unit (LSU), etc., and / or their sub-components. Additionally, although Figure 1 a single control circuit 120 and a single functional unit 130 are illustrated, in other embodiments, the processor 110 includes multiple different iterations of each, as further described herein.
[0030] When the processor 110 operates, the processor 110 powers on and powers off the functional unit 130 as needed to improve power efficiency and reduce heat generation. For example, as the workload demand increases, the processor 110 typically powers on and utilizes the functional unit 130 accordingly. As the workload demand decreases, particularly if the functional unit 130 is not needed, the processor 110 powers off the functional unit 130, places the functional unit 130 in a low-power mode, or otherwise throttles the functional unit 130. A control scheme determines when to throttle the functional unit 130. Some control schemes rely on software such as optimized code that more efficiently utilizes the processor 110 and the functional unit 130. Other control schemes observe the workload of the processor 110 and predict workload trends to throttle the functional unit 130. Still other control schemes observe the performance of the processor 110 to determine when to throttle the functional unit 130. However, such control schemes for throttling the functional unit 130 may be relatively slow, requiring hundreds to thousands of processing cycles to react. For example, such control schemes may be limited in terms of the granularity of which components are powered on / off, such that the power on / off itself may be a relatively slow process.
[0031] The control circuit 120 resides in the hardware of the processor 110 to provide a control scheme that is more restricted to the functional unit 130 than the previously described control scheme. The physical proximity and direct coupling to the functional unit 130 reduce the response time latency to the scale of dozens of cycles. Additionally, the control circuit 120 is configured to dynamically configure the functional unit 130 and / or one or more sub-components of the functional unit 130 (e.g., via power-on / off, throttling, etc.), and further configure one or more pipelines leading to the functional unit 130. Further, the control circuit 120 is coupled to the functional units that implement parallelism in the processor 110. When the control circuit 120 powers off the functional unit 130, the processor 110 can still operate on a pipeline that is parallel to the functional unit 130. Thus, the control circuit 120 provides a level of granularity for observation and configuration that may not be achievable with a higher-level control scheme.
[0032] Figure 2 An example processor 210 corresponding to the processor 110 is illustrated. The processor 210 includes a functional unit 232 corresponding to the functional unit 130, a functional unit 234 corresponding to another iteration of the functional unit 130, and a control circuit 226 corresponding to the control circuit 120. The functional unit 232 includes a sub-component 233 and a control circuit 222 that corresponds to the control circuit 120. The functional unit 234 includes a sub-component 235 and a control circuit 224 that corresponds to the control circuit 120.
[0033] The control circuit 222 observes the functional unit 232 during operation. The physical proximity of the control circuit 222 allows for a granular level of observation of the functional unit 232. In some examples, the control circuit 222 observes the utilization of the sub-component 233 and dynamically configures the sub-component 233 accordingly. For example, if the utilization is outside the desired utilization range, the control circuit 222 can throttle the sub-component 233 and / or the functional unit 232. The desired utilization range pertains to the type of the functional unit 232 and / or the sub-component 233. For example, if the sub-component 233 corresponds to a queue, the desired utilization range includes an upper limit related to how full the queue can be before new entries must be rejected and a lower limit related to how empty the queue can be before the usefulness of the queue is reduced.
[0034] If the utilization is below the desired utilization range, the underutilization may indicate that the current workload type does not require the functional unit 232. For example, if the current workload corresponds to integer arithmetic and the functional unit 232 corresponds to an FPU, the current workload does not require the FPU or otherwise provides a minimal workload to the FPU, as indicated by the underutilization. In response, the control circuit 222 throttles the sub-component 235 and / or the functional unit 232 to reduce the power consumption of the processor 210.
[0035] If the utilization rate is higher than the desired utilization rate range, overutilization may indicate that the functional unit 232 cannot execute the current workload. The control circuit 222 throttles the sub-component 233 to reduce the workload input to the functional unit 232. For example, the sub-component 233 corresponds to a pipeline that inputs the workload to the functional unit 232. The control circuit 222 throttles or shuts down the pipeline to reduce the workload burden on the functional unit 232. In some examples, the control circuit 222 queues up portions of the workload and bursts the queued workload once the functional unit 232 returns to the desired utilization rate.
[0036] In other examples, the control circuit 222 takes preventive and / or predictive actions based on detected trigger actions. For example, the control circuit 222 detects a first stall event and a second stall event in the functional unit 234, which indicates a high probability of a third stall event. To prevent the third stall event, the control circuit 222 throttles the sub-component 233 accordingly.
[0037] When the control circuit 222 detects that the utilization rate is within the desired utilization rate range, the control circuit 222 unthrottles the sub-component 233 or otherwise reverses any previously executed throttling actions. Due to the available latency caused by physical proximity, the control circuit 222 is able to quickly respond to such changes detected in the utilization rate of the sub-component 233 and / or the functional unit 232.
[0038] In some examples, the control circuit 224 dynamically configures the functional unit 234 and / or the sub-component 235, similar to the control circuit 222 described above, and in other examples, the control circuit is customized based on the type of the functional unit 234. For example, the control circuit 224 observes and / or responds to different aspects of the utilization rate and / or workload of the functional unit 234 and / or the sub-component 235.
[0039] In some embodiments, control circuit 224 coordinates with control circuit 222. For example, control circuit 226 is coupled to control circuit 222 and control circuit 224 to facilitate coordination. Control circuit 226 detects the utilization trends of functional units 232 and 234 based on feedback from control circuit 222 and / or control circuit 224. If the utilization tends to be outside the desired utilization range (e.g., if the utilization is outside the desired utilization range at an increasing rate), then control circuit 226 instructs control circuit 222 and / or control circuit 224 to be more aggressive when throttling their respective functional units. If the utilization tends to be within the desired utilization range (e.g., if the utilization is outside the desired utilization range at a decreasing rate), then control circuit 226 instructs control circuit 222 and / or control circuit 224 to be less aggressive when throttling their respective functional units.
[0040] Additionally and / or alternatively, control circuit 226 instructs one of control circuit 222 and control circuit 224 to act in response to the other. For example, if control circuit 222 throttles sub-component 233, then control circuit 226 may instruct control circuit 224 to also throttle sub-component 235. Further, control circuit 226 allows for smooth throttling and un-throttling of functional units and sub-components. Powering on multiple functional units simultaneously may cause an undesired power spike in processor 210. Control circuit 226 coordinates the staggered powering on of functional units to prevent such power spikes. Similarly, control circuit 226 coordinates the staggered throttling of functional units to prevent an excessive voltage drop in processor 210. Thus, control circuit 226 provides additional information to control circuit 222 and / or control circuit 224 for dynamically configuring their respective functional units.
[0041] Figure 3 An example multi-core processor 310 corresponding to processor 110 is illustrated. Processor 310 includes a core 332 corresponding to functional unit 130, a core 334 that may correspond to another iteration of functional unit 130, and a control circuit 326 corresponding to control circuit 120. Core 332 includes a control circuit 322 that corresponds to control circuit 120. Core 334 includes a control circuit 324 that corresponds to control circuit 120.
[0042] In some examples, dynamic configuration allows throttling of one or more cores (e.g., core 332 and / or core 334) of a multi-core processor (e.g., processor 310). Control circuit 322 throttles core 332 and / or its sub-components based on observing the utilization of core 332 and / or the workload of core 332, as described herein. Similarly, control circuit 324 throttles core 334 and / or its sub-components based on observing the utilization of core 334 and / or the workload of core 334, as described herein. As described herein, control circuit 326 coordinates the dynamic configuration between control circuit 322 and control circuit 324.
[0043] The physical proximity of control circuit 322 to core 332 and the physical proximity of control circuit 324 to core 334 advantageously allows for a relatively low latency when dynamically configuring the respective cores. Additionally, the control circuits allow a large core to be virtualized into smaller cores that are transparent to software and / or the operating system (OS) (e.g., with respect to power consumption and heat generation). For example, when the full performance of core 332 is not needed, control circuit 322 throttles one or more sub-components of core 332. Alternatively, if core 334 can adequately handle the current workload, control circuit 322 and / or control circuit 326 throttle core 332 itself. In some examples, software and / or the OS may signal or request a reduction in the performance of processor 310.
[0044] Figure 4 Illustrated is a simplified example dual pipeline 400 (e.g., CPU core front-end subsystem) of a processor (such as processor 110) corresponding to a micro-operation cache 430 to a micro-operation queue 434 of a functional unit 130. Figure 4 Includes a first pipeline 433A from the micro-operation cache 430 to the micro-operation queue 434 and a second pipeline 433B from the micro-operation cache 430 to the micro-operation queue 434. Figure 4 Also includes a control circuit 420 corresponding to control circuit 120, a cache 415 (e.g., L1 cache or other cache of the processor / core, such as L2, L3, etc.), an address translation 431A, an address translation 431B, a decoder 432A, a decoder 432B, and a dispatcher 436.
[0045] The instruction cycle of a processor includes a fetch stage, a decode stage, and an execute stage. During the fetch stage, the next encoded instruction is fetched from memory. During the decode stage, the instruction is decoded into micro-operations. During the execute stage, the decoded micro-operations are executed. Figure 4 Illustrated are the components used during the decode stage.
[0046] During the decode stage, the processor fetches instructions from cache 415 for decoder 432A and / or decoder 432B to decode. Fetching instructions for the requested memory address requires translation to fetch from cache 415. Each decoder uses a corresponding address translation unit (e.g., address translation 431A for decoder 432A and address translation 431B for decoder 432B). After fetching the instructions, the decoder (e.g., decoder 432A and / or decoder 432B) decodes the instructions. The decoded instructions (e.g., micro-operations) are stored in a micro-operation queue (such as micro-operation queue 434) until the processor is ready to execute more micro-operations. The dispatcher (such as dispatcher 436) selects micro-operations from the micro-operation queue 434 that are ready to be executed.
[0047] During the instruction cycle, particularly during the decode stage, the process of decoding instructions can be relatively time-consuming. To reduce the time required to decode instructions, certain instructions can be pre-decoded and stored in a micro-operation cache (such as micro-operation cache 430). For example, certain instructions predicted to be repeatedly accessed are pre-decoded and stored in the micro-operation cache 430. The decoded instructions are provided to the micro-operation queue 434 from the micro-operation cache 430 (e.g., via the first pipeline 433A and / or the second pipeline 433B) rather than directly from decoder 432A and / or decoder 432B.
[0048] In some examples, the processor has a high enough throughput (e.g., a large micro-operation cache and / or a large micro-operation queue) such that the micro-operation cache 430 supports receiving instructions from more than one instruction pipeline or workload source. Figure 4 A dual pipeline is illustrated where two independent pipelines (e.g., the first pipeline 433A and the second pipeline 433B) access and share the micro-operation cache 430 (although in other examples, data / requests can be fed to the micro-operation queue 434 from multiple workload sources such as decoder 432A and / or address translation 431A and decoder 432B and / or address translation 431B).
[0049] However, in some scenarios, the micro-operation cache 430 is not being optimally utilized. In some examples, the dual pipeline has led to over-utilization of the micro-operation queue 434. The micro-operation queue 434 does not have enough free entries such that the micro-operation cache 430 can throttle the throughput. For example, the micro-operation cache 430 can immediately accept a read request from address translation 431A (alternatively the first pipeline 433A), but delays servicing a request from address translation 431B (alternatively the second pipeline 433B). Operating on the request from address translation 431B may consume power unnecessarily under these conditions.
[0050] In some examples, control circuit 420 is configured to detect and respond to these conditions. Control circuit 420 is coupled to micro-operation cache 430 and / or micro-operation queue 434 to observe its workload. Control circuit 420 observes the workload over a period of time (such as 30 cycles). Control circuit 420 may observe that the utilization of micro-operation queue 434 is higher than the desired range, which, in some examples, is based on the workload outputs of first pipeline 433A and second pipeline 433B, and in other examples, additionally or alternatively based on workload outputs from workload sources such as address translation 431A, address translation 431B, decoder 432A, decoder 432B, etc. For example, each of decoder 432A and decoder 432B outputs 6 entries per cycle. Control circuit 420 may (e.g., based on tokens) observe that the number of idle entries during the time period has remained low (such as 6 to 10 idle entries), which would not support the outputs of the two pipelines (e.g., 12 entries). During this time period, micro-operation queue 434 has rejected new entries from decoder 432B, resulting in a stall in second pipeline 433B. Because micro-operation queue 434 cannot support both first pipeline 433A and second pipeline 433B, control circuit 420 throttles second pipeline 433B to save power. Since second pipeline 433B has stalled, throttling second pipeline 433B does not result in a significant reduction in performance or throughput. When control circuit 420 detects that the utilization is within the desired range (e.g., 12 or more idle entries), micro-operation queue 434 may be able to support both first pipeline 433A and second pipeline 433B. Control circuit 420 accordingly un-throttles second pipeline 433B (e.g., by accepting entries from decoder 432B).
[0051] Figure 5 is a flowchart of an exemplary computer-implemented method 500 for dynamically configuring a processor sub-component. Figure 5 The steps shown in Figure 1 and / or Figure 2 can be performed by any suitable computer-executable code and / or computing system, including Figure 5 and / or the systems illustrated in
[0052] As Figure 5As illustrated, at step 502, one or more of the systems described herein use control circuitry of a target sub-component of a functional unit coupled to a processor to observe the utilization of the target sub-component. For example, control circuitry 120 observes the utilization of functional unit 130 and / or its sub-components.
[0053] The systems described herein can perform step 502 in various ways. In one example, control circuitry 120 observes the workload input to the target sub-component. In some examples, the target sub-component corresponds to a micro-operation cache (e.g., micro-operation cache 430) for storing pre-decoded instructions and / or a micro-operation queue (e.g., micro-operation queue 434). The functional unit includes at least a first sub-component (e.g., decoder 432A, address translation 431A, and / or first pipeline 433A) and a second sub-component (e.g., decoder 432B, address translation 431B, and / or second pipeline 433B), where the first sub-component and the second sub-component each provide pre-decoded instructions and / or requests to the micro-operation cache and / or the micro-operation queue. At least one sub-component corresponds to at least one of the first sub-component and the second sub-component.
[0054] In some examples, control circuitry 420 observes the micro-operation queue 434 and its utilization, whether the micro-operation queue 434 is blocked by downstream stalls, and / or the rate of output / dispatch of the micro-operation queue 434 via a token available for the micro-operation cache 430 (e.g., at the token level as described herein).
[0055] In some examples, the functional unit corresponds to at least one of an arithmetic logic unit (ALU), a floating-point unit (FPU), or a load-store unit (LSU). In some examples, the processor corresponds to a multi-core processor, and the functional unit corresponds to a core of the multi-core processor (e.g., see Figure 4 ).
[0056] At step 504, one or more of the systems described herein detect that the utilization is outside the desired utilization range. For example, control circuitry 120 detects that the utilization of functional unit 130 is outside the desired utilization range. The desired utilization range is based on the type of the functional unit and / or the type of the workload as described herein.
[0057] At step 506, one or more of the systems described herein throttle at least one sub-component of the functional unit in response to the detection to reduce the power consumption of the functional unit. For example, control circuitry 120 throttles functional unit 130 and / or its sub-components.
[0058] The systems described herein may perform step 506 in various ways. In one example, throttling at least one sub-component includes placing at least one sub-component in a low power state. In some examples, throttling at least one sub-component includes reducing the output rate of at least one sub-component. In some examples, throttling at least one sub-component includes throttling the functional unit itself.
[0059] In some examples, control circuit 120 selects a particular sub-component of functional unit 130 for throttling, the particular sub-component being different from the target sub-component (e.g., the observed sub-component). In some examples, at least one sub-component of the functional unit is selected for throttling based on reducing the workload input to the target sub-component. Control circuit 120 selects a sub-component from one of a plurality of pipelines or workload sources for the micro-operation cache and / or micro-operation queue to reduce the workload input to the micro-operation cache and / or micro-operation queue. For example, control circuit 420 selects address translation 431B, decoder 432B, and / or micro-operation cache 430 (e.g., to throttle one of the first pipeline 433A and the second pipeline 433B) to reduce the workload input to micro-operation queue 434. Control circuit 420 may select a sub-component that has stalled or may select a sub-component that causes another sub-component to stall. In other examples, control circuit 420 selects the target sub-component itself (e.g., the observed micro-operation queue 434) for throttling.
[0060] In some examples, at least one sub-component of the functional unit is selected for throttling based on the workload type of the workload. For example, if the workload corresponds to a floating-point operation, control circuit 120 selects the integer unit for throttling.
[0061] In some examples, the workload corresponds to a minimum workload and at least one sub-component corresponds to the target sub-component. For example, control circuit 120 detects that the current stage of the workload of the FPU does not require a floating-point operation and throttles the FPU accordingly.
[0062] In some examples, control circuit 120 also detects that the utilization is within the desired utilization range (e.g., in response to a previous throttling of at least one sub-component), and un-throttles at least one sub-component in response to detecting that the utilization is within the desired utilization range. For example, control circuit 120 previously turned off a selected sub-component from one of a plurality of pipelines to the micro-operation cache and / or micro-operation queue. After detecting that the micro-operation cache and / or micro-operation queue is within the desired utilization range, control circuit 120 re-enables the turned-off sub-component to restore the corresponding pipeline to the micro-operation cache and / or micro-operation queue.
[0063] As Figure 4 illustrated, in some examples, a target sub-component corresponds to a micro-operation queue for storing pre-decoded instructions, and the functional unit further includes at least a first sub-component and a second sub-component, where the first sub-component and the second sub-component each provide pre-decoded instructions and / or entries to the micro-operation queue. At least one sub-component corresponds to at least one of the first sub-component and the second sub-component.
[0064] In some examples, a higher-level control circuit (e.g., control circuit 226) coupled to a control circuit (e.g., control circuit 222) detects throttling of at least one sub-component (e.g., sub-component 233) of a functional unit (e.g., functional unit 232). The higher-level control circuit is coupled to a second control circuit (e.g., control circuit 224), which is coupled to a second target sub-component (e.g., sub-component 235) of a second functional unit (e.g., functional unit 234) of the processor. The higher-level control circuit coordinates the throttling of the sub-components by the control circuit and the second control circuit.
[0065] As described herein, the present disclosure provides systems and methods for dynamically configuring processor sub-components. A program may have various stages during execution, such as integer arithmetic, floating-point arithmetic, loops, etc. The control circuit may identify the programming stage and which sub-components are required to execute the programming stage. The control circuit may power off or throttle sub-components that are not required for the identified programming stage. Because the control circuit may reside physically near the sub-components it observes and throttles, the control circuit can respond with a higher granularity and higher frequency than other techniques may allow. For example, the control circuit can observe the workload of the sub-components more directly rather than relying on feedback provided by the functional unit.
[0066] For example, a sub-component may receive inputs from multiple pipelines. The control circuit may observe the current workload of the sub-component and determine that operating multiple pipelines would not provide a performance benefit to the current workload. The control circuit may accordingly maintain a single pipeline and shut down the other pipelines (e.g., by turning off one or more sub-components along the respective pipelines) to reduce power consumption. Thus, the systems and methods described herein may advantageously allow for maximum performance of high-throughput workloads using multiple pipelines while minimizing the power consumption of low-throughput workloads.
[0067] As detailed above, the computing devices and systems described and / or illustrated herein broadly represent any type or form of computing device or system capable of executing computer-readable instructions (such as those included in the units or circuits described herein). In their most basic configuration, these computing devices each include at least one memory device and at least one physical processor.
[0068] In some examples, the term "memory device" generally refers to any type or form of volatile or non-volatile storage device or medium capable of storing data and / or computer-readable instructions. In one example, the memory device stores, loads, and / or maintains one or more of the units described herein. Examples of storage devices include, but are not limited to, random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), optical disk drive, cache, variations or combinations of one or more of the foregoing, or any other suitable memory.
[0069] In some examples, the term "physical processor" generally refers to any type or form of hardware-implemented processing unit capable of interpreting and / or executing computer-readable instructions. In one example, the physical processor accesses and / or modifies one or more of the units stored in the memory device described above. Examples of physical processors include, but are not limited to, microprocessors, microcontrollers, central processing units (CPUs), field programmable gate arrays (FPGAs) implementing soft-core processors, application specific integrated circuits (ASICs), system on a chip (SOCs), portions of one or more of the foregoing, variations or combinations of one or more of the foregoing, or any other suitable physical processor.
[0070] In some embodiments, the term "computer-readable medium" generally refers to any form of device, carrier, or medium capable of storing or carrying computer-readable instructions. Examples of computer-readable media include, but are not limited to, transmission-type media such as carrier waves, and non-transitory media such as magnetic storage media (e.g., hard disk drives, tape drives, and floppy disks), optical storage media (e.g., compact discs (CDs), digital video discs (DVDs), and Blu-ray discs), electronic storage media (e.g., solid state drives and flash media), and other distribution systems.
[0071] The order of process parameters and steps described or illustrated herein is given by way of example only and may vary as needed. For example, although the steps illustrated or described herein are shown or discussed in a particular order, these steps do not necessarily need to be performed in the order illustrated or discussed. The various exemplary methods described or illustrated herein may also omit one or more of the steps described or illustrated herein, or include additional steps other than those disclosed.
[0072] The foregoing description has been provided so that others skilled in the art may best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or limited to any precise form. Many modifications and variations are possible without departing from the spirit and scope of the disclosure. The embodiments disclosed herein should be considered illustrative in all respects and not restrictive. In determining the scope of the disclosure, reference should be made to the appended claims and their equivalents.
[0073] Unless otherwise indicated, the terms "connected to" and "coupled to" (and their derivatives) as used in the specification and claims will be taken to allow both direct and indirect (i.e., via other elements or components) connections. Further, the term "a" or "an" as used in the specification and claims will be taken to mean "at least one". Finally, for convenience, the terms "comprising" and "having" (and their derivatives) as used in the specification and claims may be interchanged with the word "including" and have the same meaning.
Claims
1. A device, the device comprising: Functional units of a processor, the functional units including at least one sub-component, the at least one sub-component including a target sub-component; And A control circuit, the control circuit being coupled to the target sub-component and configured to: Observe the utilization rate of the target sub-component; Detect that the utilization rate is outside a desired utilization rate range; And In response to the detection, throttle at least one sub-component of the functional unit to reduce the power consumption of the functional unit.
2. The device according to claim 1, wherein the control circuit is further configured to observe the workload input to the target sub-component.
3. The device according to claim 2, wherein the at least one sub-component of the functional unit is selected for throttling based on reducing the workload input to the target sub-component.
4. The device according to claim 2, wherein the at least one sub-component of the functional unit is selected for throttling based on the workload type of the workload.
5. The device according to claim 2, wherein the workload corresponds to a minimum workload, and the at least one sub-component corresponds to the target sub-component.
6. The device according to claim 1, wherein throttling the at least one sub-component includes placing the at least one sub-component in a low power state.
7. The device according to claim 1, wherein throttling the at least one sub-component includes throttling the functional unit.
8. The device according to claim 1, wherein the control circuit is further configured to: Detect that the utilization rate is within the desired utilization rate range; and In response to detecting that the utilization rate is within the desired utilization rate range, un-throttle the at least one sub-component.
9. The device according to claim 1, wherein: The target sub-component corresponds to a micro-operation queue for storing pre-decoded instructions; The functional unit further includes a first sub-component and a second sub-component, the first sub-component and the second sub-component each providing pre-decoded instructions to the micro-operation queue; and The at least one sub-component corresponds to at least one of the first sub-component and the second sub-component.
10. The device according to claim 1, wherein the functional unit corresponds to at least one of an arithmetic logic unit (ALU), a floating point unit (FPU), or a load-store unit (LSU).
11. The device according to claim 1, wherein the processor corresponds to a multi-core processor, and the functional unit corresponds to a core of the multi-core processor.
12. The device according to claim 1, the device further comprising: A second functional unit of the processor, the second functional unit including a second target sub-component; A second control circuit, the second control circuit being coupled to the second target sub-component; And A higher-level control circuit, the higher-level control circuit being coupled to the control circuit and the second control circuit; And Wherein the higher-level control circuit is configured to: Detect the throttling of the at least one sub-component of the functional unit; and Coordinate throttling of the sub-components by the control circuit and the second control circuit.
13. A method, the method comprising: Based on a workload in a target sub-component input to a functional unit of a processor, use a control circuit coupled to the target sub-component to observe the utilization of the target sub-component; Detect that the utilization is outside a desired utilization range of the workload; And Throttle at least one sub-component of the functional unit to reduce power consumption of the functional unit.
14. The method according to claim 13, wherein the at least one sub-component of the functional unit is selected for throttling based on reducing the workload input to the target sub-component.
15. The method according to claim 13, wherein the workload corresponds to a minimum workload, and the at least one sub-component corresponds to the target sub-component.
16. The method according to claim 13, wherein: The target sub-component corresponds to a micro-operation queue for storing pre-decoded instructions; The functional unit further includes at least a first sub-component and a second sub-component, and the first sub-component and the second sub-component each provide pre-decoded instructions to the micro-operation queue; and The at least one sub-component corresponds to at least one of the first sub-component and the second sub-component.
17. The method according to claim 13, the method further comprising: Use a higher-level control circuit coupled to the control circuit to detect the throttling of the at least one sub-component of the functional unit, wherein the higher-level control circuit is further coupled to a second control circuit, and the second control circuit is coupled to a second target sub-component of a second functional unit of the processor; And Use the higher-level control circuit to coordinate throttling of sub-components by the control circuit and the second control circuit.
18. A system, the system comprising: Physical memory; And At least one physical processor, the at least one physical processor comprising: A functional unit, the functional unit including at least one sub-component, the at least one sub-component including a target sub-component; and A control circuit, the control circuit coupled to the target sub-component and configured to: Observe the workload and utilization of the target sub-component; Detect that the utilization is outside a desired utilization range of the workload; and Throttle at least one sub-component of the functional unit to reduce power consumption of the physical processor.
19. The system according to claim 18, wherein the at least one sub-component of the functional unit is selected for throttling based on reducing the workload input to the target sub-component.
20. The system according to claim 18, wherein: The target sub-component corresponds to a micro-operation queue for storing pre-decoded instructions; The functional unit further includes at least a first sub-component and a second sub-component, and the first sub-component and the second sub-component each provide pre-decoded instructions to the micro-operation queue; and The at least one sub-component corresponds to at least one of the first sub-component and the second sub-component.
Citation Information
Patent Citations
Power consumption controller, a power consumption control system, a power consumption control method and a program thereof
US20120226923A1
Reducing power consumption of a compute node experiencing a bottleneck
US20170168540A1
Hardware processors and methods for extended microcode patching
US20200210196A1
Performance and power optimization via block oriented performance measurement and control
US6895520B1