Balanced throughput of partitions replicated in the presence of a non-functional computing unit

By using a power manager to adjust power domains with static and dynamic scaling factors, the throughput imbalance caused by manufacturing defects in replicated partitions is addressed, ensuring efficient task completion across integrated circuits.

JP2025522497APending Publication Date: 2025-07-15ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024574604
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-20
Filing Date
2023-05-03
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Integrated circuits with replicated partitions suffer from imbalanced throughput due to manufacturing defects, leading to reduced performance and inefficiency, particularly in highly parallel data microarchitectures.

Method used

A power manager generates static and dynamic scaling factors based on the number of operating computing units to adjust the operating parameters of individual power domains, ensuring balanced throughput across replicated partitions.

Benefits of technology

The solution ensures that partitions with fewer functional units complete tasks at similar times to other partitions, maintaining overall processing unit performance and reducing dependence on lock-step execution, thus enhancing throughput efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522497000001_ABST
    Figure 2025522497000001_ABST
Patent Text Reader

Abstract

Despite a loss of functionality due to manufacturing defects, an apparatus and method are provided for efficiently managing balanced performance among replicated partitions of an integrated circuit. A processing unit includes at least two replicated partitions, each partition being assigned to operating parameters of a respective power domain. The partitions include a plurality of computing units. The computing units include a plurality of execution lanes. Due to various types of manufacturing defects, one or more of the partitions of the processing unit have less than a predetermined number of operating computing units. To balance the throughput of the plurality of partitions, a power manager generates both static and dynamic scaling factors based at least on the corresponding number of operating computing units. Using these scaling factors, the power manager adjusts the operating parameters of the power domains of the partitions relative to each other.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] (Description of Related Art) Both planar transistors and non-planar transistors are manufactured for use within integrated circuits in semiconductor chips. In system packaging for integrating multiple types of integrated circuits, there are various options for placing processing circuits. Some examples are system-on-a-chip (SOC), multi-chip module (MCM), and system-in-package (SiP). Mobile devices, desktop systems, and servers use these packages. Regardless of the system packaging choice, during the assembly of semiconductor chips, one or more semiconductor dies (or dice) are placed on a single substrate or package, and these dies are susceptible to the effects of electrostatic discharge events. An electrostatic discharge event provides an accidental charge that can cause a current density exceeding a safe threshold to flow through metal wires and transistors (devices). Thus, one or more processing units and other functional blocks on the die can fail, thereby reducing the manufacturing yield.

[0002] Prior to packaging and during semiconductor manufacturing process steps for the die, one or more processing units and other functional blocks on the die can also fail. These failures result from manufacturing defects that unintentionally cause open circuits, stuck-at faults, etc. During the testing of the die and later during the testing of the package, some defect is found. In some cases, the defect occurs in a functional block that is replicated in a partition of the processing unit. A particular functional block is no longer operable, and the overall throughput of the processing unit decreases, but the partitions within the processing unit remain operable.

[0003] By using a fuse array and a fuse read-only memory (ROM), access to specific functional blocks without defects within a partition can be restricted on a die. The semiconductor die is still used, but the resulting package is placed in a reduced performance category or bin. However, in the case of a die using a highly parallel data microarchitecture, a partition that uses all of its replicated functional blocks will complete its task before another partition that uses fewer replicated functional blocks. There is an imbalance in throughput between partitions, which further degrades performance. In some cases, packages with reduced performance still function, but are not acceptable because there is a high demand in the market for running certain applications at a relatively high minimum performance level.

[0004] In view of the above, there is a desire for an efficient method and apparatus for managing balanced performance among replicated partitions of an integrated circuit despite a loss of functionality due to manufacturing defects.

Brief Description of the Drawings

[0005]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Best Mode for Carrying Out the Invention

[0006] Although the present invention has room for various modifications and alternative forms, specific embodiments are shown in the drawings by way of example and are described in detail herein. However, the drawings and their detailed description are not intended to limit the present invention to the particular forms disclosed, but on the contrary, the present invention is intended to cover all modifications, equivalents, and alternatives included within the scope of the present invention as defined by the appended claims.

[0007] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, those skilled in the art should recognize that the present invention may be practiced without these specific details. In some instances, well-known circuits, structures, and techniques have not been shown in detail in order to avoid obscuring the present invention. Further, it should be understood that the elements shown in the figures are not necessarily drawn to scale for the sake of simplicity and clarity of the description. For example, the dimensions of some elements are exaggerated relative to other elements.

[0008] Apparatus and methods are contemplated for efficiently managing balanced performance across replicated partitions of an integrated circuit despite loss of functionality due to manufacturing defects. In some embodiments, a processing unit includes at least two replicated partitions. As used herein, the term "replicated" is used to refer to identical instantiations of hardware such as a particular functional block, a particular type of unit, or another particular type of circuit. For example, a "replicated partition" refers to two or more partitions where each partition is an identical instantiation of a particular type of partition. In one embodiment, each of the partitions is a shader engine of a graphics processing unit (GPU). Similarly, a "replicated computational unit" refers to two or more computational (or compute) units where each computational unit is an instantiation of a particular type of computational unit. In one embodiment, a particular type of computational unit includes a plurality of execution lanes that support a parallel data microarchitecture for processing a workload. Thus, in one embodiment, each of the replicated (instantiated) partitions is a shader engine of a GPU, and each of the shader engines includes a plurality of replicated (instantiated) computational units.

[0009] In various embodiments, a processing unit includes at least two replicated (instantiated) partitions, and each partition is assigned to operating parameters of a respective power domain. Each of the power domains includes operating parameters such as at least an operating supply voltage and an operating clock frequency. Also, each of the power domains includes control signals for enabling and disabling a connection to a clock generation circuit and a power reference. Thus, the at least two replicated partitions do not share the same clock generation circuit and the same connection to a power reference.

[0010] Due to various types of manufacturing defects, one or more of the partitions of the processing unit have less than a predetermined number of operation calculation units. As used herein, an operation calculation unit is also referred to as a functional calculation unit. An operation calculation unit (functional calculation unit) can successfully process a task because there is no manufacturing defect. In one example, the processing unit includes four partitions, and each partition has eight calculation units. However, due to manufacturing defects, any one of the four partitions has seven operation calculation units instead of the predetermined number of eight. To balance the throughput of multiple partitions, the power manager generates corresponding static scaling factors for each of the multiple partitions relative to each other.

[0011] The power manager generates a static scaling factor for each of the plurality of replicated partitions based on the number of corresponding operation calculation units. The power manager uses each power domain to generate a static scaling factor to balance the throughput of the plurality of replicated partitions in each partition. In other words, for any one of the plurality of partitions, the power manager selects the individual operation parameters for that partition using the corresponding static scaling factor. The difference in throughput between any two partitions is less than a threshold value. Using at least the static scaling factor, a first partition having six out of eight functional calculation units achieves approximately the same throughput as a second partition having eight out of eight functional calculation units. The difference in throughput between the first partition and the second partition is less than the throughput threshold value. Therefore, the difference in task completion time between the first partition and the second partition is less than the time threshold value. The static scaling factor of the first partition having six out of eight functional calculation units causes the power manager to select the operation parameters of the first power domain that provide a higher transistor switching speed than the operation parameters of the second power domain used by the second partition having eight out of eight functional calculation units.

[0012] If a task with a workload is executed by a plurality of replicated partitions in a lock-step format and any one partition completes significantly later than other partitions, the overall throughput of the processing unit decreases. If a checkpoint is used to synchronize the execution of a task with a workload across a plurality of replicated partitions and any one partition completes significantly later than other partitions, the overall throughput of the processing unit decreases. For example, if the first partition has seven out of eight functional computing units instead of eight out of eight functional computing units and each partition uses the same power domain, the first partition will complete later than other partitions that have eight out of eight functional computing units. Accordingly, the throughput of the entire processing unit decreases. However, if the first partition uses the operation parameters of an individual power domain as described above, the first partition will have the same or approximately the same completion time, so the dependence on lock-step execution or synchronization checkpoints does not degrade the performance of the processing unit. The difference in the completion time of tasks between the first partition and other partitions is less than a time threshold. In addition, the power manager can dynamically adjust the operation parameters of individual power domains at the granularity of partitions rather than at the granularity of the entire processing unit. This dynamic adjustment is based on performance metrics monitored during the processing of the workload. For example, the power manager receives performance metrics from performance counters distributed across the computing units. Further details for efficiently managing balanced performance among replicated partitions of an integrated circuit, despite functional loss due to manufacturing defects, are provided in the following description.

[0013] Referring to FIG. 1, a schematic block diagram of an apparatus 100 for managing balanced performance between replicated partitions of an integrated circuit is shown, despite a loss of functionality due to manufacturing defects. In the illustrated embodiment, apparatus 100 includes a power manager 170 and at least two partitions, such as partition 110 and partition 150, each of which is assigned by power manager 170 to a respective power domain. Each of the power domains includes at least operating parameters such as at least an operating supply voltage and an operating clock frequency. Also, each of the power domains includes control signals for enabling and disabling connections to a clock generation circuit and a power reference. Only two partitions 110 and 150 are shown, but other numbers of partitions used by apparatus 100 are possible and contemplated, and that number is based on design requirements. Other components of apparatus 100 are not shown for ease of explanation. For example, a memory controller, one or more input / output (I / O) interface units, an interrupt controller, one or more phased locked loop (PLL) or other clock generation circuits, one or more levels of cache memory subsystems, and various other functional blocks are not shown, but they may be used by apparatus 100.

[0014] In some embodiments, the functionality of apparatus 100 is included as components on a single die, such as a single integrated circuit. In one embodiment, the functionality of apparatus 100 is included as one of multiple dies on a system on chip (SOC). In various embodiments, apparatus 100 is used in a desktop, portable computer, mobile device, server, peripheral device, etc. Apparatus 100 can communicate with an external general-purpose central processing unit (CPU) that includes circuitry for executing instructions according to a given general-purpose instruction set architecture (ISA). Also, apparatus 100 can communicate with various other external circuits, such as one or more of a digital signal processor (DSP), a display controller, various application specific integrated circuits (ASICs), a multimedia engine, etc.

[0015] Power manager 170 decreases (or increases) power consumption when apparatus 100 is operating above (or below) a threshold limit. In some embodiments, power manager 170 selects a respective power management state for each of partitions 110 and 150. As used herein, "power management state" is either any one of a plurality of "P states" or any one of a plurality of power performance states that includes a set of operating parameters such as operating clock frequency and operating supply voltage. In various embodiments, apparatus 100 uses a parallel data microarchitecture that provides high instruction throughput for compute-intensive tasks. In one embodiment, apparatus 100 uses one or more processor cores having a relatively wide single instruction multiple data (SIMD) microarchitecture to achieve high throughput in high data parallel applications. Each object is processed independently of other objects, but the same operation sequence is used.

[0016] In one embodiment, device 100 is a graphics processing unit (GPU). Modern GPUs are efficient for data parallel computing found within the loops of applications such as applications for manipulating and displaying computer graphics, molecular dynamics simulations, financial calculations, etc. The highly parallel structure of the GPU makes the GPU more effective than a general-purpose central processing unit (CPU) for a wide range of complex algorithms. In various embodiments, partitions 110 and 150 are partitions replicated within device 100, so partition 150 includes the same components as partition 110. In one embodiment, each of partitions 110 and 150 is a shader engine of the GPU, and each of the shader engines includes a plurality of computing units 140A - 140C for processing data parallel applications such as graphics shader tasks.

[0017] Each of computing units 140A - 140C includes a plurality of lanes 142. Each lane is also referred to as a SIMD unit or SIMD lane. In some embodiments, lanes 142 operate in lockstep. In other embodiments, the processing of tasks by lanes 142 uses synchronization checkpoints. In various embodiments, the data flow within each of lanes 142 is pipelined. Pipeline registers are used to store intermediate results, and the circuitry for the arithmetic logic unit (ALU) performs integer arithmetic, floating-point arithmetic, Boolean logic operations, branch condition comparisons, etc. These components are not shown for ease of explanation. Each of the computing units within a given row across lanes 142 is the same computing unit. Each of these computing units operates on different data associated with different threads but with the same instruction.

[0018] As shown, each of the computing units 140A - 140C includes a respective register file 144, a local data store 146, and a local cache memory 148. In some embodiments, the local data store 146 is shared among the lanes 142 within each of the computing units 140A - 140C. In other embodiments, the local data store is shared among the computing units 140A - 140C. Thus, one or more of the lanes 142 within the computing unit 140A can share result data with one or more of the lanes 142 within the computing unit 140B based on the operating mode. The high parallelism provided by the hardware of the computing units 140A - 140C is used to render multiple pixels simultaneously, but it is also possible to process multiple data elements of science, medical, financial, encryption / decryption, and other computations simultaneously. For example, the partition 110 is used for real-time data processing such as rendering of multiple pixels, image blending, pixel shading, vertex shading, and geometry shading. The circuitry of a controller (not shown) receives tasks via a memory controller (not shown). In some embodiments, the controller is the command processor of the GPU, and the tasks are a sequence of commands (instructions) of function calls of an application.

[0019] The partition 110 receives the operating parameters 160 of the first power domain from the power manager 170, and the computing units 140A - 140C use the operating parameters 160 to process tasks. The partition 150 receives the operating parameters 164 of the second power domain from the power manager 170. The first power domain and the second power domain can be, and are contemplated to be, different power domains. Thus, the power manager 170 can dynamically adjust the power domains at the granularity of partitions such as the partitions 110 and 150, rather than at the granularity of the entire device 100. In some embodiments, the power manager 170 is an integrated controller as shown, but in other embodiments, the power manager 170 is an external unit.

[0020] Due to various types of manufacturing defects, one or more of partitions 110 and 150 have less than a predetermined number of operating computing units, such as computing units 140A - 140C. In one example, each of partitions 110 and 150 includes eight computing units. However, due to a manufacturing defect, partition 110 has seven operating computing units instead of the predetermined number of eight. To balance the throughput of partitions 110 and 150, power manager 170 generates corresponding static scaling factors for each of the plurality of partitions relative to each other.

[0021] Power manager 170 generates static scaling factors for each of partitions 110 and 150 based on the number of corresponding operating computing units. For partition 110 having seven operating computing units out of eight computing units 140A - 140C, power manager 170 generates a static scaling factor for partition 110 indicating that partition 110 uses a set of operating parameters of a power domain that provides a higher transistor switching speed than another set of operating parameters of another power domain used by partition 150 having eight operating computing units out of eight computing units 140A - 140C. Thus, the difference in completion times between partition 110 and partition 150 is reduced, especially when compared to the case where each of partition 110 and partition 150 uses the same set of operating parameters of the same power domain. Despite the fact that the number of operating computing units of computing units 140A - 140C in partition 110 is less than that in partition 150, when partition 110 uses the operating parameters of the power domain based on the static scaling factor, the dependence on lockstep execution or synchronous checkpoints does not degrade the performance of device 100. For example, partitions 110 and 150 have the same or approximately the same completion time. The difference in completion times of tasks in partitions 110 and 150 is less than a time threshold.

[0022] In addition, the power manager 170 can dynamically adjust the power domains of partitions 110 and 150 based on the number of corresponding operation calculation units monitored during the processing of the workload and performance metrics 162 and 166. For example, the power manager 170 receives performance metrics 162 and 166 from performance counters such as performance counter 149 distributed across calculation units 140A-140C and other components (not shown) of partitions 110 and 150. In some embodiments, the collected data includes a predetermined sampled signal. The switching of the sampled signal indicates the amount of switched capacitance. Examples of signals selected for sampling include a clock gating enable signal, a bus driver enable signal, a mismatch in a content-addressable memory (CAM), a CAM word-line (WL) driver, and the like. Also, the collected data can include data indicating the performance or throughput of each of partitions 110 and 150, such as the number of retired instructions, the number of cache accesses, the monitored latency of cache accesses, the number of cache hits, the count of issued instructions or issued threads, and the like.

[0023] In one embodiment, the power manager 170 collects data to characterize the power consumption and throughput of partitions 110 and 150 during a specific sampling interval. If one or more of the estimated power consumption and estimated throughput of partitions 110 and 150 change significantly, the power manager 170 updates the operation parameters 160 and 164 of the individual power domains of partitions 110 and 150. The operation parameters 160 and 164 can also be referred to as a set of operation parameters 160 and 164. The updated values of the operation parameters 160 and 164 cause partitions 110 and 150 to achieve approximately the same throughput. The difference in throughput between partitions 110 and 150 is less than a throughput threshold. Therefore, the difference in task completion time for partitions 110 and 150 is less than a time threshold.

[0024] Referring to FIG. 2, a schematic block diagram of a method 200 for efficiently managing balanced performance between replicated partitions of an integrated circuit is shown, despite a loss of functionality due to manufacturing defects. For purposes of explanation, the steps in this embodiment (as well as in FIG. 4) are presented in order. However, in other embodiments, some steps are performed in a different order than shown, some steps are performed simultaneously, some steps are combined with other steps, and some steps do not exist.

[0025] The parallel data processing unit includes at least two partitions, and each partition is assigned to a respective power domain. Each of the power domains includes operating parameters such as at least an operating power supply voltage and an operating clock frequency. In some embodiments, the circuitry of the first partition includes a plurality of computing units each having a plurality of execution lanes. Hardware such as the circuitry of the power manager of the parallel data processing unit determines the same throughput level expected for each of the plurality of partitions during the execution of the workload (block 202). The power manager determines the number of computing units operable in each of the plurality of partitions (block 204). As used herein, a computing unit is considered to be "operable" if the computing unit can process a task. An "operable" computing unit is also referred to as an "operating" computing unit or a "functional" computing unit. In contrast, a "non-functional" computing unit, also referred to as a "non-operating" computing unit, cannot process a task. For example, a non-operating computing unit has one or more manufacturing defects that prevent it from processing a task. In one embodiment, a fuse array or a fuse ROM is accessed to identify which computing units are non-operating computing units, thereby determining the number of operating computing units operable in each of the plurality of partitions.

[0026] The power manager generates static scaling factors for each of a plurality of partitions relative to one another based on the number of corresponding operation calculation units (block 206). The power manager converts the scaling factors into specific operation parameters of individual power domains for the plurality of partitions to achieve a balanced throughput level (block 208). The power manager assigns the operation parameters to the plurality of partitions (block 210). The parallel data processing units process the workload using the assigned operation parameters (block 212).

[0027] Referring to FIG. 3, a schematic block diagram of a power manager 300 that manages balanced performance among replicated partitions of an integrated circuit is shown, despite a loss of functionality due to manufacturing defects. As shown, the power manager 300 includes a table 310 and a control unit 330. The control unit 330 includes a plurality of components 332-338 that are used to generate operation parameters 350 for a plurality of power domains that are transmitted to a plurality of replicated partitions. The table 310 includes a plurality of table entries (or entries), and each entry stores information in a plurality of fields, such as at least fields 312-318.

[0028] Table 310 is implemented using any of a flip-flop circuit, a random access memory (RAM), a content addressable memory (CAM), etc. Although it is shown that specific information is stored in fields 312 to 318 in a specific consecutive order, in other embodiments, different orders are used and different numbers and types of information are stored. As shown, field 312 stores a partition identifier (ID) that designates a specific partition among a plurality of partitions used in a parallel data processing unit. In one embodiment, the partition is any one of a plurality of shader engines of a GPU. Field 314 stores the number of computing units within the identified partition. In one embodiment, this information is found in a fuse read-only memory (ROM) that is set after the manufacture and testing of the semiconductor package.

[0029] Field 316 stores the static scaling factor for the identified partition. The static scaling factor is set based on the number of corresponding operation calculation units among multiple partitions. Therefore, the value of the static scaling factor is a relational value. In various embodiments, this value is set during or immediately after the boot-up operation and does not change from this value. Field 318 stores the dynamic scaling factor for the identified partition. This value is based on both the number of corresponding operation calculation units among multiple partitions and the measured performance metric. This dynamic scaling factor changes over time. For example, the dynamic scaling factor is updated based on one or more of determining that a specific time interval has elapsed, determining that a new workload has been assigned for execution, and the power manager determining that the throughput level has changed beyond a threshold amount. In some embodiments, the corresponding weight values are associated with each of the static scaling factor and the dynamic scaling factor. In one embodiment, each of the multiple replicated partitions has its own pair of weight values. These weight values change over time and it is possible and contemplated that either the static scaling factor or the dynamic scaling factor may significantly influence the selection of the operating parameters over time.

[0030] Control unit 330 receives at least activity level measurement values or usage measurement values 320 representing data from multiple partitions. One example is the sampled signal as described above. Control unit 330 receives sensor input 322 representing measured temperature values from analog or digital thermal sensors arranged across the die. Control unit 330 receives performance metric 324 representing values read from performance counters arranged across multiple partitions. Also, control unit 330 receives data from table 310 and control unit 330 can update the information stored in table 310.

[0031] The power reporting unit 332 calculates the power value from the usage measurement value 320. Also, the power reporting unit 332 calculates the leakage power value included in the total power value. The leakage power value depends on the calculated temperature. In some embodiments, the power reporting unit 332 associates the total number of power credits of the parallel data processing unit with the thermal design power (TDP) value of the processing unit. The power reporting unit 332 allocates a predetermined number of power credits to each of the partitions of the parallel data processing unit. The sum of the associated power credits is equal to the total number of power credits of the die 202. The power reporting unit 332 adjusts the number of power credits for each of the external partitions over time.

[0032] The calculated temperature is determined by the temperature reporting unit 334 and uses the worst-case ambient temperature value. In an embodiment, if the sensor measured temperature is significantly different from the calculated temperature, the calculated power value does not change. The balance throughput manager 336 (or manager 336) has the functions of the balance throughput manager 174 (FIG. 1). For example, the manager 336 determines the dynamic scaling factor stored in the field 318 of the table 310 for a plurality of partitions. The manager 336 calculates these dynamic scaling factors based on the corresponding number of operation calculation units and the performance metric 324. The manager 336 determines when a performance bottleneck occurs in any of the plurality of partitions during the execution of the workload and recalculates the dynamic scaling factor used by the operation parameter selector 338 to generate a new power domain for the plurality of partitions.

[0033] The operation parameter selector 338 receives temperature-related values from the temperature reporting unit 334, receives both the power values calculated for each partition and the current power credit number and the updated power credit number from the power reporting unit 332, and receives the updated dynamic scaling factor from the manager 336. Based on these inputs, the operation parameter selector 338 generates updated operation parameters for individual power domains for a plurality of partitions. The updated operation parameters include the operation parameters 350. The operation parameter selector 338 receives various input values, but the performance metric 324, the predetermined static scaling factor stored in the field 314, and the updated dynamic scaling factor from the manager 336 are values that adjust the operation parameters 350 so that the partitions have approximately equal throughput.

[0034] Referring to FIG. 4, a schematic block diagram of a method 400 for efficiently managing balanced performance among replicated partitions of an integrated circuit is shown, despite a loss of functionality due to manufacturing defects. A plurality of partitions of a parallel data processing unit process a workload using corresponding assigned operation parameters of separate power domains (block 402). In one embodiment, each of the partitions is a shader engine of a graphics processing unit (GPU), and each of the shader engines includes a plurality of computing units. Hardware such as a circuit of a power manager of the parallel data processing unit monitors performance metrics of the plurality of partitions (block 404). For example, when a specific sampling interval elapses, values stored in performance counters arranged across the plurality of partitions are read and reported to the power manager.

[0035] When the power manager determines that the conditions for updating the operating parameters of an individual power domain are not met (condition block 406: "no"), the multiple partitions of the parallel data processing unit continue to process the workload using the corresponding assigned operating parameters of the individual power domain (block 408). In some embodiments, the conditions for updating a power domain include the power manager determining that one or more specific time intervals have elapsed and the power manager determining that the throughput level has changed beyond a threshold amount.

[0036] When the power manager determines that the conditions for updating the operating parameters of an individual power domain are met (condition block 406: "yes"), the power manager determines the dynamic scaling factor for each of the multiple partitions based on the number of corresponding operation calculation units and the monitored performance metrics (block 410). Based at least on the dynamic scaling factor, the power manager assigns the updated operating parameters of the individual power domain to the multiple partitions (block 412). In some embodiments, when updating the operating parameters of an individual power domain, the power manager further uses a static scaling factor and a weight value corresponding to both the static scaling factor and the dynamic scaling factor. The power manager resets one or more performance metric measurements eligible for reset (block 414).

[0037] Referring to FIG. 5, one embodiment of a computing system 500 is shown. As shown, the computing system 500 includes a processing unit 510, a memory 520, and a parallel data processing unit 530. In some embodiments, the functions of the computing system 500 are included as components on a single die, such as a single integrated circuit. In other embodiments, the functions of the computing system 500 are included as multiple dies on a system-on-chip (SOC). In various embodiments, the computing system 500 is used in desktop computers, portable computers, mobile devices, servers, peripheral devices, and the like.

[0038] The circuitry of the processing unit 510 processes instructions of a given algorithm. The processing unit includes fetching instructions and data, decoding the instructions, executing the instructions, and storing the results. In one embodiment, the processing unit 510 uses one or more processor cores having circuitry for executing instructions according to a predefined general-purpose instruction set. In various embodiments, the processor 510 is a central processing unit (CPU). The parallel data processing unit 530 includes the circuitry and functions of the device 100 (of FIG. 1).

[0039] The balance throughput manager 532 (or manager 532) has the functions of the balance throughput manager 174 (FIG. 1) and the balance throughput manager 336 (FIG. 3). For example, the manager 532 determines dynamic scaling factors used to dynamically update the power domains of multiple partitions of the parallel data processing unit 530. The manager 532 calculates these dynamic scaling factors based on the number of corresponding operation calculation units of the multiple partitions and performance metrics monitored over time during the processing of one or more workloads. The manager 532 determines when a performance bottleneck occurs in any of the multiple partitions during the execution of a workload and recalculates the dynamic scaling factors used to generate a new power domain for the multiple partitions.

[0040] In various embodiments, a thread is scheduled on one of the processing unit 510 and the parallel data processing unit 530 based at least in part on the runtime hardware resources of the processing unit 510 and the parallel data processing unit 530 such that each thread has the highest instruction throughput. In some embodiments, some threads are associated with general-purpose algorithms scheduled on the processing unit 510, and other threads are associated with parallel data computation-intensive algorithms such as a video graphics rendering algorithm scheduled on the parallel data processing unit 530. Applications using these algorithms have a copy stored on the memory 520.

[0041] Some threads that are not video graphics rendering algorithms still exhibit parallel data and intensive throughput. These threads have instructions that can operate on a relatively large number of different data elements simultaneously. Examples of these threads are threads for scientific, medical, financial, and encryption / decryption computations. The high parallelism provided by the hardware of the parallel data processing unit 530 and used to render multiple pixels simultaneously can also process multiple data elements of scientific, medical, financial, encryption / decryption, and other computations simultaneously.

[0042] Function calls within the application are converted into commands by a predetermined application programming interface (API). The processing unit 510 sends the converted commands to the memory 520 to store them in the ring buffer 522. The commands are placed into groups called command groups. In some embodiments, the processing units 510 and 530 use a producer-consumer relationship, also known as a client-server relationship. The processing unit 510 writes commands to the ring buffer 522. Next, the parallel data processing unit 530 reads commands from the ring buffer 522, processes the commands, and writes the resulting data to the buffer 524. The processing unit 510 is configured to update the write pointer to the ring buffer 522 and provide a size for each command group. The parallel data processing unit 530 updates the read pointer of the ring buffer 522, indicating the entry within the ring buffer 522 that the next read operation will use.

[0043] It should be noted that one or more of the above-described embodiments include software. In such embodiments, the program instructions implementing the method and / or mechanism are carried or stored on a computer-readable medium. A number of types of media configured to store program instructions are available, including hard disks, floppy (registered trademark) disks, CD-ROMs, DVDs, flash memories, programmable ROM (Programmable ROM, PROM), random access memory (RAM), and various other forms of volatile or non-volatile storage devices. Generally speaking, computer-accessible storage media includes any storage media that is accessible by a computer during use to provide instructions and / or data to the computer. For example, computer-accessible storage media includes magnetic or optical media, such as disks (fixed or removable), tapes, CD-ROMs, DVD-ROMs, CD-Rs, CD-RWs, DVD-Rs, DVD-RWs, or storage media such as Blu-Ray (registered trademark). Storage media further includes volatile or non-volatile memory media such as RAM (e.g., synchronous dynamic RAM (synchronous dynamic RAM, SDRAM), double data rate (double data rate, DDR, DDR2, DDR3, etc.) SDRAM, low-power DDR (low-power DDR, LPDDR2, etc.) SDRAM, Rambus DRAM (RDRAM), static RAM (static RAM, SRAM), etc.), ROM, flash memory, and non-volatile memory (e.g., flash memory) accessible via a peripheral interface such as a Universal Serial Bus (Universal Serial Bus, USB) interface. Storage media includes microelectromechanical systems (microelectromechanical system, MEMS), as well as storage media accessible via communication media such as networks and / or wireless links.

[0044] Additionally, in various embodiments, the program instructions include an operational level description of a hardware function or a register-transfer level (RTL) description in a high-level programming language such as C, a design language (HDL) such as Verilog or VHDL, or a database format such as GDSII stream format (GDSII). In some cases, the description is read by a synthesis tool that synthesizes the description to generate a netlist that includes a list of gates from a synthesis library. The netlist includes a set of gates that also represent the functionality of the hardware including the system. The netlist is then placed and routed to generate a data set that describes the geometric shapes to be applied to the mask. Next, the mask is used in various semiconductor manufacturing steps to generate one or more semiconductor circuits corresponding to the system. Alternatively, the instructions on a computer-accessible storage medium are, optionally, a netlist (with or without a synthesis library) or a data set. Additionally, the instructions are utilized for emulation by hardware-based types of emulators from vendors such as Cadence®, EVE®, and Mentor Graphics®.

[0045] Although the above embodiments have been described in considerable detail, numerous variations and modifications will become apparent to those skilled in the art upon a full understanding of the above disclosure. It is intended that the following claims be interpreted to embrace all such variations and modifications.

Claims

1. A plurality of partitions, each including a plurality of replicated computing units, and a power manager, wherein the power manager is configured to allocate a first set of operating parameters to the first partition, at least partially based on the number of operable replicated computing units within the first partition among the plurality of partitions; and allocate a second set of operating parameters to the second partition, at least partially based on the number of operable replicated computing units within the second partition among the plurality of partitions. The apparatus is configured to perform the above. Apparatus.

2. The first partition and the second partition are configured to process task of workload using the first set of operating parameters and the second set of operating parameters respectively. The apparatus according to Claim 1.

3. Based on the first set of operating parameters and the second set of operating parameters, the difference between the throughput of the first partition and the throughput of the second partition is less than a threshold value. The apparatus according to Claim 2.

4. Each of the plurality of partitions is configured to process the workload using a parallel data microarchitecture. The apparatus according to Claim 2.

5. The power manager is configured to allocate updated values to the first set of operating parameters and the second set of operating parameters in response to determining that conditions for updating the operating parameters of the plurality of partitions are met. The apparatus according to Claim 2.

6. The updated values for the first set of operating parameters and the second set of operating parameters are at least partially based on the power manager receiving a plurality of performance metrics monitored during the processing of the workload. The apparatus according to Claim 5.

7. The conditions for updating the power domain include the power manager determining that a time interval has elapsed; and the power manager determining that the throughput of the plurality of partitions has changed by more than a threshold amount. The conditions include one or more of the above. The apparatus according to Claim 5.

8. A plurality of partitions, each including a plurality of replicated computing units, process tasks. The power manager assigns a first set of operating parameters to the first partition, at least in part based on the number of operable replicated computing units within the first partition of the plurality of partitions. The power manager assigns a second set of operating parameters to the second partition, at least in part based on the number of operable replicated computing units within the second partition of the plurality of partitions. A method.

9. The first partition and the second partition each process a task of the workload using the first set of operating parameters and the second set of operating parameters, respectively. The method of claim 8.

10. Based on the first set of operating parameters and the second set of operating parameters, the difference between the throughput of the first partition and the throughput of the second partition is less than a threshold value. The method of claim 9.

11. Each of the plurality of partitions processes the workload using a parallel data microarchitecture. The method of claim 9.

12. The power manager assigns updated values to the first set of operating parameters and the second set of operating parameters in response to determining that a condition for updating the operating parameters of the plurality of partitions is satisfied. The method of claim 9.

13. The updated values for the first set of operating parameters and the second set of operating parameters are at least in part based on the power manager receiving a plurality of performance metrics monitored during the processing of the workload. The method of claim 12.

14. The condition for updating the power domain is The power manager determines that a time interval has elapsed. The power manager determines that the throughput of the plurality of partitions has changed beyond a threshold amount. including one or more of The method of claim 12.

15. A memory configured to store one or more applications of a workload, a processing unit, The processing unit a plurality of partitions each including a plurality of replicated computing units, a power manager, The power manager Allocating a first set of operating parameters to the first partition, at least in part based on the number of operable replicated computing units within the first partition of the plurality of partitions; Allocating a second set of operating parameters to the second partition, at least in part based on the number of operable replicated computing units within the second partition of the plurality of partitions; configured to perform; A computing system. **Claim 16** The first partition and the second partition are configured to process tasks of the workload using the first set of operating parameters and the second set of operating parameters, respectively. The computing system of claim 15. **Claim 17** Based on the first set of operating parameters and the second set of operating parameters, the difference between the throughput of the first partition and the throughput of the second partition is less than a threshold value. The computing system of claim 16. **Claim 18** Each of the plurality of partitions is configured to process the workload using a parallel data microarchitecture. The computing system of claim 16. **Claim 19** The power manager is configured to allocate updated values to the first set of operating parameters and the second set of operating parameters in response to determining that conditions for updating the operating parameters of the plurality of partitions are met. The computing system of claim 16. **Claim 20** The updated values for the first set of operating parameters and the second set of operating parameters are at least in part based on the power manager receiving a plurality of performance metrics monitored during the processing of the workload. The computing system of claim 19.