Scalable graphics processing using dynamic shader engine allocation

Dynamic scaling of shader engines based on application profiles and power configurations addresses inefficiencies in APUs, enhancing performance and power efficiency by optimizing shader engine usage.

JP2026520124APending Publication Date: 2026-06-22ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ADVANCED MICRO DEVICES INC
Filing Date
2024-05-29
Publication Date
2026-06-22

AI Technical Summary

Technical Problem

Conventional APUs face inefficiencies in power consumption and performance due to static shader engine resource allocation, particularly in low concurrent active context workloads, leading to suboptimal performance-to-power ratios and reduced battery life.

Method used

Dynamic scaling of shader engine resources based on application profiles and power configurations, using a command processor to activate or deactivate shader engines as needed, optimizing performance and power efficiency through software-controlled allocation and deallocation.

Benefits of technology

Optimizes performance per watt by dynamically adjusting shader engines according to application demands and power sources, improving user experience and extending battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026520124000001_ABST
    Figure 2026520124000001_ABST
Patent Text Reader

Abstract

A technique is described for performing the selective activation and deactivation of a dynamically allocated subset of shader engines (160) based on application-based profile information (210) and / or the active system power configuration (230), etc. Instructions to be executed are received from the application associated with the first application profile. Based on the application profile, the number of shader engines to be activated among multiple shader engines is changed. The number of activated shader engines is further changed in response to receiving additional instructions from a second application and / or receiving one or more indicators of the modified active system power configuration.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] An integrated processing unit (Accelerated Processing Unit, APU) typically combines the functions of a central processing unit (CPU) and a graphics processing unit (GPU) within a single package such as a chip or die. The APU improves system performance and power efficiency in a computing system by eliminating the need for a separate graphics card that generally consumes significant power and can generate additional heat. The APU is generally used in various portable computing devices (e.g., laptop computers, tablet computers, mobile computing, etc.) where power consumption and size are important factors for improving the user experience.

Summary of the Invention

Means for Solving the Problems

[0002] The techniques and systems described herein are directed to enabling scaling of shader engine resources, for example, based on an application profile associated with an application that generates instructions to be executed, to change the number of shader engines activated among a larger plurality of shader engines, based on a particular application for which instructions are to be provided. According to an exemplary embodiment, the system includes a command processor communicatively coupled to a plurality of shader engines. The command processor is configured to receive one or more instructions to be executed in place of a first application, change the number of shader engines activated among the plurality of shader engines based on profile information associated with the first application, and initiate execution of the one or more instructions for the first application on one or more processors using the changed number of activated shader engines.

[0003] In some embodiments, the command processor is further configured to receive one or more additional instructions to be executed on behalf of a second application, to dynamically increase the number of shader engines to be activated from among multiple shader engines in response to the one or more additional instructions and based on second profile information associated with the second application, and to execute one or more additional instructions on one or more processors using the increased number of activated shader engines.

[0004] In some embodiments, dynamically increasing the number of shader engines to be activated involves initializing a first set of one or more shader engines using state information associated with a second set of one or more shader engines, such that one or more of the second set of shader engines are activated before receiving one or more additional instructions to be executed.

[0005] The command processor may be further configured to receive one or more additional instructions to be executed on behalf of the second application, dynamically reduce the number of shader engines to be activated from among multiple shader engines in response to the one or more additional instructions and based on second profile information associated with the second application, and then execute one or more additional instructions on one or more processors using the reduced number of activated shader engines.

[0006] In some embodiments, dynamically reducing the number of activated shader engines involves clearing state information from a first set of one or more shader engines before deactivating one or more of those shader engines. Changing the number of activated shader engines from a plurality of shader engines may also be based on the system's power configuration, or on whether the system is currently coupled to an alternating current (AC) power supply or a direct current (DC) power supply.

[0007] The command processor may be further configured to receive profile information from the graphics driver, which may include any of several application profiles maintained by the graphics driver.

[0008] In another exemplary embodiment, the method includes receiving one or more instructions to be executed on behalf of a first application; changing the number of shader engines to be activated from among a plurality of shader engines of the processor based on profile information associated with the first application; and executing one or more instructions for the first application using the changed number of activated shader engines. According to some embodiments, the method further includes receiving one or more additional instructions to be executed on behalf of a second application; dynamically increasing the number of shader engines to be activated from among a plurality of shader engines in response to the receipt of one or more additional instructions and based on second profile information associated with the second application; and executing one or more additional instructions using the increased number of activated shader engines.

[0009] Dynamically increasing the number of shader engines to be activated may include initializing a first set of one or more shader engines using state information associated with a second set of shader engines, the second set of shader engines being activated before receiving one or more additional instructions to be executed. In some embodiments, the method further includes receiving one or more additional instructions to be executed on behalf of a second application, dynamically decreasing the number of shader engines to be activated from among the multiple shader engines in response to the receipt of one or more additional instructions and based on second profile information associated with the second application, and executing one or more additional instructions using the reduced number of shader engines.

[0010] Dynamically reducing the number of activated shader engines may include clearing state information from a first set of one or more shader engines before deactivating one or more of those shader engines. Changing the number of activated shader engines among multiple shader engines may further depend on the power configuration of the computing system containing multiple shader engines, and may also depend on whether the computing system is currently coupled to an alternating current (AC) power supply or a direct current (DC) power supply.

[0011] In some embodiments, the method further includes determining profile information associated with a first application based on heuristic analysis of one or more analyzed applications. One or more analyzed applications may include the first application. The method may further include receiving profile information from a software driver, the profile information including any application profile from a plurality of application profiles maintained by the software driver. In some embodiments, the method also includes selecting any application profile from the plurality of application profiles based on the application type of the first application.

[0012] In another exemplary embodiment, the command processor is configured to receive one or more instructions from a graphics driver to be executed on behalf of a first application, to change the number of shader engines to be activated from among a plurality of shader engines coupled to the command processor based on profile information associated with the first application, and to start executing one or more instructions for the first application using the changed number of activated shader engines.

[0013] This disclosure will be better understood by referring to the accompanying drawings, and many of its features and advantages may become apparent to those skilled in the art. The use of the same reference numerals in different drawings indicates similar or identical items. [Brief explanation of the drawing]

[0014] [Figure 1] This is a block diagram of a processing system 100 that performs selective activation and deactivation of dynamically allocated subsets of shader engines according to several embodiments. [Figure 2] This figure shows application-based activation and deactivation of dynamically allocated subsets of shader engines according to several embodiments. [Figure 3] This figure shows operational routines for selectively activating and deactivating dynamically allocated subsets of shader engines based on an application profile, according to several embodiments. [Figure 4] This figure shows operational routines for selectively activating and deactivating dynamically allocated subsets of shader engines based on application profiles and active power configurations, according to several embodiments. [Modes for carrying out the invention]

[0015] Larger APUs typically include many Work Group Processors (WGPs) across multiple Shader Engines (SEs). This architecture offers various performance advantages. However, making more hardware resources available introduces power consumption issues, for example, when running workloads associated with relatively low concurrent active contexts (CACs). Such workloads typically utilize very few graphics processing resources to efficiently accomplish their tasks, often using only a few WGPs within a single SE. The resulting power utilization causes the processor to operate at a suboptimal performance-to-power ratio, at least in part, due to relatively large power leaks consumed in the idle portions of the graphics pipeline and power wasted on the clock distribution paths to those portions.

[0016] Conventional solutions involve throttling one or more system clock signals or system voltages according to application needs. However, simply running at a slower frequency does not enable operation under minimum power constraints, thereby reducing battery life and contributing to a degraded user experience. In addition, such solutions statically enable or disable shader resources, for example, through hardware fuzing methods that are only implemented during system initialization (boot time), thereby preventing any runtime modifications to scale the shader engine resources available to the APU.

[0017] Embodiments of the technology described herein enable scaling of SE resources based on an application profile associated with an application that generates instructions to be executed, for example, to change the number of shader engines activated from a larger group of shader engines based on a particular application that provides instructions to be executed. In certain embodiments, the allocation and deallocation of shader engines are performed, for example, by being dynamically software-controlled by a user-mode driver (UMD) and / or a kernel-mode driver, and implemented by a run list controller (RLC) and a command processor (CP).

[0018] For example, in certain embodiments, dynamic SE activation is performed using application heuristics to analyze and profile numerous SE allocation configurations for various popular applications (e.g., gaming applications, productivity applications, visual production applications, etc.). In certain embodiments, information on such configurations is incorporated into one or more software drivers to selectively activate (e.g., power) and / or deactivate (e.g., substantially deplete power) certain shader engines (e.g., a subset of more shader engines) to achieve an optimal performance-to-power operating point. By scaling graphics pipeline resources (e.g., the shader engines to be activated) based on individual application requirements, the APU can dynamically enable or disable SEs based on the requirements indicated by these software components, keeping the graphics pipeline operating with substantially optimal power efficiency.

[0019] In certain embodiments, the number of shader engines activated for use in the indicated application is further determined by the APU based on the power configuration of the computing system. For example, in embodiments and scenarios where sufficient power is available, the APU may be configured to optimize the GPU for performance by allowing the APU to use more internal resources to achieve higher frame rates at the expense of additional power. More generally, when operating under AC power, the APU may be optimized for performance, and under DC power, it may be optimized for power consumption, for example, to extend battery life. In both scenarios and any power configuration, performance per watt is optimized or improved by the APU.

[0020] As used herein, the power of a shader configuration refers to the relative number of activated (powered) shader engines among multiple shader engines; therefore, a high-power shader configuration includes more activated shader engines than a low-power shader configuration. Accordingly, in at least some embodiments, shader engines referred to as deactivated herein are substantially unpowered, for example, to mitigate or avoid leakage power consumed in idle portions of the graphics pipeline and power wasted in any relevant portions of the clock distribution path.

[0021] In certain embodiments, switching from a low-power shader configuration to a high-power shader configuration involves restoring previously saved states to all activated SEs, thereby initializing and programming one or more newly added SEs in the high-power shader configuration using information from the shader engines activated in the low-power shader configuration. For example, in various embodiments, the RLC and CP initialize and program newly added shader engines without additional software assistance from software drivers or the application itself by provisioning newly activated shader engines using, for example, state information from one or more previously activated shader engines.

[0022] While the various embodiments discussed herein employ techniques described in the context of a particular APU processing system having specific components, it should be understood that such techniques described may be utilized in additional contexts and situations, for example, within and / or by a GPU, including discrete graphics processing units (GPUs) (one or more GPUs contained in a separate package and coupled to one or more CPUs via a hardware interface) and integrated GPUs (one or more GPUs integrated with one or more CPUs in a single package).

[0023] Figure 1 is a block diagram of a processing system 100 that performs selective activation and deactivation of a dynamically allocated subset of shader engines, according to several embodiments. The processing system 100 includes, or has access to, memory 105 or other storage components implemented using a non-temporary computer-readable storage medium such as dynamic random-access memory (DRAM). However, in embodiments, memory 105 is implemented using other types of memory, such as static random-access memory (SRAM), non-volatile RAM, etc. According to embodiments, memory 105 includes external memory implemented outside of the processing units implemented within the processing system 100. The processing system 100 also includes a bus 110 that facilitates communication between entities implemented within the processing system 100, such as memory 105. Some embodiments of the processing system 100 include other buses, bridges, switches, routers, etc., which are not shown in Figure 1 for clarity.

[0024] The techniques described herein are at least partially employed in an integrated processing unit (APU) 115, also referred to as an integrated processor, in various embodiments. The APU 115 includes, for example, any of various parallel processors, vector processors, coprocessors, graphics processing units (GPUs), general-purpose GPUs (GPGPUs), non-scalar processors, high parallel processors, artificial intelligence (AI) processors, inference engines, machine learning processors, other multi-threaded processing units, scalar processors, serial processors, or any combination thereof. In some embodiments, the APU 115 renders images according to one or more applications 135 (e.g., shader programs) for presentation on the display 190. For example, the APU 115 renders objects (e.g., groups of primitives) according to one or more shader programs to generate pixel values provided to the display 190, and this display uses the pixel values to display an image representing the rendered objects.

[0025] To render objects, the APU115 implements multiple processor cores 121-123 that execute instructions from one or more applications 135 simultaneously or in parallel. For example, the APU115 uses multiple processor cores 121-123 to execute instructions from shader programs, ray tracing programs, graphics pipelines, or both to render one or more objects. In the exemplary embodiment shown in Figure 1, three processor cores (121-123) representing N cores are shown, but the number of processor cores 121-123 implemented within the APU115 is a matter of design choice. Therefore, in other embodiments, the APU115 can contain any number of processor cores 121-123. Some embodiments of the APU115 are used for general-purpose computing. The APU 115 executes instructions such as program code 125 (e.g., shader code, ray tracing code) for one or more applications 135 (e.g., shader program, ray tracing program) stored in memory 105, and the APU 115 stores information such as the results of the executed instructions in memory 105. In the shown embodiment, memory 105 further includes, for example, part or all of an operating system (OS) 126 to provide an interface between application 135 and graphics driver 128.

[0026] Each of the processor cores 121-123 is communicatively coupled to one or more respective sets of computational unit resources (resource, RES) 141. For example, each computational unit of the processor cores 121-123 includes or is otherwise coupled to a respective set of computational unit resources within RES 141. RES 141 is configured to store values, register files, operands, instructions, variables, result data (e.g., data obtained from the performance of one or more operations), flags, or any combination thereof that are necessary to perform, assist in performing, or are useful in performing one or more operations indicated in one or more instructions from application 135. In various embodiments, the processing system 100 includes any number of sets of computational unit resources 141 for use by the processor cores 121-123.

[0027] The APU 115 further includes a plurality of shader engines 160, which, in the illustrated embodiment, include shader engines 161, 162, 163, 164, 165, 166. In various embodiments, the shader engines 160 may include any number of shader engines, and the number of shader engines 160 implemented within the APU 115 is a matter of design choice. Each of the shader engines 160 includes one or more workgroup processors (WGP), which are omitted here for clarity.

[0028] The APU115 includes a command processor (CP) 140 (also known as a scheduler) and a run list controller (RLC) 144, both of which, in various embodiments, include hardware-based circuitry, software-based circuitry, or both. The RLC 144 is responsible for managing and scheduling the execution of a list of commands sent to the APU115. These commands, also known as a "run list," are typically sequences of low-level instructions specifying various actions (e.g., drawing a triangle, setting a color, or updating a texture). The RLC ensures that the commands in the run list are executed in the correct order and that any necessary resources of RES141 are available, while the CP 140 is responsible for interpreting and executing the individual commands in the run list, for example, by decoding the commands and converting them into appropriate hardware instructions to be executed by one or more shader engines of the shader engine 160.

[0029] The processing system 100 also includes a central processing unit (CPU) 130 connected to a bus 110 and therefore communicating with an APU 115 and memory 105 via a bus 112. The CPU 130 implements a plurality of processor cores 131-133 that execute instructions simultaneously or in parallel. In some embodiments, each of one or more of the processor cores 131-133 operates as one or more computing units (e.g., single instruction multiple data or SIMD units) that perform the same operation on different datasets. In the exemplary embodiment shown in Figure 1, three processor cores (131-133) representing M cores are shown, but the number of processor cores 131-133 implemented in the CPU 130 is a matter of design choice. Therefore, in other embodiments, the CPU 130 can contain any number of processor cores 131-133. In some embodiments, the CPU 130 and APU 115 have the same number of processor cores, but in other embodiments, the CPU 130 and APU 115 have different numbers of processor cores. Processor cores 131 to 133 execute instructions such as program code 125 stored in memory 105, and the CPU 130 stores information such as the results of the executed instructions in memory 105. The CPU 130 can also start graphics processing by issuing a draw call to the APU 115. In this embodiment, the CPU 130 implements multiple processor cores (not shown in Figure 1 for clarity) that execute instructions simultaneously or in parallel.

[0030] The input / output (I / O) engine 145 includes hardware and software that handles input or output operations associated with the display 190, as well as other elements of the processing system 100 such as a keyboard, mouse, printer, and external disk. The I / O engine 145 is coupled to a bus 110 so that it can communicate with memory 105, APU 115, or CPU 130.

[0031] Figure 2 illustrates application-based activation and deactivation of a dynamically allocated subset of shader engines in several embodiments. Continuing to refer to the processing system 100 in Figure 1, in the shown embodiments, the graphics driver 128 includes a plurality of application profiles 210, individually identified as application profiles 211, 212, ..., 213. In various embodiments, the application profiles 210 may include any number of application profiles. The graphics driver 128 further includes a kernel mode driver (KMD) 220.

[0032] In the first time T1, the APU 115 executes instructions on behalf of the application associated with application profile 212. In this example, application profile 212 is associated with a text-based application that uses very few graphics rendering resources. Based on information about its effect within application profile 212, the command processor 140 instructs the runlist controller 144 to activate (provide operating power to) only a single shader engine 161, leaving shader engines 162, 163, 164, 165, and 166 deactivated and therefore effectively unpowered in the first shader engine activation profile 250. Thus, instructions received from the text-based application associated with application profile 212 are executed using only the single activated shader engine 161.

[0033] In a later second time T2, the APU 115 receives one or more instructions on behalf of a second application associated with application profile 211. In this example, application profile 211 is associated with a gaming application that heavily utilizes 3D rendering during gameplay. Based on information about its effects within application profile 211, the command processor 140 instructs the runlist controller 144 to utilize all shader engines 160 in the new shader engine activation profile 260. As a result, each of the shader engines 162, 163, 164, 165, and 166 that were deactivated in shader engine activation profile 250 are initialized and activated (provided operating power) for use in executing instructions received from the gaming application associated with application profile 211, or on behalf of that gaming application.

[0034] In certain embodiments, switching from a low-power shader engine activation profile 250 to a high-power shader engine activation profile 260 involves providing state information to each of the newly activated SE162, 163, 164, 165, and 166 from the already activated SE161. For example, in one embodiment, after the RLC has finished enabling SE162, 163, 164, 165, and 166, the RLC sends a command to the CP to instruct it to reinitialize the entire system state using the state information from SE161, including the newly activated shader engines. In this way, the CP140 and RLC144 initialize and program the newly added SE162, 163, 164, 165, and 166 without any additional software assistance from the graphics driver 128 or the application associated with the application profile 211.

[0035] In a later third time T3, while the APU 115 is still executing instructions on behalf of the gaming application associated with application profile 211, the APU 115 receives notification of a modification to the active system power configuration 230. In various embodiments, the notification of the active system power configuration 230 may be proactively transmitted by one or more power monitoring components communicably coupled to the APU, and may be received by polling from one or more registers or memory locations or in some other way.

[0036] For example, in one embodiment, KMD220 sends a message to CP140 instructing it that SE reconfiguration is required. In response, CP140 instructs RLC144 to unmap the SE hardware queue and reconfigure the activated shader engines SE161, 162, 163, 164, 165, and 166. Following the reconfiguration, RLC sends a completion response, causing CP140 to remap the previous SE queue and restart the reconfigured system.

[0037] In this example, at time T3, the APU 115 receives a notification (not shown) indicating that the active system power configuration 230 has transitioned from a first configuration in which multiple shader engines 160 are coupled to an AC power source to a second configuration in which multiple shader engines 160 are coupled to a DC power source. Based on the currently active application profile 211 and the active system power configuration, the CP 140 instructs the RLC 144 to deactivate shader engines 165 and 166 in the shader engine activation profile 270, while leaving shader engines 161, 162, 163, and 164 active. In this way, the APU 115 optimizes or improves system performance per watt based on both the active application and the active system power configuration.

[0038] In certain embodiments, switching from a high-power shader engine activation profile 260 to a low-power shader engine activation profile 270 includes clearing state information from SE165, 166 before deactivating those shader engines. For example, in the shown embodiment, a drain command is issued to RLC144 by the command processor 140 to ensure that shader waves or events being processed by SE165, 166 are not stored as part of their respective state information.

[0039] Figure 3 shows operational routines for selectively activating and deactivating dynamically allocated subsets of shader engines based on application profiles, according to several embodiments. Routine 300 may be performed, for example, by an APU (e.g., APU 115 in Figure 1) when it receives instructions to be executed on behalf of one or more applications (e.g., application 135 in Figure 1) based on one or more application profiles (e.g., application profile 210 in Figure 2), for example, by an APU (e.g., APU 115 in Figure 1).

[0040] Routine 300 begins in block 305, where the APU receives instructions for execution on behalf of the first application. Routine 300 then proceeds to block 310.

[0041] In block 310, the APU determines profile information (first profile information) associated with the first application. In certain embodiments and as described elsewhere in this specification, the profile information may be stored as part of a software driver (e.g., the graphics driver 128 in Figures 1 and 2). In various embodiments, the first profile information may be directly associated with the first application, or it may be indirectly associated with the first application, for example, if the first application is identified as having an application type corresponding to one or more additional applications associated with the determined first profile information. For example, the APU may determine that the application is a text-based application (e.g., a word processor, text editor), a 2D graphical application that presents purely graphical content or a combination of graphical and text content (e.g., a web browser), a gaming application that presents rendered 3D content, etc. Once the profile information associated with the first application is determined, routine 300 proceeds to block 315.

[0042] In block 315, the APU modifies the number of shader engines to be activated from among multiple shader engines based on the determined first profile information. In various embodiments, modifying the number of activated shader engines may involve one or more additional processes to appropriately save or release state information associated with the shader engines being activated or deactivated. For example, as described elsewhere in this specification, in certain embodiments, increasing the number of activated shader engines may include, for example, provisioning the newly activated shader engines with state information from one or more previously activated shader engines in order to initialize them. In contrast, in various embodiments, decreasing the number of activated shader engines may include, for example, clearing state information from a set of one or more shader engines before deactivating them by executing a drain command to ensure that shader waves or events are not saved as part of the state information of those deactivated shader engines. Routine 300 proceeds to block 320.

[0043] In block 320, the APU executes instructions on behalf of the first application using a modified number of activated shader engines. Routine 300 then proceeds to block 325.

[0044] In block 325, the APU receives instructions to be executed on behalf of the second application. Routine 300 proceeds to block 330.

[0045] In block 330, the APU determines profile information associated with a second application (second profile information). Similar to the profile information associated with the first application determined in block 310, the second profile information may be stored as part of a software driver (e.g., the graphics driver 128 in Figures 1 and 2). Also, in a similar manner to that described above for determining the first profile information in block 310, the second profile information may be directly associated with the second application, or indirectly associated with the second application based on the application type associated with the second application (e.g., a text-based application, a 2D graphical application, a gaming application, or another application that presents rendered 3D content). Once the second profile information is determined, routine 300 proceeds to block 335.

[0046] In block 335, the APU modifies the number of shader engines to be activated to a second modified number, which is, for example, more or less than the number of shader engines to be activated selected in block 315, based on the determined second profile information. Modifying the number of shader engines to be activated according to the second profile information in a manner similar to that described above for block 315 may involve one or more additional processes to appropriately save or release state information associated with the shader engines being activated or deactivated. Routine 300 proceeds to block 340.

[0047] In block 340, the APU uses a second modified number of activated shader engines to execute instructions on behalf of a second application.

[0048] Figure 4 shows operational routines for selectively activating and deactivating a dynamically allocated subset of shader engines based on application profiles and active power configurations, according to several embodiments. Routine 400 may be performed, for example, by an APU (e.g., APU 115 in Figure 1) when it receives an instruction to be executed (e.g., part or all of program code 125 in Figure 1) based on an application profile (e.g., any of the application profiles 210 in Figure 2) and an active power configuration (e.g., power configuration 230 in Figure 2).

[0049] Routine 400 begins in block 405, where the APU receives instructions to be executed on behalf of the first application. Routine 400 then proceeds to block 410.

[0050] In block 410, the APU determines profile information associated with a first application (first profile information), such as profile information stored as part of a software driver (e.g., the graphics driver 128 in Figures 1 and 2). As described above with respect to the operation routine 300 in Figure 3, the profile information may be directly or indirectly associated with the first application. Once the profile information associated with the first application is determined, routine 400 proceeds to block 415.

[0051] In block 415, the APU changes (modifies) the number of shader engines to be activated from among the multiple shader engines based on the determined first profile information. As described above with respect to the operation routine 300 in Figure 3, in various embodiments, changing (modifying) the number of activated shader engines may involve one or more additional processes to appropriately save or release the state information associated with the activated or deactivated shader engines. Routine 400 proceeds to block 420.

[0052] In block 420, the APU executes instructions on behalf of the first application using a modified number of activated shader engines. Routine 400 then proceeds to block 425.

[0053] In block 425, the APU receives notification of the active system power configuration. In various embodiments, the notification of the active system power configuration may be actively transmitted, for example, by one or more power monitoring components communicably coupled to the APU, and may be polled from one or more registers or memory locations. Routine 400 proceeds to block 430.

[0054] In block 430, the APU modifies the number of shader engines to be activated based on the determined profile information and the active system power configuration. For example, in certain scenarios and embodiments, the number of shader engines to be activated is modified based on whether multiple shader engines are currently coupled to an AC (alternating current) power supply or a DC (direct current) power supply.

[0055] In some embodiments, the apparatus and techniques described above are implemented in a system including one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips), such as the integrated processing units and other devices described above with reference to Figures 1 to 4. Electronic design automation (EDA) and computer-aided design (CAD) software tools can be used to design and manufacture these IC devices. These design tools are typically represented as one or more software programs. One or more software programs include computer-executable code for operating a computer system to operate with code representing the circuit of one or more IC devices in order to perform at least part of the process of designing or adapting a manufacturing system for manufacturing the circuit. This code may include instructions, data, or combinations of instructions and data. Software instructions representing the design or manufacturing tools are typically stored in a computer-readable storage medium accessible to the computing system. Similarly, code representing one or more stages of designing or manufacturing an IC device is stored in and accessed from the same or different computer-readable storage medium.

[0056] Computer-readable storage media include any non-temporary storage media or combination of non-temporary storage media that are accessible by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray® discs), magnetic media (e.g., floppy disks, magnetic tapes, magnetic hard drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical system (MEMS) based storage media. Computer-readable storage media (e.g., system RAM or ROM) may be built into the computing system, computer-readable storage media (e.g., magnetic hard drives) may be permanently mounted to the computing system, computer-readable storage media (e.g., optical disks or Universal Serial Bus (USB) based flash memory) may be detachably mounted to the computing system, and computer-readable storage media (e.g., network-accessible storage (NAS)) may be connected to the computer system via a wired or wireless network.

[0057] In some embodiments, certain aspects of the technology described above are implemented by one or more processors of a processing system that executes the software. The software includes one or more sets of executable instructions, which are stored in a non-temporary computer-readable storage medium or otherwise clearly embodied. The software may also include instructions and specific data, which, when executed by one or more processors, operate the one or more processors to execute one or more aspects of the technology described above. Non-temporary computer-readable storage mediums may include, for example, magnetic or optical disk storage devices, solid-state storage devices such as flash memory, caches, random-access memory (RAM), or other non-volatile memory devices (one or more). Executable instructions stored in a non-temporary computer-readable storage medium can be implemented as source code, assembly language code, object code, or other instruction forms that can be interpreted or otherwise executed by one or more processors.

[0058] In addition to the foregoing, it should be noted that not all activities or elements described in the summary are required, and certain activities or parts of devices may not be required, and one or more additional activities may be performed, and one or more additional elements may be included. Furthermore, the order in which the activities are listed does not necessarily indicate the order in which they are performed. Also, the concepts have been described with reference to specific embodiments. However, those skilled in the art will understand that various modifications and variations can be made without departing from the scope of the invention as described in the claims. Therefore, the specification and drawings should be considered illustrative rather than restrictive, and all of these variations are intended to fall within the scope of the invention.

[0059] Benefits, other advantages, and solutions to problems have been described above with respect to specific embodiments. However, benefits, advantages, solutions to problems, and features that may give rise to or manifest any benefits, advantages, or solutions are not to be construed as essential, necessary, or indispensable features to any or all of the claims. Furthermore, the disclosed invention can be modified and implemented in different but similar ways, in ways that are obvious to those skilled in the art who are interested in the teachings of this specification; therefore, the specific embodiments described above are merely illustrative. There are no limitations to the details of the configuration or design shown herein beyond those described in the appended claims. Accordingly, the specific embodiments described above may be modified or altered, and it is clear that all such modifications are within the scope of the disclosed invention. Accordingly, the protection sought herein is described in the appended claims.

Claims

1. It is a system, It features a command processor that is communicatively coupled to multiple shader engines, The aforementioned command processor, To receive one or more instructions to be executed on behalf of the first application, Based on the profile information associated with the first application, the number of shader engines to be activated among the plurality of shader engines is changed, Using a modified number of activated shader engines, to initiate the execution of the one or more instructions for the first application on one or more processors, It is configured to do, system.

2. The aforementioned command processor, Receiving one or more additional instructions to be executed on behalf of the second application, In response to one or more additional instructions, and based on second profile information associated with the second application, the number of shader engines to be activated among the plurality of shader engines is to be dynamically increased. Using an increased number of activated shader engines, execute the one or more additional instructions on the one or more processors, It is configured to do, The system according to claim 1.

3. Dynamically increasing the number of shader engines to be activated includes initializing a first set of one or more shader engines using state information associated with a second set of one or more shader engines. One or more shader engines from the second set of shader engines are activated before receiving the one or more additional instructions to be executed. The system according to claim 2.

4. The aforementioned command processor, Receiving one or more additional instructions to be executed on behalf of the second application, In response to one or more additional instructions, and based on second profile information associated with the second application, the number of shader engines to be activated among the plurality of shader engines is dynamically reduced. Using a reduced number of activated shader engines, execute the one or more additional instructions on the one or more processors, It is configured to do, The system according to claim 1.

5. Dynamically reducing the number of activated shader engines includes clearing state information from a first set of shader engines before deactivating one or more shader engines from a first set of one or more shader engines. The system according to claim 4.

6. The number of shader engines to be activated among the aforementioned plurality of shader engines is changed based on the power configuration of the system. A system according to any one of claims 1 to 5.

7. Changing the number of shader engines to be activated is done based on whether the system is currently coupled to an alternating current (AC) power supply or a direct current (DC) power supply. The system according to claim 6.

8. The command processor is configured to receive the profile information from the graphics driver. The profile information includes one of the application profiles among the multiple application profiles maintained by the graphics driver. The system according to claim 1.

9. It is a method, To receive one or more instructions to be executed on behalf of the first application, Based on the profile information associated with the first application, the number of shader engines to be activated among the processor's multiple shader engines is changed, This includes executing one or more instructions for the first application using a modified number of activated shader engines, method.

10. Receiving one or more additional instructions to be executed on behalf of the second application, In response to receiving one or more additional instructions, and based on the second profile information associated with the second application, the number of shader engines to be activated among the plurality of shader engines is dynamically increased. This includes executing one or more of the additional instructions using an increased number of activated shader engines, The method of claim 9.

11. Dynamically increasing the number of shader engines to be activated includes initializing a first set of one or more shader engines using state information associated with a second set of shader engines. The second set of shader engines is activated before receiving the one or more additional instructions to be executed. The method of claim 10.

12. Receiving one or more additional instructions to be executed on behalf of the second application, In response to receiving one or more additional instructions, and based on the second profile information associated with the second application, the number of shader engines to be activated among the plurality of shader engines is dynamically reduced. This includes using a reduced number of shader engines to execute one or more of the additional instructions, The method of claim 9.

13. Dynamically reducing the number of activated shader engines includes clearing state information from a first set of shader engines before deactivating one or more shader engines from a first set of one or more shader engines. The method according to claim 12.

14. The number of shader engines to be activated among the aforementioned plurality of shader engines is changed based on the power configuration of the computing system including the plurality of shader engines. The method according to any one of claims 9 to 13.

15. Changing the number of shader engines to be activated is done based on whether the computing system is currently coupled to an alternating current (AC) power supply or a direct current (DC) power supply. The method according to claim 14.

16. This includes determining profile information associated with the first application based on heuristic analysis of one or more analyzed applications, The method of claim 9.

17. The one or more analyzed applications include the first application. The method according to claim 16.

18. This includes receiving the profile information from the software driver, The profile information includes one of the application profiles among the multiple application profiles maintained by the software driver. The method of claim 9.

19. This includes selecting one of the application profiles from the plurality of application profiles based on the application type of the first application, The method of claim 18.

20. A command processor, The aforementioned command processor, Receiving one or more instructions from the graphics driver to be executed on behalf of the first application, Based on the profile information associated with the first application, the number of shader engines to be activated among the multiple shader engines coupled to the command processor is changed, Using a modified number of activated shader engines, the execution of one or more instructions for the first application is initiated, It is configured to do, Command processor.