Execution of software-controlled variable wavefront size on GPU
By implementing different operational modes in GPUs, the challenge of scheduling instructions for wavefronts with more work items than execution units is addressed, resulting in improved performance and resource utilization.
Patent Information
- Application Number
- JP2021570312
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-05-29
- Filing Date
- 2020-05-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-05-29
AI Technical Summary
The challenge in graphics processing units (GPUs) is efficiently scheduling instructions for wavefronts with more work items than execution units, leading to difficulties in determining how to execute instructions effectively.
The GPU operates in different modes to optimize performance and resource utilization. In one mode, instructions are executed across multiple parts of the wavefront before proceeding to the next instruction, while in another mode, instructions are executed on one portion of the wavefront before moving to the next portion, allowing for dynamic control of the operating mode based on indicators or cache miss rates.
This approach enables efficient execution of instructions for wavefronts with varying sizes, improving performance and power management by adapting the execution strategy based on workload and resource availability.
Smart Images

Figure 0007675660000002 
Figure 0007675660000003 
Figure 0007675660000004
Abstract
Description
[Background technology]
[0001] A graphics processing unit (GPU) is a complex integrated circuit configured to perform graphics processing tasks. For example, a GPU may perform graphics processing tasks required for end-user applications, such as video game applications. Increasingly, GPUs also perform other tasks unrelated to graphics. A GPU may be a separate device or may be included in the same device as another processor, such as a central processing unit (CPU).
[0002] In many applications, such as graphics processing on a GPU, a sequence of work-items, also called threads, are processed to produce a final result. For example, in many modern parallel processors, a processor within a Single Instruction, Multiple Data (SIMD) core executes a sequence of work-items synchronously. Multiple identical synchronous work-items processed by separate processors are called a wavefront or warp.
[0003] During processing, one or more SIMD cores execute multiple wavefronts simultaneously. Execution of a wavefront ends when all work items in the wavefront have completed processing. Each wavefront contains multiple work items that are processed in parallel using the same instruction set. In some cases, the number of work items in a wavefront does not match the number of execution units of a SIMD core. In one embodiment, each execution unit of a SIMD core is an arithmetic logic unit (ALU). If the number of work items in a wavefront does not match the number of execution units of a SIMD core, it may be difficult to determine how to schedule instructions for execution.
[0004] Advantages of the methods and mechanisms described herein may be better understood by referring to the following description in conjunction with the accompanying drawings. [Brief description of the drawings]
[0005] [Figure 1] FIG. 1 is a block diagram of an embodiment of a computing system. [Diagram 2] FIG. 2 is a block diagram of one embodiment of a GPU. [Diagram 3] FIG. 2 is a block diagram of one embodiment of a set of Vector General Purpose Registers (VGPR). [Figure 4] FIG. 2 illustrates one embodiment of an exemplary wavefront and an exemplary instruction sequence. [Diagram 5] FIG. 2 illustrates an embodiment of a first operating mode of a processor. [Figure 6] FIG. 2 illustrates an embodiment of a second operating mode of the processor. [Figure 7] 1 is a generalized flow diagram illustrating one embodiment of a method for scheduling instructions on a processor. [Figure 8] FIG. 2 is a generalized flow diagram illustrating one embodiment of a method for determining an operating mode to use in a parallel processor. [Figure 9] FIG. 2 is a generalized flow diagram illustrating one embodiment of a method for using different operating modes for a parallel processor. [Figure 10] FIG. 2 is a block diagram of one embodiment of the software architecture of the computing system of FIG. 1. [Figure 11] FIG. 1 is a generalized flow diagram illustrating one embodiment of a method for generating modified shader programs for dynamically controlling a processor's operating mode for each section of the program. [Figure 12] FIG. 2 illustrates one embodiment of a register usage analysis of a shader program. [Figure 13] FIG. 1 illustrates one embodiment of a modification of a shader program to implement a loop of sub-vectors over a section of a shader program. [Figure 14]FIG. 13 is another embodiment of a modification of a shader program to implement a loop of sub-vectors over a section of a shader program. [Figure 15] FIG. 13 is another embodiment of a modification of a shader program to implement a loop of sub-vectors over a section of a shader program. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0006] In the following description, numerous specific details are set forth to provide a thorough understanding of the methods and mechanisms described herein. However, those skilled in the art will recognize that various embodiments may be practiced without the use of such specific details. In some instances, well-known structures, components, signals, computer program instructions and techniques have not been shown in detail to avoid obscuring the approaches described herein. It should be understood that for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements.
[0007] Various systems, apparatus, methods and computer readable media are disclosed for processing variable wave front sizes on a processor. When operating in a first mode, the processor executes the same instructions on multiple portions of the wave front before proceeding to the next instruction of the shader program. When operating in a second mode, the processor executes a set of instructions on a first portion of the wave front, and when the processor finishes executing the set of instructions on the first portion of the wave front, the processor executes the set of instructions on a second portion of the wave front, and so on until all portions of the wave front have been processed. The processor then continues executing subsequent instructions of the shader program.
[0008] In one embodiment, an indicator is declared in a code sequence to specify which mode to utilize for a given region of the program. In another embodiment, a compiler generates an indicator when generating executable code, the indicator specifying an operating mode of the processor. In another embodiment, the processor includes a control unit that determines the operating mode of the processor.
[0009] Referring to FIG. 1, a block diagram of one embodiment of a computing system 100 is shown. In one embodiment, the computing system 100 includes a system on chip (SoC) 105 coupled to a memory 150. The SoC 105 may also be referred to as an integrated circuit (IC). In one embodiment, the SoC 105 includes processing units 115A-115N, an input / output (I / O) interface 110, a shared cache 120A-120B, a fabric 125, a graphics processing unit (GPU) 130, and a memory controller 140(s). The SoC 105 may include other components not shown in FIG. 1 to avoid obscuring the diagram. The processing units 115A-115N represent any number and type of processing units. In one embodiment, the processing units 115A-115N are central processing unit (CPU) cores. In another embodiment, one or more of processing units 115A-115N are other types of processing units (e.g., application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs)). Processing units 115A-115N are coupled to shared caches 120A-120B and fabric 125.
[0010] In one embodiment, the processing units 115A-115N are configured to execute instructions of a particular instruction set architecture (ISA). Each processing unit 115A-115N includes one or more execution units, a cache memory, a scheduler, branch prediction circuitry, etc. In one embodiment, the processing units 115A-115N are configured to execute main control software of the system 100, such as an operating system. Generally, during use, the software executed by the processing units 115A-115N may control other components of the system 100 to achieve desired functionality of the system 100. The processing units 115A-115N may execute other software, such as application programs.
[0011] GPU 130 includes computation units 145A-145N, which represent any number and type of computation units used for graphics or general purpose processing. Computation units 145A-145N may be referred to as "shader arrays," "shader engines," "single instruction multiple data (SIMD) units," or "SIMD cores." Each computation unit 145A-145N includes multiple execution units. GPU 130 is coupled to shared caches 120A-120B and fabric 125. In one embodiment, GPU 130 is configured to execute graphics pipeline operations, such as drawing commands, pixel operations, geometric calculations, and other operations, to render an image to a display. In another embodiment, GPU 130 is configured to execute non-graphics related operations. In a further embodiment, GPU 130 is configured to execute both graphics and non-graphics related operations.
[0012] GPU 130 is configured to receive shader programs and wavefront instructions for execution. In one embodiment, GPU 130 is configured to operate in different modes. In one embodiment, the number of work items in each wavefront is greater than the number of execution units in GPU 130.
[0013] In one embodiment, GPU 130, in response to detecting the first indication, schedules a first instruction for execution in the first and second portions of the first wavefront before scheduling a second instruction for execution in the first portion of the first wavefront. GPU 130 follows this pattern for other instructions of the shader program and other wavefronts as long as the first indication is detected. Note that "scheduling an instruction" is also referred to as "issuing an instruction." Depending on the embodiment, the first indication may be specified in software or may be generated by GPU 130 based on one or more operating conditions. In one embodiment, the first indication is a command for GPU 130 to operate in the first mode.
[0014] In one embodiment, in response to not detecting the first indication, GPU 130 schedules the first instruction and the second instruction for execution in the first portion of the first wavefront before scheduling the first instruction for execution in the second portion of the first wavefront. GPU 130 follows this pattern for other instructions of the shader program and other wavefronts until the first indication is detected.
[0015] I / O interface 110 is coupled to fabric 125, and I / O interface 110 may represent any number and type of interface (e.g., Peripheral Component Interconnect (PCI) bus, PCI-Extended (PCI-X), PCI Express (PCIE) bus, Gigabit Ethernet (GBE) bus, Universal Serial Bus (USB)). Various types of peripheral devices may be coupled to I / O interface 110. Such peripheral devices may include, but are not limited to, displays, keyboards, mice, printers, scanners, joysticks or other types of game controllers, media recording devices, external storage devices, network interface cards, and the like.
[0016] The SoC 105 is coupled to a memory 150 that includes one or more memory modules. Each of the memory modules includes one or more memory devices mounted thereon. In some embodiments, the memory 150 comprises one or more memory devices mounted on a motherboard or other carrier on which the SoC 105 is also mounted. The implemented RAM can be static RAM (SRAM), dynamic RAM (DRAM), resistive RAM (ReRAM), phase change RAM (PCRAM), or other volatile or non-volatile RAM. Types of DRAM used to implement the memory 150 include (but are not limited to) double data rate (DDR) DRAM, DDR2 DRAM, DDR3 DRAM, etc. Although not explicitly shown in FIG. 1, the SoC 105 can include one or more cache memories internal to the processing units 115A-115N and / or the computing units 145A-145N. In some embodiments, SoC 105 includes shared caches 120A-120B utilized by processing units 115A-115N and computational units 145A-145N. In one embodiment, caches 120A-120B are part of a cache subsystem that includes a cache controller.
[0017] In various embodiments, computing system 100 may be a computer, a laptop, a mobile device, a server, or any of various other types of computing systems or devices. It should be noted that the number of components of computing system 100 and / or SoC 105 may vary from embodiment to embodiment. Each component / subcomponent may be more or less than that shown in FIG. 1. For example, in another embodiment, SoC 105 may include multiple memory controllers coupled to multiple memories. It should also be noted that computing system 100 and / or SoC 105 may include other components not shown in FIG. 1. Furthermore, in other embodiments, computing system 100 and SoC 105 may be configured in a manner other than that shown in FIG. 1.
[0018] Referring to Figure 2, a block diagram of one embodiment of a graphics processing unit (GPU) 200 is shown. In one embodiment, GPU 200 includes at least SIMDs 210A-210N, a branch and message unit 240, a scheduler unit 245, an instruction buffer 255, and a cache 260. Note that GPU 200 may also include other logic that is not shown in Figure 2 to avoid obscuring the diagram. Note also that other processors (e.g., FPGAs, ASICs, DSPs) may include the circuitry shown in GPU 200.
[0019] In one embodiment, GPU 200 is configured to operate in different modes to process shader program instructions on wavefronts of different sizes. GPU 200 utilizes a given mode to optimize performance, power consumption, and / or other factors depending on the type of workload being processed and / or the number of work items in each wavefront. In one embodiment, each wavefront includes a number of work items that is greater than the number of lanes 215A-215N, 220A-220N, and 225A-225N of SIMD 210A-210N. In this embodiment, GPU 200 processes the wavefront differently based on the operating mode of GPU 200. In another embodiment, GPU 200 processes the wavefront differently based on one or more detected conditions. Each lane 215A-215N, 220A-220N, and 225A-225N of SIMD 210A-210N may also be referred to as an "execution unit."
[0020] In one embodiment, GPU 200 receives instructions for a wavefront that has a number of work items that is greater than the total number of lanes of SIMD 210A-210N. In this embodiment, GPU 200 executes a first instruction on multiple portions of the wavefront before proceeding to a second instruction when GPU 200 is in a first mode. GPU 200 continues this pattern of execution for subsequent instructions as long as GPU 200 is in the first mode. In one embodiment, the first mode may be specified by a software generated declaration. When GPU 200 is in a second mode, GPU 200 executes instructions on a first portion of the wavefront before proceeding to a second portion of the wavefront. When GPU 200 is in the second mode, GPU 200 shares some of the vector general purpose registers (VGPRs) 230A-230N between different portions of the wavefront. Further, when GPU 200 is in the second mode, if the execution mask of mask(s) 250 indicates that a given portion of the wavefront is temporarily masked out, GPU 200 does not execute instructions for the given portion of the wavefront.
[0021] In another embodiment, if the wavefront size is greater than the number of SIMDs, GPU 200 determines a cache miss rate for the program's cache 260. If the cache miss rate is less than a threshold, GPU 200 executes a first instruction in multiple portions of the wavefront before proceeding to a second instruction. GPU 200 continues this execution pattern for subsequent instructions as long as the cache miss rate is determined or predicted to be less than the threshold. The threshold may be specified as a number of bytes, a percentage of cache 260, or other suitable metric. If the cache miss rate is greater than or equal to the threshold, GPU 200 executes instructions in a first portion of the wavefront before executing instructions in a second portion of the wavefront. Additionally, if the cache miss rate is greater than or equal to the threshold, GPU 200 shares portions of vector general purpose registers (VGPRs) 230A-230N between different portions of the wavefront, and if the execution mask indicates that a given portion is masked out, GPU 200 skips instructions for that given portion of the wavefront.
[0022] It should be noted that the letter "N," when it appears next to various structures herein, is meant to generally refer to any number of elements of that structure (e.g., any number of SIMDs 210A-210N). Additionally, the different references in FIG. 2 that use the letter "N" (e.g., SIMDs 210A-210D and lanes 215A-215N) are not intended to imply that an equal number of different elements are provided (e.g., the number of SIMDs 210A-210D may differ from the number of lanes 215A-215N).
[0023] 3, there is shown a block diagram of one embodiment of a set of vector general purpose registers (VGPR) 300. In one embodiment, VGPR 300 is included in SIMDs 210A-210N of GPU 200 (of FIG. 2). VGPR 300 may include any number of registers, depending on the embodiment.
[0024] As shown in Figure 3, VGPR 300 includes VGPR 305 for a first wavefront, VGPR 310 for a second wavefront, and any number of other VGPRs for other numbers of wavefronts. For purposes of this description, it is assumed that the first wavefront and the second wavefront have 2xN work items, where N is a positive integer and N varies from embodiment to embodiment. In one embodiment, N is equal to 32. VGPR 305 includes a region with dedicated VGPRs and a shared VGPR region 315. Similarly, VGPR 310 includes a region with dedicated VGPRs and a shared VGPR region 320.
[0025] In one embodiment, when the host GPU (e.g., GPU 200) is in a first mode, each of the shared VGPR 315 and the shared VGPR 310 is not shared between different portions of the first and second wavefronts. However, when the host GPU is in a second mode, each of the shared VGPR 315 and the shared VGPR 310 is shared between different portions of the first and second wavefronts. In other embodiments, the enforcement of sharing or non-sharing is based not on the first or second mode, but on detection of a first indicator. The first indicator may be generated by software, may be generated based on a cache miss rate, or may be generated based on one or more other operating conditions.
[0026] 4, an embodiment of an exemplary wavefront 405 and an exemplary instruction sequence 410 are shown. The wavefront 405 is intended to illustrate an example of a wavefront according to one embodiment. The wavefront 405 includes 2×N work items, where “N” is a positive integer and is the number of lanes 425A-425N in the vector unit 420. The vector unit 420 may also be referred to as a SIMD unit or parallel processor. A first portion of the wavefront 405 includes work items W0-WN-1, and a second portion of the wavefront 405 includes work items WN-W2N-1. A single portion of the wavefront 405 is intended to execute on lanes 425A-425N of the vector unit 420 in a given instruction cycle. In other embodiments, the wavefront 405 may include other numbers of portions.
[0027] In one embodiment, N is 32 and the number of work items per wavefront is 64. In other embodiments, N may be other values. In an embodiment where N is 32, vector unit 420 includes 32 lanes, shown as lanes 425A-425N. In other embodiments, vector unit 420 may include other numbers of lanes.
[0028] Instruction sequence 410 illustrates an example of an instruction sequence. As shown in FIG. 4, instruction sequence 410 includes instructions 415A-415D, which represent any number and type of instructions of a shader program. For purposes of this description, it should be assumed that instruction 415A is the first instruction of instruction sequence 410, instruction 415B is the second instruction, instruction 415C is the third instruction, and instruction 415D is the fourth instruction. In other embodiments, an instruction sequence can have other numbers of instructions. Wavefront 405, instruction sequence 410, and vector unit 420 are reused in the discussion following FIGS. 5 and 6 below.
[0029] Referring to Figure 5, a diagram of one embodiment of a first operating mode of a processor is shown. The description of Figure 5 is a continuation of the description of Figure 4. In one embodiment, the size of the wave front is twice the size of the vector units 420. In another embodiment, the size of the wave front is an integer multiple of the size of the vector units 420. In these embodiments, the processor may implement different operating modes for determining how to use the vector units 420 to execute work items of the wave front.
[0030] In a first mode of operation, each instruction is executed on a different subset of the wavefront before the next instruction is executed on a different subset. For example, instruction 415A is executed for the first half of the wavefront (i.e., work items W0 through WN-1) during a first instruction cycle on lanes 425A through 425N of vector unit 420, and then instruction 415A is executed for the second half of the first wavefront (i.e., work items WN through W2N-1) during a second instruction cycle on lanes 425A through 425N. For example, during the first instruction cycle, work item W0 may be executed on lane 425A and work item W1 may be executed on lane 425B.
[0031] Next, instructions 415B are executed for the first half of the wavefront during the third instruction cycle of lanes 425A-425N, then instructions 415B are executed for the second half of the wavefront during the fourth instruction cycle of lanes 425A-425N. Next, instructions 415C are executed for the first half of the wavefront during the fifth instruction cycle of lanes 425A-425N, then instructions 415C are executed for the second half of the wavefront during the sixth instruction cycle of lanes 425A-425N. Next, instructions 415D are executed for the first half of the wavefront of lanes 425A-425N of vector unit 420 during the seventh instruction cycle of lanes 425A-425N, then instructions 415D are executed for the second half of the wavefront of lanes 425A-425N of vector unit 420 during the eighth instruction cycle of lanes 425A-425N. For the purposes of this description, it can be assumed that a second instruction cycle follows a first instruction cycle, and a third instruction cycle follows the second instruction cycle.
[0032] Referring now to Figure 6, a diagram of one embodiment of an operating mode of a second operation of a processor is shown. The description of Figure 6 continues from the description of Figure 5. In the second operating mode, the entire sequence of instructions 410 is executed on one portion of the wavefront before the entire sequence of instructions 410 is executed on the next portion of the wavefront.
[0033] For example, in a first instruction cycle, instruction 415A is executed for the first half of the wavefront (i.e., work items W0 through WN-1) on lanes 425A-425N of vector unit 420. Then, in a second instruction cycle, instruction 415B is executed for the first half of the wavefront of lanes 425A-425N. Then, in a third instruction cycle, instruction 415C is executed for the first half of the wavefront of lanes 425A-425N. Then, in a fourth instruction cycle, instruction 415D is executed for the first half of the wavefront of lanes 425A-425N.
[0034] Next, instruction sequence 410 is executed on the second half of the wavefront. Thus, in the fifth instruction cycle, instruction 415A is executed for the second half of the wavefront on lanes 425A-425N of vector unit 420 (i.e., work items WN-W2N-1). Next, in the sixth instruction cycle, instruction 415B is executed for the second half of the wavefront on lanes 425A-425N. Next, in the seventh instruction cycle, instruction 415C is executed for the second half of the wavefront on lanes 425A-425N. Next, in the eighth instruction cycle, instruction 415D is executed for the second half of the wavefront on lanes 425A-425N.
[0035] In another embodiment, if a wavefront has 4×N work items, instruction sequence 410 may be executed in a first quarter of the wavefront, then instruction sequence 410 may be executed in a second quarter of the wavefront, followed by a third quarter of the wavefront, a fourth quarter of the wavefront, etc. Other wavefronts of vector units having other sizes and / or other numbers of lanes may be utilized in a similar manner for the second mode of operation.
[0036] Referring to FIG. 7, one embodiment of a method 700 for scheduling instructions on a processor is shown. For purposes of illustration, the steps in this embodiment and in FIGS. 8-9 are shown in a sequential order. However, it should be noted that in various embodiments of the described method, one or more of the described elements may be performed simultaneously, in a different order than shown, or omitted entirely. Other additional elements may also be performed, as desired. Any of the various systems, apparatus, or computing devices described herein may be configured to perform the method 700.
[0037] The processor receives the wavefront and instructions of the shader program for execution (block 705). In one embodiment, the processor includes at least a number of execution units, a scheduler, a cache, and a number of GPRs. In one embodiment, the processing units are GPUs. In other embodiments, the processor is any of a variety of other types of processors (e.g., DSPs, FPGAs, ASICs, multi-core processors). In one embodiment, the number of work items of the wavefront is greater than the number of execution units of the processor. For example, in one embodiment, the wavefront includes 64 work items and the processor includes 32 execution units. In this embodiment, the number of work items of the wavefront is equal to twice the number of execution units. In other embodiments, the wavefront can include other numbers of work items and / or the processor can include other numbers of execution units. In some cases, the processor receives multiple wavefronts for execution. In these cases, method 700 can be performed multiple times for multiple wavefronts.
[0038] The processor then determines whether a first indicator is detected (conditional block 710). In one embodiment, the first indicator is a setting or parameter declared within the software instructions, which setting or parameter specifies the operating mode that the processor utilizes. In another embodiment, the first indicator is generated based on a cache miss rate of the wavefront. In other embodiments, other types of indicators are possible and contemplated.
[0039] If the first indicator is detected (conditional block 710: "yes"), the processor schedules the execution units to execute the first instruction on the first and second portions of the wavefront, and then schedules the execution units to execute the second instruction on the first portion of the wavefront (block 715). The processor may follow this same pattern of scheduling instructions for the remainder of the plurality of instructions as long as the first indicator is detected. If the first indicator is not detected (conditional block 710: "no"), the processor schedules the execution units to execute the first and second instructions on the first portion of the wavefront, and then schedules the execution units to execute the first instruction on the second portion of the wavefront (block 720). The processor may follow this same pattern of scheduling instructions for the remainder of the plurality of instructions as long as the first indicator is not detected. The processor may also share a portion of the GPR between the first portion of the wavefront and the second portion of the wavefront if the first indicator is not detected (block 725). After blocks 715 and 725, the method 700 ends.
[0040] Referring to FIG. 8, one embodiment of a method 800 for determining an operating mode to use in a parallel processor is illustrated. A control unit of the processor determines a cache miss rate for a wavefront (block 805). Depending on the embodiment, the control unit may determine a cache miss rate for a portion of the wavefront or for the entire wavefront at block 805. Depending on the embodiment, the control unit may be implemented utilizing any suitable combination of hardware and / or software. In one embodiment, the control unit predicts a cache miss rate for the wavefront. In another embodiment, the control unit receives a software-generated indicator, the indicator specifying a cache miss rate for the wavefront. The control unit then determines whether the cache miss rate for the wavefront is less than a threshold (conditional block 810). Alternatively, if the control unit receives a software-generated indicator, the indicator may specify whether the cache miss rate is less than a threshold. In one embodiment, the threshold is programmable. In another embodiment, the threshold is pre-determined.
[0041] If the cache miss rate for the wavefront is less than the threshold (conditional block 810: "yes"), the processor utilizes a first mode of operation in processing the wavefront (block 815). In one embodiment, the first mode of operation includes issuing each instruction in all parts of the wavefront before moving on to the next instruction of the shader program. If the cache miss rate for the wavefront is equal to or greater than the threshold (conditional block 810: "no"), the processor utilizes a second mode of operation in processing the wavefront (block 820). In one embodiment, the second mode of operation includes executing a set of instructions in a first part of the wavefront, then executing the same set of instructions in a second part of the wavefront, etc., until all parts of the wavefront have been executed. After blocks 815 and 820, the method 800 ends.
[0042] Referring to FIG. 9, another embodiment of a method 900 for utilizing different operating modes for a parallel processor is shown. A control unit of the processor determines the operating mode of the processor (block 905). Depending on the embodiment, the control unit may be implemented using any suitable combination of hardware and / or software. The criteria utilized by the control unit to determine the selected operating mode may vary from embodiment to embodiment. An example of the criteria that may be utilized is described in FIG. 8 in the description of method 800. Other examples of criteria that may be utilized to select the operating mode of the processor are possible and contemplated.
[0043] If the control unit selects the first mode of operation (conditional block 910: "first"), the processor does not share registers between different subsets of wavefronts being processed by the processor (block 915). If the control unit instead selects the second mode of operation (conditional block 910: "second"), the control unit shares one or more registers between different subsets of wavefronts being processed by the processor (block 920). For example, in one embodiment, sharing registers includes the processor using a shared portion of a register file for a first portion of a wavefront for a first set of instructions. The processor then reuses the shared portion of the register file for a second portion of the wavefront. If there are more than two portions of the wavefront, the processor reuses the shared portion of the register file for additional portions of the wavefront. After block 920, the method 900 ends.
[0044] It should be noted that in some embodiments, the processor may have more than two operating modes. In these embodiments, the conditional block 910 may be applied such that a first subset of the operating modes (e.g., first mode, third mode, seventh mode) follows the "first" leg, and a second subset of the operating modes (e.g., second mode, fourth mode, fifth mode, sixth mode) follows the "second" leg shown in FIG. 7. Alternatively, in another embodiment, the size of the portion of the register file that is shared may vary according to different operating modes. For example, for the second mode, a first number of registers are shared, for the third mode, a second number of registers are shared, for the fourth mode, a third number of GPRs are shared, etc.
[0045] 10-15, embodiments of a computing system are disclosed in which a shader program is analyzed for register usage for each program section and then modified to control an operating mode of the computing system's GPU for each program section based on the register usage. In at least one embodiment, the computing system uses a compiler in a CPU or other processing unit to analyze some or all of each section of a shader program to determine the number of VGPRs expected to be required to execute the corresponding section for an entire wavefront (i.e., for all work items of the wavefront). For each section identified as expected to require no more than a certain threshold number of VGPRs, the compiler refrains from modifying the shader program such that the section is executed by the computing system's GPU in a mode in which the GPU schedules execution units to execute each instruction for all portions of the wavefront (i.e., the entire wavefront) before executing the next instruction for all portions of the wavefront, as described above. For each section identified as expected to require a number of VGPRs exceeding a certain threshold, the compiler modifies the shader program so that the section is executed in a different mode in which the GPU schedules an execution unit to execute a set of instructions of the shader program for a first portion of the wavefront (i.e., two or more instructions) before the GPU schedules an execution unit to execute this same set of instructions for a second portion of the same wavefront, as described above. If the wavefront is composed of three or more portions, in this mode the GPU schedules an execution unit to execute a set of instructions for a third portion of the wavefront after scheduling an execution unit for the second portion of the wavefront, and so on. For ease of reference, the former mode is referred to hereinafter as "normal mode" (identified as "first mode" in the above description of Figures 1-9) and the latter mode is referred to hereinafter as "sub-vector loop mode" (identified as "second mode" in the above description of Figures 1-9).The sections of a shader program used in this technique may comprise individual instructions, logical sequences or blocks of instructions (e.g., instruction loops or subroutines), etc. In one embodiment, the particular threshold applied to the use of registers in each section to determine the operating mode used to execute the section is based on the number of VGPRs available in each execution unit.
[0046] 10, an exemplary software architecture 1000 for use in the computing system 100 of FIG. 1 for dynamic sub-vector loops is shown. The software architecture 1000 includes an operating system (OS) 1002 that supports the execution of one or more software applications 1004 in cooperation with one or more processing units 115 (e.g., CPUs) and a GPU 130. The OS 1002 and the software applications 1004, as well as much of the data utilized by the processing units 115 and some of the data utilized by the GPU 130, typically reside in the system memory 150.
[0047] The software application 1004 includes one or more sets of executable instructions 1006 and one or more shader programs 1008. The sets of executable instructions 1006 represent one or more programs compiled into machine code suitable for execution on the processing unit 115. Each shader program 1008 (also commonly known as a "compute kernel" or simply a "shader") is a program that represents a task or workload intended to be performed at least in part by the GPU 130, and typically multiple instances of a shader program 1008 are executed in parallel by two or more compute units 145 (FIG. 1) of the GPU 130. Such shader programs may be graphics-related, such as pixel shaders, vertex shaders, geometry shaders, tessellation shaders, etc., or may be machine learning (ML) shaders or other general computation shaders.
[0048] The OS 1002 includes an OS kernel 1010, one or more kernel mode drivers 1012, one or more application programming interfaces (APIs) 1014, and one or more user mode drivers 1016. The OS kernel 1010 represents the functional core of the OS 1002 and is responsible for boot initialization, memory allocation / deallocation, input / output control, and other basic hardware control, as well as facilitating the execution of software applications 1004. The kernel mode drivers 1012 manage the general operation of the GPU 130 hardware in system memory 150 (not shown) which facilitates task management of commands from the processing unit 115 to the GPU 130, including initializing the GPU 130, setting display modes, managing mouse hardware, managing allocation / deallocation of physical memory for the GPU 130, and a command buffer (not shown).
[0049] A user mode driver 1016 acts as an interface to the GPU 130 for one or more shader programs 1008 of a software application 1004. However, to facilitate hardware abstraction, the shader programs 1008 are typically not implemented in the software application 1004 as machine-readable code (i.e., "native" code), but rather as OpenGL™ Shading Language (GLSL) or High-Level Shading Language (HLSL) syntax, or partially compiled bytecode, such as the Standard Portable Intermediate Representation (SPIR) bytecode format, or source code (i.e., human-readable syntax) that depends on one or more APIs 1014, such as the OpenCL™ API, OpenGL™ API, Direct3D™ API, CUDA™ API, and / or associated libraries. Because the shader programs 1008 are not in a native code format, the user mode driver 1016 employs a shader compiler 1018 that operates to perform run-time compilation (also known as real-time compilation or just-in-time (JIT) compilation) of a source code or byte code representation of the shader programs 1008 into machine-readable code executable by the GPU 130. In other embodiments, an offline compiler is used to compile the code representing the shader programs 1008 into executable native code. The compiled executable code representation of the shader programs 1008 is then provided by the user mode driver 1016 to the GPU 130 via a command buffer (not shown) implemented in the system memory 150 and managed, for example, by the scheduler unit 245 (FIG. 2).
[0050] 11, an example of an embodiment of a shader compiler 1018 and its method of operation 1100 for modifying a shader program 1008 to implement a software-controlled dynamic sub-vector loop is shown. In the illustrated embodiment, the shader compiler 1018 includes a set of instructions for operating the processing unit 115 to perform a set of tasks at run-time, which are logically organized as a front-end stage 1102, a dynamic sub-vector modification stage 1104, and a back-end stage 1106. The shader program 1008 is provided to the front-end stage 1102 in the form of human readable source code or in the form of partially compiled bytecode, depending on the implementation. The front-end stage 1102 may perform one or more initial preparation processes, such as lexical analysis, syntactic analysis, and semantic analysis, and then generate the shader program 1008, which may include, for example, converting human readable source code to bytes, generating an intermediate representation 1108 of the shader program 1008, or converting from a high-level shader language to a low-level shader language. In the sub-vector modification stage 1104, the shader compiler 1018 applies one or more register usage analysis techniques to identify register usage per section, and then modifies the intermediate representation 1108 of the shader program 1008 to implement software-controlled selective implementation of the sub-vector loops per section, thereby generating a modified representation 1110 of the shader program 1008. This process is described in more detail below with reference to the method 1100. The modified representation 1110 is then processed by a back-end stage 1106, which converts the modified representation 1110 into one or more shader objects represented in machine language for the GPU 130 and links the one or more objects to generate an executable machine language representation of the modified shader program, identified herein as a "native-code shader 1112."The native-coded shaders 1112 are then passed to the GPU 130 via a command buffer in memory 150, so that one or more compute units 145 execute the executable shader representations, i.e., the native-coded shaders 1112, in parallel.
[0051] As mentioned above, the number of work items for a wavefront may exceed the number of execution units of the SIMD cores of GPU 130, making scheduling of the wavefront's instructions difficult. Similarly, in some embodiments, a wavefront may operate on data structures, such as vertex buffers or arrays of other structures, that are larger than the caches allocated to store the data for the data structures, resulting in inefficient execution performance if GPU 130 were to attempt to schedule the entire wavefront for execution in parallel. Thus, as mentioned above and further described below, GPU 130 employs at least two distinct modes of operation, one of which involves GPU 130 executing each instruction in parallel for the entire wavefront (i.e., a "normal mode" or "first mode") and another mode in which GPU 130 executes a set of instructions in parallel on one portion of the wavefront, which then executes this same set of instructions on another portion of the wavefront (i.e., a "sub-vector loop mode" or "second mode"). In some embodiments, GPU 130 includes a controller or other hardware-based mechanism for independently identifying when it is advantageous to switch from a normal mode of operation to a sub-vector loop mode of operation or vice versa. However, in other embodiments, this mode selection process is instead implemented at compile time, thus providing software-implemented dynamic mode switching to better optimize the execution of different sections of a shader program. Method 1100 of FIG. 11 illustrates an exemplary implementation of such a software-based technique.
[0052] As described below, in some embodiments, the method 1100 relies on comparing the VGPR usage of each section of the shader program 1008 to a particular VGPR usage to determine whether to implement a sub-vector loop mode to execute the corresponding section. Thus, as part of the initialization, at block 1120, the computing system 100 identifies a particular VGPR usage threshold to be used. In one embodiment, the VGPR usage threshold is pre-identified and fixed, such as by programming a fuse or one-time programmable element, by setting the VGPR usage threshold as a constant in the code implementing the shader compiler 1018. In other embodiments, the VGPR usage threshold is user programmable or otherwise programmable, such as through programming a guest-accessible register. Additionally, in other embodiments, the computing system 100 dynamically identifies or calculates the VGPR usage threshold for each iteration of the method 1100. In any of these approaches, the VGPR usage may be set based on the maximum number of VGPRs available to each execution unit. For example, assume that each execution unit has 12 dedicated VGPRs. In this example, the VGPR usage threshold may be set to be equal to the number of dedicated VGPRs (ie, VGPR usage threshold=12VGPRs), or may be set to a fixed percentage.
[0053] In block 1122, the dynamic subvector modification state 1104 of the shader compiler 1018 analyzes the intermediate representation 1108 of the shader program 1008 (or, in some cases, the shader program 1008 itself) to identify, on a section-by-section basis, the expected VGPR usage for each section of the shader program 1008. To this end, the shader program 1008 (or its intermediate representation 1108) may be logically segmented into multiple sections using any of a variety of criteria or combinations of criteria. For example, all instructions of a loop may be designated collectively as a section, as well as instructions of a subroutine or function call. Similarly, for instructions not contained in a loop or subroutine, sections may be designated in order based on a particular number of instructions. Additionally, in other embodiments, a section may be identified as all instructions that occur between an instance where VGPR usage by the shader program 1008 exceeds a VGPR usage threshold and the next instance where VGPR usage falls below the VGPR usage threshold (i.e., the section is defined in terms of VGPR usage and the VGPR usage threshold). VGPR usage by the shader program 1008 for any given point in time or any given section of an instruction sequence may be determined using any of a variety of techniques. For example, in some embodiments, the shader compiler 1018 determines VGPR usage using any of a variety of well-known or proprietary register "liveness" analysis techniques (also commonly referred to as "live variable analysis").
[0054] At block 1124, the shader compiler 1018 compares the VGPR usage of the selected section of the shader program 1008 to the VGPR usage threshold determined at block 1120. If the VGPR usage of the selected section exceeds the VGPR usage threshold, then at block 1126, the shader compiler 1018 modifies the shader program 1008 such that when the GPU 130 executes the resulting modified shader program (e.g., the modified representation 1110 of the shader program 1008), the GPU 130 is configured to operate in a sub-vector loop mode to execute instructions of the selected section.
[0055] As will be described in more detail below with reference to Figures 13-15. This modification can take any of a variety of forms. If GPU 130 provides direct hardware support for sub-vector loop mode, the modification of shader program 1008 can include inserting a first instruction or other command before the instructions of the selected section and inserting a second instruction or other command after the instructions of the selected section, the first instruction and the second instruction being part of the ISA of GPU 130, which when executed by GPU 130 configures GPU 130 to switch to operating in sub-vector loop mode and to switch to normal operation. In implementations where GPU 130 does not provide direct hardware support for sub-vector loops, or where utilizing such hardware support is undesirable, the modification of shader program 1008 can instead take the form of replacing the instructions of the selected section with a set of instructions or code that introduces either a loop that causes the instructions of the selected section to be executed sequentially for each portion of the wavefront, or by explicitly duplicating the instructions of the selected section for each portion of the wavefront. Returning to block 1124, if the VGPR usage of the selected section does not exceed the VGPR usage threshold, the shader compiler 1018 refrains from modifying the shader program 1008 of the selected section (and thus instructs the GPU 130 to execute the selected section in normal mode since there are no modifications).
[0056] At block 1128, the shader compiler 1018 determines whether there are any further sections to consider. If so, the shader compiler 1018 selects the next section of the shader program 1008 and repeats the process of blocks 1124 and 1126 for the selected section. If not, at block 1130, the shader compiler 1018 issues the resulting modified shader program (e.g., as a modified representation 1110) from the dynamic subvector modification stage 1104 to the back-end stage 1106 for further processing and generation of a native code shader 1112 from the modified shader program.
[0057] 12, a diagram 1200 is shown illustrating a simplified embodiment of the method 1100. The horizontal axis of the diagram 1200 represents a sequence of ten sections of the shader program 1008, designated section 0 through section 9, with section 0 being executed first and section 9 being executed last in the program order of the shader program 1008. The vertical axis of the diagram 1200 represents VGPR usage of each section, with the maximum VGPR usage designated as "K" and the VGPR usage threshold in this example designated as horizontal line 1202 (and thus also referred to herein as "VGPR usage threshold 1202"). Thus, line 1204 represents the VGPR usage of sections 0-9 as determined by the shader compiler 1018 via a variable liveness analysis or other type of register usage analysis of the shader program 1008 (block 1122 of the method 1100). In this example, sections 0, 1, 2, 6, 8, and 9 do not exceed the VGPR usage threshold 1202, but sections 3, 4, 5, and 7 do exceed the VGPR usage threshold 1202. Thus, during compilation, the shader compiler 1018 modifies the shader program 1008 such that sections 3, 4, 5, and 7 are executed by the GPU 130 using sub-vector loop mode and refrains from making any modifications to the shader program 1008 for sections 0, 1, 2, 6, 8, and 9 such that these sections are executed by the GPU 130 using normal mode.
[0058] As discussed above, modifying the shader program 1008 such that sections having VGPR usage above a threshold are executed by the GPU 130 using the sub-vector loop mode may be implemented in any of a variety of ways. With reference to FIG. 13, an ISA-based command modification scheme is shown. In some embodiments, the GPU 130 provides a hardware implementation of the sub-vector loop mode (i.e., hard-coded logic in the GPU 130 can implement the multi-pass approach of the sub-vector loop mode without reconfiguring the shader program 1008). In such an implementation, the shader program 1008 may be modified to switch between normal mode and sub-vector loop mode through the insertion of mode switch instructions into the shader program 1008 by the shader compiler 1018, such that these mode switch instructions become part of the ISA of the GPU 130 and, when executed by the GPU 130, trigger the GPU 130 to switch modes. For illustrative purposes, block 1300 of FIG. 13 shows a portion of the code of the shader program 1008 representing sections 6, 7, and 8 from the example of FIG. 12. Section 6 is represented by a set of instructions 1302 having instructions INST0-INST2, section 7 is represented by a set of instructions 1304 having instructions INST3-INST6, and section 8 is represented by a set of instructions 1306 having instructions INST7-INST9.
[0059] 12, sections 6 and 8 have VGPR usage that does not exceed a particular VGPR usage threshold 1202 (FIG. 12), while section 7 has VGPR usage that exceeds a particular VGPR usage threshold 1202. Thus, in this example, the shader compiler 1018 modifies the shader program 1008 so that a set of instructions 1304 representing section 7 is executed in subvector loop mode by inserting a subvector loop mode begin instruction 1308 (labeled “S_SUBVECTOR_START”) before the first instruction of set 1304 in program order (INST3), and inserting a subvector loop mode end instruction 1310 (labeled “S_SUBVECTOR_END”) after the last instruction of set 1304 (INST6), to generate a modified set of instructions 1314 in the resulting modified shader program represented by block 1312. Enter subvector loop mode instruction 1308 is an ISA support instruction that, when executed by GPU 130, triggers GPU 130 to operate in subvector loop mode. Conversely, exit subvector loop mode instruction 1310 is an ISA support instruction that, when executed by GPU 130, triggers GPU 130 to exit subvector loop mode and begin execution in normal mode. Thus, set of instructions 1302 representing section 6 of the modified shader program is executed by GPU 130 in normal mode, GPU 130 switches to subvector loop mode and executes set of instructions 1304 representing section 7, GPU 130 returns to normal mode and executes set of instructions 1306 representing section 8.
[0060] This example illustrates the modifications that are performed when a single section whose VGPR usage exceeds the VGPR usage threshold is executed between two sections whose VGPR usage does not exceed the VGPR usage threshold. If multiple sections are executed consecutively while the GPU 130 is in the subvector loop mode, rather than inserting a subvector loop mode start instruction 1308 and a subvector loop mode end instruction 1310 separately for each such section, the shader compiler 1018 can instead insert a single subvector loop mode start instruction 1308 before the first instruction of the first section of the sequence and a single subvector loop mode end instruction 1310 after the last instruction of the last section of the sequence. For example, when sections 3-5 each exceed the VGPR usage threshold 1202, the shader compiler 1018 can insert a single subvector loop mode start instruction 1308 before the first instruction of section 3 and a single subvector loop mode end instruction 1310 after the last instruction of section 5 in generating the modified shader program. Further, while FIG. 13 was discussed above, in FIGS. 14 and 15 described below, based on modifications to the shader program 1008, sections identified as having expected register usage exceeding a threshold are modified to reconfigure the GPU 130 from normal mode as the default mode to sub-vector loop mode; in other embodiments, the sub-vector loop mode is the default mode, and the shader compiler 1018 instead modifies sections having register usage that does not exceed a threshold to trigger the GPU 130 to switch from sub-vector loop mode as the default mode to normal mode for executing such sections.
[0061] Referring to FIG. 14, a replication-based modification scheme for software-implemented sub-vector loops is shown. If the GPU 130 does not provide a hardware implementation for the sub-vector loop mode, or if it is not desirable to utilize such a hardware implementation, the GPU 130 may be configured to provide a sub-vector loop mode via software reconfiguration of the shader program 1008, in which instructions of a section executed in the sub-vector loop mode are replicated once for each portion of the wavefront in which the section is executed separately. For illustrative purposes, block 1400 of FIG. 14 shows a portion of code of the shader program 1008 representing sections 6, 7, and 8 from the example of FIG. 13. Section 6 is represented by a set of instructions 1402 having instructions INST0-INST2, section 7 is represented by a set of instructions 1404 having instructions INST3-INST6, and section 8 is represented by a set of instructions 1406 having instructions INST7-INST9. In the example of FIG. 12, sections 6 and 8 have VGPR usage that does not exceed a particular VGPR usage threshold 1202 (FIG. 12), while section 7 has VGPR usage that exceeds a particular VGPR usage threshold 1202.
[0062] Thus, in this example, shader compiler 1018 modifies shader program 1008 to effect execution of set 1304 in only one portion of the wavefront at a time by replicating set 1304 in the modified shader program for each portion of the wavefront, and by inserting instructions that restrict execution of each replicated set 1404 of instructions to a corresponding wavefront. To illustrate, as shown by block 1412 representing a resulting portion of the modified shader program, set 1404 of instructions in section 7 is replaced with a replacement set 1414 of instructions, which includes two replicas of set 1404 of instructions in the resulting modified shader program, shown as set 1404_L and set 1404_H. Replicated set 1404_H is inserted in the bottom half of the wavefront, and replicated set 1404_H is inserted in the top half of the wavefront. As shown, replacement set of instructions 1414 further includes instructions that restrict execution of instructions in replicated set 1404_L to the first half of the wavefront and instructions in replicated set 1404_H to the second half of the wavefront.
[0063] Referring to FIG. 15, a loop-based modification scheme for a software-implemented sub-vector loop is shown. Rather than duplicating a set of instructions for a section that executes in a software-based sub-vector loop mode, the GPU 130 may instead be configured to provide the sub-vector loop mode via a software reconfiguration of the shader program 1008 in which instructions are added to the section that executes in the sub-vector loop mode, and an iteration of the section is performed for each portion of the wavefront. For illustrative purposes, block 1500 of FIG. 15 shows a portion of code for the shader program 1008 representing sections 6, 7, and 8 from the example of FIG. 12. Section 6 is represented by a set of instructions 1402 having instructions INST0-INST2, section 7 is represented by a set of instructions 1404 having instructions INST3-INST6, and section 8 is represented by a set of instructions 1506 having instructions INST7-INST9. As discussed above, sections 6 and 8 have VGPR usage that does not exceed a particular VGPR usage threshold 1202 (FIG. 12), while section 7 has VGPR usage that exceeds a particular VGPR usage threshold 1202.
[0064] In this example, the shader compiler 1018 modifies the shader program 1008 by including instructions that effectively place the set of instructions 1504 in a loop of the modified shader program to provide for execution of the set of instructions 1504 on only one portion of the wavefront at a time, such that each iteration of the loop executes sequentially on a different portion of the same wavefront. To illustrate, as shown by block 1512 representing the resulting portion of the modified shader program, the set of instructions 1504 of section 7 has been replaced from the replacement set of instructions 1414, which includes instructions 1516 that place the set of instructions 1504 in a loop, resulting in each iteration of the loop executing on a separate portion of the wavefront. As shown, these instructions 1516 include function calls "S_SUBVECTOR_LOOP_END" and "S_SUBVECTOR_LOOP_BEGIN" that represent the following functions: [Table 1]
[0065] Thus, if EXEC_HI=0, set 1504 of instructions is executed only once, and EXEC_LO is stored in S0 and restored at the end, which is zero in either case. If EXEC_LO is zero at the start, the same result occurs. If both halves of the EXEC mask are non-zero, the low pass is executed first (storing EXECHI in S0), then EXECHI is restored and saved from EXECLO, and the pass is executed again. Then EXECLO is restored at the end of the second pass. The "pass #" is encoded by noting which half of EXEC is zero.
[0066] In some embodiments, certain aspects of the techniques described above may be implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions stored or otherwise tangibly embodied in a non-transitory computer-readable storage medium. The software may include instructions and specific data that, when executed by one or more processors, operate the one or more processors to perform one or more aspects of the techniques described above. Such non-transitory computer-readable storage media may include, for example, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs), magnetic media (e.g., floppy disks, magnetic tapes, magnetic hard drives), volatile memory (e.g., random access memory (RAM), cache), non-volatile memory (e.g., read-only memory (ROM), flash memory), or microelectromechanical systems (MEMS)-based storage media. The computer readable storage medium may be embedded in a computing system (e.g., system RAM or ROM), fixedly attached to a computing system (e.g., a magnetic hard drive), removably attached to a computing system (e.g., an optical disk or Universal Serial Bus (USB)-based flash memory), or coupled to a computer system via a wired or wireless network (e.g., Network Accessible Storage (NAS)). The executable instructions stored on the non-transitory computer readable storage medium may be source code, assembly language code, object code, or other instruction formats that are interpreted or otherwise executable by one or more processors.
[0067] It should be noted that not all of the activities or elements described above in the overall description are required, that some of the specific activities or devices may not be required, and that one or more additional activities may be performed or elements may be included in addition to those described. Furthermore, the order in which the activities are listed is not necessarily the order in which the activities are performed. Also, the concepts have been described with reference to specific embodiments. However, those skilled in the art will appreciate that various modifications and changes can be made without departing from the scope of the invention as set forth in the following claims. Thus, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of the present disclosure.
[0068] Benefits, other advantages, and solutions to problems have been described above with respect to specific embodiments. However, any feature or features that may cause or make more pronounced any benefit, advantage, or solution should not be construed as critical, essential, or essential features of any or all of the claims. Moreover, the specific embodiments disclosed above are merely illustrative, and the disclosed invention may be modified and embodied in different but similar manners apparent to those skilled in the art having the benefit of the teachings herein. No limitations are intended to the details of construction or design shown herein, other than as set forth in the following claims. It is therefore apparent that the particular embodiments described above may be altered or modified, and all such variations are considered to be within the scope of the disclosed invention. Accordingly, the protection sought herein is as set forth in the following claims.
Claims
1. 1. A computer-implemented method comprising: providing a first processor configurable to operate in at least a first mode and a second mode, in which in the first mode the first processor is operative to execute instructions for an entire wavefront before executing a next instruction for the entire wavefront, and in the second mode the first processor is operative to execute a set of instructions for a portion of a wavefront before executing a set of instructions for another portion of the same wavefront; Configuring a shader program executed by the first processor to include at least one indicator associated with either the first mode or the second mode; executing the shader program on the first processor, wherein an operating mode of the first processor during execution of the shader program corresponds to the at least one indication present in the shader program. method.
2. the at least one indication includes a command for the first processor to operate in the second mode.
2. The method of claim 1.
3. Configuring the shader program includes: analyzing the shader program at a second processor to identify register usage of different sections of the shader program; and selectively associating an index within the shader program with each section of the shader program based on the section's identified register usage. The method of claim 1 or 2.
4. Selectively associating the indicator includes selectively associating the indicator based on a comparison of usage of the identified register to a particular threshold. The method of claim 3.
5. the particular threshold is set based on the number of registers available in each execution unit of the first processor; The method of claim 4.
6. 1. A computer-implemented method comprising: modifying a shader program in a compiler of a processing system such that a first section of the shader program that is predicted to require a large number of registers to execute across a wavefront that exceeds a particular threshold is configured to execute in a first mode, and a second section of the shader program that is predicted to require a large number of registers to execute across a wavefront that does not exceed the particular threshold is configured to execute in a second mode; executing the modified shader program in a first processor of the processing system; In the first mode, the first processor schedules execution units of the first processor to execute a set of instructions of the shader program for a first portion of the entire wavefront before scheduling the execution units of the first processor to execute a set of instructions for a second portion of the entire wavefront; In the second mode, the first processor schedules a plurality of execution units to execute a single instruction of the shader program on the first and second portions of the entire wavefront before scheduling the plurality of execution units to execute a next single instruction of the shader program on the first and second portions of the entire wavefront. method.
7. Modifying the shader program includes: In response to identifying a first section of the shader program, inserting a first command into the shader program before the identified first section, the first command configuring the first processor to operate in the first mode; and inserting a second command into the shader program after the identified first section, the second command configuring the first processor to operate in the second mode. The method of claim 6.
8. Modifying the shader program includes: in response to identifying a first section of the shader program, modifying the shader program to replace instructions of the identified first section with a first set of instructions representing a first iteration of instructions of the identified first section that are executed on the first portion of the entire wavefront, and a second set of instructions representing a second iteration of instructions of the identified first section that are executed on the second portion of the entire wavefront. The method of claim 6 or 7.
9. Modifying the shader program includes: in response to identifying a first section of the shader program, modifying the shader program to perform a first loop of instructions of the identified first section that are executed on the first portion of the entire wavefront, and a second loop of instructions of the identified first section that are executed on the second portion of the entire wavefront. The method of claim 6.
10. analyzing, in a compiler, the shader program to identify an expected register usage of each section of the shader program, where the expected register usage of a section indicates a number of registers expected to be required to execute a corresponding section on the first processor across a wavefront; The method of claim 6.
11. analyzing the shader program includes performing a register liveness analysis of the shader program. The method of claim 10.
12. the first portion of the entire wave-front includes a first half of the wave-front, and the second portion of the entire wave-front includes a second half of the wave-front; The method according to any one of claims 6 to 11.
13. the particular threshold is set based on the number of registers available in each execution unit of the first processor; The method according to any one of claims 6 to 12.
14. 1. A system comprising: a first processor configurable to operate in at least a first mode and a second mode, in which in the first mode the first processor operates to execute instructions for an entire wavefront before executing a next instruction for the entire wavefront, and in the second mode the first processor operates to execute a set of instructions for a portion of a wavefront before executing a set of instructions for another portion of the same wavefront; the first processor is configured to execute either the first mode or the second mode while executing the first shader program in response to at least one indicator present in a first shader program specifying either the first mode or the second mode. system.
15. the at least one indication includes a command for the first processor to operate in the second mode.
15. The system of claim 14.
16. a second processor coupled to the first processor and a memory; The second processor comprises: Analyzing a second shader program to identify register usage of different sections of the second shader program; modifying the second shader program by selectively associating an indication of the second shader program with each section of the second shader program based on usage of a register identified in the section; providing a modified second shader program for storage in the memory as the first shader program; [0023] 15. The system of claim 14.
17. the second processor is configured to selectively associate the indicator based on a comparison of the identified register usage to a particular threshold.
17. The system of claim 16.
18. 1. A system comprising: a first processor implementing a compiler configured to modify a shader program such that a first section of the shader program that is predicted to require a large number of registers to execute across a wavefront that exceeds a particular threshold is configured to execute in a first mode, and a second section of the shader program that is predicted to require a large number of registers to execute across a wavefront that does not exceed the particular threshold is configured to execute in a second mode; a second processor configured to execute the modified shader program, wherein in the first mode, the second processor schedules the execution units to execute a set of instructions of the shader program for a first portion of the entire wavefront before scheduling the execution units of the second processor to execute a set of instructions for a second portion of the entire wavefront, and in the second mode, the second processor schedules the execution units to execute a single instruction of the shader program for the first and second portions of the entire wavefront before scheduling the execution units to execute a next single instruction of the shader program for the first and second portions of the entire wavefront. system.
19. the first processor includes a central processing unit and the second processor includes a graphics processing unit; 20. The system of claim 18.
20. The first processor, In response to identifying a first section of the shader program, inserting a first command into the shader program before the identified first section, the first command configuring the second processor to operate in the first mode; inserting a second command into the shader program after the identified first section, the second command configuring the second processor to operate in the second mode; and modifying the shader program by 20. The system of claim 18.
21. The first processor, in response to identifying a first section of the shader program, modifying the shader program by replacing instructions of the identified first section with a first set of instructions representing a first iteration of instructions of the identified first section that are executed on the first portion of the entire wavefront and a second set of instructions representing a second iteration of instructions of the identified first section that are executed on the second portion of the entire wavefront.
20. The system of claim 18.
22. The first processor, in response to identifying a first section of the shader program, modifying the shader program by executing a first loop of instructions of the identified first section that are executed on the first portion of the entire wavefront and a second loop of instructions of the identified first section that are executed on the second portion of the entire wavefront.
20. The system of claim 18.
23. the compiler is configured to analyze the shader program to identify an expected register usage for each section of the shader program, the expected register usage for a section indicating a number of registers expected to be required to execute a corresponding section on the first processor across a wavefront; A system according to any one of claims 18 to 22.
24. the compiler is configured to analyze the shader program based on a register liveness analysis of the shader program; 24. The system of claim 23.
25. The particular threshold is set based on the number of registers available in each computing unit of the first processor. A system according to any one of claims 18 to 24.
26. In the first mode, the execution units of the second processor are configured to share a subset of registers. A system according to any one of claims 18 to 24.
Citation Information
Patent Citations
Method and apparatus for subdividing shader workloads in a graphics processor for efficient machine configuration
US20190206110A1
Variable wavefront size
WO2018156635A1