Method and system for partial wavefront combining

By adopting partial wavefront merging technology in vector processing machines, the problem of low resource utilization caused by inactive work items in some wavefronts is solved, and higher computing resource utilization and computing efficiency are achieved.

CN110716750BActive Publication Date: 2025-05-30ADVANCED MICRO DEVICES INC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201810758486.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2018-07-11
Publication Date
2025-05-30
Estimated Expiration
2038-07-11

AI Technical Summary

Technical Problem

In a vector processing machine, some wavefronts contain inactive work items, resulting in a decrease in utilization of computing resources.

Method used

Using partial wavefront merging technology, wavefronts containing inactive work items are detected and merged through partial wavefront managers, and merged them into one or more new wavefronts, thereby improving the resource utilization of SIMD units.

Benefits of technology

The utilization rate of arithmetic logic unit (ALU) in SIMD units is improved, resource waste is reduced, and computing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0001727314630000011
    Figure HDA0001727314630000011
  • Figure HDA0001727314630000021
    Figure HDA0001727314630000021
  • Figure HDA0001727314630000031
    Figure HDA0001727314630000031
Patent Text Reader

Abstract

Methods and systems for partial wavefront merging are described. A vector processing machine employs the partial wavefront merging to merge partial wavefronts into one or more wavefronts. The system includes a partial wavefront manager and a unified register. The partial wavefront manager detects wavefronts (hereinafter referred to as "partial wavefronts") including inactive work items and active work items in different single instruction multiple data ("SIMD") units, moves the partial wavefronts into one or more SIMD units and merges the partial wavefronts into one or more wavefronts. The unified register allows each active work item in the one or more merged wavefronts to access a previously allocated register in the original SIMD unit. Accordingly, the content of the unified register does not have to be copied to the SIMD unit executing the one or more merged wavefronts.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Vector processing machines, such as graphics processing units (GPUs), general-purpose graphics processing units (GPGPUs), and similar machines, use or include one or more parallel processing units configured to perform computations according to the single instruction multiple data (“SIMD“) paradigm. In these types of machines, a certain number of work items form a wavefront that runs in one SIMD unit.

[0002] A partial wavefront is a wavefront that includes some inactive work items. Partial wavefronts are common in applications and result in reduced utilization of the resources that make up the SIMD unit. For example, an OpenCl kernel can have a complex branching pattern. In some cases, some work items execute under one branch while the remaining work items are inactive. In another case, some work items execute under one branch while some other work items execute under another branch. The branches can have different execution times, which causes the work items with shorter execution times to be inactive while the work items with longer execution times are executing. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] A more detailed understanding can be obtained from the following description given by way of example in conjunction with the accompanying drawings, in which:

[0004] Figure 1 is a block diagram of an example apparatus according to certain implementations;

[0005] Figure 2 is according to certain implementations Figure 1 of the apparatus;

[0006] Figure 3 is a block diagram of a system with unified registers and a partial wavefront manager according to certain implementations;

[0007] Figure 4 is according to certain implementations for Figure 3 partial wavefront merging of the system shown in; and

[0008] Figure 5 is a block diagram of another system with unified registers and a partial wavefront manager according to certain implementations; and

[0009] Figure 6 is a block diagram of a system 600 that performs partial wavefront merging according to certain implementations.

[0010] DETAILED IMPLEMENTATIONS

[0011] Methods and systems for partial wavefront merging are described herein. A vector processing machine employs partial wavefront merging to merge partial wavefronts into one wavefront, where the partial wavefronts include some inactive work items. This results in higher utilization of computational resources in a single instruction multiple data ("SIMD") unit, such as arithmetic logic unit (ALU) utilization. The system includes a partial wavefront manager and a unified register structure. In an implementation, the unified register structure is a single register structure. In an implementation, the single register structure is composed of multiple registers. In an implementation, the unified register structure includes multiple registers, where each register is associated with an ALU or SIMD. In an implementation, each such register is composed of multiple registers. By way of example and not limitation, the unified register structure may be a general purpose register (GPR). The partial wavefront manager detects wavefronts (hereinafter referred to as "partial wavefronts") including inactive work items and active work items in different single instruction multiple data ("SIMD") units, moves the partial wavefronts to an appropriate number of SIMDs and merges the partial wavefronts into one or more wavefronts. The unified register structure allows each active work item in the partial wavefronts of the merged wavefronts to access the registers previously allocated in the original SIMD units. Thus, the contents of the previously allocated registers do not have to be copied to the SIMDs of the wavefronts performing the merge. This is in contrast to software solutions where the state of active work has to be moved from one thread to another.

[0012] Figure 1 FIG. is a block diagram of an example apparatus 100 that may implement one or more features of the present disclosure. Apparatus 100 may include, for example, a computer, a gaming device, a handheld device, a set-top box, a television, a mobile phone, or a tablet computer. Apparatus 100 includes a processor 102, a memory 104, a storage device 106, one or more input devices 108, and one or more output devices 110. Apparatus 100 may also optionally include an input driver 112 and an output driver 114. It should be understood that apparatus 100 may include Figure 1 additional components not shown in FIG.

[0013] In various alternatives, processor 102 includes a central processing unit (CPU), a graphics processing unit (GPU), a CPU and GPU located on the same die, or one or more processor cores, where each processor core may be a CPU or GPU. In various alternatives, memory 104 is located on the same die as processor 102, or is located separately from processor 102. Memory 104 includes volatile and / or non-volatile memory, such as random access memory (RAM), dynamic RAM, or a cache.

[0014] The storage device 106 includes fixed or removable storage devices such as hard disk drives, solid state drives, optical discs, or flash drives. The input device 108 includes, but is not limited to, a keyboard, keypad, touch screen, touchpad, detector, microphone, accelerometer, gyroscope, biometric scanner, or network connector (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals). The output device 110 includes, but is not limited to, a display, speaker, printer, haptic feedback device, one or more lights, antenna, or network connector (e.g., a wireless local area network card for transmitting and / or receiving wireless IEEE 802 signals).

[0015] The input driver 112 communicates with the processor 102 and the input device 108 and permits the processor 102 to receive input from the input device 108. The output driver 114 communicates with the processor 102 and the output device 110 and permits the processor 102 to send output to the output device 110. It should be noted that the input driver 112 and the output driver 114 are optional components, and in the absence of the input driver 112 and the output driver 114, the device 100 will operate in the same manner. The output driver 114 includes an acceleration processing device (“APD”) 116 coupled to the display device 118. The APD 116 is configured to receive compute commands and graphics rendering commands from the processor 102, process those compute and graphics rendering commands, and provide pixel output to the display device 118 for display. As described in further detail below, the APD 116 includes one or more parallel processing units configured to execute computations according to the single instruction multiple data (“SIMD”) paradigm. Thus, although various functions are described herein as being performed by or in conjunction with the APD 116, in various alternatives, the functions described as being performed by the APD 116 are additionally or alternatively performed by other computing devices having similar capabilities, which in some cases are not driven by a host processor (e.g., the processor 102) and which in some implementations are configured to provide (graphical) output to the display device 118. For example, it is contemplated that any processing system that executes processing tasks according to the SIMD paradigm can be configured to perform the functions described herein. Alternatively, it is contemplated that a computing system that does not execute processing tasks according to the SIMD paradigm performs the functions described herein.

[0016] Figure 2FIG. 0 is a block diagram of apparatus 100, showing additional details related to the execution of processing tasks on APD 116. Processor 102 maintains in system memory 104 one or more control logic modules for execution by processor 102. The control logic modules include operating system 120, kernel mode driver 122, and application 126. These control logic modules control various features of the operation of processor 102 and APD 116. For example, operating system 120 communicates directly with the hardware and provides a hardware interface for other software executing on processor 102. Kernel mode driver 122 controls the operation of APD 116 by accessing various functions of APD 116, for example, by providing an application programming interface ("API") to software (e.g., application 126) executing on processor 102. Kernel mode driver 122 also includes a just-in-time compiler that compiles programs for execution by processing components of APD 116, such as SIMD unit 138 discussed in more detail below.

[0017] APD 116 executes commands and programs for selected functions, such as graphics operations and non-graphics operations that may be suitable for parallel processing and / or out-of-order processing. APD 116 can be used to perform graphics pipeline operations such as pixel operations, geometric calculations, and rendering images to display device 118 based on commands received from processor 102. APD 116 also performs computational processing operations not directly related to graphics operations based on commands received from processor 102, such as operations related to video, physics simulation, computational fluid dynamics, or other tasks.

[0018] APD 116 includes computational unit 132, which includes one or more SIMD units 138 configured to perform operations in parallel according to the SIMD paradigm at the request of processor 102. The SIMD paradigm is a paradigm in which multiple processing elements share a single program control flow unit and program counter and thus execute the same program but can execute the program with different data. Although the term SIMD unit is used herein, a SIMD unit is a type of processing unit in which the processing unit includes multiple processing elements. In one example, each SIMD unit 138 includes sixteen lanes, where each lane executes the same instruction simultaneously with other lanes in SIMD unit 138 but can execute the instruction with different data. If not all lanes are needed to execute a given instruction, lanes can be predicted off. Prediction can also be used to execute programs with divergent control flow. More specifically, for programs with conditional branches or other instructions where the control flow is based on calculations performed by individual lanes, prediction of lanes corresponding to currently unexecuted control flow paths and serial execution of different control flow paths allow arbitrary control flow.

[0019] The basic execution unit in the compute unit 132 is a work item. Each work item represents a single instantiation of a program to be executed in parallel in a particular lane. Work items can be executed simultaneously as a "wavefront" on a single SIMD processing unit 138. One or more wavefronts are included in a "workgroup", which includes a set of work items designated to execute the same program. A workgroup can be executed by executing each of the wavefronts that make up the workgroup. In an alternative, wavefronts are executed sequentially on a single SIMD unit 138 or partially or fully in parallel on different SIMD units 138. A wavefront can be considered the largest set of work items that can be executed simultaneously on a single SIMD unit 138. Thus, if a command received from the processor 102 indicates that a particular program is to be parallelized to an extent that the program cannot be executed simultaneously on a single SIMD unit 138, then the program is decomposed into two or more wavefronts, parallelized across two or more SIMD units 138 or serialized (or parallelized and serialized as needed) on the same SIMD unit 138. The scheduler 136 is configured to perform operations related to scheduling the respective wavefronts on different compute units 132 and SIMD units 138.

[0020] The parallelism provided by the compute unit 132 is suitable for graphics-related operations such as pixel value calculations, vertex transformations, and other graphics operations. Thus, in some cases, the graphics pipeline 134 that receives graphics processing commands from the processor 102 provides compute tasks to the compute unit 132 for parallel execution. For example, the graphics processing pipeline 134 includes stages each performing a particular function. A stage represents a breakdown of the functionality of the graphics processing pipeline 134. Each stage is partially or fully implemented as a shader program executed in a programmable processing unit, or partially or fully implemented as fixed-function, non-programmable hardware external to the programmable processing unit.

[0021] The compute unit 132 is also used to execute compute tasks that are not graphics-related or not executed as part of the "normal" operation of the graphics pipeline 134 (e.g., custom operations that are executed to supplement the processing performed for the operations of the graphics pipeline 134). An application 126 or other software executing on the processor 102 transfers a program defining such compute tasks to the APD 116 for execution.

[0022] Figure 3FIG. 0 is a block diagram of a system 300 that performs partial wavefront merging according to some implementations. The system 300 includes SIMD units 305, e.g., SIMD0, SIMD1, SIMD2, …, SIMDN, where each SIMD unit 305 includes at least one arithmetic logic unit (ALU) 310. Each of SIMD0, SIMD1, SIMD2, …, SIMDN also includes unified GPRs 315, e.g., unified GPR0, unified GPR1, unified GPR2, …, unified GPRN. Each of unified GPR0, unified GPR1, unified GPR2, …, unified GPRN and the ALU 310 are connected or communicate (collectively referred to as “connected” hereinafter) via a network 320. In an implementation, the network 320 is a mesh network. In another implementation, the unified GPRs 315 are connected via a shared bus. In yet another implementation, the unified GPRs 315 are daisy-chained together. A partial wavefront manager 325 is connected to each of SIMD0, SIMD1, SIMD2, …, SIMDN. Figure 3 The system 300 shown in Figure 1 and Figure 2 can be implemented using the system shown and the elements presented therein.

[0023] The partial wavefront merging feature in the system 300 can be enabled or disabled via one or more register settings. In an implementation, built-in functions are added to the instruction set to set the register settings. The register settings can be set by a driver, a shader program, a kernel, etc.

[0024] As described below with respect to Figure 4 and Figure 5 the partial wavefront manager 325 detects partial wavefronts from the SIMD units 305 by determining which of the work items (corresponding to SIMD lanes) are active or inactive. The partial wavefront manager 325 can look at any number of SIMD units 305 for the purpose of merging wavefronts. In an implementation, the partial wavefront manager 325 moves the partial wavefronts into one SIMD unit 305 and merges the partial wavefronts into one wavefront, where the SIMD unit 305 can have any combination of active and inactive lanes after the merge. The unified GPRs 315 allow each active work item in the merged wavefront to access the previously allocated GPRs in the original SIMD units. Although the description below herein is with respect to the foregoing implementation, the description applies to other implementations without departing from the scope of this specification. In another implementation, e.g., the partial wavefront manager 325 can merge the partial wavefronts into multiple SIMD units.

[0025] Specifically, the partial wavefront manager 325 records the stateId, execution mask, and program counter for each wavefront in the SIMD unit 305. The stateId indicates the wavefronts sharing a particular or specific shader program. The execution mask for each wavefront indicates which work items are active. The program counter indicates which instruction of the shader program is being executed.

[0026] Figure 4 is a flowchart of a method 400 for partial wavefront merging used in conjunction with the Figure 3 system 300 as described herein. This description can be extended to other implementations without departing from the scope of this specification. The partial wavefront manager 325 determines which partial wavefronts are candidates for merging. This determination includes determining whether there are partial wavefronts using the same application, such as a shader program (405). If not, no further action is required. If there are partial wavefronts using the same application, then the total number of active work items of the partial wavefronts using the same application is determined (410), and the program counter of each of the partial wavefronts using the same application is determined (415). If the total number of active work items of the partial wavefronts using the same application does not exceed N, where N is the number of active work items that can be executed in a wavefront (i.e., the maximum number of lanes in the SIMD unit) and the program counter of each of the partial wavefronts using the same application is within a certain threshold, i.e., within a certain distance, where the program counter threshold is set based on application type, processing optimizations, and other considerations, then the partial wavefront manager 325 marks these partial wavefronts as merge candidates (420). If both conditions are met, then the partial wavefront manager 325 then moves the partial wavefronts to one SIMD unit (425). As described herein, each in the partial wavefront accesses its originally allocated unified GPR via the network 320.

[0027] The partial wavefront manager 325 then attempts to synchronize the program counters of the partial wavefronts (430). The arbitration algorithm allows one partial wavefront to execute more frequently in order to synchronize the program counters within a certain synchronization threshold. Then the synchronization of the program counters is checked (435). If the program counters fail to synchronize within a certain synchronization threshold, then the partial wavefronts are moved back to their original SIMD units (440). The criteria for setting the synchronization threshold are set based on application type, processing optimizations, and other considerations. In an implementation, synchronization can be attempted before any merging, as described with respect to Figure 5 as described.

[0028] If the program counter is successfully synchronized within a synchronization threshold, then partial wavefronts are merged into one wavefront (445). Work items within the merged wavefront run within a single SIMD unit. As described herein, each of the partial wavefronts, and specifically the work items, access their originally allocated GPRs via network 320. The merged wavefront is executed (450) when appropriate. Once the merged wavefront is complete, the GPRs allocated to it within the SIMD unit are freed (455).

[0029] Figure 5 is a flow chart of method 500 for partial wavefront merging for use in conjunction with system 300 of Figure 3 As described herein, this description may be extended to other implementations without departing from the scope of this specification. The partial wavefront manager 325 determines which partial wavefronts are candidates for merging. This determination includes determining whether there are partial wavefronts using the same application, such as a shader program (505). If not, no further action is required. In an illustrative example, if there are M partial wavefronts using the same application, then the total active work items of the M partial wavefronts using the same application are determined (510), and the program counter of each of the M partial wavefronts is determined (515), where the M partial wavefronts are executed in M different SIMD units. If the total active work items of the M partial wavefronts using the same application do not exceed N*K and K<M, where N is the number of active work items that can be executed within a wavefront (i.e., the maximum number of lanes within a SIMD unit), the partial wavefronts have a number of active work items less than N, and the program counter of each of the M partial wavefronts using the same application is within a certain threshold, i.e., within a certain distance, where the program counter threshold is set based on application type, processing optimizations, and other considerations, then the partial wavefront manager 325 marks these M partial wavefronts as merge candidates (520). Note that N*M represents the maximum number of work items that can be executed, and if K<M, there are partial wavefronts. In other words, the total active work items of the M partial wavefronts are less than the maximum number of work items that a given number of SIMD units can execute.

[0030] The partial wavefront manager 325 then attempts to synchronize the program counters of the M partial wavefronts (525). The arbitration algorithm allows some of the M partial wavefronts to execute more frequently in order to synchronize their program counters with the other partial wavefronts within a synchronization threshold. Synchronization can be achieved using a variety of methods. In an implementation, the wavefront with the highest-ranked program counter is selected (denoted as "wavefront x"), and the remaining wavefronts are given higher priority during scheduling so that their program counters will eventually all synchronize with wavefront x. The criteria for setting the synchronization threshold are based on application type, processing optimizations, and other considerations.

[0031] Then, the synchronization of the program counter is checked (530). If the program counter fails to synchronize within a certain synchronization threshold, no further action is required. If the program counter successfully synchronizes within a certain synchronization threshold, then partial wavefronts are merged (535). In an illustrative example, if L (L ≤ M) out of M partial wavefronts can have the same program counter within a certain time threshold, and the total active work items of the wavefronts do not exceed N*P (P < L), then these L partial wavefronts are merge candidates. Specifically, the L partial wavefronts are merged into P wavefronts. The P wavefronts are executed in P different SIMD units. As described herein, each in a partial wavefront, and specifically a work item, accesses its originally assigned GPR via (e.g.) network 320. The partial wavefronts, such as the P wavefronts, are executed (540) when appropriate. Once the merged wavefronts are completed, the GPRs assigned to them in the SIMD units are released (545).

[0032] There are multiple ways to merge partial wavefronts. In an illustrative example, the work items of (L - P) partial wavefronts are moved into the other P partial wavefronts. There are many ways to divide the L partial wavefronts into P and L - P. For example, the first P partial wavefronts among the L partial wavefronts are selected. In another example, P partial wavefronts are arbitrarily selected from the L partial wavefronts. In another example, the P partial wavefronts with the most active work items are selected. For each partial wavefront among the (L - P) wavefronts, the active items are moved into the selected P partial wavefronts to be merged into new P wavefronts. In an implementation, among the P (P > 2) wavefronts, after merging, the wavefronts can be partial. In an implementation, the work items of a partial wavefront can be split and merged with multiple partial wavefronts. For example, if N = 8, L is 3, and each wavefront has 5 active work items, then the original wavefronts can be merged into 2 new wavefronts. The 5 work items in one original partial wavefront are split and merged into 2 new wavefronts. In this case, one new wavefront has 8 active items, and the other wavefront has 7 active items.

[0033] Figure 6FIG. 600 is a block diagram of a system 600 that performs partial wavefront merging according to some implementations. As described herein, this description can be extended to other implementations without departing from the scope of this specification. System 600 includes SIMD units 605, e.g., SIMD0, SIMD1, SIMD2, …, SIMDN, where each SIMD unit 605 includes at least one arithmetic logic unit (ALU) 610. System 600 also includes a unified GPR 615 connected to each ALU 610 via network 620. In an implementation, network 620 is a mesh network. In another implementation, unified GPR 615 is connected via a shared bus. Partial wavefront manager 625 is connected to each of SIMD0, SIMD1, SIMD2, …, SIMDN. Figure 6 The system 600 shown in Figure 1 and Figure 2 can be implemented using the system shown therein and the elements presented therein. The partial wavefront merging feature in system 600 can be enabled or disabled via register settings. Built-in functions are added to the instruction set to set this register. The register settings can be set by a driver, a shader program, or a kernel. Generally, system 600 operates as described with respect to Figure 4 method 400 of Figure 5 and method 500 of

[0034] In an illustrative example, a shader unit or program uses a built-in register function to enable the partial wavefront merging feature. In the example, a wavefront contains 8 work items and thus the execution mask is 8 bits. Assume that wavefront 0 in SIMD0 has an execution mask of 0x11110000 (where 1 indicates an active channel and 0 indicates an inactive channel), stateId is 0 and the program counter is equal to 10, and wavefront 1 in SIMDN has an execution mask of 0x00001111 and the program counter is equal to 12. In this illustrative example, wavefront 0 has 4 active channels and wavefront 1 has 4 active channels. The program counter distance between wavefront 0 and wavefront 1 is 2, within the program counter threshold of 4. The partial wavefront manager detects both wavefronts as merge candidates and moves wavefront 1 from SIMDN to SIMD0 because the total number of combined active channels is 8 and fits within a single SIMD unit. When wavefront 1 is moved from SIMDN to SIMD0, wavefront 0 is still executing. Thus, once the move is complete, the program counter of wavefront 0 becomes 14 and the program counter of wavefront 1 remains 12.

[0035] The arbitration algorithm enables wavefront 1 to run in SIMD0 to attempt to synchronize the two program counters at 14. If the program counter of wavefront 1 skips 14 due to a branch or similar program execution, the partial wavefront manager returns wavefront 1 to SIMDN. If the two program counters are successfully synchronized, the partial wavefront manager merges wavefront 0 and wavefront 1 into a merged wavefront. The execution mask of the merged wavefront is 0x11111111, indicating all active lanes. As described herein, work items from the original wavefront 0 access the unified GPRs within SIMD0, while work items from the original wavefront 1 access the unified GPRs in SIMDN via the network. Once the merged wavefront is complete, the GPRs allocated to it in the SIMD are released. As described herein, this description can be extended to other implementations without departing from the scope of this specification.

[0036] In another illustrative example, the driver sets a register to enable the partial wavefront merge feature. In the example, the wavefront contains 16 work items and thus the execution mask is 16 bits. Assume that wavefront 0 in SIMD0 has an execution mask of 0x1111111100000000, a stateId of 0, and a program counter equal to 10, and wavefront 1 in SIMD2 has an execution mask of 0x1111111100000000 and a program counter equal to 12. The program counter distance between wavefront 0 and wavefront 1 is 2, within the program counter threshold of 4. The partial wavefront manager detects both wavefronts as merge candidates and moves wavefront 1 from SIMD2 to SIMD0 because the total number of active lanes is 16 and fits within one SIMD unit. When wavefront 1 is moved from SIMD2 to SIMD0, wavefront 0 is still executing. Thus, once the move is complete, the program counter of wavefront 0 becomes 14, and the program counter of wavefront 1 remains 12.

[0037] The arbitration algorithm enables wavefront 1 to run in SIMD0 to attempt to synchronize the two program counters at 14. If the program counter of wavefront 1 is not equal to 14, the timeout counter is incremented. If the timeout counter is greater than the timeout threshold, the partial wavefront manager moves wavefront 1 back to SIMD2. If the two program counters are successfully synchronized, the partial wavefront manager merges wavefront 0 and wavefront 1 into a merged wavefront. The execution mask of the merged wavefront is 0x1111111111111111. As described herein, work items from the original wavefront 0 access the unified GPRs within SIMD0, while work items from the original wavefront 1 access the unified GPRs in SIMD2 via the network. Once the merged wavefront is complete, the GPRs allocated to it in the SIMD unit are released. As described herein, this description can be extended to other implementations without departing from the scope of this specification.

[0038] It should be understood that many variations are possible based on the disclosure herein. Although the features and elements have been described above in specific combinations, each feature or element can be used alone without the other features and elements or can be used in various combinations with or without the other features and elements.

[0039] The provided methods can be implemented in a general-purpose computer, a processor, or a processor core. Suitable processors include, for example, a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and / or a state machine. Such a processor can be manufactured by configuring a manufacturing process using the results of processed hardware description language (HDL) instructions and other intermediate data including a netlist, which instructions can be stored on a computer-readable medium. The result of such processing can be a mask work, which is then used in a semiconductor manufacturing process to fabricate a processor that implements aspects of the embodiments.

[0040] The methods or flowcharts provided herein can be implemented by a computer program, software, or firmware incorporated into a non-transitory computer-readable storage medium for execution by a general-purpose computer or a processor. Examples of non-transitory computer-readable storage media include read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVD).

Claims

1. A processing system, which comprises: a plurality of processing units, each processing unit executing a plurality of work items configured as a wavefront; a partial wavefront manager that communicates with each of the plurality of processing units; and a unified register structure that communicates with each of the plurality of processing units, wherein the partial wavefront manager is configured to: detect wavefronts that are partial wavefronts in different processing units, wherein the partial wavefronts include inactive work items and active work items; move the partial wavefronts from the different processing units to a processing unit from the plurality of processing units; and in response to determining that the execution synchronization of each of the partial wavefronts is within a threshold, merge the partial wavefronts into a merged wavefront executed by the processing unit, and wherein each work item in the merged wavefront accesses a register space in the unified register structure previously allocated to the different processing units during the execution of the merged wavefront by the processing unit.

2. The processing system according to claim 1, which further comprises: a communication network that communicates with the unified register structure and the plurality of processing units.

3. The processing system according to claim 2, wherein the unified register structure is a plurality of registers, and each processing unit includes a register connected to the communication network.

4. The processing system according to claim 1, wherein the unified register structure is a plurality of registers, and each processing unit includes at least one register connected to a communication network that is connected to the plurality of processing units.

5. The processing system according to claim 1, wherein the partial wavefront manager records for each wavefront: which wavefronts are sharing a given program; an execution mask to identify the active work items; and a program counter to indicate which instruction is being executed.

6. The processing system according to claim 5, wherein the partial wavefront manager further determines whether the total number of the active work items exceeds the maximum number of work items that can be executed relative to the wavefront.

7. The processing system according to claim 5, wherein the partial wavefront manager further determines whether program counters associated with the wavefronts sharing the given program are within a certain threshold.

8. The processing system according to claim 5, wherein the partial wavefront manager further synchronizes the program counters of the partial wavefronts within a certain synchronization threshold.

9. A method for improving wavefront processing in a processing system, the method comprises: determining partial wavefronts from a plurality of different processing units, wherein each processing unit is executing a plurality of work items configured as a wavefront, and wherein the partial wavefronts include inactive work items and active work items; moving the partial wavefronts from the different processing units to a processing unit from the plurality of processing units; and in response to determining that the execution synchronization of each of the partial wavefronts is within a threshold, merging the partial wavefronts into a merged wavefront executed by the processing unit, wherein each work item in the merged wavefront accesses a register space in the unified register structure previously allocated to the different processing units during the execution of the merged wavefront.

10. The method according to claim 9, wherein the method further comprises: communicating between the plurality of processing units and the unified register structure via a communication network.

11. The method according to claim 10, wherein the unified register structure is a plurality of registers, and each processing unit includes at least one register connected to the communication network.

12. The method according to claim 9, wherein the determining the partial wavefront further comprises: recording for each wavefront: which wavefronts are sharing a given program; an execution mask to identify the active work items; and a program counter to indicate which instruction is being executed.

13. The method according to claim 12, wherein the determining the partial wavefront further comprises: determining whether the total number of the active work items exceeds a maximum number of work items that can be executed relative to the wavefront.

14. The method according to claim 13, wherein the determining the partial wavefront further comprises: determining whether the program counters associated with the wavefronts sharing the given program are within a certain threshold.

15. The method according to claim 14, wherein the method further comprises: synchronizing the program counters of the partial wavefronts to within a certain synchronization threshold.

16. The method according to claim 15, wherein the method further comprises: if the synchronization fails, moving the partial wavefront back to the original processing unit.

17. The method according to claim 16, wherein the method further comprises: releasing the previously allocated register space in the unified register structure after completion.

18. A method for improving wavefront processing in a processing system, the method comprises: parallelly executing a plurality of wavefronts using an associated number of processing units, wherein each wavefront includes a plurality of work items; detecting whether any of the plurality of wavefronts is a partial wavefront, wherein a partial wavefront includes inactive work items; moving the partial wavefront from different processing units to a processing unit from among the plurality of processing units; and in response to determining that the execution synchronization of each of the partial wavefronts is within a threshold, merging the partial wavefronts into a merged wavefront to be executed by the processing unit, wherein each work item in the merged wavefront accesses a register space in the unified register structure previously allocated to the different processing units during the execution of the merged wavefront.

19. The method according to claim 18, wherein the detecting further comprises: recording for each wavefront: which wavefronts are sharing a given program; an execution mask to identify active work items; and a program counter to indicate which instruction is being executed; determining whether the total number of the active work items exceeds a maximum number of work items that can be executed relative to the wavefront; and determining whether the program counters associated with the wavefronts sharing the given program are within a certain threshold.

20. The method according to claim 19, wherein the method further comprises: synchronizing the program counters of the partial wavefronts to within a certain synchronization threshold; if the synchronization fails, moving the partial wavefront back to the original processing unit; and releasing the previously allocated register space in the unified register structure after completion.

Citation Information

Patent Citations

  • System and method for efficiently executing single program multiple data (SPMD) programs

    US20050108720A1

  • Method and System for Synchronization of Workitems with Divergent Control Flow

    US20130326524A1