Graphics workflow techniques for distributed architectures

Through logical launch slot technology and streaming launch slot manager (S-KSM), graphics work is distributed across multiple GPU subunits (mGPUs), solving the problem of inefficient work distribution and scheduling in GPUs, improving performance and reducing power consumption.

CN120693601APending Publication Date: 2025-09-23APPLE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202480012608.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-08-16
Filing Date
2024-02-09
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Prior Art In Graphics Processing Units (GPUs), as the number of shader cores increases, work distribution and scheduling techniques significantly impact performance and power consumption, resulting in low efficiency.

Method used

It uses logical launch slot technology to virtualize graphics work sets and distribute them across multiple GPU sub-units (mGPUs). It also combines fine-grained virtual launch slot scheduling and streaming launch slot manager (S-KSM) technology to optimize work distribution and scheduling.

Benefits of technology

It improves the processing performance of the GPU, especially for workloads with a large number of short tasks, reduces power consumption, and provides a more flexible work streaming mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120693601A_ABST
    Figure CN120693601A_ABST
Patent Text Reader

Abstract

The disclosed technology relates to scheduling a set of graphics jobs using queues. In some embodiments, a tracking circuit implements entries for a plurality of tracking slots of a graphics processor. The queue access circuitry may access a data structure in the memory that specifies a plurality of queues, where each queue enters the queue using a plurality of sets of control information for graphics work. Queue selection circuitry may select a set of graphical operations from the data structure based on one or more selection parameters and store control information for the selected set of graphical operations in a trace slot of trace slot circuitry. The distribution circuitry may assign portions of the respective sets of graphics jobs from the tracking slots to the graphics processor circuitry for execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to computer graphics processors and, more particularly, to techniques for distributing processing work to graphics subunits. Background Art

[0002] Given the ever-increasing computational power of graphics processing units (GPUs), they are now being widely used for large-scale workloads. For example, a workload may include vertex shaders, fragment shaders, and compute tasks. APIs such as Metal and OpenCL provide an interface for software developers to access the computational power of GPUs for their applications. Recently, software developers have been shifting a large portion of their applications to use GPUs. As processing technology shrinks and GPUs become more powerful, they may contain a large number of shader cores. Software or firmware can provide a unit of work to be performed, called a "kick." Data controller circuits (e.g., a compute data controller, a vertex data controller, and a pixel data controller) can distribute the work of these kicks to multiple replicated shader cores, for example, via a communication fabric. As the number of shaders is scaled, work distribution and scheduling techniques can significantly affect performance and power consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Figure 1A is a diagram illustrating an overview of example graphics processing operations according to some embodiments.

[0004] Figure 1B is a block diagram illustrating an example graphics unit according to some embodiments.

[0005] Figure 2 is a block diagram illustrating example master control circuitry configured to map logical slots to distributed hardware slots according to some embodiments.

[0006] Figure 3 is a block diagram illustrating an example set of main control circuitry and GPU hardware subunits according to some embodiments.

[0007] Figure 4 is a diagram illustrating three example allocation patterns for mapping logical slots to distributed hardware slots according to some embodiments.

[0008] Figure 5 is a diagram illustrating example mappings of multiple logical sockets to distributed hardware using different allocation patterns according to some embodiments.

[0009] Figure 6 is a block diagram illustrating detailed example elements of a main control circuit according to some embodiments.

[0010] Figure 7 is a diagram illustrating example distributed socket state and kernel residency information according to some embodiments.

[0011] Figure 8 is a flow chart illustrating an example method for mapping logical slots to distributed mGPU hardware slots according to some embodiments.

[0012] Figure 9 is a diagram illustrating example software coverage fields according to some embodiments.

[0013] Figures 10A to 10C is a flow chart illustrating an example technique for selecting a hardware slot based on hardware slot status for different example allocation modes, according to some embodiments.

[0014] Figure 11A is a diagram illustrating example logical slot retention status values ​​according to some embodiments.

[0015] Figure 11B is a flow chart illustrating an example technique for reclaiming a hardware slot according to some embodiments.

[0016] Figure 12 is a flow diagram illustrating an example software-controlled hardware slot reservation process according to some embodiments.

[0017] Figure 13 is a diagram illustrating example reserved hardware slots for higher priority logic slots in an mGPU according to some embodiments.

[0018] Figure 14A illustrates an example cache flush invalidate command encoding an unconditional field according to some embodiments, and Figure 14B Example refresh control circuits according to some embodiments are illustrated.

[0019] Figure 15 is a block diagram illustrating an example affinity map indicated by a collection of graphical tasks in accordance with some embodiments.

[0020] Figure 16 is a block diagram illustrating example kernel walker circuitry for affinity-based scheduling according to some embodiments.

[0021] Figure 17 is a diagram illustrating an example iteration of a kernel based on software-indicated affinity according to some embodiments.

[0022] Figure 18 is a block diagram illustrating an example work sharing control circuit according to some embodiments.

[0023] Figure 19Ais a block diagram illustrating an example logical slot manager with dependency tracking and status circuitry according to some embodiments, and Figure 19B Example tracking and status fields are illustrated according to some embodiments.

[0024] Figure 20 is a diagram illustrating an example register prefetch buffer for booting a socket manager according to some embodiments.

[0025] Figure 21 is a diagram illustrating an example state machine for a top slot handled by a boot slot manager according to some embodiments.

[0026] Figures 22 to 25 is a flowchart illustrating an example method according to some embodiments.

[0027] Figure 26 is a diagram illustrating an example setup scenario for the startup of six of the graphics processors for eight mGPUs.

[0028] Figure 27 is a diagram illustrating example allocations of launches that trap launches to single-mGPU, single-group, or full-machine allocation modes, according to some embodiments.

[0029] Figure 28 is a diagram illustrating example mGPU granularity allocation for startup according to some embodiments.

[0030] Figure 29 is a diagram illustrating an example "start when ready" assignment for startup according to some embodiments.

[0031] Figure 30 is a block diagram illustrating an example initiation queue technique according to some embodiments.

[0032] Figure 31 is a block diagram illustrating a detailed example implementation of a launch queue, a launch slot manager, a master control circuit, and a distributed mGPU according to some embodiments.

[0033] Figure 32 is a block diagram illustrating example initiation queue information according to some embodiments.

[0034] Figure 33 is a block diagram illustrating an example master controller circuit implementing a logic boot slot technique for geometry booting, according to some embodiments.

[0035] Figure 34 is a diagram illustrating an example of geometry-enabled execution using the disclosed logical socket technology for execution and splicing, according to some embodiments.

[0036] Figure 35A is a block diagram illustrating an example dependency graph including system dependencies and startup dependencies according to some embodiments.

[0037] Figure 35B Shown is an event flag dependency mask stored for a given trace slot according to some embodiments.

[0038] Figure 35C A set of example event flag fields is shown according to some embodiments.

[0039] Figure 36 is a block diagram illustrating an example organization of frame portions according to some embodiments.

[0040] Figure 37 is a diagram illustrating example row and column buffer mappings according to some embodiments.

[0041] Figure 38A and Figure 38B is a block diagram illustrating an example inter-block notification routing technique for event flags according to some embodiments.

[0042] Figure 39 is a diagram illustrating an example pipelined execution of dependent launches according to some embodiments.

[0043] Figure 40 is a diagram illustrating an example pipelined execution of dependent launches from different host controllers according to some embodiments.

[0044] Figure 41 Illustrated is a table with example dependency states of a launch on another launch, according to some embodiments.

[0045] Figure 42 A more detailed example dependency scenario according to some embodiments is illustrated.

[0046] Figures 43 to 47 is a flowchart illustrating additional example methods according to some embodiments.

[0047] Figure 48 is a block diagram illustrating an example initiation queue technique according to some embodiments.

[0048] Figure 49 is a block diagram illustrating example graphics control circuitry configured to map logical slots to distributed hardware slots according to some embodiments.

[0049] Figure 50 is a block diagram illustrating an example computing device according to some embodiments.

[0050] Figure 51is a diagram illustrating an example application of the disclosed systems and devices according to some embodiments.

[0051] Figure 52 is a block diagram illustrating an example computer-readable medium storing circuit design information according to some embodiments. DETAILED DESCRIPTION

[0052] Section I of this disclosure covers various techniques involving logical boot slots and the distribution of work from the logical boot slots to distributed hardware. Figure 1A to Figure 1B Provides an overview of graphics processors. Figures 2 to 8 An overview of techniques for mapping logical slots to distributed hardware slots for graphics processing is provided. Figure 9 Example software overlays that may affect the mapping are shown. Figures 10-14 illustrate example techniques for implementing allocation modes, reclaiming hardware slots, reserving hardware slots, providing logical slot priorities, and handling cache flush operations in the context of a logical slot map. Figures 15 to 18 An example technique for affinity-based scheduling is shown in FIG. Figure 21 A boot slot manager is shown that interacts with software and configures logical slots. The remaining figures illustrate example methods, systems, applications, etc. In various embodiments, the disclosed technology can advantageously improve the performance of a graphics processor or reduce the power consumption of a graphics processor relative to conventional technologies, as explained in detail below.

[0053] In general, Section I describes techniques related to virtualization of sets of graphics work, referred to as "launches." Software can provide launches to a graphics processing unit (GPU), which can utilize virtual / logical launch slot technology to distribute the launches across all or a subset of multiple GPU subunits for distributed processing (these subunits may be referred to as "mGPUs").

[0054] Section II describes various implementations with various alternatives or additions to the techniques of Section I. As one example, fine-grained virtual launch slot scheduling can assign a given launch to any number of mGPUs. As another example, a "start when ready" implementation can allow portions of a launch to start when one or more mGPUs are available (e.g., before all mGPUs to be used by the launch are available).

[0055] Additionally, while Section 1 describes top slot techniques for allowing software to program launches, the disclosed Streaming Launch Slot Manager (S-KSM) technique allows software to program various launch queues in memory. The launch slot manager can then pull launches from the queue into the top slot using queue selection logic (which can be configured to only pull work that is likely to be able to run in the near future). These techniques can reduce the software overhead of scheduling GPU work, pack more work into the GPU faster, and provide the hardware with a wider pool of work from which to select launches. Additionally, the GPU can utilize multiple completion queues with different priorities, which can allow software to quickly perform completion processing for higher priority work. This can advantageously improve GPU performance on certain processing workloads, particularly workloads with a large number of relatively short launches.

[0056] Additionally, the disclosed embodiments arbitrate to select a launch from a queue based on various parameters, such as queue priority, queue dependencies, the master controller associated with the launch (e.g., compute, vertex, or pixel controller), queue deadlines, available resources, etc. This may also allow for further decoupling of software from hardware implementation details (e.g., top socket details) relative to the embodiments described in Section I.

[0057] The disclosed S-KSM techniques may enable various additional functionalities, such as allowing the GPU to queue work for itself, allowing playback of pre-encoded workflows (e.g., from long-term storage such as a hard drive), and so on.

[0058] Section II also discusses interrupt suppression techniques for completed work (e.g., allowing software to control what happens when the GPU completes a launch), launch queue remapping (e.g., pausing selection from a first queue and redirecting it to another queue that the first queue depends on), and event flags (e.g., specifying various conditions that should be met before a given launch begins). These techniques can allow software to stream work into the GPU in a flexible manner while providing the GPU hardware with information useful for efficient scheduling. More generally, event flags can allow fine-grained streaming of various components across a device (e.g., fine-grained streaming of image portions), such as in a system-on-chip implementation.

[0059] The disclosed techniques also provide efficient and scalable processing of geometry launches. Geometry launches can be pre-parsed, split into multiple parts based on the pre-parse, then processed individually, and then stitched together for completion. For example, a slot-based technique can be used to distribute pre-processing work, launch execution, stitching work, or some combination thereof.

[0060] Finally, Section II also discusses techniques for pipelining portions of a launch even in the presence of inter-launch dependencies. The disclosed techniques can improve performance by overriding all or part of the spin-up, spin-down, or both for certain launches.

[0061] Part I

[0062] Graphics Processing Overview

[0063] refer to Figure 1A , shows a flowchart illustrating an example process flow 100 for processing graphics data. In some embodiments, the transformation and lighting process 110 may involve processing lighting information for vertices received from an application based on defined light source positions, reflectivity, etc., assembling the vertices into polygons (e.g., triangles), and transforming the polygons to the correct size and orientation based on their positioning in three-dimensional space. The clipping process 115 may involve discarding polygons or vertices outside the visible area. In some embodiments, geometry processing may utilize object shaders and mesh shaders before rasterization for flexibility and efficiency. The rasterization process 120 may involve defining fragments within each polygon and assigning initial color values ​​to each fragment, for example, based on the texture coordinates of the polygon vertices. Fragments may specify attributes of the pixels they overlap, but the actual pixel attributes may be determined by combining multiple fragments (e.g., in a frame buffer), ignoring one or more fragments (e.g., if they are covered by other objects), or both. The shading process 130 may involve altering pixel components based on lighting, shading, elevation mapping, translucency, etc. The shaded pixels may be assembled in a frame buffer 135. Modern GPUs typically include programmable shaders that allow application developers to customize shading and other processing. Thus, in various embodiments, Figure 1A The example elements of the present invention may be performed in various orders, in parallel, or omitted. Additional processing procedures may also be implemented.

[0064] Now refer to Figure 1B , shows a simplified block diagram illustrating a graphics unit 150 according to some embodiments. In the illustrated embodiment, graphics unit 150 includes a programmable shader 160, a vertex pipe 185, a fragment pipe 175, a texture processing unit (TPU) 165, an image write buffer 170, and a memory interface 180. In some embodiments, graphics unit 150 is configured to process both vertex data and fragment data using programmable shader 160, which can be configured to process graphics data in parallel using multiple execution pipelines or instances.

[0065] In the illustrated embodiment, vertex pipe 185 may include various fixed-function hardware configured to process vertex data. Vertex pipe 185 may be configured to communicate with programmable shader 160 to coordinate vertex processing. In the illustrated embodiment, vertex pipe 185 is configured to send processed data to fragment pipe 175 or programmable shader 160 for further processing.

[0066] In the illustrated embodiment, the fragment pipe 175 may include various fixed-function hardware configured to process pixel data. The fragment pipe 175 may be configured to communicate with the programmable shader 160 to coordinate fragment processing. The fragment pipe 175 may be configured to perform rasterization on polygons from the vertex pipe 185 or the programmable shader 160 to generate fragment data. The vertex pipe 185 and the fragment pipe 175 may be coupled to a memory interface 180 (coupling not shown) to access graphics data.

[0067] In the illustrated embodiment, the programmable shader 160 is configured to receive vertex data from the vertex pipe 185 and fragment data from the fragment pipe 175 and the TPU 165. The programmable shader 160 can be configured to perform vertex processing tasks on the vertex data, which can include various transformations and adjustments to the vertex data. For example, in the illustrated embodiment, the programmable shader 160 is also configured to perform fragment processing tasks, such as texturing and shading, on the pixel data. The programmable shader 160 can include multiple sets of multiple execution pipelines for processing data in parallel.

[0068] In some embodiments, the programmable shader includes a pipeline configured to execute one or more different SIMD groups in parallel. Each pipeline may include various stages configured to perform operations (such as fetch, decode, issue, execute, etc.) in a given clock cycle. The concept of a processor "pipeline" is well understood and refers to the concept of dividing the "work" of the processor performing an instruction into multiple stages. In some embodiments, decoding, dispatching, executing (i.e., performing), and retiring an instruction may be examples of different pipeline stages. Many different pipeline architectures may have different element / partial orderings. Various pipeline stages perform such steps on an instruction during one or more processor clock cycles and then pass the instruction or the operation associated with the instruction to other stages for further processing.

[0069] The term "SIMD group" is intended to be interpreted according to its well-known meaning, comprising a group of threads for which the processing hardware processes the same instruction in parallel using different input data for different threads. A SIMD group may also be referred to as a SIMT (Single Instruction Multiple Thread) group, a Single Instruction Parallel Thread (SIPT), or a channel stack thread. Various types of computer processors may include multiple groups of pipelines configured to execute SIMD instructions. For example, graphics processors typically include programmable shader cores that are configured to execute instructions for a group of related threads in a SIMD manner. Other examples of names that may be used for SIMD groups include: wavefront, blob, or warp. A SIMD group may be part of a larger thread group, which may be split into multiple SIMD groups based on the computer's parallel processing capabilities. In some embodiments, each thread is assigned to a hardware pipeline (which may be referred to as a "channel") that extracts operands for that thread and performs the specified operation in parallel with the other pipelines in the group. Note that a processor may have a large number of pipelines, allowing multiple separate SIMD groups to execute in parallel. In some embodiments, each thread has private operand storage, such as in a register file. Thus, reading a particular register from the register file may provide a version of the register for each thread in the SIMD group.

[0070] As used herein, the term "thread" includes its meaning as is well known in the art and refers to a sequence of program instructions that can be scheduled to execute independently of other threads. A SIMD group may include multiple threads to execute in lockstep. Multiple threads may be included in a task or process (which may correspond to a computer program). The threads of a given task may or may not share resources such as registers and memory. Therefore, when switching between threads of the same task, a context switch may or may not be performed.

[0071] In some embodiments, multiple programmable shaders 160 are included in a GPU. In these embodiments, global control circuitry can assign work to different subsections of the GPU, which in turn can assign work to shader cores for processing by the shader pipeline.

[0072] In the illustrated embodiment, TPU 165 is configured to dispatch fragment processing tasks from programmable shader 160. In some embodiments, TPU 165 is configured to pre-fetch texture data and assign initial colors to fragments for further processing by programmable shader 160 (e.g., via memory interface 180). TPU 165 can be configured to provide fragment components, for example, in normalized integer format or floating point format. In some embodiments, TPU 165 is configured to provide fragments in a 2x2 format of four ("fragment quads") that are processed by a set of four execution pipelines in programmable shader 160.

[0073] In some embodiments, image write buffer 170 is configured to store processed tiles of an image and can perform operations on the rendered image before transferring it for display or transferring it to memory for storage. In some embodiments, graphics unit 150 is configured to perform tile-based deferred rendering (TBDR). In tile-based rendering, different portions of screen space (e.g., squares or rectangles of pixels) can be processed separately. In various embodiments, memory interface 180 can facilitate communication with one or more of various memory hierarchies.

[0074] As discussed above, a graphics processor typically includes dedicated circuitry configured to perform certain graphics processing operations requested by a computing system. For example, this may include fixed-function vertex processing circuitry, pixel processing circuitry, or texture sampling circuitry. A graphics processor may also perform non-graphics computing tasks that may use GPU shader cores but may not use fixed-function graphics hardware. As an example, machine learning workloads (which may include inference, training, or both) are typically assigned to the GPU due to the GPU's parallel processing capabilities. Thus, the compute kernels executed by the GPU may include program instructions that specify machine learning tasks, such as implementing a neural network layer or other aspects of a machine learning model to be executed by the GPU shader. In some scenarios, non-graphics workloads may also use dedicated graphics circuitry, for example, for purposes different from those originally intended.

[0075] In addition, in other embodiments, the various circuits and techniques discussed herein with reference to graphics processors may be implemented in other types of processors. Other types of processors may include general-purpose processors such as CPUs or machine learning or artificial intelligence accelerators with dedicated parallel processing capabilities. These other types of processors may not be configured to execute graphics instructions or perform graphics operations. For example, other types of processors may not include the fixed-function hardware included in a typical GPU. A machine learning accelerator may include dedicated hardware for certain operations (such as implementing neural network layers or other aspects of a machine learning model). Generally speaking, there may be design trade-offs between memory requirements, computing power, power consumption, and programmability of a machine learning accelerator. Therefore, different specific implementations may focus on different performance targets. Developers can choose from a plurality of potential hardware targets for a given machine learning application, such as choosing from a general-purpose processor, a GPU, and different dedicated machine learning accelerators.

[0076] Overview of work distribution and logical slots

[0077] Figure 2 is a block diagram illustrating example main control circuitry and graphics processor subunits according to some embodiments. In the illustrated embodiment, the graphics processor includes a main unit 210 and subunits 220A through 220N.

[0078] For example, the master control circuit 210 can be a compute data master, a vertex data master, or a pixel data master. Thus, in some embodiments, the graphics processor includes multiple instances of the master control circuit 210 that send different types of work to the same set of subunits. The master control circuit 210 can receive launches from software, firmware, or both via an interface. As used herein, the term "software" refers broadly to executable program instructions and covers, for example, firmware, operating systems, and third-party applications. Therefore, it should be understood that the various references to software herein may apply alternatively or additionally to firmware. In the illustrated embodiment, the master control circuit 210 includes a logical slot to distributed hardware slot mapping control circuit 215. The control circuit 215 can distribute work from logical slots (which may be referred to as "launch slots") to distributed hardware slots on all or a portion of the graphics processor (e.g., according to the following references). Figure 4 different allocation models discussed).

[0079] Various circuits are described herein as controlling logical slots. The term "logical" refers to graphics instructions that assign work to logical slots without implying which hardware will actually execute the assigned work. The control circuitry may include hardware that maintains information about logical slots and assigns work from logical slots to hardware slots for actual execution. Therefore, when initially assigned to a logical slot, the hardware slot that will execute the set of work is unknown. As discussed in detail below, logical slots can provide various advantages in terms of performance and power consumption when scheduling graphics work, particularly in a graphics processor with multiple shader cores.

[0080] Multiple "launches" can be executed to render a frame of graphics data. In some embodiments, a launch is a unit of work from a single context that can include multiple threads to be executed (and potentially other types of graphics work not performed by shaders). A launch may not provide any guarantees regarding memory synchronization between threads (other than specified by the threads themselves), concurrency between threads, or the order in which threads are launched. In some embodiments, a launch can be identified based on the results of another launch, which can allow memory synchronization without requiring hardware memory consistency support. Typically, graphics firmware or hardware programs configure each launch to register before sending work to the pipeline for processing. Typically, once a launch is started, it will not access more than a certain level of the memory hierarchy until the launch is complete (at which point the results can be written to another level in the hierarchy). Information for a given launch may include state information required to complete the corresponding graphics operation, the location of the shader program to be executed, buffer information, the location of texture data, available address space, etc. For example, when a launch is complete, the graphics firmware or hardware may schedule the launch and detect interrupts. In some embodiments, various portions of the graphics unit are configured to work on a single launch at a time. As discussed in detail herein, trace slots (also known as "top slots") and logical launch slots can be used to control launches before they are assigned to shader hardware. A launch can include a collection of one or more rendering commands, which may include commands for drawing procedural geometry, commands for setting shadow sampling methods, commands for drawing meshes, commands for retrieving textures, commands for performing render calculations, and the like. Launches can be executed at one of the various stages during frame rendering. Examples of rendering stages include, but are not limited to, camera rendering, light rendering, projection, texturing, fragment shading, and the like. For example, a launch can be scheduled for compute work, vertex work, or pixel work.

[0081] In some embodiments, subunit 220 is a scaling unit that can be replicated to increase the processing power of the GPU. Each GPU subunit 220 can be capable of independently processing instructions of a graphics program. In the illustrated embodiment, subunit 220 includes circuitry that implements corresponding distributed hardware slots 230. These hardware slots may also be referred to as "dSlots" herein. Each subunit may include multiple hardware slots 230. A subunit may also be referred to as an "mGPU" herein. In some embodiments, the main control circuit 210 assigns work from logical slots to at most one distributed hardware slot in each subunit 220. In some embodiments, each subunit includes fragment generator circuitry, shader core circuitry configured to execute shader programs, memory system circuitry (which may include one or more caches and a memory management unit), geometry processing circuitry, and distributed workload distribution circuitry (which may coordinate with the main control circuit 210 to distribute work to the shader pipeline).

[0082] Each distributed hardware slot may include various circuits configured to handle the assigned launch or portion thereof, including configuration registers, a work queue, circuitry configured to iterate through the work in the queue (e.g., batches of computational work items), circuitry to order context loads / stores, and work allocation tracking circuitry. Each subunit 220 may include multiple shaders that accept work from the distributed slots in the subunit and execute the work using a pipeline. For example, each shader may include a queue for each distributed hardware slot and may select work from the queue based on the priority of the work.

[0083] In some embodiments, a given subunit 220 includes a plurality of programmable shaders 160 of FIG. 1 .

[0084] As discussed in detail below, the logical slot to distributed hardware slot mapping control circuitry 215 may distribute activations across the subunits 220 based on various parameters, software control inputs, and the like.

[0085] Figure 33 is a block diagram illustrating a more detailed example of a main control circuit and a grouped processor subunit according to some embodiments. In the illustrated embodiment, the main control circuit 210 communicates with the boot slot manager (KSM) 350 and includes configuration registers 312. These configuration registers may include both setup registers and execution registers. The setup phase registers may be a global structure that is independent of the distributed hardware used to perform the boot, while the execution registers may be per-subunit structures. Generally speaking, although shown in the main control circuit 210, the configuration registers may be included in various appropriate circuits (e.g., in the distributed control circuit 340) and may have different scopes (e.g., some registers may be boot scoped, some registers may be associated with logical slots, and some registers may be associated with distributed slots). Some configuration registers may be shared and the same value may be programmed into both the global register circuit and the per-subunit register circuit. The configuration register data may be stored in a memory in a defined format and retrieved and unpacked to fill the physical configuration registers for a given boot.

[0086] In the illustrated embodiment, mGPUs 320A through 320N are grouped, and master control circuitry 210 communicates with multiple such groups. An mGPU is an example of a subunit 220. In the illustrated embodiment, each group 305 of mGPUs shares a cache 360. For example, in an embodiment where each mGPU 320 maintains a Level 1 cache, this may be a Level 2 cache. This shared cache may be used to store instructions, data, or both. As discussed in detail below, scheduling work with data affinity attributes to the same group 305 can be beneficial for cache efficiency. In some embodiments, such as in multi-die implementations, each group 305 of mGPUs is implemented on the same die or semiconductor substrate.

[0087] In the illustrated embodiment, each mGPU 320 includes distributed control circuitry that can receive work from the main control circuitry 210, dispatch the work within the mGPU, and report the completion of the work back to the main control circuitry 210 (e.g., via a communication fabric). The signal dispatching the work may not include the actual instructions to be executed or the data to be operated on, but may identify the location of the program instructions to be executed.

[0088] In the illustrated embodiment, the boot slot manager 350 is configured to receive boots from the software / firmware interface and pass the boots to the main control circuit 210 for assignment to logical slots. Figure 6 Example communications between the boot slot manager 350 and the control circuitry are discussed in detail, and a detailed example implementation of the boot slot manager 350 is discussed below with reference to FIG. 19 .

[0089] Discussed in detail below Figure 4 and Figure 5 Examples are provided of techniques implemented by embodiments of control circuitry 215 to allocate work from logical slots according to some embodiments.

[0090] Figure 4 is a diagram illustrating three example allocation modes according to some embodiments. Generally speaking, the allocation mode indicates the width of the allocation. In the illustrated example, each mGPU implements three distributed hardware slots (DS0 to DS2) and two groups (Group 0 and Group 1) each include two mGPUs (mGPU0 and mGPU1). It should be noted that in various embodiments, various numbers of hardware slots per mGPU, various numbers of mGPUs per group, and various numbers of groups per graphics processor can be implemented. Embodiments with different specific example numbers of elements are discussed herein for the purpose of explanation, but these examples are not intended to limit the scope of this disclosure.

[0091] As discussed above, in all three example modes, a logical slot is allowed to occupy at most one hardware slot of a given mGPU. Figure 4 The hardware slots to which work is assigned from the logical slots are shown using diagonal shading. It should be further noted that in some cases, the control circuitry 215 can dynamically adjust the logical-to-hardware slot mapping. The logical slots assign work to the distributed control circuitry in the mGPU, where it is assigned the hardware slot. The distributed control circuitry can then assign the work to its shader within the mGPU.

[0092] In the illustrated example, mode A is a single mGPU allocation mode. In this mode, the control circuit 215 assigns work from the logical slots to a single hardware slot on a single mGPU.

[0093] In the illustrated example, mode B is a single-bank assignment mode. In this mode, the control circuitry 215 assigns work from a logical slot to a slot on each mGPU in the bank of mGPUs (bank 0 in this example).

[0094] In the illustrated example, mode C is a larger multi-bank allocation mode in which the control circuitry 215 assigns work from logical slots to slots in each mGPU in multiple banks of mGPUs (e.g., in some embodiments, each bank on a graphics processor).

[0095] The control circuitry 215 may determine the allocation pattern for logical slots (or, for example, for a kernel, the portion of a boot that is assigned to a logical slot) based on various considerations discussed in detail below. Generally speaking, the control circuitry 215 may select an allocation pattern based on the workload being managed by the main control circuitry at a particular time, the size of the working set, or both. Additionally, a software override function may allow software or firmware to adjust the allocation of work during a boot. Furthermore, prioritization, dynamic remapping, and reclamation techniques may influence the logical-to-hardware slot mapping.

[0096] The control circuitry 215 may report hardware slot allocations and deallocations to the boot slot manager 350 , which may allow software or firmware to query information about the current logical slot mapping (eg, allocation mode, specific mapping, etc.).

[0097] Note that the group / mGPU / hardware slot level of organization is included for purposes of explanation and is not intended to limit the scope of this disclosure. In some embodiments, the "group" level of organization can be omitted, which can result in only two allocation modes: single mGPU or multi-mGPU. In some embodiments, additional levels of organization can be implemented, which can be associated with additional allocation modes (e.g., groups of groups, which can result in single-group group mode and multi-group group mode in addition to single-mGPU mode and single-group mode).

[0098] Figure 5 is a diagram illustrating an example population of available hardware slots from multiple logical slots according to some embodiments. In the illustrated example, the control circuitry 215 uses multiple allocation patterns to map nine logical slots 510A to 510I to forty-eight distributed slots (in four groups of four mGPUs).

[0099] In the illustrated example, circuitry 215 uses a single mGPU allocation mode for logical slots 510A, 510D, 510G, and 5101. For example, logical slot 510A receives a single distributed slot DS0 in mGPU 0 of group 0.

[0100] Circuitry 215 uses a single-group allocation mode for logical sockets 510B, 510E, and 510H. For example, logical socket 510B receives distributed socket DS0 on each mGPU in group 1.

[0101] Circuitry 215 uses a multi-group allocation mode for logical sockets 510C and 510F. For example, logical socket 510C receives a distributed socket on each instantiated mGPU.

[0102] Note that not all hardware slots may always be assigned, but generally filling available slots can improve performance. When a boot assigned to a logical slot has completed, another boot can be assigned to that logical slot, and the logical slot can be remapped to a physical slot.

[0103] Example Control Circuit

[0104] Figure 6 is a block diagram illustrating detailed example control circuitry according to some embodiments. In the illustrated embodiment, the boot slot manager communicates with mapping control circuitry 215, which in the illustrated embodiment includes: a dSlot resource allocator 620, control stream processors (CSPs) 630A to 630N, core processors 640A to 640N, mGPU assignment circuitry 650A to 650N, and a boot slot arbiter 660. In some embodiments, each logical slot supported by a processor has an assigned set of elements 630, 640, and 650. Note that while Figure 6 Some of the details are related to compute work, but similar techniques can be used for other types of work, such as vertex and pixel shading.

[0105] In the illustrated embodiment, the boot slot manager 350 assigns the boot to a logical slot and sends the boot information to the corresponding control flow processor 630. The control flow processor 630 may notify the boot slot manager 350 when the boot has completed processing.

[0106] In the illustrated embodiment, the control flow processor 630 manages the sequencing of its launch slots, extracts and executes control flows for launch, and tracks launch completion. The control flow processor 630 can operate at the kernel granularity (kernels can be extracted from the control flow for launch). The control flow processor 630 can communicate with the dSlot resource allocator 620 to obtain dSlot resources for its logical slots. The control flow processor 630 is configured to determine the allocation mode of the kernel and send the kernel with its allocation mode and distributed slot assignment to the kernel processor 640.

[0107] In some embodiments, a dSlot resource allocator (DRA) 620 includes circuitry configured to receive requests from multiple logical slots and process these requests to assign dSlots to kernels. In some embodiments, the dSlot resource allocator 620 selects an allocation pattern and assigns dSlots to various parts of a boot (e.g., at a kernel granularity), although other granularities are also contemplated. In some embodiments, the dSlot resource allocator 620 first assigns logical slots based on boot priority, and then assigns based on boot age, as discussed in further detail below. For example, the DRA 620 may reserve some distributed slots for kernels from boots with a priority level greater than a threshold.

[0108] In the illustrated embodiment, the core processor 640 is included in the master compute data master. The core processor 640 is configured to create batches from the kernel's workgroup and send the batches with their allocation patterns and distributed slot assignments to the mGPU assignment circuit 650. The core processor 640 may select batches for allocation based on affinity, load balancing, or both, as discussed in detail below. The core processor 640 may receive an indication of the assigned dSlot and a target mask indicating which mGPUs are allowed to be targets of the kernel.

[0109] As used herein, the term "compute kernel" in the context of graphics is intended to be interpreted according to its well-known meaning, which includes routines compiled for acceleration hardware such as a graphics processor. For example, a kernel can be specified by a separate programming language such as OpenCL C, written as a compute shader in a shading language such as OpenGL, or embedded in application code in a high-level language. A compute kernel typically includes multiple workgroups, which in turn include multiple work items (also called threads). It should be noted that the various techniques discussed herein with reference to compute kernels can be applied to other types of work, such as vertex or pixel processing tasks.

[0110] In the illustrated embodiment, mGPU assignment circuit 650 receives a batch and sends it to a target mGPU. Circuit 650 may receive a batch and a mask of allowable mGPU targets, which may vary depending on the allocation mode. Using this mask, circuit 650 may select an mGPU target based on load balancing.

[0111] In the illustrated embodiment, the launch slot arbiter 660 selects from the available batches to send to the target mGPU. For example, the launch slot arbiter 660 may select one or more logical launch slots to send a batch each cycle. The selected batch (and return information associated with the execution status) may be transmitted via a communication fabric. This fabric may be dedicated to control signaling, for example, as discussed in U.S. patent application Ser. No. 17 / 158,943, filed on Jan. 26, 2021, entitled "Shared Control Bus for Graphics Processors."

[0112] The following discusses in detail the Figure 6 Various additional functions are performed by the circuitry of the , for example, in the sections discussing specific functions such as dynamic mapping, software overlay, priority, retention techniques, eviction techniques, cache flushing, and affinity.

[0113] According to the kernel mapping technology

[0114] In some embodiments, per-kernel mapping during execution of a compute launch can provide dynamic allocation, which would be difficult at the launch granularity (it may be difficult to determine how many distributed slots a launch should occupy before executing the launch). As briefly discussed above, control flow processor 630 and dSlot resource allocator 620 can facilitate these techniques.

[0115] Figure 7 is a diagram illustrating example distributed socket state and core residency information according to some embodiments. This information can facilitate dynamic mapping.

[0116] In the illustrated example, a dslot_status is maintained for each dSlot and indicates whether the dSlot is invalid, running, cleared, refreshing, or reserved. The invalid state indicates that the dSlot is not owned by any logical slot. The running state indicates that the dSlot is owned by a logical slot and is currently executing. The cleared state indicates that the dSlot is owned by a logical slot and has completed execution. The refreshing state indicates that the dSlot is owned by a logical slot while in the process of a cache refresh (e.g., a refresh-invalidate with the memory hierarchy). The reserved state indicates that the dSlot is owned by a logical slot and is retained after the kernel completes (e.g., after the kernel ends cache refresh invalidation), for example to save performance data. Note that these states are included for explanation purposes, but in other embodiments, other states may be implemented, states may be omitted, or both situations may be implemented.

[0117] In the illustrated example, a dslot_owner state is maintained for each dSlot and indicates the logical slot that owns the dSlot. This field is irrelevant to the invalid state because no logical slot owns an invalid dSlot.

[0118] In the illustrated example, a per_kernel_residency state is maintained for each kernel and each mGPU, and the per_kernel_residency state indicates whether the kernel is assigned to an mGPU. Note that the various information maintained per kernel for compute work can similarly be maintained for other types of work that do not utilize kernels, for launches or portions of launches.

[0119] Figure 8 is a flowchart illustrating an example method for mapping a launched kernel according to some embodiments.

[0120] At 810, in the illustrated embodiment, the control circuitry waits until kernel dependencies have cleared and the logical socket assigned to the launch has a free kernel processor. This allows the previous kernel to complete its iteration before starting the next kernel for launch.

[0121] At 820 , in the illustrated embodiment, the CSP determines an allocation mode and sends a request with the allocation mode to the DRA 620 .

[0122] The DRA 620 responds with distributed slot allocation at 830. An example DRA arbitration algorithm is discussed in detail below.

[0123] At 840, in the illustrated example, the master control circuitry performs several activities. First, it sends a distributed slot start message to all mGPUs on which a dSlot has been allocated for the kernel. Next, it sends a register write command to the register replication unit, which includes a dSlot mask indicating which dSlots are affected. The register replication unit writes the distributed slot-wide control register for the kernel. (The register replication unit may have already written to the logical slot-wide control register for startup.) Finally, the master control circuitry sends the work to the indicated mGPU. Note that the work can be isolated until all register writes for the register replication unit are complete.

[0124] The master control circuitry can also track the completion status of each kernel it assigns. For example, it can detect when the dSlots on which the kernel is executing have all transitioned from running to empty.

[0125] Example software coverage technique

[0126] In some embodiments, software can provide various instructions that override the default kernel allocation mode. For example, this can allow software to parallelize important work rather than risk assigning it to a single mGPU. Additionally, this can allow software to assign kernels to specific groups of mGPUs.

[0127] Figure 9 is a diagram illustrating example software override fields. Software or firmware can adjust these fields to control kernel allocation. In the illustrated embodiment, the mGPU mask field indicates which mGPUs can be used for the launch. For example, the mask may include one bit per mGPU. This may allow software to indicate that certain mGPUs should be avoided or targeted for the launch. The allocation mode field allows software to select the allocation mode. The default value may allow the control flow processor 630 of the logical slot to select the allocation mode. Other values ​​may specify an allocation mode that the control flow processor 630 may implement regardless of the mode it would otherwise select (at least in the operating mode with software override enabled). In the default mode, the mGPU assignment circuitry 650 may flexibly select dSlots based on load balancing according to the allocation mode selected by the CSP, while in other modes, the mGPU assignment circuitry may adhere to restrictions specified by the software override.

[0128] The force group field allows software to select the group on which to execute the launch. For example, this can be specified in conjunction with a single mGPU or single group allocation mode. The policy field allows software to specify a scheduling policy for single mGPU or single group allocation. In the illustrated example, the software can specify a "select first option" policy (which can be the default) or a round-robin policy. The select first option policy selects the first available element based on its index (e.g., mGPU or group), which avoids fragmentation and leaves more consecutive dSlots for other slices. The round-robin policy randomizes the use of resources, which avoids significant performance differences caused by the location of the selected resources, but can spread small slices over multiple groups. In other embodiments, various policies can be specified. A detailed example of arbitration considering software override fields is discussed below.

[0129] Example arbitration techniques

[0130] Figures 10A to 10C is a flow chart illustrating an example technique for hardware slot arbitration corresponding to different allocation modes according to some embodiments. Note that the disclosed techniques can generally allocate consecutive kernels to the same logical slot (e.g., such that if kernel A is a single mGPU kernel assigned to a dSlot in mGPU0, kernel B, also a single mGPU kernel, will be assigned to a dSlot in mGPU1, which can prioritize the completion of logical slot execution while allowing fewer logical slots to run simultaneously).

[0131] In some embodiments, DRA 620 keeps a dSlot in the clear state as long as possible, for example, to allow subsequent cores from the same logical socket to use the dSlot. This can reduce cache-flush invalidations and writes to the execution configuration registers of the newly allocated dSlot. In some embodiments, a dSlot in the clear state owned by another logical socket must undergo a reclamation process (discussed below with reference to FIG. 11 ) and become invalid before being assigned to a new logical socket.

[0132] Generally speaking, as described in detail below, the DRA 620 uses the following priority scheme to select a dSlot for a kernel. The highest priority is an empty dSlot already owned by a logical slot. These dSlots have their control registers written to and are free for immediate execution. The middle priority is an invalid dSlot, which is newly allocated and may require a control register write, but is free for immediate execution. The lowest priority is a running dSlot already owned by a distributed slot. These dSlots have their control registers written to, but may need to wait behind another kernel.

[0133] Figure 10A The arbitration method for single mGPU allocation mode is shown. At 1010, in the illustrated embodiment, DRA 620 determines the set of allowable mGPUs for the kernel based on its force group and mGPU mask fields. This set may omit any set of mGPUs that are not selected by software.

[0134] At 1012, in the illustrated embodiment, DRA 620 selects the mGPU for which the kernel's logical slot already has a dSlot in the clear state. Note that in the case of elements 1012, 1016, and 1018 in parallel, DRA 620 uses a determined policy (e.g., in some embodiments, a default, software-specified, or a single type of policy) to select the hardware resource. For example, if there are multiple mGPUs that meet element 1012, DRA 620 can apply the policy to select the mGPU. If one or more mGPUs meet these criteria, one of them is selected and arbitration for that logical slot ends until the kernel completes. Otherwise, the process continues.

[0135] At 1014, in the illustrated embodiment, DRA 620 selects an mGPU with at least one invalid dSlot, where the logical slot does not already have a dSlot. If one or more mGPUs meet these criteria, one of them is selected and arbitration for that logical slot ends until the kernel completes. Otherwise, the process continues.

[0136] At 1016, in the illustrated embodiment, DRA 620 selects the mGPU with the most invalid slots. If one or more mGPUs meet these criteria, one of them is selected and arbitration for that logical slot ends until the kernel completes. Otherwise, the process continues.

[0137] At 1018, in the illustrated embodiment, the DRA 620 selects the mGPU where the logical slot already has a hardware slot in the running state. If one or more mGPUs meet these criteria, one of them is selected and arbitration for that logical slot ends until the kernel completes. Otherwise, the process continues.

[0138] At 1020, in the illustrated embodiment, the DRA 620 attempts a reclaim process. An example of such a process is discussed in more detail below with reference to Figure 11. If the reclaim process is unsuccessful, the process continues.

[0139] At 1022, in the illustrated embodiment, the DRA 620 restarts the provisioning machine and re-arbitrates. For various allocation modes, re-arbitration may occur until a sufficient number of hardware slots are available to satisfy the allocation mode.

[0140] Figure 10B The arbitration method for the single group allocation mode is shown. At 1030, in the illustrated embodiment, DRA 620 determines the set of allowable mGPUs, similar to Figure 10A Element 1010.

[0141] At 1032, in the illustrated embodiment, the DRA 620 selects a group in which all mGPUs in the group have dSlots owned by the kernel's logical slot in a clear or invalid state. In the event of a tie, the DRA 620 selects the group with the fewest invalid dSlots. If one or more groups meet these criteria, one of them is selected and arbitration for that logical slot ends until the kernel completes. Otherwise, the process continues.

[0142] At 1034, in the illustrated embodiment, the DRA 620 selects a group in which all mGPUs in the group have dSlots owned by logical slots in the Running, Invalid, or Clear states. In the event of a tie, the DRA 620 selects the group with the fewest mGPUs with slots in the Running state. If a tie still exists, the DRA 620 can apply a policy. If one or more groups meet these criteria, one of them is selected and arbitration for that logical slot ends until the kernel completes. Otherwise, the process continues.

[0143] At 1038 and 1040, the dRA 620 attempts to reclaim and then restart the dispenser and re-arbitrate, similar to elements 1020 and 1022 discussed above.

[0144] Figure 10C An arbitration method for a multi-group allocation mode is shown. At 1050 , in the illustrated embodiment, the DRA 620 determines a set of allowable mGPUs based on the mGPU mask (rather than based on the force group command, as all groups are used in this example).

[0145] At 1052, in the illustrated embodiment, the DRA 620 performs the operations of elements 1054 through 1058 for each target mGPU in the set of allowable mGPUs. At 1054, the DRA selects a dSlot that is already owned by the kernel's logical slot and is either empty or running. If one or more dSlots meet these criteria, one of them is selected and arbitration for that logical slot ends until the kernel completes. Otherwise, the process continues.

[0146] At 1056, in the illustrated embodiment, DRA 620 selects an invalid dSlot. If one or more dSlots meet these criteria, one of them is selected and arbitration for that logical slot ends until the kernel completes. Otherwise, the process continues. At 1058, in the illustrated embodiment, DRA 620 attempts to reclaim.

[0147] If the operation of element 1052 is unsuccessful in allocating a dSlot in each mGPU to the kernel, flow proceeds to 1060 and the DRA 620 restarts the allocation engine and re-arbitrates.

[0148] Note that while the various techniques discussed above consider software overlay fields, in other embodiments, in certain operating modes, software overlays may not be implemented or may be disabled. In this case, DRA may operate as discussed above, but omitting software overlay considerations.

[0149] Slot Recycling

[0150] In some embodiments, the control circuitry is configured to allow a logical slot to reclaim a hardware slot that has been assigned to another logical slot. In some embodiments, only higher-priority logical slots are allowed to reclaim a hardware slot from another logical slot. Example techniques for implementing logical slot priorities are discussed below, but priority levels can generally be dictated by software. In some embodiments, only hardware slots in an empty state are eligible for reclaim by another logical slot.

[0151] Generally speaking, the control circuitry may attempt to keep a hardware slot in a clear state for as long as possible. This avoids the overhead of performing cache flush invalidations and writing configuration registers when switching the hardware slot to a new logical slot (since keeping the hardware slot clear allows the same logical slot to send another core from the same boot to use the hardware slot, thus avoiding this overhead). However, due to this, allowing other important logical slots to use such a hardware slot may improve performance.

[0152] Figure 11A is a diagram illustrating example values ​​of a hold signal for hardware slot reclamation according to some embodiments. The hold signal may also be referred to as a persist signal. Each CSP 630 may send a hold signal to the DRA 620 indicating the extent to which it wants to hold onto its hardware slot (e.g., according to the extent to which the CSP 630 performs its startup).

[0153] In the illustrated example, the hold signal has one of three values, but other sets of values ​​are contemplated in other embodiments. A low value indicates that the logical slot has reached the control flow termination signal for launch and does not have any kernels remaining in the kernel queue to be processed. In this case, the logical slot will not require another hardware slot for launch. A medium value indicates that the logical slot has not reached the control flow termination signal, but there are currently no kernels ready to request a hardware slot for execution. A high value indicates that the logical slot has a kernel requesting a hardware slot for execution.

[0154] In some embodiments, DRA 620 is configured to reclaim slots only if enough hardware slots can be reclaimed to satisfy the request. Otherwise, the reclaim attempt may fail. Once the reclaim is successful, DRA 620 restarts its state machine and re-arbitrates the logical slots. DRA 620 may initiate a cache flush invalidation with the memory hierarchy for any reclaimed slots. This may transition those slots to a refreshed state, but once they have completed the flush and transitioned to an invalidated state, those slots may become available for arbitration.

[0155] Figure 11B is a flow chart illustrating an example technique for reclaiming one or more hardware slots currently assigned to another logical slot, according to some embodiments. At 1110, in the illustrated embodiment, the DRA 620 finds all dSlots in the flushing state. It may generate a data structure indicating the set of dSlots in the flushing_set. If these dSlots are sufficient to service the kernel's request, the DRA 620 cancels the reclamation and waits for the flush to complete. If not, the process continues.

[0156] At 1120, in the illustrated embodiment, the DRA 620 finds all dSlots that are emptied and owned by logical slots that (a) do not have context stored and (b) do not have any flushing dSlots. It can generate a data structure indicating the set of dSlots in the allowed_set. If the dSlots in the allowed_set with low retention values ​​are combined with the dSlots in the flushing_set and are sufficient to service the request, the DRA 620 reclaims those dSlots and begins a cache flush invalidation for those dSlots. If not, the process continues.

[0157] At 1130, in the illustrated embodiment, DRA 620 first determines whether the request is for a low-priority logical slot and acts accordingly, or for a high-priority logical slot and acts accordingly. Note that other granularities of priority may be supported in other embodiments. For a low-priority requester, DRA 620 generates a do_set of slots owned by the low-priority logical slots that have a medium retention value in allowed_set. DRA 620 finds dSlot in both flushing_set and do_set. If these dSlots are sufficient to service the request, DRA 620 reclaims those dSlots and initiates a cache flush invalidation for those dSlots. If not, the process continues.

[0158] For high-priority requesters, DRA 620 generates a do_set of slots owned by high-priority logical slots that have medium retention values ​​in allowed_set. DRA 620 finds dSlots in both flushing_set and do_set. If these dSlots are sufficient to service the request, DRA 620 reclaims those dSlots and initiates cache flush invalidation for those dSlots. If not, the process continues.

[0159] At 1140, in the illustrated embodiment, the DRA 620 adds the slots in the allowed_set that have a high retention value and belong to logical slots with lower priority and lower age to the do_set. The DRA 620 finds the dSlots in both the flushing_set and the updated do_set. If these dSlots are sufficient to service the request, the DRA 620 reclaims those dSlots and begins cache flush invalidation for those dSlots. If not, it may cancel the reclamation and restart arbitration.

[0160] In various embodiments, the disclosed techniques may advantageously provide a balance between keeping hardware slots empty for the current logical slot (to avoid overhead) while still allowing those hardware slots to be reclaimed by other logical slots in certain scenarios.

[0161] Slot Reservation

[0162] In some embodiments, the control circuitry is configured to reserve a hardware slot for a logical slot until instructed to release the slot (e.g., by software). This may allow software to query various boot information, such as performance registers, memory, or other data affected by the boot execution. In some embodiments, each boot includes a retain_slots field (e.g., a bit) that indicates whether the hardware slot mapped to the logical slot should wait to be de-allocated.

[0163] In some embodiments, if a boot that reserves a slot is assigned to a logical slot, other slots cannot reclaim resources from the logical slot regardless of priority.

[0164] Figure 12 is a flow chart illustrating an example method performed by the master control circuitry to handle a boot with reserved slots, according to some embodiments. The process may be performed while communicating with the KSM 350 to allow software communication. At 1210, in the illustrated example, the master control circuitry 210 initiates a boot with the retain_slots field set, indicating that the hardware slots should be reserved.

[0165] At 1220, in the illustrated example, the boot completes its work and the device performs a kernel end flash process.The hardware slots remain mapped.

[0166] At 1230, the master control circuit 210 sends a kick_done signal to the KSM 350. It also transitions dSlot to the Reserve state.

[0167] At 1240 , software or firmware may query performance registers, memory, etc. affected by the startup. At 1250 , the KSM 350 sends a release_slots signal (eg, based on a software instruction indicating that the query is complete).

[0168] At 1260, the master control circuit 210 completes the process of de-allocating the hardware slots, which are converted to an invalid state and are now available for another logical slot. At 1270, the master control circuit 210 sends a de-allocation message to the KSM 350, notifying it that the de-allocation is complete.

[0169] In some embodiments, to avoid hang conditions, boots with reserved slots always use multi-group allocation mode and cannot be blocked from completion. Therefore, when arbitrating between logical slots with reservations and logical slots without reservations, logical slots with reservations may always have priority. In addition, the KSM 350 may only schedule up to a threshold number of logical slots with reservations (e.g., corresponding to the number of dSlots per mGPU). In some embodiments, all logical slots with reservations are promoted to high priority.

[0170] Reserved slot for high-priority boot

[0171] As briefly discussed above, different logical slots may have different priority levels, for example, specified by software. In some embodiments, on a given mGPU, a subset of hardware slots is reserved for logical slots that meet a threshold priority (e.g., the higher priority slots in a system with two priority levels).

[0172] Figure 13 is a block diagram illustrating multiple hardware slots for an mGPU. In some embodiments, one or more dSlots (in Figure 13 ) are reserved for high priority logic slots, and one or more dSlots (shown in solid black in Figure 13 ) is available to all logical slots (and is the only hardware slot available to low-priority logical slots).

[0173] In some embodiments, the high-priority logical slot first attempts to use the reserved hardware slots of the mGPU before attempting to use other slots. In other embodiments, the high-priority logical slot may, for example, use a round-robin technique to try to use all hardware slots of the mGPU equally.

[0174] In some embodiments, low priority logical slots are not allowed to reclaim hardware slots from high priority logical slots unless there is no chance that the high priority logical slot will use them.

[0175] In various embodiments, the disclosed prioritization techniques may advantageously allow software to influence the allocation of important work to reduce obstruction from less important work.

[0176] Refreshing technology

[0177] As discussed above, whenever a hardware slot is assigned to a new logical slot, a cache flush invalidation (CFI) may be performed. Additionally, the main control circuitry 210 must execute any CFI included in the control flow for compute startup. However, because hardware slots can be dynamically mapped at the kernel level, the set of hardware slots to be flushed for control flow CFIs may not be deterministic. The following discussion provides techniques for addressing this phenomenon. Specifically, an "unconditional" CFI is introduced that flushes all relevant mGPUs (e.g., in some implementations, all mGPUs in a graphics processor).

[0178] Figure 14A is a diagram illustrating example cache flush invalidate commands with an unconditional field according to some embodiments. In this example, each cache flush invalidate command 1410 includes an "unconditional" field. A standard (non-unconditional) CFI applies to all hardware slots owned by the logical slot at the time the standard CFI is issued. An unconditional CFI is sent to all mGPUs, even if the logical slot does not own any hardware slots in some mGPUs.

[0179] Figure 14B is a block diagram illustrating an embodiment of a dSlot resource allocator configured to handle unconditional CFIs according to some embodiments. In the illustrated example, DRA 620 includes a core end flush control register 1430 and a de-allocation flush control register 1440. In some embodiments, master control circuitry 210 implements a state machine such that at most one unconditional CFI can be outstanding at any given time. Logical slots can arbitrate for this resource.

[0180] The kernel end flush control register 1430 may maintain a set of bits indicating which mGPUs are to be flushed when a kernel ends. The deallocation flush control register 1440 may maintain a set of bits indicating which mGPUs are to be flushed when a dSlot deallocation occurs in the middle of a boot (note that this may be a subset of the bits specified by the kernel end flush).

[0181] When a dSlot is deallocated, the DRA 620 may implement the following process. First, if the dSlot is not in the last mGPU in the group that has a dSlot allocated to the logical slot, the DRA 620 uses the deallocation flush control register 1440, which can potentially invalidate flushing a smaller number of caches (e.g., one or more L1 caches shared by the group, but not the L2 cache). If the dSlot is in the last mGPU in the group, the DRA 620 uses the core end flush control register 1430 to determine which cache to flush.

[0182] In various embodiments, the disclosed techniques can advantageously avoid non-deterministic flushing behavior, improve cache efficiency, or both.

[0183] Affinity-based allocation

[0184] In embodiments where multiple GPU sub-units (e.g., mGPUs 320A through 320N of group 305) share a cache, the control circuitry can dispatch portions of kernels that access the same memory region to the sub-units that share the cache. This can improve cache efficiency, particularly between kernels of the same launch.

[0185] In some embodiments, the master control circuitry 210 defines a set of affinity regions, which may correspond to a set of hardware that shares resources (such as caches). In some embodiments, there is a fixed relationship between affinity regions and target groups of the mGPU (although the relationship may vary depending on the dimensions of the kernel). The master control circuitry 210 may include control registers that store multiple affinity maps. Each affinity map may specify a relationship between a kernel portion and an affinity region. In this way, each kernel may reference an affinity map that reflects its memory accesses (e.g., as determined and determined by software, which may configure the affinity map and specify an affinity map for each kernel). Thus, software may program potential affinity modes using configuration registers, which may also be shared among multiple data masters. Within a boot, different kernels may be assigned according to different affinity maps.

[0186] Figure 15 is a diagram illustrating an example affinity technique for a collection of graphics work (e.g., compute kernels) according to some embodiments. In the illustrated embodiment, the collection of graphics work (e.g., kernels) includes an affinity map indicator 1515 that specifies an affinity map 1520. For example, the indicator can be a pointer or index into an affinity map table. The affinity map indicates the corresponding target group 305 of the mGPU for N parts of the kernel. Note that the "part" of the kernel may not actually be a field in the affinity map, but may be implied based on the index of the entry. For example, the third entry in the affinity map may correspond to the 3 / Nth part of the kernel. The device may include configuration registers that can be configured to specify multiple different affinity maps. In addition, a given affinity map can be referenced by multiple kernels.

[0187] In some embodiments, rather than mapping portions of a set of graphics work directly to a target group, affinity mapping may use an indirect mapping that maps portions of a set of graphics work to affinity regions and then maps the affinity regions to sets of hardware (e.g., to groups of mGPUs).

[0188] The control circuitry may allocate the set of graphics work based on the indicated affinity mapping. Portions of the set of graphics work 1510 that target the same group may be assigned to the same group / affinity region (and thus may share caches shared by the mGPUs of the group, which may improve cache efficiency).

[0189] Note that while the disclosed embodiments specify affinity at the granularity of a group of mGPUs, affinity can be specified and implemented at any of a variety of suitable granularities, for example, utilizing shared caches at various levels in the memory hierarchy. The disclosed embodiments are included for purposes of illustration and are not intended to limit the scope of the present disclosure.

[0190] Figure 16 is a block diagram illustrating example circuitry configured to allocate batches of workgroups from a kernel based on affinity in accordance with some embodiments. In the illustrated embodiment, the control circuitry for one logical slot includes: a control flow processor 630, a master kernel walker 1610, group walkers 1620A through 1620N, a group walker arbiter 1630, an mGPU dispatch circuit 650, a boot slot arbiter 660, and a communication fabric 1660. Similar circuitry may be instantiated for each logical slot supported by the device. Note that elements 1610, 1630, and 1640 may be included in the kernel processor 640 discussed above, and similarly numbered elements may be referenced as above. Figure 6 Configured as described.

[0191] Each kernel may be organized into workgroups in multiple dimensions (typically three). These workgroups may in turn include multiple threads (also referred to as work-items). In the illustrated embodiment, the main kernel traverser 1610 is configured to iterate through the kernel according to the specified affinity map to provide affinity sub-kernels that include portions of the kernel that target the mGPU's group. The main kernel traverser 1610 may use the coordinates of the sub-kernel's initial workgroup to indicate the sub-kernel assigned to a given group traverser 1620. Note that in Figure 16 The various kernel data sent between elements may not include actual work, but may be control signaling, for example using kernel coordinates to indicate the location of the work to be assigned.

[0192] For kernels with different dimensions, the main kernel traverser 1610 may divide the kernel into N affinity regions. For example, in an embodiment where each affinity map has N affinity regions, the main kernel traverser 1610 may use all N regions for a one-dimensional kernel. For a two-dimensional kernel, the main kernel traverser 1610 may divide the kernel into N affinity regions. take For a three-dimensional kernel, the main kernel traverser 1610 may divide the kernel into rectangular affinity regions (as an example, an affinity region spanning the entire z dimension). take grid).

[0193] In the illustrated embodiment, the group traverser 1620 is configured to independently traverse the corresponding affinity sub-kernels and generate batches, where each batch includes one or more work groups. A batch can be the granularity at which computational work is dispatched to the mGPU. Note that a given affinity sub-kernel can be divided into multiple thread-restricted traversal order sub-kernels, as described below with reference to Figure 17 Various techniques for controlling the order of kernel traversal are discussed in U.S. patent application Ser. No. 17 / 018,913, filed Sep. 11, 2020, and group traverser 1620 may use these techniques to traverse affinity subkernels.

[0194] In the illustrated embodiment, the group walker arbiter 1630 is configured to arbitrate among the available batches, and the mGPU assignment circuit 650 is configured to assign the selected batches to the mGPUs.

[0195] The dispatch circuitry 650 can dispatch mGPUs using mGPU masks and load balancing, subject to any software overrides. The boot slot arbiter 660 arbitrates between the prepared batches and sends them to the target mGPU via the communication fabric 1660. The communication fabric 1660 can be a workload dispatch shared bus (WDSB) configured to send control signaling indicating the attributes of the dispatched work and tracking signaling indicating completion of the work, for example, as discussed in the '943 patent application cited above.

[0196] In some embodiments, the device may, for example, use control circuitry to turn off affinity-based scheduling based on software control or under certain conditions. In this case, the main kernel traverser 1610 may assign the entire kernel to a single set of traversers 1620.

[0197] Each instance of distributed control circuitry 340 in an mGPU may include an input queue and a batch execution queue to store received batches before assigning workgroups to shader pipelines for execution.

[0198] Figure 17 is a diagram illustrating an example kernel iteration according to some embodiments. In the illustrated embodiment, kernel 1710 includes multiple parts (M parts in one dimension and X parts in another dimension). Each of these parts can be referred to as an affinity sub-kernel and can be mapped to an affinity region (note that multiple affinity sub-kernels can be mapped to the same affinity region).

[0199] In the illustrated example, section A0 includes multiple thread-restricted sub-kernel sections A through N. Within each affinity sub-kernel, group walker 1620A may use restricted iteration as described in the '913 application. As shown, thread-restricted sub-kernel section A is divided into multiple batches, which may be allocated via communication fabric 1660 (where each block in a batch represents a work group). In the disclosed embodiment, all batches from section A0 may be assigned to the same group of the mGPU (and it should be noted that other sections of kernel 1710 may also target that group of the mGPU). In various embodiments, the disclosed affinity techniques may advantageously improve cache efficiency.

[0200] In some embodiments, affinity-based scheduling can temporarily reduce performance in certain situations (e.g., for non-homogeneous kernels). For example, some groups of mGPUs may still be working on complex portions of a kernel while other groups have already completed less complex portions. Therefore, in some embodiments, the graphics processor implements a work-stealing technique to override affinity-based scheduling, for example, at the end of a kernel. In these embodiments, groups of mGPUs that are idle for use by a kernel can take work from groups still working on that kernel, which can advantageously reduce the kernel's overall execution time.

[0201] In some embodiments, the control circuitry selects one or more donor groups of mGPUs (e.g., the group with the most remaining work) and selects other groups of mGPUs that are in certain states (e.g., have completed all of their work for the kernel or at least a threshold amount of their work) as work acceptor groups. The work acceptor groups can receive batches from the affinity child kernels assigned to the donor groups, thereby overriding the affinity technique in some cases.

[0202] Figure 18 is a block diagram illustrating example circuitry configured to facilitate work sharing according to some embodiments. In the illustrated embodiment, master kernel walker 1610 includes circuits 1810A through 1810N configured to track the remaining portion of kernels (e.g., affinity sub-kernels) targeted for each group of the mGPU. For example, if a given group is targeted by seven affinity sub-kernels and has received four affinity sub-kernels, then three affinity sub-kernels remain for that group.

[0203] In the illustrated embodiment, work sharing control circuit 1820 is configured to select work donor groups and acceptor groups based on information maintained by circuit 1810. In the illustrated embodiment, information identifying these groups is maintained in circuits 1830 and 1840. In some embodiments, a group is eligible to work only if it is associated with an affinity region in the affinity map of the core. In some embodiments, once a group has assigned all of its assigned work (assigned via the affinity map) to the core, the group becomes eligible to work for the core.

[0204] In some embodiments, the working donor group is the group that is furthest along (has the largest number of remaining parts to be dispatched). As groups become eligible to receive work, they can lock onto the donor group. As shown, the main kernel traverser 1610 can send state information (e.g., coordinate basis information for the affinity sub-kernel) for synchronization of such acceptor groups.

[0205] The group kernel walker for the donor (1620A in this example) generates batches of work groups to be sent to the mGPUs in its corresponding group or the mGPUs of any work recipient group in the work recipient group. For example, the set of eligible mGPUs may be specified by the mGPU mask from the group walker 1620A, so that the mGPU assignment circuit 650 may select from the set of eligible mGPUs based on load balancing.

[0206] In some embodiments, once a donor set completes dispatch for its current portion (e.g., affinity subkernel), the acceptor becomes unlocked and a new donor can be selected, and the process can continue until the entire kernel is dispatched.

[0207] Example Boot Slot Manager Circuit

[0208] Figure 19A is a block diagram illustrating an example boot slot manager according to some embodiments. In the illustrated embodiment, the boot slot manager 350 implements a software interface and includes a register replication engine 1910 and dependency tracking and status circuitry 1920 (e.g., a scoreboard). In the illustrated embodiment, the boot slot manager 350 communicates with a memory interface 1930, a control register interface 1940, and the main control circuitry 210.

[0209] In some embodiments, the boot slot manager 350 implements multiple "top slots" to which software can assign boots. These top slots are also referred to herein as "tracking slots." The boot slot manager 350 can then handle software-specified dependencies between boots, map boots from tracking slots to logical slots in the main control circuitry 210, track boot execution status, and provide status information to the software. In some embodiments, a dedicated boot slot manager circuit can advantageously reduce boot-to-boot transition time relative to a software-controlled implementation.

[0210] In some embodiments, register copy engine 1910 is configured to retrieve register data (e.g., for startup configuration registers) from memory via memory interface 1930 and program configuration registers for startup via interface 1940. In some embodiments, register copy engine 1910 is configured to prefetch configuration register data into an internal buffer (e.g., a memory buffer) before allocating shader resources for startup. Figure 19A In various embodiments, this can reduce the boot-to-boot transition time when initiating a new boot. Register copy engine 1910 can access control register data via memory interface 1930 and can write to control registers via control register interface 1940.

[0211] In some embodiments, the register copy engine 1910 is configured to prefetch data for startup in priority order and may not wait to retrieve the initially requested register data before requesting additional data (this may absorb memory latency associated with reading register data). In some embodiments, the register copy engine 1910 supports mask broadcast register programming, for example, based on an mGPU mask, so that the appropriate distributed slots are programmed. In some embodiments, programming control registers using the register copy engine 1910 may offload work from the main firmware processor.

[0212] In some embodiments, the boot slot manager 350 is configured to schedule the boot and send work assignment information to the master control circuit 210 before programming all configuration registers for the boot. Generally speaking, the initial boot scheduling can be pipelined. This may include setting up the stage register programming, the master control circuit identifying the distributed slots, the register replication engine 1910 and the master control circuit queuing work to program the control registers in parallel, and starting the queued work once the final control register has been written. In some embodiments, this allows downstream circuits to receive and queue work assignments and quickly begin processing once the configuration registers are written, thereby further reducing the boot-to-start transition time. Specifically, this can save latency associated with multiple control bus traversals, compared to waiting to queue work until all control registers are programmed.

[0213] Dependency tracking and status circuitry 1920 can store information received from software and provide status information to the software via a software interface, as discussed in detail below. In some embodiments, tracking slots are shared by multiple types of master control circuits (e.g., compute, pixel, and vertex control circuits). In other embodiments, certain tracking slots can be reserved for certain types of master control circuits.

[0214] Figure 19B is a diagram illustrating example tracking and status data for each tracking slot according to some embodiments. In the illustrated embodiment, circuitry 1920 maintains the following information for each tracking slot: identifier, status, data identification, dependencies, operational data, and configuration. Each of these example fields is discussed in detail below. In some embodiments, the status and operational data fields are read-only by software, while the other fields are software-configurable.

[0215] Each trace slot can be assigned a unique ID. Therefore, the boot slot manager 350 can support a maximum number of trace slots. In various embodiments, the number of supported trace slots can be selected so that it is relatively rare for independent boots that are small enough to use all available trace slots to be scheduled in parallel. In some embodiments, the number of supported trace slots is greater than the number of supported logical slots.

[0216] In some embodiments, the status field indicates the current state of the slot and whether the slot is active. If applicable, the field may also indicate the logical slot and any distributed slots assigned to the tracking slot. In some embodiments, the status field supports the following status values: clear, programming complete, register fetch start, waiting for parent process, waiting for resource, waiting for distributed slot, running, stop request, deallocated, dequeued by boot slot manager, dequeued by master control circuit, context stored, and completed. In other embodiments, the status field may support other states, subsets of the described states, etc. Figure 21 The example states are discussed in detail in the State Machine section.

[0217] In some embodiments, the data identification field indicates the location of the control register data for the startup. For example, this can be specified as an initial register address and a number of configuration registers. It can also include a register context identifier. In some embodiments, the data identification field also indicates other resources used by the startup, such as samplers or memory holes. In some cases, some of these resources can be hard resources, so that the startup cannot proceed until they are available, while other resources can be soft resources, and the startup can proceed without them or with only a portion of the requested resources. As an example, memory holes can be considered soft resources and can allow the startup to proceed even if their soft resources are not available (potentially with a notification sent to the requesting software).

[0218] In some embodiments, the dependency field indicates any dependencies a slot has on boots in other slots. As an example, circuit 1920 may implement an NxN matrix (where N is the number of tracking slots), where each slot includes an entry for each other slot, indicating whether the slot is dependent on the other slot. When boots from other slots complete, the entry may be cleared. In other embodiments, other techniques may be used to encode dependencies. The boot slot manager 350 may assign a tracking slot to a logical slot based on the indicated dependencies (e.g., by waiting to assign a boot to a logical slot until all the tracking slots it depends on have completed). Moving dependency tracking from software / firmware control to dedicated hardware may allow for more efficient use of logical slots and reduce boot-to-boot transitions.

[0219] In some embodiments, the Run Data field provides information about the operational status of the boot. For example, this field can provide a timestamp for assigning the boot to a logical slot when the boot began running on a distributed slot and when the boot completed. Various other performance or debugging information can also be indicated. In some embodiments, various tracking slot information is retained for the slots for which the Reserved field is set, and their mapped hardware resources are not released (potentially allowing access to the status register at the logical slot level, the distributed slot level, or both).

[0220] In some embodiments, the configuration field indicates the type of main control circuit that controls the slot (e.g., compute, pixel, or vertex), the priority of the slot, a reserved slot indication, a force start end interrupt indication, or any combination thereof. For example, the configuration field may be programmable by software to indicate the configuration of the slot and provide certain software overlay information. The kernel end interrupt may be set globally, or may be set to trigger each startup (or after a threshold number of startups). This can advantageously reduce the firmware time spent processing interrupts (by omitting interrupts in some cases) while still retaining interrupt functionality when needed.

[0221] In various embodiments, the disclosed tracing circuitry may allow software to handle multiple launches in parallel (eg, be able to start, stop, query, and modify the execution of those launches).

[0222] Figure 20 is a diagram illustrating an example register prefetch buffer organization according to some embodiments. In the illustrated embodiment, registers are organized by type (e.g., in this example, all setup registers are at the beginning of the buffer and execution registers are at the end of the buffer). Generally speaking, setup registers are used to configure a boot before it starts, and execution registers are used for distributed execution of a boot. In the illustrated embodiment, the buffer indicates the offset within the configuration register space where the register is located and its payload.

[0223] This organization of prefetched register data can advantageously allow rewriting of previous registers, e.g., for startup to startup buffer reuse, while still allowing new registers to be saved at the beginning or end of a block of registers of a given type. In various embodiments, two or more registers of different types can be grouped together by type to facilitate such techniques. In some embodiments, the register prefetch buffer is an SRAM. In other embodiments, the register prefetch buffer is a cache and can evict entries when additional space is needed (e.g., according to a least recently used algorithm or another appropriate eviction algorithm).

[0224] Figure 21is a state machine diagram illustrating an example boot slot manager state according to some embodiments. From the clear state 2110, the control circuit is configured to enable the slot to allocate the slot for boot. Figure 19B When programming is complete (as discussed in the dependencies and configurations discussed above), the state transitions to the "Programming Complete" state 2112. After the register copy engine 1910 has accepted the fetch request, the state transitions to the register fetch start 2114 (note that in the illustrated embodiment, this is a prefetch before resources are allocated to the trace slot). After the register copy engine 1910 has indicated the fetch is complete, the state transitions to the "Waiting for Parent" state 2116. Once all dependencies for the trace slot are satisfied, the state transitions to the "Waiting for Resources" state 2118.

[0225] As shown, if a stop is requested in any of states 2110 to 2118, the state transitions to "Queue from KSM" 2126. Once the slots are reset, the state transitions back to the empty state 2110. Note that state 2116 may require significantly fewer deallocation operations than the other stop states discussed in detail below, for example, because resources have not yet been allocated to the slots.

[0226] Once the resource is allocated, the state transitions to the "waiting for dSlot state" 2120, and the KSM waits for a control response (e.g., from the main control circuit) at 2124. Once the dSlot is allocated, the state transitions to the running state 2122. If a stop is requested in these states (shown at 2128), the KSM waits for a control response at 2130. If a start is made after a stop request or from the running state 2122, the slot is de-allocated at 2132, and the start is completed at 2138.

[0227] If a stop is requested in state 2120 or 2122 and the control response 2130 indicates that the logical slot is stored, the state transitions to the deallocation state 2134 and waits for the context to be stored at 2140 before resetting the slot. If the control response at 2130 indicates to exit the queue, the state transitions to deallocation 2136 and then "exit the queue from the main control circuit" 2142 before resetting the slot (relative to states 2134 and 2140, this can be a more elegant exit from the queue that does not require context storage of the logical slot). Generally speaking, the disclosed technology can advantageously allow the main control circuit to suspend the scheduling of multiple levels of work and allow firmware to interact with hardware in a safe manner.

[0228] Once the slot is reset from state 2138, 2140, or 2142, the startup slot manager determines whether the reserved field is set and, if not, transitions back to the clear state 2110. If the reserved field is set, the KSM waits for any assigned logical slots to be deassigned (e.g., based on software control) at 2148. Generally, tracking slots may be recycled automatically unless they are explicitly reserved.

[0229] As discussed above, dependency tracking and status circuitry 1920 may provide the current state of each socket to the software.

[0230] In some embodiments, the boot socket manager 350 can scale across multiple GPU sizes, for example, by allowing variations in the number of supported trace sockets. The disclosed dynamic hierarchical scheduling of trace sockets (via firmware or software), then logical sockets (via master control circuitry), and then distributed sockets can advantageously provide efficient allocation with scheduling intelligence allocated across hierarchical levels.

[0231] In some embodiments, the boot slot manager 350 is configured to perform one or more power control operations based on the tracking slots. For example, the control circuitry can reduce the power state of one or more circuits (e.g., through clock gating, power gating, etc.). In some embodiments with a large number of tracking slots, the control circuitry can reduce the power state of other circuits even when the other circuits have work queued in the tracking slots. For example, the control circuitry can reduce the power state of the pixel data master even when it has boots in the tracking slots.

[0232] In some embodiments, for a scheduled trace slot, if it is in a lower than desired power state, the first action of the scheduled trace slot is to increase the power state of any associated circuitry. For example, the control circuitry may initiate pixel activation by writing to a power-on register for the pixel data master. Generally, a device may power gate various types of logic (e.g., caches, filtering logic, ray tracing circuitry, etc.) and turn on those logic blocks when a trace slot is about to use that logic. In some embodiments, the boot slot manager 350 maintains one or more flags for each trace slot that indicate whether the activation assigned to the trace slot uses one or more types of circuitry. The boot slot manager 350 may cause those types of circuitry to meet the required power state in response to the scheduling of those trace slots.

[0233] Example Method

[0234] Figure 22 is a flow chart illustrating an example method for allocating graphics work using logical slots according to some embodiments. Figure 22The methods shown can be used in conjunction with any of the computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown can be performed concurrently in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.

[0235] At 2210, in the illustrated embodiment, the control circuitry assigns a first set of graphics work and a second set of graphics work to a first logical slot and a second logical slot. In some embodiments, the circuitry implements multiple logical slots, and the set of graphics processor subunits each implements multiple distributed hardware slots. In some embodiments, the graphics processor subunits are organized into multiple groups of multiple subunits, wherein subunits in the same group share a cache. In some embodiments, the subunits of a given group are implemented on the same physical die. In some embodiments, the subunits each include: a fragment generator circuit, a shader core circuit, a memory system circuit including a data cache and a memory management unit, a geometry processing circuit, and a distributed workload distribution circuit. In some embodiments, the distributed hardware slots each include: a configuration register, a batch queue circuit, and a batch iteration circuit. In various embodiments, the shader circuitry in the subunit is configured to receive work from its multiple distributed hardware slots and execute the work.

[0236] The expression "the set of graphics processor subunits each implements a plurality of distributed hardware slots" means that the set of graphics processor subunits includes at least two subunits, each of the subunits implementing a plurality of distributed hardware slots. In some embodiments, a device may have additional graphics processor subunits (not in the set) that do not necessarily implement a plurality of distributed hardware slots. Thus, the phrase "the set of graphics processor subunits each implements a plurality of distributed hardware slots" should not be interpreted to mean that in all cases, all subunits in a device implement a plurality of distributed hardware slots, it merely provides the possibility that this may be true in some cases and not in others. Similar interpretations are intended for other expressions in this text using the term "each."

[0237] At 2220, in the illustrated embodiment, the control circuitry determines an allocation rule for the first set of graphics work, the allocation rule indicating allocation to all graphics processor subunits in the set.

[0238] At 2230, in the illustrated embodiment, the control circuitry determines an allocation rule for the second set of graphics jobs that indicates allocation to fewer than all of the graphics processor sub-units in the set. In some embodiments, the determined allocation rule for the second set of graphics jobs indicates allocation of the first set of graphics jobs to a single group of sub-units. Alternatively, the determined allocation rule for the second set of graphics jobs may indicate allocation of the second set of graphics jobs to a single sub-unit.

[0239] The control circuitry may select a first allocation rule and a second allocation rule based on the amount of work in the first set of graphics work and the second set of graphics work. The control circuitry may determine the first allocation rule based on one or more software overrides signaled by the graphics program being executed. These software overrides may include any appropriate combination of the following types of example software overrides: mask information indicating which subunits are available for the first set of work; specifying allocation rules; group information indicating the groups of subunits of the first set of work that should be deployed; and policy information indicating a scheduling policy. In some embodiments, the control circuitry determines corresponding retention values ​​for slots in a plurality of logical slots, wherein the retention values ​​indicate a condition of a core of the logical slot. The control circuitry may allow a logical slot having a first priority level to reclaim a hardware slot assigned to a logical slot having a lower second priority level based on one or more of the corresponding retention values.

[0240] The first set of graphics work and the second set of graphics work may be launches. The first set of graphics work and the second set of graphics work may be compute kernels in the same launch or in different launches. Thus, in some embodiments, the first set of graphics work is a first kernel of a compute launch assigned to a first logical socket, wherein the compute launch includes at least one other kernel, and wherein the apparatus is configured to select a different allocation rule for the at least one other kernel than for the first kernel.

[0241] At 2240, in the illustrated embodiment, the control circuitry determines a mapping between a first logical slot and a first set of one or more distributed hardware slots based on a first allocation rule.

[0242] At 2250, in the illustrated embodiment, the control circuitry determines a mapping between a second logical slot and a second set of one or more distributed hardware slots based on a second allocation rule.

[0243] At 2260, in the illustrated embodiment, the control circuitry distributes the first set of graphics work and the second set of graphics work to one or more of the graphics processor subunits according to the determined mapping.

[0244] In some embodiments, the control circuitry for a logical socket includes: a control flow processor (e.g., CSP 630) configured to determine a first allocation rule and a second allocation rule; a core processor (e.g., circuit 640) configured to generate batches of computational workgroups; and subunit assignment circuitry (e.g., circuit 650) configured to assign batches of computational workgroups to subunits. In some embodiments, the control circuitry includes: a hardware socket resource allocator circuit (e.g., circuit 620) configured to allocate hardware sockets to the control flow processor based on the indicated allocation rule; and a logical socket arbiter circuit (e.g., circuit 660) configured to arbitrate between batches from different logical sockets for allocation to the assigned subunits. In some embodiments, the hardware socket resource allocator circuit is configured to allocate hardware sockets based on the state of the hardware sockets. For example, the states of different hardware sockets may include at least: invalid, running, clear, and refreshing.

[0245] In some embodiments, the device is configured to perform multiple types of cache flush invalidation operations, which may include: a first type of cache flush invalidation operation that flushes and invalidates the cache only for one or more sub-units to which the core is assigned; and an unconditional type of cache flush invalidation operation that flushes and invalidates all caches for a set of graphics processor sub-units at one or more cache levels.

[0246] Figure 23 is a flow chart illustrating an example method for prioritizing logical slots according to some embodiments. Figure 23 The methods shown can be used in conjunction with any of the computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown can be performed concurrently in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.

[0247] At 2310, in the illustrated embodiment, the control circuitry receives a first set of software-specified graphics jobs and software-indicated priority information for the first set of graphics jobs.

[0248] At 2320, in the illustrated embodiment, the control circuitry assigns a first set of graphics jobs to a first logical slot of a plurality of logical slots implemented by the device.

[0249] At 2330, in the illustrated embodiment, the control circuitry determines a mapping between logical slots and distributed hardware slots implemented by the graphics subunit of the device, wherein the mapping reserves a threshold number of hardware slots in each subunit for logical slots having a priority exceeding a threshold priority level. In some embodiments, a first subset of the logical slots are high priority slots, and the remaining logical slots are low priority slots. In these embodiments, the control circuitry may assign the first set of graphics jobs to the first logical slot based on priority information indicated by software. In other embodiments, various other techniques may be used to encode and track priorities.

[0250] At 2340, in the illustrated embodiment, the control circuitry distributes the first set of graphics work to one or more of the graphics processor subunits according to one of the mappings.

[0251] In some embodiments, control circuitry (e.g., distributed slot resource allocator circuitry) is configured to execute a reclamation process that allows logical slots having a first software-indicated priority level to reclaim hardware slots that were assigned to logical slots having a lower second priority level.

[0252] In some embodiments, based on software input (e.g., a reserve slot command) for the first set of graphics work, the control circuitry is configured to maintain the mapping of the distributed hardware slot for the first logical slot after completing processing of the first set of graphics work. In some embodiments, the control circuitry assigns the mapped distributed hardware slot for the first set of graphics work to another logical slot only after software input instructs to release the mapped distributed slot.

[0253] In some embodiments, the control circuitry provides status information to the software for the first set of graphics work. The control circuitry can support various status states, including, but not limited to, waiting for a dependency, waiting for configuration data for the first set of graphics work, waiting for assigned distributed slots, waiting for hardware resources, flushing, program completion, waiting for logical slots, deallocation, and context storage. For example, the status information can identify the first logical slot, identify the assigned distributed hardware slot, or indicate timestamp information associated with the execution of the first set of graphics work.

[0254] In addition to or in lieu of the priority information, the control circuitry may support various software control or override functions, including, but not limited to: specifying allocation rules indicating whether to allocate to only a portion of the graphics processor sub-units in the set or to allocate to all of the graphics processor sub-units in the set; group information indicating the group of sub-units of the first set on which graphics work should be deployed; mask information indicating which sub-units may be used for the first set of graphics work; and policy information indicating a scheduling policy.

[0255] In some embodiments, the device includes: a control flow processor circuit configured to determine an allocation rule for a mapping; and a distributed slot resource allocator circuit configured to determine the mapping based on: software input, the determined allocation rule from the control flow processor circuit, and distributed slot status information.

[0256] Figure 24 is a flow chart illustrating an example method for affinity-based scheduling according to some embodiments. Figure 24 The methods shown can be used in conjunction with any of the computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown can be performed concurrently in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.

[0257] At 2410, in the illustrated embodiment, control circuitry (e.g., kernel walker circuitry) receives a software-specified set of graphics work (e.g., compute kernels) and a software-indicated mapping of portions of the set of graphics work to groups of graphics processor subunits. A first group of subunits may share a first cache and a second group of subunits may share a second cache. Note that the mapping may or may not identify a specific group of graphics subunits. Conversely, the mapping may specify that multiple portions of a compute kernel should be assigned to the same group of graphics processor subunits, but may allow hardware to determine which group of graphics processor subunits to actually assign.

[0258] At 2420 , in the illustrated embodiment, the control circuitry assigns a first subset of the set of graphics work to the first group of graphics subunits and a second subset of the set of graphics work to the second group of graphics subunits based on the mapping.

[0259] The control circuitry may be configured to store in configuration registers a plurality of mappings of portions of a collection of graphics work to groups of graphics processor sub-units.

[0260] The kernel walker circuit may include: a main kernel walker circuit (e.g., Figure 161610), the main kernel traverser circuit is configured to determine a portion of the computation kernel; a first set of traverser circuits (e.g., Figure 16 The kernel walker circuits may further include: a group walker arbitration circuit (e.g., Figure 16 element 1630), the group visitor arbitration circuit being configured to select from batches of the work group determined by the first group visitor circuit and the second group visitor circuit; and a sub-unit assignment circuit (e.g., mGPU assignment circuit 650), the sub-unit assignment circuit being configured to assign the batch selected by the group visitor arbitration circuit to one or more graphics sub-units in the group of sub-units corresponding to the selected group visitor circuit.

[0261] In some embodiments, the device includes a work sharing control circuit that is configured to: determine a set of one or more other groups of sub-units that have assigned all of their assigned portions to a computing kernel; and assign at least a first portion of the computing kernel that is indicated by a mapping to target a first group of sub-units to a group of the one or more other groups of sub-units.

[0262] In some embodiments, the control circuitry disables affinity-based work allocation in one or more operating modes. The control circuitry may support mapping portions of compute kernels to sets of graphics processor subunit affinity maps for compute kernels of multiple dimensions, including single-dimensional kernels, two-dimensional kernels, and three-dimensional kernels.

[0263] In some embodiments, a non-transitory computer-readable medium stores thereon instructions executable by a computing device to perform operations comprising: receiving compute kernels and corresponding mappings of portions of the compute kernels to groups of graphics processor subunits, wherein the compute kernels and the mappings are specified by the instructions and the mappings indicate cache affinities of a set of portions of the compute kernels mapped to a given group of graphics processor subunits; and based on the mappings, assigning a first subset of the compute kernels to a first group of graphics subunits and assigning a second subset of the compute kernels to a second group of graphics subunits.

[0264] Figure 25 is a flow chart illustrating an example method for initiating slot manager operations according to some embodiments. Figure 25The methods shown can be used in conjunction with any of the computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown can be performed concurrently in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.

[0265] At 2510, in the illustrated embodiment, control circuitry (e.g., socket manager circuitry) uses an entry of a tracking socket circuitry to store software-specified information for a collection of graphics jobs, wherein the information includes: the type of job, dependencies on other collections of graphics jobs, and the location of data for the collection of graphics jobs.

[0266] In some embodiments, the tracking slot circuitry is software-accessible for querying various information associated with a collection of graphics jobs. This may include, for example, the status of the collection of graphics jobs, timestamp information associated with the execution of the collection of graphics jobs, information indicating a logical home slot, and information indicating one or more distributed hardware slots. In some embodiments, the tracking slot circuitry supports status values ​​indicating at least the following status states of a collection of graphics jobs: flushed, register fetch initiated, waiting for one or more other collections of graphics jobs, waiting for a logical slot resource, waiting for a distributed hardware slot resource, and running.

[0267] At 2520, in the illustrated embodiment, the control circuitry pre-fetches configuration register data for the set of graphics work from this location and before allocating shader core resources for the set of graphics work. Note that the pre-fetch may occur after the tracking slots for the set of graphics work are configured but before the control circuitry determines to start the set of graphics work (e.g., before all of its dependencies have been satisfied). The control circuitry may utilize various criteria to determine when to begin the pre-fetch. The pre-fetch may be performed from shared memory (which may be shared between multiple instances of the control circuitry, shared with non-GPU processors, or both) into an SRAM memory element of the socket manager circuitry.

[0268] In some embodiments, before completing programming of the configuration registers, the control circuitry sends the portion of the set of graphics work to the hardware slot assigned to the set of graphics work. The hardware slot may include queue circuitry for the received portion of the set of graphics work.

[0269] At 2530, in the illustrated embodiment, the control circuitry uses the prefetched data to program configuration registers for the set of graphics jobs. The configuration registers may specify attributes of the set of graphics jobs, the location of data for the set of graphics jobs, parameters for processing the set of graphics jobs, etc. The configuration registers may be distinct from the data registers that store data to be processed by the set of graphics jobs.

[0270] At 2540, in the illustrated embodiment, the control circuitry initiates processing of the set of graphics jobs by the graphics processor circuitry according to the dependencies. The control circuitry may assign the set of graphics jobs to a logical master socket (and at least a portion of the configuration register data is available to the configuration registers of the logical master socket) and assign the logical socket to one or more distributed hardware sockets (and at least a portion of the configuration register data is available to the configuration registers of the one or more distributed hardware sockets).

[0271] In some embodiments, the control circuitry is configured to initiate an increase from a lower power mode to a higher power mode of one or more circuits associated with the set of graphics work in conjunction with initiating the set of graphics work from an entry of the tracking slot circuit and based on information about the set of graphics work.

[0272] In some embodiments, the graphics instruction specifies information for storing a set of graphics work (e.g., indicating the type of work, dependencies on other sets of graphics work, and the location of data for the set of graphics work) and queries the tracking socket circuitry to determine status information for the set of graphics work (e.g., status from among: flushing, initiating register fetch, waiting for one or more other sets of graphics work, waiting for logical socket resources, waiting for distributed hardware socket resources, and running, timestamp information associated with execution of the set of graphics work, information indicating the assigned logical primary socket, and information indicating the assigned distributed hardware socket).

[0273] In some embodiments, in response to a stop command for a collection of graphics jobs, the control circuitry is configured to perform different operations depending on the current status of the tracking slots. For example, the control circuitry may reset entries in a tracking slot circuitry in response to determining that a logical master slot has not been assigned. As another example, the control circuitry may de-allocate the logical master slot and reset entries in a tracking slot circuitry in response to determining that a logical master slot has been assigned. As yet another example, the control circuitry may perform one or more context switch operations, de-allocate one or more distributed hardware slots, de-allocate the logical master slot, and reset entries in a tracking slot circuitry in response to determining that one or more distributed hardware slots have been assigned.

[0274] Part II

[0275] The following sections discuss improved or alternative launch scheduling and allocation techniques, streaming launch slot manager techniques with launch queues in memory (e.g., DRAM), logical launch slots in a geometry launch context, event flags for dependencies that are independent of other launches, and techniques for pipelining dependent launches.

[0276] mGPU-level rationing

[0277] In some embodiments, the launch slot manager 350 is configured to allocate logical launch slots so that the logical launch slots use the minimum number of mGPUs required for a given launch (e.g., as opposed to capturing the number of mGPUs allocated to one, the number of mGPUs in a group, or all mGPUs in a machine). In these embodiments, a launch can use any combination of mGPUs (e.g., in an eight-mGPU system, a launch can use 1, 2, 3, 4, 5, 6, 7, or 8 mGPUs). In various embodiments, this can advantageously allow more small or medium-sized launches to run in parallel (and for small launches, to launch earlier), which can increase GPU hardware utilization.

[0278] Note that small launch performance can be important in various contexts, such as games written for immediate-mode GPUs that do not minimize launch overhead and include a large number of small launches. Additionally, fine-grained mGPU-level provisioning can be combined with the "start when ready" technique discussed below to enable large launches to start earlier.

[0279] In some embodiments using tile-based rendering, the fragment / pixel master control circuitry is configured to calculate the number of mGPUs to launch for a fragment based on the number of tiles in the launch. The following pseudo-code provides an example technique that the fragment generator circuitry can use to determine the number of mGPUs (and therefore distributed slots) to allocate to a fragment launch:

[0280]

[0281]

[0282] In some embodiments, the compute control circuitry is configured to calculate the number of mGPUs to allocate to a compute launch based on the number of workgroups and / or workitems in the launch or kernel. The compute control circuitry can dynamically readjust the number of distributed slots at different dispatch boundaries within a launch. The following pseudocode provides an example technique that the compute control circuitry can use to determine the number of mGPUs (and therefore distributed slots) to allocate to a compute launch or kernel:

[0283]

[0284]

[0285] Note that the above allocations may be affected by software overrides, such as forcing the use of a single mGPU, a specific mGPU group, using all mGPUs, etc. Distributed slot assignments can take into account mGPU-level allocations, as described below. Figure 26 This logic can be implemented as a loop over a single mGPU slot dispatch mechanism, which can in some embodiments facilitate work being sent immediately.

[0286] Figure 26 FIG2 is a diagram illustrating an example setup scenario for starting six of eight mGPUs in a graphics processor. In this example, each group includes four mGPUs, and each mGPU includes two allocated slots for the type of work being allocated. Note that this example is included for illustrative purposes, but a variety of numbers of groups, mGPUs per group, and distributed slots per mGPU for a given type of work can be implemented.

[0287] In this example, launch 0 targets six mGPUs, launch 1 targets two mGPUs, launch 2 targets one mGPU, launch 3 targets eight mGPUs, launch 4 targets three mGPUs, and launch 5 targets four mGPUs.

[0288] Figure 27 is a diagram illustrating example allocations of launches that trap launches to single-mGPU, single-group, or full-machine allocation modes, according to some embodiments.

[0289] At time T0, launch 0 is allocated using the full-machine allocation model. Therefore, at T1, launch 1 uses the distributed slots on mGPU0 and mGPU1 (even though the two upper distributed slots on mGPU6 and mGPU7 are not used). Launch 1 uses group-level allocation, which also blocks the second slot on mGPU 2 and mGPU 3.

[0290] At time T2, launch 0 completes. At time T3, launch 2 is scheduled on mGPU 0. At time T4, launch 1 has completed. At time T5, launch 3 is scheduled across all mGPUs. At time T6, launch 4 is scheduled across the three mGPUs in group 1 (using group-level rationing). At time T7, launch 2 and launch 4 have completed. At time T8, launch 5 is scheduled. As discussed in detail below, Figure 28 In the example of mGPU-level provisioning, the Figure 27 Example of higher utilization.

[0291] Figure 28 is a diagram illustrating example mGPU granularity allocation for startup according to some embodiments, relative to Figure 27 The example can provide better utilization of execution resources.

[0292] At time T0, launch 1 is assigned to the six mGPUs. Therefore, at T1, launch 1 can use the upper distributed slots in mGPUs 6 and 7. At time T2, launch 2 is scheduled in the lower slot of mGPU 0. At time T3, launch 0 has completed. At time T4, launch 3 is assigned to the eight mGPUs, including the lower slots on mGPUs 6 and 7. At time T5, launch 1 has completed. At time T6, launch 4 has been assigned to mGPUs 1 through 3. At time T7, launch 5 is assigned to mGPUs 4 through 7.

[0293] In this example, relative to Figure 27 For the example, the processor distributed all launches to the machine in 120 time units within 160 time units. This suggests that for some programs, resource utilization can be improved through mGPU-level rationing / distribution.

[0294] Start when ready scheduling for the startup portion

[0295] Note that in some embodiments discussed above, a given launch may wait until the total number of its desired mGPUs is available before launching any portion of the launch. This may reduce the total execution time between the start and end execution of the launch across N mGPUs. However, running the same launch on a single mGPU as a launch on N consecutive portions may utilize the same amount of compute resources. Thus, in some embodiments, it may be desirable to launch portions of a launch as long as distributed slots are available, regardless of whether the full number of desired mGPUs is available. For some workloads, this may advantageously reduce GPU idle time (and may eliminate idle time when there are non-dependent launches available to run). Thus, in these embodiments, a launch may be started as soon as any mGPU is available, and portions of a given launch may be executed sequentially on the same distributed slots. Note that by reference below Figure 30 The launch queue and streaming launch slot manager techniques discussed in detail can significantly increase the probability of finding a launch that can be started.

[0296] Figure 29 is an example according to some embodiments Figure 26 Illustration of an example "start when ready" allocation of a startup that allows an appropriate subset of a startup's parts to be allocated when all of its allocated slots are unavailable. Figure 29 It is shown that these techniques can provide better utilization than mGPU-level provisioning alone.

[0297] In this example, from T0 to T2, the allocation is Figure 28At time T3, the seven parts of launch 3 are distributed to mGPU1 to mGPU7. At time T4, part of launch 0 is completed, and launch 3 uses the free space in mGPU0.

[0298] At time T5, launch 4 obtains three distributed slots. At time T6, launch 5 has been allocated two slots and receives all the slots it requested at time T7. In this example, all launches are launched into the machine within 80 time units.

[0299] Note that a given launch may never receive a slot in its desired number of mGPUs. For example, a seven-part launch, two of which can execute on the same mGPU, will only utilize six mGPUs. In general, N parts of a given launch can execute on a given mGPU, and the launch can be allocated any number (1 to N) of its desired number of mGPUs.

[0300] In some embodiments, various performance tracking techniques can be used in the context of "start when ready" allocations. Without this feature, a given launch may have a well-defined execution starting point (on a distributed slot allocation), but the parts may complete at different times. In these embodiments, a launch may be considered complete when the last distributed slot is released. These start and end timestamps can be provided to software for profiling and performance analysis.

[0301] In a "start when ready" implementation, the graphics processor can generate multiple types of performance indicators. For example, the launch slot manager can track how long a given launch runs on a given distributed slot (this can include the execution of multiple launch parts). The launch slot manager can also aggregate these counts into a single runtime counter for the launch (these counts indicate the total resources in space and time utilized by the launch), and can also report per-distributed slot counts. The launch slot manager can provide this information in tracking slot registers, write this information to the launch completion buffer (discussed in detail below), or do both. This can advantageously provide developers with tools for profiling their code.

[0302] In some embodiments, for a context repository, the master control circuitry for a given launch is configured to store state information specifying which distributed slots were in use at the time the launch was context-stored. Similarly, the master control circuitry can load state information on context load and wait for those distributed slots to be allocated before continuing with the launch. Note that the "start when ready" technique can operate in conjunction with affinity scheduling and can utilize work stealing mechanisms to redistribute work assigned to lagging mGPUs as needed.

[0303] Launch Queue and Streaming Launch Slot Manager Technology

[0304] The various techniques discussed above can allow work from a boot to be quickly streamed to the distributed mGPUs in a fine-grained manner. Depending on the number of trace slots implemented (which can have circuit area implications), it can be difficult for software to configure enough trace slots in the boot slot manager 350 to fully utilize these features, especially without knowledge of hardware resource utilization.

[0305] Thus, in some embodiments, the launch queue is stored external to the launch slot manager (e.g., in a data structure in DRAM), and the launch slot manager is configured to select launches from the queue to fill tracking slots, logical slots, or both. As discussed in detail below, these embodiments can utilize various techniques to indicate and track dependencies, select from the launch queue to fill tracking slots, store data for completed work in a completion queue, handle interrupts for completed launches, support partial rendering operations, and the like.

[0306] Figure 30 is a block diagram illustrating an example initiation queue technique according to some embodiments. In the illustrated example, initiation slot manager 350 is configured to receive work from initiation queues 3010A through 3010N and output initiation completion data to completion queues 3040A through 3040M.

[0307] In some embodiments, the boot slot manager 350 is configured to read work from a number of launch queues located in DRAM (although other types of memory are also contemplated.) For example, in different embodiments or operating modes, various numbers of launch queues may be implemented, such as 16, 32, 64, 128, 256, 512, etc. Similarly, each launch queue may allow up to a threshold number of entries to enter the queue (such as 16, 32, 64, 128, 256, 512, 1024, etc.).

[0308] In some embodiments, each boot queue entry is a fixed-size structure in memory and includes various information discussed above in the context of trace slot entries. Thus, software can write boot configuration data into a queue entry instead of configuring the top slot. The boot slot manager 350 can then select a queue entry to fill a trace slot. A given queue entry can indicate dependencies on other boots and can also indicate dependencies on system events, as described below with reference to Figures 35B to 35C Discussed in detail.

[0309] In some embodiments, the launch queue is implemented as a circular buffer. Figure 32Provides more details about an example launch queue structure. A queue can allow software to program a larger number of launches than would be possible by programming the launches directly into the trace slots.

[0310] In some embodiments, a given launch queue is configured via a set of configuration registers that may implement the following fields: fields: valid, skip, stop, pause, address, queue size, launch count, launch position, wrap count, context identifier, priority, add launch, and one or more timestamp fields.

[0311] The skip field can indicate that all starts in the queue should be skipped, so that all starts are scheduled normally, but once unblocked to run in a trace slot, they should immediately complete without performing any work. The stall field can indicate that the launch queue is unable to schedule work into the trace slot. For example, stalling can occur based on software control (e.g., for context switching) or based on a dependency on another stall queue. The pause field can indicate that the launch is paused and can include separate flags for different pause reasons.

[0312] The address, queue size, start count, start position, and wrap count fields can specify the size of the queue structure in a ring buffer embodiment. The context identifier can specify the context to use when retrieving a start entry. The priority field indicates the priority of the queue. For example, the timestamp can indicate the timing of the last start selected from the queue and the earliest stopped start seen from the queue. In some embodiments, software reads and writes these registers using masked register reads and writes.

[0313] In some embodiments, each startup is tracked using a startup identifier that can be programmed by software. Each startup identifier can include a queue identifier, a wrap count provided by software (which increments when a startup queue wraps around or is reused), and a queue position. In other embodiments, the queue identifier and position can be implied by hardware based on the position of an entry in a given queue.

[0314] In the illustrated embodiment, the boot slot manager 350 includes queue selection logic 3020 (which in turn includes a queue remapping table 3025), a register copy engine 1910, and a top slot circuit 3030 (which may implement the status circuit 1920 discussed above with reference to FIG. 19).

[0315] Example launch dispatch from launch queue to tracking slot

[0316] In some embodiments, queue selection logic 3020 is configured to select launches based on various combinations of parameters such as, but not limited to, queue priority, queue deadlines, dependencies between launches, the master controller associated with the launch (e.g., compute, vertex, or pixel controller), available hardware resources, etc. In some embodiments, queue selection logic 3020 is triggered to select a launch if there is an empty tracking slot and there is at least one available launch queue that could potentially schedule into that top slot.

[0317] In some embodiments, each queue has a programmable queue priority. Generally speaking, queue selection logic 3020 can select work from higher priority queues first. In some embodiments, queue selection logic 3020 implements a fallback arbitration mechanism (e.g., round-robin) between queues with the same priority. Note that in some embodiments, the queue priority used to select launches from a queue is separate from the tracking slot priority used to select tracking slots to be allocated to logical slots. In some cases, when a higher priority queue is unable to provide a launch, queue selection logic 3020 is configured to select a launch from a relatively lower priority queue in order to increase the utilization of tracking slot resources.

[0318] In some embodiments, for queues with the same priority, queue selection logic 3020 is configured to implement a deadline-based schedule before falling back to the loop. In these embodiments, a given launch queue entry implements a programmable timestamp value. The timestamp value may include, for example, a wrap count and a queue position field, and may be compared using an integer subtraction circuit and a sign bit checker. In some embodiments, when a launch is programmed into a tracking slot, the last_into_KSM timestamp is updated to match the timestamp of the launch entry. The control circuitry may then compare the timestamp of a given launch with the last_into_KSM timestamp to identify which time is earlier (this indicates whether a given launch has been programmed into a tracking slot or is still in the launch queue).

[0319] In some embodiments, given two or more queues (e.g., two queues with the same priority) to which deadline-aware selection is to be applied, queue selection logic 3020 is configured to select the queue with the closer deadline. In these embodiments, each launch queue can maintain a deadline_timestamp indicating the target completion time of the launches in that queue. Each launch entry in the queue can also have a deadline_timestamp that can be used to update the queue deadline_timestamp when the launch slot manager 350 reads a given launch entry. The updated queue deadline_timestamp can then be used to arbitrate subsequent selections from the queue. In other embodiments, the launch entry deadlines can be listened to or pre-fetched to provide more synchronization for the queue timestamps.

[0320] Priority-based scheduling, deadline-based scheduling, or both may be subject to constraints corresponding to dependencies. For dependencies between launches, queue selection logic 3020 may redirect attempts to pull work from a blocked queue to pull work from the blocked queue's queue (e.g., using queue remapping table 3025 to remap selections based on priority). For queues blocked by system dependencies (also known as event flags), queue selection logic 3020 may remove these queues from selection consideration until they are unblocked.

[0321] In some embodiments, each queue entry includes a list of parent launch identifiers that specify the launches that the launch depends on. The queue entry may also include a valid parent mask that indicates whether some of the entries in the list should be ignored, which can facilitate a fixed queue entry size. Note that the parent launch identifier can be specified as a launch queue identifier and a timestamp to allow resolution of whether dependencies are blocking, as described below.

[0322] If the parent launch identifier is in a different queue, the launch slot manager 350 compares the parent launch's timestamp with the parent launch queue's last_into_KSM timestamp (and oldest_halted if the parent's queue is halted). If the parent launch timestamp is older than these timestamps, the dependency has been resolved. Otherwise, the launch slot manager 350 can update the appropriate entry in the queue remapping table 3025.

[0323] In some embodiments, entries in the queue remapping table 3025 include: a launch queue identifier field, a remapped launch queue identifier field, and a reset timestamp. In some embodiments, all requests pass through the queue remapping table 3025, which initially sets the launch queue identifier and the remapped launch queue identifier to the same launch. When a dependency is detected, the launch slot manager 350 updates the remapped launch queue identifier to specify the launch that the current queue depends on and sets the reset timestamp to the timestamp of the parent launch. When the current queue wins arbitration, the queue selection logic 3020 can select from the remapped queue instead. The remapping table entry remains valid until the last_into_KSM timestamp reaches the reset timestamp.

[0324] Note that these dependency techniques differ from those discussed in Section I because software sets up launch queues rather than programming launch-to-launch dependencies directly in the trace slots. This allows software to remain agnostic about which trace slots will be running launches or which launches will reside on the GPU at a given time. Instead, software can freely specify dependencies between launches across various launch queues.

[0325] In some embodiments, the selection by queue selection logic 3020 is also constrained by tracking slot allocations. For example, in some embodiments, software provides limits on one or more of the following: the maximum number of tracking slots allocated to each launch queue, the maximum number of tracking slots allocated to each master controller (e.g., for pixel, vertex, or compute work), the maximum number of default priority tracking slots per master controller (which, combined with the previous limits, can determine the number of tracking slots reserved for high priority launches), and the maximum number of tracking slots that are dependent on another tracking slot (e.g., this can allow software to ensure that there is always at least one tracking slot available for partial rendering).

[0326] When queue selection logic 3020 normally selects from the startup queue, but when the selection of the startup queue will violate a restriction in the restriction, the startup queue can be suspended. In some embodiments, the startup queue in a paused state (this paused state can be different from the valid state and the invalid state) can reduce the excessive acquisition of startup entries. The paused state can indicate that the queue cannot move forward due to temporary restrictions. The example of this type of restriction includes: waiting for the event flag to be cleared, waiting for unavailable tracking slots (for example, based on the number of dependencies required for the main controller type, priority level and startup, etc.), the startup at the head of the queue in some scenarios depends on an invalid or empty queue, etc. The pause flag can indicate the reason why the queue is suspended. When the startup slot manager 350 determines to clear all flags for the queue, the queue is restored. In some embodiments, software is also allowed to reset the pause flag.

[0327] Note that in some cases, reserving trace slots for different master controllers (rather than allowing one master controller to use all trace slots at a given time) can facilitate finding non-dependent startups for scheduling.

[0328] In some embodiments, the queue selection logic 3020 is configured to take into account the availability of downstream processing resources when selecting to start. The disclosed embodiments discussed above in relation to the limitation of trace slots for each master controller are an example of this feature. For example, if a master controller meets the limitation on the number of its trace slots, the corresponding downstream resources are likely busy (e.g., the geometry hardware of the vertex controller). In some embodiments, additional flag fields are assigned to queues to indicate what resources they are targeting. For example, a given flag may indicate whether a queue uses dedicated resources, such as: ray tracing acceleration hardware, texture sampling hardware, fragment processing hardware, matrix multiplication pipelines, and so on. The queue selection logic 3020 can then limit the number of trace slots allocated to queues that target a given resource (e.g., N trace slots to access ray tracing hardware, M trace slots to access texture sampling hardware, etc.).

[0329] In some embodiments, hardware can provide more detailed feedback about resource availability as input to queue selection logic 3020. While tracking slot usage can roughly correspond to resource usage, more granular information can allow queue selection logic 3020 to use more complex logic (based on which launch's resources are more likely to be available) to select launches from the queue, which can further improve performance. Similarly, queues can include fields with more detailed information about resource utilization (e.g., encoding information such as the specific resources being targeted, the amount of workload targeted at those resources, etc.).

[0330] In general, resource availability can be considered at various levels of provisioning, including selecting a queue for a tracking slot, assigning a tracking slot to a master slot, allocating work from a master slot to distributed slots, and so on.

[0331] In some embodiments, each boot entry includes a field for specifying up to N pointers to register values ​​that should be programmed at the start of the boot (e.g., by register copy engine 1910). Additionally, a boot phase value can allow software to specify which of the N pointers should be used. In some embodiments, all boots prefetch and program registers using data located in memory identified by the first register pointer, but optionally append additional register data located in memory identified by one of the other pointers based on the boot phase value. In some embodiments, this can provide hardware partial rendering support.

[0332] Additionally, the boot can optionally program a set of registers (directed by software in DRAM) at the boot completion time (e.g., in addition to generating data for the completion queue). This can facilitate tasks that wish to update certain data without waiting for software to process the completion queue. Examples of such tasks include, but are not limited to, providing memory cache discard notifications, updating event flags, and the like.

[0333] Example Completion Queue

[0334] When the startup is complete, the startup slot manager can evict it from its tracking slot and move its data to a completion queue. In some embodiments, the completion queue is stored in the same memory space as the startup queue. Generally speaking, the completion queue can allow software to delay processing of completed startups, rather than immediately processing each startup when it is completed.

[0335] In some embodiments, completion queue comprises queue with different priorities and queue for different complexity.Different completion queue priorities can correspond to the priority of the startup performed (for example, wherein one or more startup priority thresholds divide startup into two or more bins of different priorities to assign to the different sets of completion buffers).About queue complexity, for example, the startup of normally completing work can be expelled to normal completion queue, and the startup completed as the result of stop request can be expelled to complex completion queue.Generally speaking, in other embodiments, various number of completion queues can be included for different categories of startup.As a fine-grained example, processor can realize completion queue for each input startup queue.

[0336] In some embodiments, a complex completion queue supports the following states for a given launch: exited the queue by the launch slot manager, exited the queue by the host controller, and stored context. These states can indicate how far the launch progressed before being stopped. Normal completion queues can support completed and skipped states.

[0337] Based on the comparison of the priority of startup and the priority threshold value configurable by software, startup can be classified as high priority or normal priority. In some embodiments, at least four completion queues are supported: normal completion default priority, normal completion high priority, complex completion default priority and complex completion high priority. In other embodiments, various other categories of completion queues can be realized, for example, based on additional priority classes, other types of completion etc. Independent completion queue can allow software to prioritize completion tasks efficiently. In some embodiments, software can disable one or more types of completion queues.

[0338] Example Stop and Context Switching Techniques

[0339] In some embodiments, the boot slot manager 350 is configured to stop the boot queue, for example, to remove a given application from the system. This can include providing a halt signal to the boot from the boot queue and tracking (directly or via other boots) the slots that depend on the boot queue. In some embodiments, any subsequent attempts to schedule a startup entry that depends on the stopped startup or boot queue will fail and its boot queue will be marked as halted (and an interrupt may be initiated). In some embodiments, when the startup is halted, the control circuit updates the oldest_halted timestamp of the boot queue to the earliest of (A) the stopped startup and (B) the current oldest_halted timestamp of the boot queue.

[0340] In some embodiments, stopping can utilize the tracking slot dependency matrix (and the dependency information of the startups still in the queue) to recursively determine which startups are dependent on the stopped queue. The recursive method of the startup slot manager 350 can determine the set of dependency tracking slots that should be stopped. The startup slot manager 350 can also set the oldest_halted timestamp to the earliest timestamp of any startup that sends a stop signal in the process. On context loading, the software can empty the complex completion buffer to determine which startups are exited from the queue or the context is stored (and can also check the normal completion buffer to determine which startups have been successfully completed). The software can then program the stopped startup as a new startup entry in the startup queue, or can rewind the queue (for example, by manipulating the startup position and startup count value of the stopped queue). The software can also appropriately modify the last_into_KSM timestamp before unstopping the queue.

[0341] Example Part Rendering Techniques

[0342] As mentioned above, the launch slot manager 350 may include hardware support for partial rendering operations. This may occur, for example, when a geometry launch encounters an out-of-memory situation. Software may check whether the previous render from the paired fragment launch queue has completed (e.g., based on the completion buffer timestamp). If not, the previous render may be allowed to complete (parameter buffer pages may be released), and the geometry launch may be restarted.

[0343] If the previous render has completed, the fragment launch required for the partial render should be the launch entry at the head of the launch queue. Software can request to stop only the launches that depend on the current launch queue (this can be called stopping the child process). This stops all launches that depend on the low-memory geometry launch (and any such later-arriving launches), and the low-memory geometry launch can remain in its tracking slot. Software can then remove the fragment launch's dependency on the geometry launch completion and remove the stall on the fragment launch queue, allowing the first copy of the fragment launch to run. The launch slot manager can schedule the first copy of the fragment launch (partial render) into the tracking slot (without updating any of the queue's launch queue pointers, launch counts, or last_into_KSM), effectively submitting the copy without removing the fragment launch from the head of the queue so that it can be scheduled again in the future. The fragment queue can then be marked as stalled to prevent a second copy of the fragment launch from being scheduled. Software can then restart the geometry launch, add back the dependency, and remove the stall on the fragment launch.

[0344] Software may also update the register programming of the fragment launch so that the fragment launch knows to resume from an existing partially rendered image.As discussed above, the launch phase value may correspond to multiple different pointers to registers that should be used by the register copy engine for different phases of the launch.

[0345] Once the geometry launch is complete (typically, for example because a partial render freed up the page), the second copy of the fragment launch can be scheduled normally and complete the rendering.

[0346] Example Interrupt Control Techniques

[0347] In some embodiments, the boot slot manager 350 can be configured to generate interrupts in various boot completion contexts. For example, the boot slot manager 350 can trigger an interrupt in response to the following events: normal boot completion, complex boot completion, default priority boot completion, high priority boot completion, boot completion that was initially stuck in retain_slot mode, a boot stalled by hardware, the boot slot manager 350 being idle due to lack of work, the boot slot manager 350 being idle due to all available boots being blocked from starting, the boot slot manager 350 being no longer idle due to an external event, etc., or some combination thereof.

[0348] Software can configure the boot slot manager 350 to generate interrupts only for a subset of completion types, for example, by suppressing other completion types. In some embodiments, a single interrupt interface is shared by multiple interrupt types, and the boot slot manager 350 includes a real-time status register with a read-only flag indicating the cause of the interrupt (the real-time status register can also be reset independently). The disclosed interrupt control techniques can allow a given launched thread to control its own behavior when it ends, for example, by configuring the boot slot manager 350.

[0349] Detailed streaming KSM block diagram

[0350] Figure 31 31 is a block diagram illustrating a detailed example implementation of a launch queue, a launch socket manager, master control circuitry, and distributed mGPUs according to some embodiments. In the illustrated example, the system includes DRAM 3110 (which may be external to the GPU), a launch socket manager 350, a master geometry control 3132, a master pixel control 3134, a master compute control 3136, a workload distribution bus 3150, and mGPUs 320A through 320N.

[0351] In the illustrated embodiment, DRAM 3110 stores data structures for the start queue 3010, data for the start setup register 3115, and data for the start execute register 3120 (e.g., register data accessible by the register copy engine 1910). DRAM 3110 also stores data structures for four completion queues 3142, 3144, 3146, 3148 having different complexity and priority attributes.

[0352] In the illustrated example, the master pixel control circuit 3134 and the master compute control circuit 3136 have separate default priority sets 3135 and 3138 for logical slots, as well as high priority sets 3137 and 3139. The high priority slots can utilize dedicated logical slots in the mGPU. In the illustrated example, the master geometry control circuit 3132 is shown as including a single set 3133 of logical slots, but multiple sets of geometry logical slots can be implemented in various embodiments. Similarly, in other embodiments, various numbers of priority classes for the logical slots in the various master control circuits can be implemented; the disclosed examples are not intended to limit the scope of this disclosure.

[0353] In some embodiments, workload distribution bus 3150 is configured to route control data for workload distribution from logical slots to distributed slots. In some embodiments, bus 3150 is an extensible serial bus that is distinct from one or more interfaces or fabrics used to access actual graphics data for operations to be performed.

[0354] In the illustrated embodiment, the mGPU 320 includes corresponding distributed pixel control circuitry 3161, distributed geometry control circuitry 3162, distributed compute control circuitry 3163, and shader / texture / pixel circuitry 3164. The distributed control circuitry can use multiple distributed hardware slots of each mGPU to handle the execution of the assigned startup portion. The execution resources of circuitry 3164 can be replicated within a given mGPU (e.g., having instances corresponding to different distributed hardware slots). In addition, some resources of circuitry 3164 can be shared by multiple types of work (e.g., shader resources), while other resources can be dedicated to work from certain master controllers or certain operations (e.g., texture sampling). In some embodiments, each distributed controller controls a subset of the distributed hardware slots of a given mGPU for its type of work (e.g., N slots for geometry work, M slots for pixel work, and P slots for compute work). As described below with reference to Figure 33 As discussed in detail, certain types of jobs may have multiple different types of distributed hardware slots (e.g., segment processing slots and stitching slots for geometry jobs).

[0355] Example Ring Buffer Initiate Queue Implementation

[0356] Figure 3232 is a block diagram illustrating example startup queue information according to some embodiments. In the illustrated embodiment, the startup queue 3210 for startup queue ABC is implemented as a ring buffer with a base address, a wrap count, a startup position, and a startup count. In this example, the startup at position 2 to position 5 is effective for selection, the startup from position 0 and position 1 has been assigned to the tracking slot in the KSM scoreboard 3220, and the startup from position N has been completed and written to the entry in the completion buffer 3230. The KSM scoreboard 3220 can also include various dependency tracking information (not shown), such as a tracking slot dependency matrix generated based on startup dependencies and an event flag dependency mask based on system dependencies.

[0357] The illustrated example shows the use of timestamps ("t=" values) to track the movement of launches through the system, and a ring buffer structure for a given launch queue. These details are included for illustrative purposes and are not intended to limit the scope of this disclosure. In other embodiments, various queue data structures, timestamp encodings, dependency encodings, etc. may be utilized.

[0358] Logical slot technology for geometry activation

[0359] U.S. patent application Ser. No. 18 / 055,111, filed Nov. 14, 2022, and entitled “Initial Object Shader Run for Graphics Workload Distribution,” describes example techniques for performing a resolve run of a geometry launch to determine segment boundaries to divide the launch into multiple segments for parallel processing on distributed sockets. According to these techniques, a pre-resolved set of work can be executed on one or more distributed sockets before the work is allocated for actual parallel execution on multiple distributed sockets. U.S. patent application Ser. No. 17,805,607, filed Jun. 6, 2022, and entitled “Distributed Geometry Processing and Tracking Closed Pages,” describes techniques for splicing one or more data structures from different segments of a geometry launch.

[0360] In the logical slots technique for geometry work discussed below, a launch can be dynamically divided into an appropriate number of segments and processed on a subset of the available distributed slots (e.g., a segment can be launched as soon as a distributed slot becomes available). In these embodiments, segments from multiple launches can be processed in parallel (e.g., there may be segments from as many launches as there are mGPUs), and multiple launches can have their segments stitched together in parallel. Additionally, a unified master parameter manager circuit can allow geometry launches from different applications to be executed and multiple launches from the same application to be mixed in parallel. The disclosed techniques discussed below may be particularly advantageous in the context of applications (e.g., game titles) that send UI elements as a large number of small draw calls appended to previous rendering. These applications may have a large number of small launches mixed with larger launches, and the disclosed logical slots technique can significantly improve geometry hardware utilization.

[0361] Figure 33 is a block diagram illustrating example master controller circuitry implementing a logical launch slot technique for geometry launches, according to some embodiments. While geometry work can utilize various techniques described above similar to fragment or shader work, certain aspects of geometry processing can benefit from dedicated techniques. Specifically, the GPU can pre-parse geometry launches to generate segments for parallel processing in distributed hardware slots and can stitch together the processing results once the segments are completed. In some embodiments, both pre-parsing and segment execution utilize distributed slots. Additionally, in some embodiments, stitching can utilize dedicated distributed slots and can be controlled by the parameter manager master controller.

[0362] In the illustrated example, the processor includes a master geometry control 3132, a master parameter manager control 3310, a workload distribution bus 3150, and mGPUs 320A through 320N. In some embodiments, master controller responsibilities for geometry work are split between the master geometry control 3132 (which may handle the front end of the geometry pipeline) and the master parameter manager control 3310 (which may handle the back end of the geometry pipeline).

[0363] In the illustrated embodiment, master geometry control 3132 includes logical slots 3320, an arbiter 3325, a command stream pre-parser 3330, and a scheduler 3335. The boot slot manager 350 can assign tracking slots to the logical slots 3320. In some embodiments, the number of logical slots 3320 corresponds to at least the number of mGPUs + 1 (where additional logical slots are available for pre-parse work). In some embodiments, arbiter 3325 is configured to arbitrate between logical slots for access to mGPU resources. Master geometry control 3132 can be aware of the occupancy and status of distributed slots in the system and, therefore, determine which distributed slots are available to accept new work.

[0364] In the illustrated embodiment, command stream pre-parser 3330 is configured to control pre-parse operations, for example, by generating pre-parse tasks for one or more distributed slots. Command stream pre-parser 3330 can strip certain operations from the command stream to generate pre-parse work. Pre-parse can be run across all mGPUs (e.g., using distributed slots in each mGPU), a single mGPU, or an appropriate subset of mGPUs. Scheduler 3335 is configured to send pre-parse or actual segment work to distributed slots 3360A to 3360N in mGPU 320 via workload distribution bus 3150. Once pre-parsed, a geometry start can correspond to a stream of primitive segments followed by a termination marker.

[0365] In the illustrated example, each mGPU includes a section distributed slot 3360 and a tile distributed slot 3365. In some embodiments, geometry launches are allowed to occupy up to one tracking slot, up to one logical slot 3320 in the master geometry control 3132, up to one section distributed slot 3360 for launch execution in a given mGPU, up to one distributed slot 3360 for pre-resolve in a given mGPU, and up to one tile distributed slot 3365 in a given mGPU.

[0366] In some embodiments, the master geometry control 3132 is configured to stream launches (e.g., single-segment launches) and segments to the mGPU serially such that the first segment of the next launch is not launched to the mGPU until the last segment of the current launch is launched.

[0367] In the illustrated embodiment, the master parameter manager control 3310 is configured to control the stitching operation using a stitching distributed slot 3365. The master parameter manager control 3310 can track the completion of segments initiated by the master geometry control 3132 and stitch the data structures of the completed segments using stitching circuitry. In these embodiments, the logical start slot technology is implemented separately for geometry tasks and stitching tasks. In the illustrated example, the master parameter manager control 3310 includes a logical slot tracker 3340, an arbiter 3345, a stitching slot 3350, and a scheduler 3355.

[0368] The logical slot tracker 3340 can track the status of logical slots 3320 and can assert a ready signal when one or more segments of a given logical slot are ready for stitching. In some embodiments, the master parameter manager control 3310 and the master geometry control 3132 utilize direct interface connections to notify stitching completion, logical slot completion, etc. The arbiter 3345 can select from the logical slots 3320 that are ready for stitching and assign the stitching work to a stitching slot 3350. For distributed stitching, the scheduler 3355 is configured to schedule stitching work across one or more stitching distributed slots 3365 of the mGPU 320.

[0369] In some embodiments, stitching is performed at different granularities for different data structures. For example, in some embodiments, the stitching slot 3350 includes hardware configured to stitch one or more data structures, and the stitching distributed slot 3365 includes hardware configured to stitch one or more other data structures. As a specific example, in some embodiments, the stitching distributed slot 3365 is configured to stitch together the tile area array (RA) headers of the boot segment. In some embodiments, the stitching slot 3350 is configured to stitch together the layer identifier cache (LIC) of the boot segment (this stitching can be performed at the main controller because the stitching may affect the area array stitching) and the list of closed pages written by the boot segment (also known as the "A list").

[0370] A layer identifier cache allows the geometry processing stage to specify the layer of the final render target. For LIC stitching, distributed slots 3360 working on segments can update the LIC base address and initialize it at the start of the startup. Control 3310 can monitor the segment stitching mask and segment start pointer to determine when a threshold number of segments are available for stitching. For example, LIC stitching by a given stitching slot 3550 may involve adding the segment identifier multiplied by a fixed size to the base address and performing an atomic OR operation on the destination address associated with the layer cache base. LIC stitching can be performed before region array stitching because when layered rendering is enabled, only modified layers may need to be tiled together across region array segments, and unmodified layers can be skipped. Software can indicate a segment stitching mask that uses the segment identifier to determine the offset from the layer base address used to read LIC data for each segment. LIC stitching can be performed in the background relative to the processing of each segment, allowing segment completion and independent of LIC stitching startup.

[0371] In some embodiments, the segment distribution slot 3360 is configured to generate a list of closed pages that the hardware for a given slot has completed writing. This can be called an allocation list or (A list). In some embodiments, this is a linked list of closed pages (LLCP). The first part of an A list page can be a header that includes a link to the next A list page. The splicing slot 3350 can wait for a threshold number of segments to be ready for A list splicing, and then can update the pointer of the last page of the segment A list to point to the first A list page of the subsequent segment.

[0372] Note that in some embodiments, splicing is performed in splice slots 3350 for relatively small data structures (such as LICs and A-lists, as examples) that are spliced ​​in memory. Generally, splicing operations can be significantly impacted by memory latency (the time to retrieve data from memory and write the spliced ​​data back to memory). Therefore, as discussed in detail below, larger structures such as area arrays can be spliced ​​in parallel via splice distributed slots 3365 to take advantage of additional memory interfaces.

[0373] In some embodiments, the master stitcher circuitry is configured to coordinate the activities of the various stitching hardware (there may be an instance of a master stitcher for each tile slot 3350). In some embodiments, each tile slot 3350 is responsible for stitching a list of closed pages, stitching LICs, and scheduling zone array stitching of the mGPU with the active segment distributed slots 3360 for startup.

[0374] Region arrays can be used for tile-based deferred rendering (TBDR), where the tiling engine circuitry can divide the geometry into tiles and generate a control flow (e.g., with multiple linked lists per tile) to be processed by the segment processing circuitry. In some embodiments, a given tile distributed slot 3365 corresponds to a region array stitcher (RAS) circuit that reads the region arrays generated by the segment distributed slot 3360 from memory and stitches them back into a single region array for processing by the fragment circuitry. The RAS may include a prefetch stage with a read memory interface, a latency-hiding memory (e.g., RAM), a decode stage, a stitching stage, and a write memory interface. In some embodiments, the RAS stitcher is triggered by a threshold number of segments completing their LIC stitching. In some embodiments, the master parameter manager control 3310 provides information to the RAS stitcher, such as a stitching command and an indication of whether the stitching is for the start of a startup or appending to a previously stitched segment (and can distinguish between an initially started start stitching segment and a partially rendered stitched segment).

[0375] The master parameter manager control 3310 may also provide a RAS instance identifier, a count of RAS instances enabled for splicing, a count of segments to be spliced, a starting segment identifier, and a previous segment identifier (from a previous spliced ​​set). This may allow each RAS circuit to determine which sets of region headers to splice.

[0376] A given region array header may include a control flow pointer pointing to the first control block of the tile, an index indicating whether the region array is empty, and a shared field indicating whether the tile shares a control block with a neighboring tile. A splice link pointer may include a control flow pointer pointing to the next control block of the tile and a link end field indicating whether to follow the pointer or whether this is the last segment of the tile list. The splice algorithm may copy the control flow pointer of the region array header into the splice link pointer of the previous segment.

[0377] Thus, the RAS circuit can read the data of its set of region array headers into its latency-hidden memory, decode the data in memory, and route it to the correct bank. The stitching stage can then execute the stitching algorithm and write the results back to memory. In some embodiments, each RAS instance operates on a complete memory cache line, for example to avoid sharing cache lines across instances.

[0378] While the master geometry control 3132 can launch work serially, the master parameter manager control 3310 can see work from multiple logical slots 3320 in parallel. Thus, in some embodiments, the master parameter manager control 3310 is configured to process logical slots corresponding to up to the number of mGPUs in parallel (and these logical slots can be any combination of segmented and single-segment launches). When the number of logical slots 3320 to be tiled is greater than the number of available tile slots 3350, the arbiter 3345 can utilize round-robin arbitration.

[0379] In some embodiments, the control circuitry does not impose a strict age-based ordering on the completion of work. Thus, single-segment starts (also referred to as non-segmented starts) can complete out of order relative to segmented starts, and segmented starts that complete faster can complete out of order relative to earlier, slower starts.

[0380] In other embodiments ( Figure 33 In some embodiments, the processor may include a single stitcher per mGPU, and the master parameter manager control 3310 may simply schedule stitching work across all mGPUs that have been assigned distributed slots for startup.

[0381] In some embodiments, the processor supports various software overrides for geometry work, such as forcing a launch to schedule only on certain groups of mGPUs, scheduling overrides (e.g., to change the scheduling type between loops and other types such as find-first, targeting the mGPU with the lowest identifier available), a maximum number of mGPUs for a given launch, and so on.

[0382] In some embodiments, the processor is configured to execute segments from different geometry launches in parallel only if the launches do not share parameter buffers.

[0383] Figure 34 is a diagram illustrating an example of a geometry-based execution using the disclosed logical socket technology for execution and tiling according to some embodiments. The mGPU may correspond to the Figure 26 Those mGPUs in the example.

[0384] In the illustrated example, at time T0, master geometry control 3132 has initiated pre-resolving tasks for launch 0 across the first distributed slot on all mGPUs. Once segments are identified for launch 0 (three segments in this example), master geometry control 3132 utilizes round-robin scheduling starting at group 0, mGPU 0. Thus, at time T0, master geometry control 3132 has initiated pre-resolving of segments for launch 1 across all mGPUs, as well as launch 0 across mGPUs 0, mGPU 1, and mGPU 2. In this example, launch 1 includes at least eight segments. At time T2, master geometry control 3132 has launched five of these segments into available mGPUs 3 through 7, while the remaining segments await available distributed slots.

[0385] At time T3, the master geometry control 3132 has launched segments from launch 1 across all mGPUs in the system (segments from launch 1 are submitted when segments from launch 0 complete execution and distributed slots become available). In this example, each launch is allowed to use at most a single distributed slot per mGPU. In other embodiments, a launch may be allowed to use multiple distributed slots within a given mGPU. At time T4, the master geometry control 3132 has launched the pre-resolved tasks for launch 2 across all mGPUs, and the three segments from launch 1 are still executing.

[0386] Example Event Flags Technique

[0387] Figure 35A is a block diagram illustrating an example dependency graph according to some embodiments. In the illustrated example, renderer C depends on renderer A to complete and buffer B to be ready. In the illustrated example, external circuit 3510 should write to buffer B before renderer C uses buffer B. External circuit 3510 can be external to the graphics processor (but the graphics processor can also utilize the disclosed event flag technology). For example, external circuit 3510 can be a machine learning accelerator or image processing circuit that operates on image data generated by the GPU, provides image data to the GPU for further processing, or performs both operations. External circuit 3510 can communicate with the GPU via a shared fabric, a dedicated communication line, a shared memory space, etc.

[0388] In some embodiments, the boot slot manager 350 supports a set of event flags (eg, in the GPU register space) that indicate dependencies on system events. Software may utilize various encodings to indicate a boot's dependency on one or more event flags.

[0389] Figure 35BFigure 1 shows event flag dependency masks stored for a given trace slot according to some embodiments. These dependency masks can be configured in a launch queue and retrieved into the trace slot upon launch selection. In the illustrated example, a given dependency mask indicates a set of one or more event flags that the corresponding launch depends on.

[0390] Figure 35C An example set of N+1 event flag fields according to some embodiments is shown. A given field can be a counter, and when the counter reaches a threshold (e.g., zero), a given dependency can be cleared. In other embodiments, a given flag field is a single bit that indicates whether an event has occurred.

[0391] In the illustrated example, renderer A is assigned to trace slot 0, and renderer C is assigned to trace slot 2. Renderer C depends on renderer A (as shown in the trace slot dependency matrix, which can be maintained in the top slot circuit 3030), and also depends on event flag field 0, as indicated by dependency mask 0x1. In this example, software has assigned event flag field 0 so that external circuitry 3510 can indicate when buffer C is ready for renderer C. In this example, renderer C is not eligible for assignment to a logical slot until both dependencies are satisfied (all trace slot dependency matrix entries in the row for renderer C are cleared, and event flag 0 is cleared).

[0392] In some embodiments, event flags can be used to stream image work across different circuits of a processor or system on a chip.The following discussion explains example macroblock-level completion tracking of image data frames and example techniques for utilizing event flags across IP blocks.

[0393] Figure 36 is a block diagram illustrating an example organization of a frame portion according to some embodiments. In the illustrated embodiment, the frame is divided into a plurality of macrotiles, each macrotile corresponding to a plurality of tiles. For example, the number of macrotiles can be fixed and can be an integer multiple of the tile size for tile-based deferred rendering. As shown, there may be overhead (between the frame and the dashed line) if the frame does not fit exactly into the number of macrotiles.

[0394] In some embodiments, certain circuits (such as fragment generator circuits) may process tiles of a macrotile before selecting the next macrotile to be scheduled. In general, it may be useful to break a frame into smaller chunks of work such as macrotiles. For example, this may allow demand-based rendering to generate only the required area of ​​a given render at a time. As another example, slip streaming may process corresponding areas in multiple renders (e.g., sequential renders) to improve cache locality. As yet another example, this may allow large launches (e.g., full-frame renders) to be broken down into smaller launches (e.g., macrotile renders), which may facilitate GPU resource utilization by generating small launches for selection to run on idle shader cores (this may be particularly useful when using the mGPU granularity scheduling and "start when ready" techniques discussed above).

[0395] Using event flags, producer or consumer circuit IP can map available data blocks (e.g., macroblocks) to the format they wish to produce or consume. As an initial example, an event flag might correspond to a single macroblock being ready. As other examples, event flags might be mapped to various granularities.

[0396] For example, Figure 37 An example row and column buffer mapping is shown. In this example, the shaded work blocks are complete, while the non-shaded blocks are incomplete. As shown, the circuitry may set or monitor event flags corresponding to whether the row buffer or column buffer of the block has been completed. This may allow consumers to use event flag signaling to start consuming data as soon as the row / column becomes available. This may be particularly useful when streaming image data between different image processing blocks in a system on chip. Note that the data itself can be accessed in a shared memory space, but the disclosed technology may allow consumers to start accessing data before the entire image is ready. For example, a graphics processor, image processing circuitry (e.g., a memory scaling / rotation unit), a machine learning accelerator, etc. may excel at different image processing tasks and therefore may share the work of processing a given image for display. In this example, pipelining subsets of the work can significantly improve performance, reduce power consumption, or achieve both.

[0397] Figure 38A and Figure 38B 38 is a block diagram illustrating example inter-block notification routing techniques for event flags according to some embodiments. In each of these figures, the circuit blocks include the graphics unit 150, the machine learning accelerator 3810, the central processing unit complex 3820 (which may include one or more CPUs), the memory scaling / rotation circuit 3830, and the image processing circuit 3840. Note that these bits are included for illustrative purposes and are not intended to limit the scope of the present disclosure. The disclosed techniques can be used with a subset of the illustrated blocks, additional blocks, etc.

[0398] exist Figure 38A In the example of , direct block-to-block communication is implemented for inter-IP event flags (e.g., using interrupt lines and real-time status registers to indicate DRAM addresses). Figure 38B In the example of , arbiter 3860 acts as a centralized directory of inter-block event flags. For example, this can provide scalable line counts and a central location for real-time status registers.

[0399] In various embodiments, the disclosed event flagging techniques can provide software with a flexible mechanism to indicate launch dependencies on system events (as well as on other launches), which can facilitate the various demand-based processing techniques discussed above.

[0400] Pipelining for dependency startup

[0401] Pipelining launches can be used to improve performance. Traditionally, dependent launches may not be pipelined (instead, the parent launch must complete before the child launch is launched) to ensure that dependencies are met. However, in the disclosed embodiments discussed below, the spin-up portion of a dependent launch can be pipelined after the spin-down portion of the parent launch. For example, a dependent kick release (DKR) signal can be routed from the distributed shader circuitry to the master control circuitry and the launch slot manager 350, which can conditionally unblock the dependent launch. In some embodiments, the graphics driver can provide finer-grained dependency information to allow a determination that it is safe to release a dependent launch, for example, up to one or more stall points. Stall points may be referred to herein as shader "launch gates."

[0402] Note that a given launch may include an execution state load (ESL) program SIMD group that may be allowed to execute in certain circumstances without violating dependencies (e.g., dependencies associated with a master shader). In some embodiments, one launch gate may correspond to a point at which the first ESL SIMD group has been allocated resources but has not yet executed instructions. This may be referred to as an ESL gate or early gate. Another launch gate may correspond to a point at which the first master shader (non-ESL) SIMD group has been allocated resources but has not yet executed instructions. This may be referred to as a working gate or delay gate.

[0403] In some embodiments, when there are no known dependencies between the first launch and the parent launch's control flow or indirection, a dependent launch is allowed to proceed to the first launch gate (although there may be dependencies on ESL data). In some embodiments, when there are no known dependencies between the dependent launch and the parent launch's control flow, indirection, or ESL data, a dependent launch is allowed to proceed to the second launch gate. Generally speaking, software can specify which one or more of the N launch gates should block the launch.

[0404] Note that these launch gates are included for purposes of explanation, but in other embodiments, other launch gates may be implemented at various resource allocation or execution stages. Additionally, the listed dependencies are merely examples; different gates may correspond to various combinations of launch dependencies depending on the location of the gates within the dependent launch.

[0405] In some embodiments, launch gates are detected and implemented by the hardware based on the dependency type of a given launch. The launch slot manager 350 can use this dependency information in conjunction with status information from the distributed shader circuits to pipeline dependent launches. In other embodiments, software can provide dependency information that indicates one of multiple launch gates supported by the hardware as the gate for a particular dependent launch.

[0406] In some embodiments, there are multiple types of DKR signals. For example, an early DKR can signal a shader that all SIMD groups have been started for startup (e.g., from a token parser circuit to a tile and thread group manager circuit). In some embodiments, the token parser circuit is configured to receive work tokens from multiple master controllers, form SIMD groups, and interact with the allocator circuit to allocate pages for private memory. In some embodiments, the tile and thread group manager circuit is configured to coordinate the execution of SIMD groups within a tile (e.g., for pixel work) or a thread group (e.g., for compute work). For example, this may include implementing various types of synchronization. The SIMD group scheduler circuit can then select from the ready SIMD groups to schedule them for execution. These specific circuit examples are included for illustrative purposes, but in other embodiments, the DKR signal can be initiated at various appropriate execution points.

[0407] The delayed DKR can signal that the shader has finished executing all SIMD groups for launch. In some embodiments, the early DKR is not used for launches that depend on compute or geometry launches. Note that these release signals are included for explanation purposes, but in other embodiments, other release points can be implemented at various resource launch or execution stages (e.g., a signal for a remaining threshold number of SIMD groups that have not yet been launched or executed, etc.). The end of launch signal can indicate that all barriers, flushes, etc. of the launch have completed.

[0408] In various embodiments, the disclosed dependency launch pipelining techniques can reduce the cycle costs corresponding to the ramp-up of the first shader execution for a launch and the ramp-down after the last shader instruction is executed (e.g., associated with fixed-function hardware and memory latency). In some embodiments, this can advantageously improve overall performance. Additionally, faster execution of dependency launches can free up resources for various disclosed launch streaming and scheduling techniques, which can further improve utilization.

[0409] Figure 39 is a diagram illustrating an example pipelined execution of dependent launches according to some embodiments. In the illustrated example, fragment launch B depends on fragment launch A. In this example, fragment launch A triggers early and delayed dependent launch release signals when certain execution points are reached.

[0410] In the illustrated example, the rotational acceleration portion of fragment launch B corresponds to the portion between the start of the launch and the delayed launch gate. This can include operations such as region array fetches, control flow and primitive block fetches, rasterization, and ESL. As shown, the launch slot manager can launch dependent launch B after an early DKR signal from fragment launch A. In this example, a portion of the rotational acceleration portion of fragment launch B is hidden, reducing the total processing time of launches A and B compared to waiting to launch launch B until launch A completes.

[0411] Note that in some embodiments, if segment launch B reaches the launch gate before the corresponding release from segment launch A, it will stall.

[0412] Figure 40 is a diagram illustrating an example pipelined execution of dependent launches from different host controllers according to some embodiments. In this example, segment launch B depends on compute launch A.

[0413] As shown, the launch slot manager 350 is configured to wait to initiate a dependent launch B until a delayed DKR signal is received from a compute launch. Note that this can allow pipelining to operate in a preempted context. In this example, the compute launch can be preempted between the early DKR and delayed DKR signals. Note that at some point in fragment launch B, it is too late to do a context switch without data dependency ordering issues. Therefore, in some embodiments, launches that are dependent on certain other types of launches may not be initiated until a delayed DKR signal is received (unlike the delay DKR signal). Figure 39 Note that in some embodiments, computation initiation can begin based on an early DKR signal from a fragment initiation, by contrast.

[0414] Figure 41 An example table illustrating four dependency states for launch B on another launch A, according to some embodiments, is shown. If no dependency exists, the launch slot manager 350 is free to schedule launch B independently of launch A. If a soft early dependency exists, the launch slot manager 350 is allowed to schedule launch B when it receives an early DKR signal based on launch A. If a soft late dependency exists, the launch slot manager 350 is allowed to schedule launch B when it receives a late DKR signal based on launch A. If a hard dependency exists, the launch slot manager 350 is allowed to schedule launch B when it receives a launch end signal from launch A. In some embodiments, soft early and soft late dependencies are used for launches that have only data dependencies visible from the shader. In contrast, other dependencies, such as command stream, indirect, index, depth / stencil / parameter buffer dependencies, etc., can be hard dependencies.

[0415] Figure 42 A more detailed example dependency scenario according to some embodiments is illustrated. In the illustrated example, fragment launch D has a soft dependency (soft early) on fragment launch A, a soft dependency (soft late) on compute launch B, and a hard dependency on fragment launch C.

[0416] In the illustrated example, signal 1 is an early DKR signal processed based on fragment launch A. This satisfies the soft dependency on launch A. Signal 2 is a delayed DKR signal from launch B, satisfying the soft dependency on launch B. Signal 3 is an end-of-core (EOK) signal from launch C, satisfying the hard dependency.

[0417] At 4, the boot slot manager 350 begins boot D due to signals 1 through 3. After a spin-up interval, boot D reaches the boot gate and stalls.

[0418] In the illustrated example, signal 5 is the EOK signal for start B, and signal 6 is the EOK signal for start A. These two signals allow the start door to unlock and allow start D to continue. Note that if the EOK from segment start A is received before the start door, segment D could continue immediately after the spin-up without the stall shown in this example.

[0419] Note that multiple launches can depend on the same parent launch (e.g., a fragment launch and a compute launch can depend on another fragment launch). Therefore, multiple launches can be allowed to spin up based on one DKR from the parent launch.

[0420] In some embodiments, the boot slot manager 350 implements a soft dependency field. In some embodiments, when programming the execution phase register for boot, the software includes a gate for dependency boot as discussed above.

[0421] In some embodiments, to generate the DKR signal, the distributed control circuitry may aggregate early DKR signals from the shaders and send them to the master control circuitry for further aggregation. When the master control circuitry detects receipt of early DKR signals for all assigned shaders for a launch, the master control circuitry may generate an early DKR signal to the launch slot manager 350. When all work assignments, tiles, or geometries for a given launch are known to be complete, the master control circuitry may generate a delayed DKR signal (and may also send launch end requests, such as cache flush invalidations, memory allocator cleanups, etc., at this time). A shader may receive a token representing the last work token in a launch, which may allow the shader to generate an early DKR signal.

[0422] Additional example methods

[0423] Figure 43 is a flow chart illustrating an example method for queue-based launch slot management according to some embodiments. Figure 43 The methods shown can be used in conjunction with any of the computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown can be performed concurrently in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.

[0424] At 4310, in the illustrated embodiment, a queue access circuit (e.g., queue access circuit 4820) accesses a data structure in a memory (e.g., in DRAM 3110) that specifies multiple queues (e.g., queue 4810), where the corresponding queues queue control information for multiple sets of graphics work.

[0425] At 4320, in the illustrated embodiment, queue selection circuitry (e.g., queue selection logic 3020) selects a set of graphics work from a data structure based on one or more selection parameters and stores control information for the selected set of graphics work in a tracking slot (e.g., top slot circuitry 3030).

[0426] In some embodiments, the selection parameter comprises a work category parameter included in a data structure for a given set of graphics jobs, wherein the work category parameter indicates whether the given set of graphics jobs is compute jobs, fragment jobs, or geometry jobs. In some embodiments, the selection parameter comprises a resource availability parameter provided by graphics processor circuitry, wherein the resource availability parameter indicates the availability of different graphics processor hardware resources. In some embodiments, the selection parameter comprises a priority parameter included in a data structure for a given queue, wherein different queues have different priority parameter values. In some embodiments, the selection parameter comprises a deadline parameter included in a data structure for a given set of graphics jobs.

[0427] In some embodiments, the selection parameters include dependency parameters included in a data structure for a given set of graphics work, wherein the dependency parameters indicate one or more other sets of graphics work on which the given set of graphics work depends. In some embodiments, the dependency parameters for the given set of graphics work include a set of parent identifiers and a valid parent mask indicating which parent identifiers are valid, and the dependency parameters are specified in a manner independent of the tracking slots to which the set of graphics work is assigned. In some embodiments, the one or more dependency parameters indicate at least one parent process that is in a different queue than the given set of graphics work. In some embodiments, the queue selection circuitry is configured to suspend selection from the first queue and select the set of graphics work from the second queue in response to the dependency parameters indicating a dependency of the set of work in the first queue on the set of work in the second queue.

[0428] In some embodiments, the selection parameters include an event flag parameter included in a data structure for a given set of graphics work, the event flag parameter indicating one or more software programmable event flags on which the set of graphics work depends, and the queue selection circuitry is configured to suspend selection from the first queue based on the set of graphics work waiting for the one or more event flags. In some embodiments, at least one of the event flags is encoded as a counter value. In some embodiments, a set of graphics work that is independent of each other, is included in at least two different queues, and targets a first hardware resource, and the initiation of a set of graphics work in the set sets the event flag to acquire the hardware resource and prevents initiation of other sets of graphics work until the event flag is cleared.

[0429] At 4330, in the illustrated embodiment, distribution circuitry (eg, master control circuitry 210, boot socket manager 305, or both) assigns portions of the respective sets of graphics work from the tracked sockets to graphics processor circuitry for execution.

[0430] In some embodiments, another processor of the device is configured to perform image processing on graphics frame data generated by the graphics processor circuit, and the device is configured to execute program instructions to utilize an event flag parameter to indicate when a subset of the graphics data frame has completed processing by the another processor and is ready for the graphics processor circuit. In some embodiments, the device is configured to control at least one event based on tasks performed by one or more other circuit components external to the graphics processing circuit of the device.

[0431] In some embodiments, the graphics processor circuitry includes control circuitry configured to control a portion of a rendering process, including: stopping all sets of graphics work that depend on a geometry set of work; scheduling a first copy of a fragment set of work that operates on data from the geometry set of work; restarting the geometry set of work after executing the first copy of the fragment set of work; and configuring a second copy of the fragment set of work to resume from a partially rendered image generated by the first copy of the fragment set of work. In some embodiments, the graphics processor circuitry includes control circuitry configured to control which sets of graphics work trigger an interrupt upon completion.

[0432] In some embodiments, the graphics processor circuitry includes control circuitry configured to write result data for a set of completed graphics jobs to a completion queue structure in memory. In some embodiments, the control circuitry is configured to write the result data to a plurality of different completion queues having different priorities. In some embodiments, the control circuitry is configured to write the result data to different completion queues based on whether a given set of graphics jobs completed normally. In some embodiments, the control circuitry is further configured to program a set of registers indicated by the set of graphics jobs in response to completion of the set of graphics jobs. In some embodiments, the graphics processor circuitry is configured to execute firmware to provide one or more of the sets of graphics jobs for queues in the data structure.

[0433] In some embodiments, the memory circuit stores a data structure, and the processor executes program instructions to add control information to a queue of the data structure. In some embodiments, the plurality of single instruction multiple data pipelines execute instructions, and the fixed function circuitry controls the single instruction multiple data pipelines to perform operations for at least one of the following types of programs: a graphics shader program and a machine learning program.

[0434] Figure 44 is a flow chart illustrating an example method for allocating graphics work from logical slots to distributed hardware slots according to some embodiments. Figure 44 The methods shown can be used in conjunction with any of the computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown can be performed concurrently in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.

[0435] At 4410, in the illustrated embodiment, control circuitry (eg, master control circuitry 210) determines that a full number of distributed hardware slots to be utilized by a set of graphics jobs are unavailable.

[0436] At 4410, the control circuitry, based on the control information for the set of graphics work, assigns corresponding portions of the set of graphics work from the logical slot (e.g., logical slot 215) to the distributed hardware slots (e.g., distributed hardware slot 230). In the illustrated embodiment, this includes assigning an appropriate subset of the portions of the set of graphics work to the available distributed hardware slots in response to the determination of element 4410.

[0437] In some embodiments, the plurality of logical slots includes one or more slots dedicated to one or more types of graphics work. In some embodiments, the control circuitry assigns a first set of geometry work to N-1 graphics processor subunits of the N graphics processor subunits of the device, and assigns a second set of work to a single remaining graphics processor subunit of the device.

[0438] In some embodiments, the controller supports allocating any determined integer number of graphics processor subunits in the range from one processor subunit to the total number of processor subunits included in the device. The range may include one or more integer numbers of graphics processor subunits that are not powers of 2.

[0439] In some embodiments, the set of graphics work is a fragment processing set of graphics work, and the control circuitry is configured to determine the number of graphics processor subunits based on the number of tiles included in the set of graphics work or based on the number of pixels included in the set of graphics work. In some embodiments, the set of graphics work is a compute set of graphics work, and the control circuitry is configured to determine the number of graphics processor subunits based on the number of workgroups, work items, or both included in the set of graphics work. In some embodiments, the set of graphics work is a geometry set of graphics work, and the control circuitry is configured to determine the number of graphics processor subunits based on one or more of the following parameters: the number of primitives included in the set of graphics work, the number of vertices included in the set of graphics work, and the determined complexity of the set of graphics work.

[0440] In some embodiments, for a portion of the set of graphics work to be assigned, the control circuitry attempts to obtain a distributed hardware slot in a graphics processor subunit that does not currently have a distributed hardware slot assigned to the logical socket, and in response to failure of the attempt, sends the portion of the set of graphics work to a distributed hardware slot that the logical socket already has.

[0441] In some embodiments, the control circuitry executes the set of graphics work using a fewer number of distributed hardware slots than the number of hardware slots determined for the set of graphics work, and in some embodiments, may execute multiple portions of the set of graphics work using the same distributed hardware slots. In some embodiments, the control circuitry implements a non-overlapping mode in which different portions of the same set of graphics work cannot be assigned to the same distributed hardware slot.

[0442] In some embodiments, the control circuitry tracks runtimes of portions of the set of graphics work in corresponding distributed hardware slots and stores the tracked runtimes using one or more of the following techniques: 1) aggregating the tracked runtimes into a single runtime count, or 2) writing the tracked runtimes to a completion buffer. In some embodiments, the control circuitry is further configured to, in response to a context store, store information indicating which distributed hardware slots are used by the logical slot, and, in response to a context load, wait for allocation of the indicated distributed hardware slots before continuing execution of the set of graphics work.

[0443] Figure 45 is a flow chart illustrating an example method for resolving a collection of geometry work and assigning it to distributed hardware slots according to some embodiments. Figure 45The methods shown can be used in conjunction with any of the computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown can be performed concurrently in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.

[0444] At 4510, in the illustrated embodiment, the control circuitry assigns a resolved version of the set of geometry work to a distributed hardware slot (e.g., distributed hardware slot 230) of one or more graphics processor subunits (e.g., subunit 220).

[0445] In some embodiments, the control circuitry is configured to assign the parsing work of the set of geometry work to at most one distributed hardware slot of a given graphics processor subunit and assign the segment execution work to at most one distributed hardware slot of a given graphics processor subunit. In some embodiments, the control circuitry assigns the parsed version to all graphics processor subunits in the set of graphics processor subunits and serially assigns the segments of the determined number of segments to the available graphics processor subunits in the set.

[0446] At 4520, in the illustrated embodiment, the control circuitry determines a number of segments for the set of geometry work based on execution of the parsed version.

[0447] In some embodiments, the control circuitry dynamically changes the number of distributed hardware slots assigned to the set of geometry work during execution of the set of geometry work.

[0448] At 4530, in the illustrated embodiment, the control circuitry assigns the determined segments to distributed hardware slots of corresponding graphics processor subunits for execution.

[0449] In some embodiments, the control circuitry allows different sets of geometry work to execute in parallel on different distributed hardware sockets only if the different sets of geometry work share a parameter buffer.

[0450] At 4540, in the illustrated embodiment, a stitching circuit (e.g., stitching circuit 4930) stitches the results of the segments processed by the assigned distributed hardware slots.

[0451] In some embodiments, the stitching circuit is configured to stitch results across multiple graphics processor subunits of segments of a set assigned for geometry work. In some embodiments, the stitching circuit includes a stitching control circuit configured to assign stitching work for one or more first data structure categories to hardware stitching slots in the main control circuit, and to assign stitching work for one or more second data structure categories to distributed hardware stitching slots in the corresponding graphics processor subunits, wherein the distributed hardware stitching slots include corresponding memory interfaces to access memory storing the data structures to be stitched. In some embodiments, the one or more first data structure categories utilize less memory space than the one or more second data structure categories. In some embodiments, the one or more first data structure categories include a layer identifier cache and a closed page list, and the one or more second data structure categories include a tile region array header.

[0452] Figure 46 is a flow chart illustrating an example method for scheduling graphics work based on dependencies between different sets of graphics work, according to some embodiments. Figure 46 The methods shown can be used in conjunction with any of the computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown can be performed concurrently in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.

[0453] At 4610, in an illustrated embodiment, control circuitry (e.g., boot socket manager 350) receives different sets of graphics work and schedules the sets of graphics work for execution on distributed hardware resources, the sets of graphics work including a first set of work that is dependent on a second set of work.

[0454] At 4620, in the illustrated embodiment, the control circuitry initiates processing of the first set of work in response to a release signal from the second set of work indicating that the second set of work has reached the first processing point.

[0455] At 4630, in the illustrated embodiment, the control circuitry stalls processing of the first set of work in response to reaching a gate point in the first set of work.

[0456] At 4640, in the illustrated embodiment, the control circuitry resumes processing of the first set of work in response to the end signal of the second set of work.

[0457] Figure 47 is a flow chart illustrating an example method for specifying soft and hard dependencies between different sets of graphics work, according to some embodiments. Figure 47The methods shown can be used in conjunction with any of the computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown can be performed concurrently in a different order than shown, or can be omitted. Additional method elements can also be performed as needed.

[0458] In some embodiments, Figure 47 The method is performed by a compiler. In other embodiments, Figure 47 The method may be performed in real time, for example, by a graphics driver or firmware based on detecting dependencies in instructions to be executed. In other embodiments, for example, Figure 47 The functionality can be split between the compiler and the graphics driver.

[0459] At 4710 , in the illustrated embodiment, the computing device determines a first dependency of the first set of graphics work on the second set of graphics work.

[0460] At 4720, in the illustrated embodiment, the device determines a second dependency of the third set of graphics work on the fourth set of graphics work.

[0461] At 4730, in the illustrated embodiment, the device specifies soft dependencies for the first set of graphics work based on determining that the initial portion of the first set of graphics work is not dependent on the second set of graphics work.

[0462] At 4740, in the illustrated embodiment, the device specifies hard dependencies for the third set of graphics work.

[0463] In some embodiments, the device generates release and gate information associated with hard dependencies and soft dependencies.

[0464] Figure 48 is a block diagram illustrating example launch queue techniques according to some embodiments. In the illustrated embodiment, the computing system includes: queue access circuitry 4820 configured to access a data structure in memory specifying a plurality of queues 4810A through 4810N; top slot circuitry 3030 implementing entries for a plurality of tracking slots for a graphics processor circuit (and which may also be referred to as a "tracking slot circuit"); queue selection logic 3020 configured to select a set of graphics work from the data structure based on one or more selection parameters and store control information for the selected set of graphics work in a tracking slot of the tracking slot circuit; and allocation circuitry 4840 configured to assign portions of a corresponding set of graphics work from the tracking slots to the graphics processor circuit for execution. Note that similarly numbered elements may be referenced above. Figure 30The various disclosed circuits may be included in a boot slot manager, a host controller, other circuit elements, or some combination thereof.

[0465] Figure 49 is a block diagram illustrating example graphics control circuitry configured to map logical slots to distributed hardware slots, according to some embodiments. In the illustrated example, a computing system includes: graphics control circuitry 4910 configured to implement multiple logical slots; a set of graphics processor subunits 220A through 220N, each of which implements multiple distributed hardware slots 230; and stitching circuitry 4930 configured to stitch the results of segments processed by the assigned distributed hardware slots. Control circuitry 4910 may also assign a parsed version of a set of geometry work to the distributed hardware slots of one or more of the graphics processor subunits, determine a number of segments for the set of geometry work based on execution of the parsed version, and assign the determined segments to the distributed hardware slots of the corresponding graphics processor subunits for execution. Note that, in different embodiments, circuitry 4930 may be included in graphics control circuitry 4910, one or more subunits 220, or both.

[0466] The following numbered clauses list various non-limiting embodiments disclosed herein:

[0467] Set A

[0468] A1. A device comprising:

[0469] A graphics processor circuit, the graphics processor circuit comprising:

[0470] a trace slot circuit implementing entries for a plurality of trace slots of the graphics processor circuit;

[0471] queue access circuitry configured to access a data structure in the memory specifying a plurality of queues, wherein a corresponding queue queues control information for a plurality of sets of graphics jobs;

[0472] a queue selection circuit configured to select a set of graphics jobs from the data structure based on one or more selection parameters and to store control information for the selected set of graphics jobs in a tracking slot of the tracking slot circuit; and

[0473] Distribution circuitry is configured to assign portions of respective sets of graphics work from the tracking slots to graphics processor circuitry for execution.

[0474] A2. An apparatus according to any of the preceding clauses within set A, wherein the selection parameter includes a work category parameter included in the data structure for a given set of graphics work, wherein the work category parameter indicates whether the given set of graphics work is compute work, fragment work, or geometry work.

[0475] A3. An apparatus as recited in any preceding clause of Set A, wherein the selection parameter comprises a resource availability parameter provided by the graphics processor circuitry, wherein the resource availability parameter indicates availability of different graphics processor hardware resources.

[0476] A4. An apparatus according to any preceding clause within Set A, wherein the selection parameter comprises a priority parameter included in the data structure for a given queue, wherein different queues have different priority parameter values.

[0477] A5. An apparatus according to any preceding clause within Set A, wherein the selection parameter comprises a deadline parameter included in a data structure for a given set of graphics jobs.

[0478] A6. An apparatus according to any preceding clause within set A, wherein the selection parameter comprises a dependency parameter included in the data structure for a given set of graphics work, wherein the dependency parameter indicates one or more other sets of graphics work on which the given set of graphics work depends.

[0479] A7. An apparatus according to clause A6, wherein:

[0480] The one or more dependency parameters indicate at least one parent process that is in a different queue than the given set of graphics work.

[0481] A8. An apparatus as described in clause A7, wherein the queue selection circuit is configured to suspend selection from the first queue and select a set of graphics work from the second queue in response to a dependency parameter indicating a dependency of the set of work in the first queue on the set of work in the second queue.

[0482] A9. An apparatus according to clause A6, wherein:

[0483] The dependency parameters for a given set of graphics jobs include a set of parent identifiers and a valid parent mask indicating which parent identifiers are valid; and

[0484] The dependency parameters are specified in a manner that is independent of the tracking slot to which the set of graphics jobs is assigned.

[0485] A10. An apparatus according to any preceding clause in Set A, wherein:

[0486] said selection parameters comprising event flag parameters included in said data structure for a given set of graphics jobs;

[0487] The event flags parameter indicates one or more software programmable event flags on which the set of graphics jobs depends; and

[0488] The queue selection circuit is configured to wait for one or more event flags to suspend selection from the first queue based on the set of graphics jobs.

[0489] A11. The apparatus of clause 10, wherein at least one of the event flags is encoded as a counter value.

[0490] A12. The apparatus of clause A10, wherein:

[0491] A set of graphics jobs are independent of each other, are included in at least two different queues, and are targeted to a first hardware resource; and

[0492] The control circuit is configured to set an event flag for the launched set of graphics jobs in the set to acquire the first hardware resource, and prevent other sets of graphics jobs from being launched until the event flag is cleared.

[0493] A13. The apparatus of clause A10, wherein:

[0494] Another processor of the apparatus is configured to perform image processing on graphics frame data generated by the graphics processor circuit; and

[0495] The apparatus is configured to execute program instructions to utilize an event flag parameter to indicate when a subset of a frame of graphics data has completed processing by the other processor and is ready for use by the graphics processor circuitry.

[0496] A14. The apparatus of clause 10, wherein the apparatus is configured to control at least one event based on tasks performed by one or more other circuit components of the apparatus external to the graphics processing circuitry.

[0497] A15. An apparatus according to any preceding clause of Set A, wherein the graphics processor circuit comprises a control circuit configured to control a portion of a rendering process, including:

[0498] Stop all collections of graphics jobs that depend on the job's geometry collection;

[0499] scheduling a first copy of the set of fragments of the job to operate on data from the set of geometry of the job;

[0500] restarting the geometry set of work after executing the first copy of the fragment set of work; and

[0501] A second copy of the working set of fragments is configured to be restored from a partially rendered image generated by the first copy of the working set of fragments.

[0502] A16. An apparatus according to any preceding clause of Set A, wherein the graphics processor circuit comprises a control circuit configured to:

[0503] Writes result data for the completed set of graphics work to a completion queue structure in memory.

[0504] A17. The apparatus of clause A16, wherein the control circuitry is configured to write the result data to a plurality of different completion queues having different priorities.

[0505] A18. The apparatus of clause A16, wherein the control circuitry is configured to write result data to different completion queues based on whether a given set of graphics work completed normally.

[0506] A19. The apparatus of clause A16, wherein the control circuitry is further configured to, in response to completion of a set of graphics work, program a set of registers indicated by the set of graphics work.

[0507] A20. The apparatus of any preceding clause within Set A, wherein the graphics processor circuitry is configured to execute firmware to provide one or more of the sets of graphics work for queues in the data structure.

[0508] A21. An apparatus according to any preceding clause in Set A, further comprising:

[0509] a memory circuit configured to store the data structure; and

[0510] A processor circuit is configured to execute program instructions to add control information to a queue of the data structure.

[0511] A22. The apparatus of any preceding clause within Set A, wherein the graphics processor circuitry comprises control circuitry configured to control which sets of graphics work trigger an interrupt upon completion.

[0512] A23. An apparatus according to any preceding clause of Set A, wherein the image processor circuit comprises:

[0513] a plurality of SIMD pipelines configured to execute instructions; and

[0514] Fixed function circuitry configured to control the single instruction multiple data pipeline to perform operations for at least one of the following types of programs:

[0515] graphics shader programs; and machine learning programs.

[0516] A24. An apparatus according to any preceding clause in Set A, wherein the apparatus is a computing device, the computing device further comprising:

[0517] monitor;

[0518] central processing unit; and

[0519] Network interface.

[0520] A25. A method comprising any combination of operations performed by an apparatus according to any preceding clause within Set A.

[0521] A26. A non-transitory computer-readable medium having stored thereon instructions in a hardware description programming language, the instructions, when processed by a computing system, programming the computing system to generate a computer simulation model, wherein the model represents a hardware circuit, the hardware circuit comprising:

[0522] Any combination of the elements described in any of the preceding clauses within Set A.

[0523] Set B

[0524] B1. A device comprising:

[0525] a graphics control circuit, wherein the graphics control circuit implements a plurality of logical slots;

[0526] a set of graphics processor sub-units, the set of graphics processor sub-units each implementing a plurality of distributed hardware slots; and

[0527] A control circuit, the control circuit being configured to:

[0528] Based on the control information for the set of graphics work, assigning corresponding portions of the set of graphics work from logical slots to distributed hardware slots, including: in response to determining that the full number of distributed hardware slots to be utilized by the set of graphics work are unavailable, assigning an appropriate subset of the portions of the set of graphics work to the available distributed hardware slots.

[0529] B2. An apparatus according to any preceding clause within Set B, wherein the control circuitry is configured to: for a portion of the set of graphics work to be assigned:

[0530] Attempting to acquire a distributed hardware slot in a graphics processor subunit that does not currently have a distributed hardware slot assigned to the logical slot; and

[0531] In response to failure of the attempt, the portion of the set of graphics work is sent to a distributed hardware slot already owned by the logical slot.

[0532] B3. An apparatus according to any preceding clause within set B, wherein the apparatus is configured to perform the set of graphics work using fewer distributed hardware slots than the determined number of hardware slots for the set of graphics work, including performing multiple portions of the set of graphics work using the same distributed hardware slots.

[0533] B4. An apparatus according to any preceding clause within Set B, wherein the control circuitry is configured to implement a non-overlapping mode in which different portions of the same set of graphics work cannot be assigned to the same distributed hardware slot.

[0534] B5. The apparatus of any preceding clause within Set B, wherein the control circuitry is configured to track runtimes of portions of the set of graphics jobs in corresponding distributed hardware slots, and to store the tracked runtimes using one or more of the following techniques:

[0535] aggregating the traced runtimes into a single runtime count; and

[0536] The traced runtime is written to a completion buffer.

[0537] B6. An apparatus according to any preceding clause in Set B, wherein the control circuit is further configured to:

[0538] In response to the context storage, storing information indicating which distributed hardware slots are used by the logical slot; and

[0539] Responsive to the context loading, allocation of the indicated distributed hardware slot is awaited before continuing execution of the set of graphics work.

[0540] B7. An apparatus according to any preceding clause within Set B, wherein the plurality of logical slots includes one or more slots dedicated to one or more types of graphics work.

[0541] B8. An apparatus according to any preceding clause in Set B, further comprising:

[0542] queue access circuitry configured to access a data structure in memory specifying a plurality of queues, wherein a respective queue queues control information for a plurality of sets of graphics jobs; and

[0543] A queue selection circuit is configured to select a set of graphics jobs from the data structure based on one or more selection parameters and to store control information for the selected set of graphics jobs in the slot tracking circuit.

[0544] B9. An apparatus according to any of the preceding clauses within set B, wherein the control circuit is configured to assign a first set of geometry work to N-1 graphics processor sub-units of N graphics processor sub-units of the apparatus, and to assign a second set of work to a single remaining graphics processor sub-unit of the apparatus.

[0545] B10. An apparatus as recited in any preceding clause of Set B, wherein the apparatus supports allocation of any determined integer number of graphics processor subunits in the range from one processor subunit to the total number of processor subunits included in the apparatus.

[0546] B11. The apparatus of clause B10, wherein the range includes one or more integer numbers of graphics processor sub-units that are not powers of two.

[0547] B12. An apparatus according to clause B10, wherein the set of graphics work is a fragment processing set of graphics work, and the control circuit is configured to determine the number of graphics processor subunits based on the number of tiles included in the set of graphics work or based on the number of pixels included in the set of graphics work.

[0548] B13. An apparatus according to clause B10, wherein the set of graphics work is a compute set of graphics work, and the control circuit is configured to determine the number of graphics processor subunits based on the number of workgroups, work items, or both included in the set of graphics work.

[0549] B14. The apparatus of clause B10, wherein the set of graphics work is a geometry set of graphics work, and the control circuitry is configured to determine the number of graphics processor subunits based on one or more of the following parameters:

[0550] the number of graphics primitives included in the set of graphics work;

[0551] the number of vertices included in the set of graphics work; and

[0552] The determined complexity of the set of graphics tasks.

[0553] B15. A non-transitory computer-readable medium having stored thereon instructions in a hardware description programming language, which, when processed by a computing system, program the computing system to generate a computer simulation model, wherein the model represents a hardware circuit, the hardware circuit comprising:

[0554] Any combination of the elements described in any of the preceding clauses within Set B.

[0555] B16. A method comprising any combination of operations performed by an apparatus according to any preceding clause within Set B.

[0556] Set C

[0557] C1. A device comprising:

[0558] a graphics control circuit, wherein the graphics control circuit implements a plurality of logical slots;

[0559] a set of graphics processor sub-units, the set of graphics processor sub-units each implementing a plurality of distributed hardware slots; and

[0560] A control circuit, the control circuit being configured to:

[0561] assigning resolved versions of the set of geometry work to distributed hardware slots of one or more of the graphics processor subunits;

[0562] determining a number of segments of the set to use for geometry work based on execution of the parsed version; and

[0563] Assigning the determined segments to distributed hardware slots of corresponding graphics processor subunits for execution; and

[0564] A stitching circuit is configured to stitch results of the segments processed by the assigned distributed hardware slots.

[0565] C2. The apparatus of any preceding clause of Set C, wherein the stitching circuit comprises:

[0566] A splicing control circuit, wherein the splicing control circuit is configured to:

[0567] Assigning splicing work for one or more first data structure categories to hardware splicing slots in the main control circuit; and

[0568] Stitching work for the one or more second data structure categories is assigned to distributed hardware tiling slots in corresponding graphics processor subunits, wherein the distributed hardware tiling slots include corresponding memory interfaces to access memory storing the data structures to be tiled.

[0569] C3. The apparatus of clause C2, wherein the one or more first data structure types utilize less memory space than the one or more second data structure types.

[0570] C4. An apparatus according to clause C2, wherein:

[0571] The one or more first data structure categories include a layer identifier cache and a closed page list; and

[0572] The one or more second data structure types include a tile region array header.

[0573] C5. An apparatus according to any preceding clause in Set C, wherein the control circuit is configured to:

[0574] assigning resolution work of the set of geometry work to at most one distributed hardware slot of a given graphics processor subunit; and

[0575] Assigns segment execution work to at most one distributed hardware slot of a given graphics processor subunit.

[0576] C6. An apparatus according to any preceding clause in Set C, wherein the control circuit is configured to:

[0577] assigning the parsed version to all graphics processor sub-units in a set of graphics processor sub-units; and

[0578] Segments of the determined number of segments are serially assigned to available graphics processor sub-units in the set.

[0579] C7. An apparatus according to any preceding clause within Set C, wherein the tiling circuitry is configured to tile results across multiple graphics processor subunits assigned to segments of the set for geometry work.

[0580] C8. An apparatus according to any preceding clause within set C, wherein the control circuit is configured to allow different sets of geometry work to be executed in parallel on different distributed hardware sockets only if the different sets of geometry work share a parameter buffer.

[0581] C9. An apparatus according to any preceding clause within Set C, wherein the control circuit is configured to dynamically change the number of distributed hardware slots assigned to the set of geometry work during execution of the set of geometry work.

[0582] C10. A non-transitory computer-readable medium having stored thereon instructions in a hardware description programming language, which, when processed by a computing system, program the computing system to generate a computer simulation model, wherein the model represents a hardware circuit, the hardware circuit comprising:

[0583] Any combination of the elements described in any of the preceding clauses within Set C.

[0584] C11. A method comprising any combination of operations performed by an apparatus according to any preceding clause within Set C.

[0585] Set D

[0586] D1. A device comprising:

[0587] A control circuit, the control circuit being configured to:

[0588] receiving a different set of graphics work and scheduling the set of graphics work for execution on the distributed hardware resources, the set of graphics work including a first set of work that is dependent on a second set of work;

[0589] in response to a release signal from the second set of work indicating that the second set of work has reached a first processing point, initiating processing of the first set of work;

[0590] In response to reaching a gate point in the first set of jobs, stalling processing of the first set of jobs; and

[0591] Responsive to an end signal for the second set of work, processing of the first set of work is resumed.

[0592] D2. The apparatus of any preceding clause within Set D, wherein the control circuit is configured to receive multiple types of release signals, the multiple types of release signals comprising:

[0593] an early release signal indicating that all SIMD groups have been started for a given set of work; and

[0594] A delayed release signal indicating that all SIMD groups for a given set of work have completed.

[0595] D3. An apparatus as described in clause D2, wherein the control circuit is configured to allow a dependent set of work to initiate processing based on an early release signal from one or more sets of work of a first type and based on a delayed release signal from one or more sets of work of a second type.

[0596] D4. An apparatus according to any preceding clause in Set D, wherein the control circuit is configured to implement:

[0597] A hard dependency indicating that a parent set of work must be completed before processing of a child set of work can be initiated; and

[0598] A soft dependency indicates initiating processing on a child set of work based on a release signal from the parent set of work before the parent set of work completes.

[0599] D5. The apparatus of clause D4, wherein the control circuitry is configured to track both hard dependencies and soft dependencies using a dependency matrix circuit.

[0600] D6. The apparatus of any preceding clause within set D, wherein the control circuitry supports multiple types of gating points, the multiple types of gating points comprising:

[0601] A first gate point class, the first gate point class corresponds to a point where one or more execution state loading SIMD groups have been allocated resources but have not yet executed instructions; and a second gate point class, the second gate point class corresponds to a point where one or more working SIMD groups have been allocated resources but have not yet executed instructions.

[0602] D7. An apparatus according to any preceding clause in Set D, wherein the processor circuit comprises:

[0603] a plurality of SIMD pipelines configured to execute instructions; and

[0604] Fixed function circuitry configured to control the single instruction multiple data pipeline to perform operations for at least one of the following types of programs:

[0605] Graphics shader programs; and

[0606] Machine learning program.

[0607] D8. An apparatus according to any preceding clause in Set D, wherein the apparatus is a computing device, the computing device further comprising:

[0608] monitor;

[0609] central processing unit; and

[0610] Network interface.

[0611] D9. A non-transitory computer-readable medium having stored thereon instructions in a hardware description programming language, which, when processed by a computing system, program the computing system to generate a computer simulation model, wherein the model represents a hardware circuit, the hardware circuit comprising:

[0612] Any combination of the elements described in any of the preceding clauses within Set D.

[0613] D10. A method comprising any combination of operations performed by an apparatus according to any preceding clause within set D.

[0614] ***

[0615] The concept of "execution" is broad and can refer to 1) the processing of instructions in the entire execution pipeline (e.g., through the fetch stage, decode stage, execute stage, and rollback stage), and 2) the processing of instructions at the execution unit or execution subsystem of such pipeline (e.g., integer execution unit or load store unit). The latter meaning can also be referred to as "perform" an instruction. Thus, "perform" an add instruction refers to adding two operands to produce a result, which in some embodiments can be implemented by circuitry at the execution stage of the pipeline (e.g., an execution unit). Conversely, "perform" an add instruction can refer to the entire operation that occurs in the entire pipeline as a result of the add instruction. Similarly, "perform" a "load" instruction can include retrieving a value (e.g., from a cache, memory, or the stored result of another instruction) and storing the retrieved value in a register or other location.

[0616] As used herein, in the context of an instruction, the term "completion" refers to the committal of the result of the instruction to the architectural state of a processor or processing element. For example, completion of an add instruction includes writing the result of the add instruction to a destination register. Similarly, completion of a load instruction includes writing a value (e.g., a value retrieved from a cache or memory) to a destination register or a representation thereof.

[0617] The concept of a processor "pipeline" is well known and refers to the concept of dividing the "work" that a processor performs on an instruction into multiple stages. In some embodiments, decoding, dispatching, executing (i.e., performing), and retiring an instruction may be examples of different pipeline stages. Many different pipeline architectures may have different element / part orderings. Various pipeline stages perform such steps on an instruction during one or more processor clock cycles and then pass the instruction or an operation associated with the instruction to other stages for further processing.

[0618] Example device

[0619] Now refer to Figure 50 , a block diagram of an example embodiment of an exemplary device 5000 is shown. In some embodiments, elements of the device 5000 may be included in a system on a chip. In some embodiments, the device 5000 may be included in a mobile device that may be battery powered. Therefore, the power consumption of the device 5000 may be an important design consideration. In the illustrated embodiment, the device 5000 includes a fabric 5010, a compute complex 5020, an input / output (I / O) bridge 5050, a cache / memory controller 5045, a graphics unit 5075, and a display unit 5065. In some embodiments, in addition to or in place of the illustrated components, the device 5000 may include other components (not shown), such as video processor encoders and decoders, image processing or recognition elements, computer vision elements, etc.

[0620] Fabric 5010 may include various interconnects, buses, MUXs, controllers, etc., and may be configured to facilitate communication between various elements of device 5000. In some embodiments, portions of fabric 5010 may be configured to implement various different communication protocols. In other embodiments, fabric 5010 may implement a single communication protocol, and elements coupled to fabric 5010 may internally convert from the single communication protocol to other communication protocols.

[0621] In the illustrated embodiment, compute complex 5020 includes a bus interface unit (BIU) 5025, a cache 5030, and cores 5035 and 5040. In various embodiments, compute complex 5020 may include various numbers of processors, processor cores, and caches. For example, compute complex 5020 may include one, two, four, or any other suitable number of processor cores. In one embodiment, cache 5030 is a set-associative L2 cache. In some embodiments, cores 5035 and 5040 may include internal instruction and data caches. In some embodiments, a coherence unit (not shown) in fabric 5010, cache 5030, or elsewhere in device 5000 may be configured to maintain coherence between the various caches of device 5000. BIU 5025 may be configured to manage communications between compute complex 5020 and other components of device 5000. Processor cores (such as core 5035 and core 5040) can be configured to execute instructions of a specific instruction set architecture (ISA), which can include operating system instructions and user application instructions. These instructions can be stored in a computer-readable medium (such as a memory coupled to a memory controller 5045 discussed below).

[0622] As used herein, the term "coupled to" may indicate one or more connections between elements, and a coupling may include intervening elements. For example, in Figure 50 In , graphics unit 5075 can be described as being "coupled to" memory via fabric 5010 and cache / memory controller 5045. In contrast, in Figure 50 In the illustrated embodiment, graphics unit 5075 is "directly coupled" to fabric 5010 because there are no intervening elements.

[0623] The cache / memory controller 5045 can be configured to manage data transfer between the fabric 5010 and one or more caches and memories. For example, the cache / memory controller 5045 can be coupled to an L3 cache, which in turn can be coupled to system memory. In other embodiments, the cache / memory controller 5045 can be directly coupled to the memory. In some embodiments, the cache / memory controller 5045 can include one or more internal caches. The memory coupled to the controller 5045 can be any type of volatile memory, such as dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate (DDR, DDR2, DDR3, etc.) SDRAM (including mobile versions of SDRAM such as mDDR3, etc., and / or low-power versions of SDRAM such as LPDDR4, etc.), RAMBUS DRAM (RDRAM), static RAM (SRAM), etc. One or more memory devices can be coupled to a circuit board to form a memory module, such as a single inline memory module (SIMM), a dual inline memory module (DIMM), etc. Alternatively, these devices may be mounted with integrated circuits in a chip-on-chip configuration, a package-on-package configuration, or a multi-chip module configuration. The memory coupled to the controller 5045 may be any type of non-volatile memory, such as NAND flash memory, NOR flash memory, nano RAM (NRAM), magnetoresistive RAM (MRAM), phase-change RAM (PRAM), racetrack memory, memristor memory, etc. As described above, the memory may store program instructions that can be executed by the computing complex 5020 to cause the computing device to perform the functionality described herein.

[0624] The graphics unit 5075 may include one or more processors, such as one or more graphics processing units (GPUs). For example, the graphics unit 5075 may receive graphics-oriented instructions such as or Instructions. The graphics unit 5075 may execute dedicated GPU instructions or perform other operations based on the received graphics-oriented instructions. The graphics unit 5075 may generally be configured to process large blocks of data in parallel and may build images in a frame buffer for output to a display, which may be included in the device or may be a separate device. The graphics unit 5075 may include transformation, lighting, triangle, and rendering engines in one or more graphics processing pipelines. The graphics unit 5075 may output pixel information for displaying an image. In various embodiments, the graphics unit 5075 may include a programmable shader circuit, which may include a highly parallel execution core configured to execute graphics programs, which may include pixel tasks, vertex tasks, and compute tasks (which may or may not be graphics-related).

[0625] The display unit 5065 may be configured to read data from the frame buffer and provide a stream of pixel values ​​for display. In some embodiments, the display unit 5065 may be configured as a display pipeline. Additionally, the display unit 5065 may be configured to blend multiple frames to produce an output frame. In addition, the display unit 5065 may include one or more interfaces (e.g., a touch screen or external display) for coupling to a user display (e.g., a touch screen or external display). or embedded DisplayPort (eDP).

[0626] The I / O bridge 5050 may include various components configured to implement, for example, Universal Serial Bus (USB) communications, security, audio, and low-power always-on functionality. The I / O bridge 5050 may also include interfaces such as, for example, pulse width modulation (PWM), general-purpose input / output (GPIO), serial peripheral interface (SPI), and inter-integrated circuit (I2C). Various types of peripherals and devices may be coupled to the device 5000 via the I / O bridge 5050.

[0627] In some embodiments, the device 5000 includes a network interface circuit (not explicitly shown) that can be connected to the fabric 5010 or the I / O bridge 5050. The network interface circuit can be configured to communicate via various networks, which can be wired networks, wireless networks, or both. For example, the network interface circuit can be configured to communicate via a wired local area network, a wireless local area network (e.g., via Wi-Fi), or a wireless local area network (e.g., via Wi-Fi). TM ) or a wide area network (e.g., the Internet or a virtual private network). In some embodiments, the network interface circuit is configured to communicate via one or more cellular networks using one or more radio access technologies. In some embodiments, the network interface circuit is configured to use device-to-device communication (e.g., or Wi-Fi TMIn various embodiments, the network interface circuit can provide connectivity for the device 5000 with various types of other devices and networks.

[0628] Sample Application

[0629] Now go to Figure 51 , shows various types of systems that may include any of the circuits, devices, or systems discussed above. A system or device 5100 that may incorporate or otherwise utilize one or more of the techniques described herein may be used in a wide range of fields. For example, the system or device 5100 may be used as part of the hardware of a system such as a desktop computer 5110, a laptop computer 5120, a tablet computer 5130, a cellular or mobile phone 5140, or a television 5150 (or a set-top box coupled to a television).

[0630] Similarly, the disclosed elements can be used in wearable devices 5160, such as smart watches or health monitoring devices. In many embodiments, a smart watch can perform a variety of different functions—for example, access to email, cellular service, calendar, health monitoring, etc. A wearable device can also be designed to perform only health monitoring functions, such as monitoring the user's vital signs, performing epidemiological functions such as contact tracing, providing communications to emergency medical services, etc. Other types of devices are also contemplated, including devices worn around the neck, devices that can be implanted in the human body, glasses or helmets designed to provide computer-generated reality experiences, such as those based on augmented reality and / or virtual reality, etc.

[0631] The system or device 5100 may also be used in various other contexts. For example, the system or device 5100 may be used in the context of a server computer system (such as a dedicated server) or on shared hardware that implements cloud-based services 5170. Furthermore, the system or device 5100 may be implemented in a wide range of dedicated everyday devices, including devices 5180 commonly found in homes, such as refrigerators, thermostats, security cameras, and the like. The interconnection of such devices is commonly referred to as the "Internet of Things" (IoT). Components may also be implemented in various modes of transport. For example, the system or device 5100 may be used for control systems, guidance systems, entertainment systems, and the like for various types of vehicles 5190.

[0632] Figure 51 The applications illustrated in the examples are merely exemplary and are not intended to limit potential future applications of the disclosed systems or devices. Other example applications include, but are not limited to, portable gaming devices, music players, data storage devices, drones, etc.

[0633] Example computer-readable media

[0634] The present disclosure has described various example circuits in detail above. It is intended that the present disclosure encompass not only embodiments including such circuits, but also computer-readable storage media including design information specifying such circuits. Accordingly, the present disclosure is intended to support claims that encompass not only apparatuses including the disclosed circuits, but also storage media that specify the circuits in a format that allows programming a computing system to generate a simulation model of a hardware circuit, programming a manufacturing system configured to produce hardware (e.g., an integrated circuit) that includes the disclosed circuits, and the like. Claims to such storage media are intended to encompass, for example, entities that generate circuit designs but do not themselves perform complete operations (such as design simulation, design synthesis, circuit fabrication, etc.).

[0635] Figure 52 is a block diagram illustrating an example non-transitory computer-readable storage medium storing circuit design information according to some embodiments. In the illustrated embodiment, computing system 5240 is configured to process the design information. This may include executing instructions included in the design information, interpreting instructions included in the design information, compiling, transforming, or otherwise updating the design information, etc. Thus, in some embodiments, the design information (e.g., by programming computing system 5240) controls computing system 5240 to perform various operations discussed below.

[0636] In the illustrated example, computing system 5240 processes the design information to generate both a computer simulation model of hardware circuit 5260 and lower-level design information 5250. In other embodiments, computing system 5240 may generate only one of these outputs, may generate the other output based on the design information, or both. With respect to the computer simulation, computing system 5240 may execute instructions of a hardware description language, which may include register transfer level (RTL) code, behavioral code, structural code, or some combination thereof. The simulation model may perform the functionality specified by the design information, facilitate verification of the functional correctness of the hardware design, generate power consumption estimates, generate timing estimates, and the like.

[0637] In the illustrated example, computing system 5240 also processes the design information to generate lower-level design information 5250 (e.g., gate-level design information, a netlist, etc.). As shown, this may include synthesis operations such as building a multi-level network, optimizing the network using technology-independent techniques, technology-dependent techniques, or both, and outputting a gate network (with potential constraints based on the available gate-to-technology library, sizing, delay, power, etc.). Based on lower-level design information 5250 (and potentially other inputs), semiconductor manufacturing system 5220 is configured to manufacture integrated circuit 5230 (which may correspond to the functionality of simulation model 5260). Note that computing system 5240 can generate different simulation models based on design information at various levels of description (including information 5250, 5215, etc.). Data representing design information 5250 and model 5260 may be stored on medium 5210 or on one or more other media.

[0638] In some embodiments, lower-level design information 5250 controls (e.g., programs) semiconductor fabrication system 5220 to fabricate integrated circuit 5230. Thus, when processed by the fabrication system, the design information can program the fabrication system to fabricate circuits including the various circuits disclosed herein.

[0639] The non-transitory computer-readable storage medium 5210 may include any of various suitable types of memory devices or storage devices. The non-transitory computer-readable storage medium 5210 may be an installation medium, such as a CD-ROM, floppy disk, or tape device; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media, such as a hard drive or optical storage device; registers, or other similar types of memory elements, etc. The non-transitory computer-readable storage medium 5210 may also include other types of non-transitory memory or combinations thereof. Therefore, the non-transitory computer-readable storage medium 5210 may include two or more memory media; such media may reside in different locations—for example, in different computer systems connected by a network.

[0640] Design information 5215 may be specified using any of a variety of suitable computer languages, including hardware description languages ​​such as, but not limited to, VHDL, Verilog, SystemC, SystemVerilog, RHDL, M, MyHDL, and the like. The formats of the various design information may be recognized by one or more applications executed by computing system 5240, semiconductor manufacturing system 5220, or both. In some embodiments, the design information may also include one or more cell libraries that specify the synthesis, layout, or both of integrated circuit 5230. In some embodiments, the design information is specified in whole or in part in the form of a netlist specifying the cell library elements and their connectivity. Taken alone, the design information discussed herein may or may not include sufficient information for manufacturing the corresponding integrated circuit. For example, the design information may specify the circuit elements to be manufactured, but not their physical layout. In such cases, the design information may need to be combined with layout information to actually manufacture the specified circuit.

[0641] In various embodiments, integrated circuit 5230 may include one or more custom macrocells, such as memory and analog or mixed-signal circuits. In this case, the design information may include information related to the included macrocells. Such information may include, but is not limited to, a schematic capture database, mask design data, behavioral models, and device or transistor-level netlists. The mask design data may be formatted according to Graphics Data System II (GDSII) or any other suitable format.

[0642] Semiconductor manufacturing system 5220 may include any of a variety of suitable elements configured to manufacture integrated circuits. This may include, for example, elements for depositing semiconductor material (e.g., on a wafer that may include a mask), removing material, changing the shape of deposited material, modifying material (e.g., by doping the material or using ultraviolet treatment to modify the dielectric constant), etc. Semiconductor manufacturing system 5220 may also be configured to perform various tests on the manufactured circuits for proper operation.

[0643] In various embodiments, integrated circuit 5230 and model 5260 are configured to operate according to the circuit design specified by design information 5215, which may include performing any of the functionality described herein. For example, integrated circuit 5230 may include Figures 30 to 31 、 Figure 33 、 Figure 36 、 Figures 38A to 38B and Figures 48 to 50 Any of the various elements shown in . In addition, integrated circuit 5230 can be configured to perform various functions described herein in conjunction with other components. In addition, the functionality described herein can be performed by multiple connected integrated circuits.

[0644] As used herein, a phrase of the form “design information specifying a design of a circuit configured to…” does not imply that the circuit in question must be manufactured in order to satisfy the element. Rather, the phrase indicates that the design information describes a circuit that, when manufactured, will be configured to perform the indicated actions or will include the specified components. Similarly, the statement that “instructions of a hardware description programming language are “executable” to program a computing system to generate a computer simulation model” does not mean that the instructions must be executed in order to satisfy the element, but rather that the characteristics of those instructions are specified. In this case, additional features associated with the model (or the circuit represented by the model) may similarly be related to the characteristics of those instructions. Thus, an entity that sells a computer-readable medium having instructions that satisfy the stated characteristics may be providing an infringing product even if another entity actually executes those instructions on the medium.

[0645] It is important to note that a given design, at least in the context of digital logic, can be implemented using a variety of different gate arrangements, circuit technologies, etc. As one example, different designs may select or connect gates based on design tradeoffs (e.g., to focus on power consumption, performance, circuit area, etc.). Additionally, different manufacturers may have proprietary libraries, gate designs, physical gate implementations, etc. Different entities may also use different tools to process design information at various levels (e.g., from behavioral specifications to physical layout of gates).

[0646] However, once a digital logic design is specified, one skilled in the art does not need to perform extensive experimentation or research to determine these implementations. Instead, one skilled in the art understands the procedures for reliably and predictably generating one or more circuit implementations that provide the functionality described by the design information. Different circuit implementations may affect the performance, area, power consumption, etc. of a given design (potentially with trade-offs between different design goals), but the logical functionality does not change between different circuit implementations of the same circuit design.

[0647] In some embodiments, the instructions included in the design information instructions provide RTL information (or other higher-level design information) and are executable by a computing system to synthesize a gate-level netlist representing a hardware circuit based on the RTL information as input. Similarly, the instructions can provide behavioral information and be executed by the computing system to synthesize a netlist or other lower-level design information. The lower-level design information can be used to program the manufacturing system 5220 to manufacture the integrated circuit 5230.

[0648] ***

[0649] The various techniques described herein may be performed by one or more computer programs. The term "program" should be interpreted broadly to encompass a sequence of instructions in a programming language that is executable by a computing device. These programs may be written in any suitable computer language, including lower-level languages ​​such as assembly and higher-level languages ​​such as Python. The program may be written in a compiled language such as C or C++ or an interpreted language such as JavaScript.

[0650] Program instructions may be stored on a "computer-readable storage medium" or "computer-readable medium" to facilitate execution of the program instructions by a computer system. Generally speaking, these phrases include any tangible or non-transitory storage medium or memory medium. The terms "tangible" and "non-transitory" are intended to exclude propagating electromagnetic signals, but do not otherwise limit the type of storage medium. Thus, the phrases "computer-readable storage medium" or "computer-readable medium" are intended to cover types of storage devices that do not necessarily store information permanently (e.g., random access memory (RAM)). Thus, the term "non-transitory" is a limitation on the nature of the medium itself (i.e., the medium cannot be a signal), as opposed to limitations on the data storage persistence of the medium (e.g., RAM vs. ROM).

[0651] The phrases "computer-readable storage medium" and "computer-readable medium" are intended to refer to storage media within a computer system as well as removable media such as a CD-ROM, memory stick, or portable hard drive. These phrases encompass any type of volatile memory within a computer system, including DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc., as well as non-volatile memory such as magnetic media (e.g., hard drives) or optical storage devices. These phrases are expressly intended to encompass the memory of a server that facilitates the download of program instructions, the memory within any intermediate computer systems involved in the download, and the memory of all destination computing devices. Still further, these phrases are intended to encompass combinations of different types of memory.

[0652] Furthermore, a computer-readable medium or storage medium may be located in a first set of one or more computer systems in which the program is executed, and in a second set of one or more computer systems connected to the first set via a network. In the latter instance, the second set of computer systems may provide program instructions to the first set of computer systems for execution. In short, the phrases "computer-readable storage medium" and "computer-readable medium" may include two or more media that may reside in different locations (e.g., in different computers connected via a network).

[0653] This disclosure includes references to "an embodiment" or groups of "embodiments" (e.g., "some embodiments" or "various embodiments"). An embodiment is a different specific implementation or example of the disclosed concepts. References to "an embodiment," "one embodiment," and "a specific embodiment," etc., do not necessarily refer to the same embodiment. Numerous possible embodiments are contemplated, including those specifically disclosed, as well as modifications or alternatives that fall within the spirit or scope of this disclosure.

[0654] This disclosure may discuss potential advantages that may result from the disclosed embodiments. Not all implementations of these embodiments will necessarily exhibit any or all of these potential advantages. Whether a particular implementation achieves an advantage depends on many factors, some of which are outside the scope of this disclosure. Indeed, there are many reasons why an implementation falling within the scope of the claims may not exhibit some or all of the disclosed advantages. For example, a particular implementation may include additional circuitry outside the scope of this disclosure that, in combination with one of the disclosed embodiments, negates or mitigates one or more of the disclosed advantages. Furthermore, suboptimal design implementation of a particular implementation (e.g., a particular implementation technique or tool) may also negate or mitigate the disclosed advantages. Even assuming a specific implementation of the technique, the realization of an advantage may still depend on other factors, such as the environmental circumstances in which the implementation is deployed. For example, inputs provided to a particular implementation may prevent one or more of the problems addressed by this disclosure from occurring in a particular situation, and as a result, the benefits of its solution may not be realized. In light of the possible existence of factors external to this disclosure, it is expressly stated that any potential advantages described herein should not be construed as claim limitations that must be met in order to prove infringement. Rather, the identification of such potential advantages is intended to illustrate the types of improvements available to designers who benefit from this disclosure. Permanently describing such advantages (eg, stating that a particular advantage "may occur") is not intended to convey a doubt as to whether such advantage can actually be achieved, but rather to recognize that achievement of such advantages often depends on technical realities of additional factors.

[0655] Unless otherwise stated, the embodiments are non-restrictive. That is, the disclosed embodiments are not intended to limit the scope of claims drafted based on this disclosure, even when only a single example is described with respect to a particular feature. The disclosed embodiments are intended to be illustrative and not restrictive, without any statement to the contrary in this disclosure. Therefore, this application is intended to allow claims covering the disclosed embodiments, as well as such alternatives, modifications, and equivalents, which will be apparent to those skilled in the art knowing the beneficial effects of this disclosure.

[0656] For example, features in this application may be combined in any suitable manner. Accordingly, new claims may be formulated during the prosecution of this application (or an application claiming priority thereto) directed to any such combination of features. Specifically, with reference to the appended claims, features of dependent claims may, where appropriate, be combined with features of other dependent claims, including claims that are dependent on other independent claims. Similarly, features from corresponding independent claims may, where appropriate, be combined.

[0657] Thus, while the appended dependent claims can be drafted so that each dependent claim is dependent on a single other claim, additional dependencies are also contemplated. Any combination of dependent features consistent with the present disclosure is contemplated and may be claimed in this or another application. In short, the combinations are not limited to those specifically recited in the appended claims.

[0658] It is also contemplated that claims drafted in one format or legal type (eg, apparatus) are intended to support corresponding claims in another format or legal type (eg, method), where appropriate.

[0659] ***

[0660] Because this disclosure is a legal document, various terms and phrases may be subject to regulatory and judicial interpretation. Notice is hereby given that the definitions provided in the following paragraphs and throughout this disclosure will be used to determine how claims drafted based on this disclosure are to be interpreted.

[0661] Unless the context clearly dictates otherwise, reference to an item in the singular (i.e., a noun or noun phrase preceded by "a," "an," or "the") is intended to mean "one or more." Thus, reference to "an item" in a claim, without accompanying context, does not exclude additional instances of that item. A "plurality" of an item refers to a collection of two or more of the items.

[0662] The word "may" is used herein in a permissive sense (ie, having the potential to, being able to), rather than the mandatory sense (ie, must).

[0663] The terms "include" and "including" and their forms are open ended and mean "including, but not limited to."

[0664] When the term "or" is used in this disclosure with respect to a list of options, unless the context provides otherwise, it will generally be understood to be used in an inclusive sense. Thus, the expression "x or y" is equivalent to "x or y, or both," and thus encompasses 1) x but not y, 2) y but not x, and 3) both x and y. On the other hand, phrases such as "either, but not both, x or y" make it clear that "or" is used in an exclusive sense.

[0665] The expressions "w, x, y, or z, or any combination thereof" or "... at least one of w, x, y, and z" are intended to cover all possibilities involving individual elements up to the total number of elements in the set. For example, given the set [w, x, y, z], these phrases cover any single element in the set (e.g., w but not x, y, or z), any two elements (e.g., w and x, but not y or z), any three elements (e.g., w, x, and y, but not z), and all four elements. The phrase "... at least one of w, x, y, and z" thus refers to at least one element in the set [w, x, y, z], thereby covering all possible combinations in that list of elements. The phrase should not be interpreted as requiring the presence of at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.

[0666] In this disclosure, various "labels" may precede a noun or noun phrase. Unless the context provides otherwise, different labels used for a feature (e.g., "first circuit," "second circuit," "particular circuit," "given circuit," etc.) refer to different instances of the feature. Additionally, unless otherwise specified, the labels "first," "second," and "third" do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) when applied to features.

[0667] The phrase "based on" or "based on" is used to describe one or more factors that influence a determination. The term does not exclude the possibility that additional factors may influence the determination. That is, a determination may be based solely on the specified factors or on the specified factors as well as other unspecified factors. Consider the phrase "A is determined based on B." The phrase specifies that B is a factor used to determine A or that B influences the determination of A. The phrase does not exclude that the determination of A may also be based on some other factor (such as C). The phrase is also intended to cover embodiments in which A is determined solely based on B. As used herein, the phrase "based on" is synonymous with the phrase "based at least in part on."

[0668] The phrases "in response to" and "in response to" describe one or more factors that trigger an effect. The phrases do not exclude the possibility that additional factors may influence or otherwise trigger the effect, either in conjunction with the specified factors or independently of the specified factors. That is, the effect may be responsive only to these factors, or may be responsive to the specified factors as well as other unspecified factors. Consider the phrase "A is performed in response to B." The phrase specifies that B is the factor that triggers the performance of A or triggers a particular result of A. The phrase does not exclude that the performance of A may also be responsive to some other factor, such as C. The phrase also does not exclude that the performance of A may be performed in response to B and C in combination. The phrase is also intended to cover embodiments in which A is performed only in response to B. As used herein, the phrase "in response to" is synonymous with the phrase "at least partially in response to." Similarly, the phrase "in response to" is synonymous with the phrase "at least partially in response to."

[0669] ***

[0670] Within the present disclosure, different entities (which may be variously referred to as "units," "circuits," other components, etc.) may be described or claimed as being "configured to" perform one or more tasks or operations. The expression "an entity configured to perform one or more tasks" is used herein to refer to a structure (i.e., a tangible thing). More specifically, the expression is used to indicate that the structure is arranged to perform one or more tasks during operation. A structure may be considered to be "configured to" perform a task even if the structure is not currently being operated. Thus, an entity described or stated as "configured to" perform a task refers to a physical thing used to implement the task, such as a device, a circuit, a system with a processor unit, a memory storing executable program instructions, etc. The phrase is not used herein to refer to an intangible thing.

[0671] In some cases, various units / circuits / components may be described herein as performing a set of tasks or operations. It should be understood that these entities are "configured to" perform those tasks / operations, even if not specifically stated.

[0672] The term "configured to" is not intended to mean "capable of being configured to." For example, an unprogrammed FPGA would not be considered "configured to" perform a particular function. However, the unprogrammed FPGA could be "configurable to" perform that function. After being appropriately programmed, the FPGA could then be considered "configured to" perform the particular function.

[0673] For purposes of a U.S. patent application based on the present disclosure, stating in a claim that a structure is “configured to” perform one or more tasks is expressly intended not to invoke 35 U.S.C. §112(f) for that claim element. If the applicant wishes to invoke section 112(f) during prosecution of a U.S. patent application based on the present disclosure, it would use the “means for [performing the function]” construct to phrase the claim element.

[0674] Different “circuits” may be described in this disclosure. These circuits or “circuits” constitute hardware that includes various types of circuit elements, such as combinational logic, clock storage devices (e.g., flip-flops, registers, latches, etc.), finite state machines, memories (e.g., random access memory, embedded dynamic random access memory), programmable logic arrays, etc. Circuits can be custom designed or taken from a standard library. In various specific implementations, circuits may include digital components, analog components, or a combination of both, as appropriate. Certain types of circuits may be generally referred to as “units” (e.g., decoding units, arithmetic logic units (ALUs), functional units, memory management units (MMUs), etc.). Such units are also referred to as circuits or circuitry.

[0675] Thus, the disclosed circuits / units / components and other elements illustrated in the accompanying drawings and described herein include hardware elements, such as those described in the preceding paragraphs. In many cases, the internal arrangement of hardware elements in a particular circuit can be specified by describing the functionality of that circuit. For example, a particular "decode unit" may be described as performing the function of "processing an instruction's opcode and routing that instruction to one or more of a plurality of functional units," meaning that the decode unit is "configured to" perform that function. For one skilled in the computer arts, this functional specification is sufficient to suggest a set of possible architectures for the circuit.

[0676] In various embodiments, as described in the preceding paragraphs, circuits, units, and other elements may be defined by the functions or operations they are configured to implement. The arrangement of such circuits / units / components relative to one another and the manner in which they interact form a microarchitecture definition of the hardware that is ultimately manufactured in an integrated circuit or programmed into an FPGA to form a physical implementation of the microarchitecture definition. Thus, a microarchitecture definition is considered by those skilled in the art to be a structure from which many physical implementations can be derived, all of which fall within the broader structure described by the microarchitecture definition. That is, a person skilled in the art having a microarchitecture definition provided in accordance with the present disclosure can, without undue experimentation and with the application of ordinary skill, implement the structure by coding a description of the circuits / units / components in a hardware description language (HDL) such as Verilog or VHDL. HDL descriptions are often expressed in a manner that can be rendered as functional. However, for those skilled in the art, the HDL description is a means for converting the structure of a circuit, unit, or component into the next level of implementation details. Such HDL descriptions may take the form of behavioral code (which is generally non-synthesizable), register transfer language (RTL) code (which is generally synthesizable compared to behavioral code), or structural code (e.g., a netlist specifying logic gates and their connectivity). The HDL description may be sequentially synthesized for a library of cells designed for a given integrated circuit manufacturing technology and may be modified for timing, power, and other reasons to obtain a final design database that is transmitted to the factory to generate masks and ultimately produce the integrated circuit. Some hardware circuits, or portions thereof, may also be custom designed in the schematic editor and captured into the integrated circuit design along with the synthesized circuits. The integrated circuit may include transistors and other circuit elements (e.g., passive elements such as capacitors, resistors, inductors, etc.), as well as interconnects between the transistors and the circuit elements. Some embodiments may implement multiple integrated circuits coupled together to implement the hardware circuit, and / or discrete elements may be used in some embodiments. Alternatively, the HDL design may be synthesized into a programmable logic array such as a field programmable gate array (FPGA) and implemented in the FPGA. This decoupling between the design of a set of circuits and the subsequent low-level implementation of those circuits often leads to scenarios where the circuit or logic designer never specifies a particular set of structures for the low-level implementation beyond a description of what the circuits are configured to do, because that process is performed at different stages of the circuit implementation process.

[0677] The fact that many different low-level combinations of circuit elements can be used to achieve the same specifications of a circuit results in a large number of identical structures for that circuit. As noted, these low-level circuit implementations can vary depending on variations in manufacturing technology, the foundry chosen to manufacture the integrated circuit, the cell libraries available for a particular project, and so on. In many cases, the selection made by different design tools or methodologies to produce these different implementations can be arbitrary.

[0678] Furthermore, for a given embodiment, a single implementation of a particular functional specification of a circuit typically includes a large number of devices (e.g., millions of transistors). Consequently, the shear volume of this information makes it impractical to provide a complete description of the low-level structure used to implement a single embodiment, let alone the large number of equivalent possible implementations. For this reason, the present disclosure describes the structure of the circuit using functional shorthand commonly used in the industry.

Claims

1. A device, comprising: A graphics processor circuit, the graphics processor circuit comprising: a trace slot circuit implementing entries for a plurality of trace slots of the graphics processor circuit; queue access circuitry configured to access a data structure in the memory specifying a plurality of queues, wherein a corresponding queue queues control information for a plurality of sets of graphics jobs; a queue selection circuit configured to select a set of graphics jobs from the data structure based on one or more selection parameters and to store control information for the selected set of graphics jobs in a tracking slot of the tracking slot circuit; and Distribution circuitry is configured to assign portions of respective sets of graphics work from the tracking slots to graphics processor circuitry for execution.

2. An apparatus according to claim 1, wherein the selection parameter comprises a work category parameter for a given set of graphics work included in the data structure, wherein the work category parameter indicates whether the given set of graphics work is compute work, fragment work, or geometry work. 3 . The apparatus of claim 1 , wherein the selection parameter comprises a resource availability parameter provided by the graphics processor circuitry, wherein the resource availability parameter indicates availability of different graphics processor hardware resources.

4. The apparatus of claim 1, wherein the selection parameter comprises a priority parameter for a given queue included in the data structure, wherein different queues have different priority parameter values.

5. The apparatus of claim 1, wherein the selection parameters include deadline parameters for a given set of graphics jobs included in the data structure.

6. The apparatus of claim 1, wherein the selection parameters include dependency parameters for a given set of graphics jobs included in the data structure, wherein the dependency parameters indicate one or more other sets of graphics jobs that the given set of graphics jobs depends on.

7. The device according to claim 6, wherein: The one or more dependency parameters indicate at least one parent job that is in a different queue than the given set of graphics jobs.

8. The apparatus of claim 7, wherein the queue selection circuitry is configured to suspend selection from a first queue and select a set of graphics work from the second queue in response to a dependency parameter indicating a dependency of a set of work in a first queue on a set of work in a second queue.

9. The apparatus according to claim 6, wherein: The dependency parameters for a given set of graphics jobs include a set of parent identifiers and a valid parent mask indicating which parent identifiers are valid; and The dependency parameters are specified in a manner that is independent of the tracking slot to which the set of graphics jobs is assigned.

10. The apparatus according to claim 1, wherein: The selection parameters include event flag parameters for a given set of graphics jobs included in the data structure; The event flags parameter indicates one or more software programmable event flags on which the set of graphics jobs depends; and The queue selection circuit is configured to wait for one or more event flags to suspend selection from the first queue based on the set of graphics jobs. The apparatus of claim 10 , wherein at least one of the event flags is encoded as a counter value.

12. The apparatus according to claim 10, wherein: A set of graphics jobs that are independent of each other, are included in at least two different queues, and are targeted to a first hardware resource; and The control circuit is configured to set an event flag for the launched set of graphics jobs in the set to acquire the first hardware resource, and prevent other sets of graphics jobs from being launched until the event flag is cleared.

13. The apparatus according to claim 10, wherein: Another processor of the apparatus is configured to perform image processing on graphics frame data generated by the graphics processor circuit; and The apparatus is configured to execute program instructions to utilize an event flag parameter to indicate when a subset of a frame of graphics data has completed processing by the other processor and is ready for use by the graphics processor circuitry.

14. The apparatus of claim 10, wherein the apparatus is configured to control at least one event based on tasks performed by one or more other circuit components of the apparatus external to the graphics processing circuit.

15. The apparatus of claim 1 , wherein the graphics processor circuit comprises a control circuit configured to control a portion of a rendering process, including: Stop all collections of graphics work that depend on the geometry work collection; scheduling a first copy of the fragment working set to operate on data from the geometry working set; restarting the geometry working set after executing the first copy of the fragment working set; as well as The second copy of the fragment working set is configured to continue from the partially rendered image generated by the first copy of the fragment working set.

16. The apparatus of claim 1 , wherein the graphics processor circuit comprises a control circuit configured to: Writes result data for the completed set of graphics work to a completion queue structure in memory. 17 . The apparatus of claim 16 , wherein the control circuitry is configured to write result data to a plurality of different completion queues having different priorities.

18. The apparatus of claim 16, wherein the control circuitry is configured to write result data to different completion queues based on whether a given set of graphics work completed normally.

19. The apparatus of claim 16, wherein the control circuit is further configured to, in response to completion of a set of graphics jobs, program a set of registers indicated by the set of graphics jobs.

20. The apparatus of claim 1, wherein the graphics processor circuit is configured to execute firmware to provide one or more of the sets of graphics work for queues in the data structure.

21. The apparatus according to claim 1, further comprising: a memory circuit configured to store the data structure; and A processor circuit is configured to execute program instructions to add control information to a queue of the data structure.

22. The apparatus of claim 1, wherein the graphics processor circuitry comprises control circuitry configured to control which sets of graphics work trigger interrupts upon completion.

23. The apparatus of claim 1 , wherein the graphics processor circuit comprises: a plurality of single instruction multiple data pipelines, the plurality of single instruction multiple data pipelines being configured to execute instructions; and Fixed function circuitry configured to control the single instruction multiple data pipeline to perform operations for at least one of the following types of programs: Graphics shader programs; and Machine learning program.

24. The apparatus of claim 1, wherein the apparatus is a computing device, the computing device further comprising: monitor; Central processing unit; and Network interface.

25. A non-transitory computer-readable storage medium having stored thereon design information, the design information specifying a design of at least a portion of a hardware integrated circuit in a format recognizable by a semiconductor manufacturing system, the semiconductor manufacturing system configured to use the design information to produce the circuit according to the design, wherein the design information specifies the circuit, the circuit comprising: A graphics processor circuit, the graphics processor circuit comprising: a trace slot circuit implementing entries for a plurality of trace slots of the graphics processor circuit; queue access circuitry configured to access a data structure in the memory specifying a plurality of queues, wherein a corresponding queue queues control information for a plurality of sets of graphics jobs; a queue selection circuit configured to select a set of graphics jobs from the data structure based on one or more selection parameters and to store control information for the selected set of graphics jobs in a tracking slot of the tracking slot circuit; and Distribution circuitry is configured to assign portions of respective sets of graphics work from the tracking slots to graphics processor circuitry for execution.

26. The non-transitory computer-readable storage medium of claim 25, wherein the selection parameters include dependency parameters for a given set of graphics jobs included in the data structure, wherein the dependency parameters indicate one or more other sets of graphics jobs that the given set of graphics jobs depends on.

27. The non-transitory computer-readable storage medium of claim 25, wherein: The selection parameters include event flag parameters for a given set of graphics jobs included in the data structure; The event flags parameter indicates one or more software programmable event flags on which the set of graphics jobs depends; and The queue selection circuit is configured to wait for one or more event flags to suspend selection from the first queue based on the set of graphics jobs.

28. The non-transitory computer-readable storage medium of claim 27, wherein: another processor configured to perform image processing on graphics frame data generated by the graphics processor circuit; and The circuitry is configured to execute program instructions to utilize an event flag parameter to indicate when a subset of a frame of graphics data has completed processing by the other processor and is ready for use by the graphics processor circuitry.

29. A method comprising: accessing, by a computing device, a data structure specifying a plurality of queues, wherein a corresponding queue queues control information for a plurality of sets of graphics jobs; selecting, by the computing device, a set of graphics jobs from the data structure based on one or more selection parameters, and storing, by the computing device, control information for the selected set of graphics jobs in a tracking slot circuit; as well as Portions of a respective set of graphics work from the trace slot circuits are assigned for execution by the computing device.

30. The method of claim 29, wherein: The selection parameters include event flag parameters for a given set of graphics jobs included in the data structure; The event flags parameter indicates one or more software programmable event flags on which the set of graphics jobs depends; and The method also includes pausing, by the computing device, selection from the first queue based on the set of graphics jobs waiting for one or more event flags.

31. The method of claim 29, further comprising: stopping, by the computing device, all sets of graphics work that are dependent on the geometry work set; scheduling, by the computing device, a first copy of the fragment working set to operate on data from the geometry working set; restarting, by the computing device, the geometry working set after executing the first copy of the fragment working set; as well as A second copy of the fragment working set is configured, by the computing device, to continue from the partially rendered image produced by the first copy of the fragment working set.

32. The method of claim 29, wherein the selection parameters include dependency parameters for a given set of graphics jobs included in the data structure, wherein the dependency parameters indicate one or more other sets of graphics jobs that the given set of graphics jobs depends on.

Citation Information

Patent Citations

  • United states graphics processor techniques with split between workload distribution control data on shared control bus and corresponding graphics data on memory interfaces

    US11847489B2

  • Initial object shader run for graphics workload distribution

    US12182926B1

  • Compute Kernel Parsing with Limits in one or more Dimensions

    US20220083377A1