Programming Model for Resource-Constrained Scheduling
By introducing a semaphore-based resource management scheduling model into the GPU computing programming model, the problem of irregular parallel computing workloads being inefficient in the existing technology is solved, and more efficient resource utilization and performance improvement is achieved.
Patent Information
- Application Number
- CN202180003525.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-20
- Filing Date
- 2021-03-17
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-03-17
AI Technical Summary
Existing GPU computing programming models are inefficient in handling irregular parallel computing workloads, resulting in insufficient resource utilization, reduced performance, reduced bandwidth and waste of power.
A new scheduling model is adopted, which is based on the availability semaphore of abstract hardware resources for resource management, allowing programmers to organize data marshalling in a more flexible way, and achieve work amplification and aggregation through data structures such as lock-free algorithms and queues to avoid deadlocks and context switching.
Improves the efficiency of GPU computing processing, reduces external memory traffic, and achieves higher performance expansion, suitable for handling irregular parallel computing workloads.
Smart Images

Figure CN113874906B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 992,872, filed Mar. 20, 2020, which is incorporated herein by reference. TECHNICAL FIELD
[0003] The technology herein relates to a Graphics Processing Unit (GPU), and more particularly, to scheduling work to be performed by a Graphics Processing Unit.
[0004] Background Art & Summary of the Invention
[0005] Figure 1 It is illustrated how many or most traditional Graphics Processing Units (GPUs) are capable of handling different kinds of workloads, such as graphics workloads and compute workloads. There are differences in how graphics workloads and compute workloads are scheduled.
[0006] A graphics workload typically includes a set of logical stages in the processing of graphics data related to coloring pixels of an image. In this context, the coloring of a pixel typically produces visualization (brightness, color, contrast, or other visual characteristics) information that affects the appearance of one or more pixels of the image.
[0007] A GPU typically uses a graphics pipeline to process graphics workloads. Figure 1 The left side shows a typical graphics pipeline, including a series of shader stages 10 that communicate using pipeline on-chip memory 11. Thus, the graphics pipeline provides a sequence of logical stages for graphics processing as a pipeline sequence of shader stages 10a, 10b, …, 10n, where each shader stage 10 in the sequence performs its respective graphics computation. In Figure 1 the graphics pipeline shown on the left side, each shader stage 10 operates on the result of the previous stage in the pipeline and provides the processed result for further processing by a subsequent stage in the pipeline. Graphics processing distributed across hardware defined by a hardware-based graphics pipeline has proven to be very effective and is the basis of most modern GPU designs.
[0008] Thus, such a graphics programming model is built around the concept of a fixed-function graphics pipeline that is finely scheduled in a producer-consumer manner. In this context, each stage 10 of the graphics pipeline can be individually scheduled in a pipeline fashion. The shader stages 10 in the pipeline are initiated based on the presence of data generated by earlier stages and the availability of output space in the fixed-function first-in-first-out memory (FIFO) for later stages. Data marshalling is provided by a set of hardware-based triggers provided by a well-defined system to ensure that data packets flow from one shader stage to the next in an orderly manner and that subsequent stages do not start processing data before the previous stage has generated the data. Additionally, data marshalling is typically supported by an explicit definition of the number of input data attributes used by a dedicated graphics pipeline hardware stage (e.g., vertex shader) and the number of output data attributes it generates.
[0009] In such graphics processing, in many (although not all) embodiments, the memory 11 used to transfer data between stages in such a graphics pipeline includes on-chip local memory that can be accessed quickly and efficiently. This pipeline scheduling approach using on-chip local memory has proven to be very effective for processing graphics workloads as it avoids using external memory bandwidth and associated memory latency to achieve an orderly, formatted data flow between pipeline stages.
[0010] By analogy, multiple batches of laundry are planned to be washed and dried in a pipeline. The washing machine cleans each new load. After the washing machine has finished cleaning a batch of laundry, the laundry worker moves the laundry to the dryer and starts the dryer, then moves a new batch of laundry into the washing machine and starts the washing machine. In this way, when the washing machine has finished washing the new laundry, the dryer will dry the laundry. After the dryer is done, the laundry worker pours the dried laundry into a basket or bin for ironing and / or folding, and irons / folds the laundry while the washing machine and dryer process earlier loads throughout the sequential process. In a modern GPU graphics pipeline, the "laundry worker" is part of the system and software developers can simply assume that it will provide the necessary data marshalling to ensure an orderly data flow and properly sequenced processing through the pipeline. It should be noted, however, that modern graphics pipelines are not limited to a linear topology, and the defining characteristic of the pipeline architecture is related to the way data is marshalled between successive processing stages rather than to any particular topology.
[0011] In the early days of GPU design, this fixed-function pipeline graphics processing was typically the only workload that the GPU could support, and all other functions were executed in software running on the CPU. The first step in GPU programming was modest: making some shaders programmable and enhancing the hardware to support floating-point operations. This opened the door to executing some non-graphical scientific applications on the GPU graphics pipeline hardware. See, e.g., Du et al., “From CUDA to OpenCL: Towards a High-Performance Portable Solution for Multi-Platform GPU Programming,” Parallel Computing. 38(8):391–407 (2012). Graphics APIs such as OpenGL and DirectX were subsequently enhanced to represent some general-purpose computing functions as graphics primitives. Additionally, more general-purpose computing APIs were developed. See, e.g., Tarditi et al., “Accelerators: Programming General-Purpose GPUs with Data Parallelism,” ACM SIGARCH Computer Architecture News, 34(5) (2006). NVIDIA's CUDA allowed programmers to ignore underlying graphics concepts and instead adopt more common high-performance computing concepts, paving the way for Microsoft's DirectCompute and Apple / Khronos Group's OpenCL. See Du et al. for further upgrades to GPU hardware to provide high-performance general-purpose computing on the graphics processing unit (“GPGPU”).
[0012] Thus, modern GPU computing processing involves operating according to general-purpose application programming interfaces (APIs) such as CUDA, OpenCL, and OpenCompute for general-purpose digital computing. These computing workloads can be workloads not defined by programmable or fixed-function graphics pipelines, e.g., defining a set of logical computational stages for processing data not directly related to shaded pixels in an image. Some example computational operations can include, but are not limited to, physical computations associated with generating animated models, analyzing large data sets from scientific or financial domains, deep learning operations, artificial intelligence, tensor computations, and pattern recognition.
[0013] Currently, unlike the graphics programming model, the GPU computing programming model is built around the concept of data parallel computing, specifically flat bulk synchronous data parallel (BSP). See, e.g., Leslie G. Valiant, “A Bridging Model for Parallel Computation,” Communications of the ACM, Vol. 33, No. 8, August 1990. Figure 1 Shown are example GPU computing processing stages 12a, 12b, …, 12m on the right. Each parallel computing processing stage can be defined by launching multiple parallel execution threads for execution by the GPU's massively parallel processing architecture. A broad collection of work 12 is launched one at a time and communicated via main memory or other global memory 13 (see Figure 2)。In one embodiment, a work set can include any general computing operation or program defined or supported by software, hardware, or both.
[0014] In such computing models, there is typically a hierarchy of execution of the computational workload. For example, in one embodiment, a "grid" can include a set of thread blocks, each thread block including a set of "warps" (using a fabric analogy), and these warps in turn include individually executing threads. The organization of the thread blocks can be used to ensure that all threads within a thread block will run concurrently, which means they can share data (scheduling guarantee), communicate with each other, and work together. However, it may not be guaranteed that all thread blocks within a grid will execute concurrently. Instead, depending on the available machine resources, one thread block may start execution and run to completion before another thread block is launched. This means that there is no guarantee of concurrent execution between thread blocks within a grid.
[0015] Generally, grids and thread blocks are not executed in conventional graphics workloads (although it should be noted that some of these general distinctions may not necessarily apply to certain advanced graphics processing techniques, such as the grid shaders described in US20190236827, which introduce a computational programming model into the graphics pipeline, as threads are used in concert to directly generate a compact grid (small grid) on the chip for use by the rasterizer). Additionally, traditional graphics processing generally does not have a mechanism to locally share data between specific selected threads and guarantee the concurrency of the thread block or grid model supported by the computational workload.
[0016] As Figure 2 shown, shared (e.g., main or global) memory 13 is typically used to transfer data between computational operations 12. Each grid can read from and write to the global memory 13. However, in addition to providing a backing memory that each grid can access, there is no system-provided marshaling or any mechanism to transfer data from one grid to the next. The system does not define or constrain the data input to or output from the grids. Instead, the computational application itself defines all other aspects of data input / output, synchronization, and marshaling of data between grids, e.g., completing the results of one grid processing (or at least those specific results on which subsequent processing depends) and storing them in memory before the start of subsequent grid processing.
[0017] Such data marshaling involves how data groups are passed from one computational process 12a to another computational process 12b and addresses various issues related to the transfer of payloads from one computational process to another. For example, data marshaling may involve transferring data from producers to consumers and ensuring that consumers can identify the data and have a consistent view of the data. Data marshaling may also involve ensuring certain functions and / or processing, such as cache coherence and scheduling, so that consumer computational processes can consistently access and operate on cached data that has been generated by producer computational processes. Data marshaling may or may not involve or require moving or copying data, depending on the application. Under modern GPU computing APIs, all of this is left to the application to handle. While this provides great flexibility, there are also some drawbacks.
[0018] Data marshaling typically involves providing some synchronization from one grid to the next. If each grid is independent of all other grids and can be processed independently, little synchronization is required. This is a bit like cars on a congested toll road passing through a row of toll booths, where each car can pass through its respective toll booth independently without waiting for any other car. However, some computational workloads (e.g., certain graphics algorithms, sparse linear algebra, and bioinformatics algorithms, etc.) exhibit "irregular parallelism", which means that the processing of some consumer grids may depend on the processing of one or more producer grids.
[0019] Under the computing API, the application itself is responsible for synchronizing with each other to organize the data flow between them. For example, if a consumer grid wishes to view and use the data generated by a producer grid, the application is responsible for inserting barriers, fences, or other synchronization events to provide a consistent view of the data generated by the producer grid to the consumer grid. For example, see U.S. Patent Application No. 16 / 712236, filed on December 12, 2019, titled "High-Performance Synchronization Mechanisms for Coordinating Operations on a Computer System", USP9223578, USP9164690, and US20140282566, which describe example ways in which grids 12 can synchronize with each other and communicate data through the global memory 13 using reweighted synchronization primitives (such as barriers or fences). Such fence / barrier synchronization techniques can, for example, provide synchronization that requires the first computational grid to finish writing data to memory before the next computational operation accesses the data. According to the bulk synchronous programming model, synchronization is done in batches per grid, which typically means that all threads of the producer grid must finish execution and write their data results to the global memory before the consumer grid begins processing. The resulting suboptimal utilization of GPU processing resources can lead to significant time waste, reduced performance, reduced bandwidth, and power waste, depending on the computational workload.
[0020] Due to these different issues, for certain types of computational workloads, the execution behavior exhibited by current computational programming models can be inefficient. There have been several attempts in the past to improve this by allowing applications to express such "irregular" parallel workloads. For example, NVIDIA introduced CUDA nested parallelism using the Kepler architecture, targeting irregular parallelism in HPC applications. See, e.g., USP8180998; and Zhang et al., "Taming Irregular Applications via Advanced Dynamic Parallelism on GPUs" CF'18 (May 8 - 10, 2018, Ischia, Italy). Another example project looked at queue-based graphics programming models. Intel's Larrabee GPU architecture (not released as a product) introduced Ct, which has the concept of "woven" parallelism. See, e.g., Morris et al., "Kite: Woven Parallelism for Heterogeneous Systems" (Computer Science 2012). All of these attempts have had relatively limited success and have not necessarily tried to capture the aspects of scheduling and data marshalling provided by the graphics pipeline system for computational workloads.
[0021] To further explain, assume one grid is the consumer and another grid is the producer. The producer grid will write data to global memory, and the consumer grid will read that data from global memory. In the current bulk synchronous computational model, these two grids must run serially - the producer grid must complete entirely before the consumer grid starts. If the producer grid is large (multiple threads), a few threads may be scattered and take a long time to complete. This will result in inefficient use of computational resources because the consumer grid must wait until the scattered threads are complete before it can start. As the threads of the producer grid gradually exit, the machine occupancy drops, leading to inefficiencies compared to the case where all threads in both grids could run concurrently. Aspects of the techniques described herein avoid such inefficiencies and achieve continuous occupancy, for example, in a typical graphics pipeline, by leveraging continuous simultaneous producer / consumer to gain occupancy benefits.
[0022] Figure 3 Shows additional expansion and contraction support (work magnification and aggregation) for common GPU pipeline graphics processing. Many GPUs replicate certain shader stages (e.g., tessellation) to enable parallel processing of system calls on the input data of the same stage as needed, to avoid bottlenecks and increase data parallelism, which is very useful for certain types of graphics workloads. This functionality of system calls has not been available in the past for computational workloads supported by the GPU system scheduler.
[0023] Accordingly, there is a need for a new model that allows computational workloads to be scheduled in a way that captures the scheduling efficiency of the graphics pipeline. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The following detailed description of exemplary non - limiting illustrative embodiments will be read in conjunction with its accompanying drawings:
[0025] Figure 1 Schematically shows the separate personalities of graphics processing units that perform computing processing and graphics processing.
[0026] Figure 2 Shows how shader stages communicate with each other using on - pipeline chip memory, and how computing processes communicate with each other using global memory and synchronization mechanisms.
[0027] Figure 3 Shows non - limiting examples of expansion and contraction support for graphics processing, with limitations on such support for computing processes.
[0028] Figure 4 Shows an example non - limiting scheduling model.
[0029] Figure 5 Shows an example non - limiting scheduling model that uses startup - guaranteed resources.
[0030] Figure 6 Shows an example of resource - constrained scheduling between producers and consumers.
[0031] Figure 7 Shows an example of using a startup threshold Figure 6 for resource - constrained scheduling.
[0032] Figure 8 Shows a resource - constrained scheduling example that includes production and consumption feedback Figure 6 of.
[0033] Figure 9 Shows an example non - limiting scheduler that starts when the input is ready and the output is available.
[0034] Figure 9A Shows a more detailed view of graphics execution startup.
[0035] Figure 10 Shows resource - constrained scheduling using a queue.
[0036] Figure 11 Shows an example of resource - constrained scheduling using data marshalling with a queue.
[0037] Figure 12 Shows an example of resource - constrained scheduling using internal and external put and get.
[0038] Figure 13 Shows an example of non - limiting resource release.
[0039] Figure 14 Shows a non - restrictive example of resource release in the case of early release if the producer does not generate output.
[0040] Figure 15 Shows a non - restrictive example of resource release if subsequent consumers require visibility.
[0041] Figure 16 Shows a non - restrictive example of resource release including more fine - grained tracking of resource usage.
[0042] Figure 17 Shows a non - restrictive example of resource release when partial output is produced.
[0043] Figure 18 Shows a non - restrictive example of resource release when the producer releases unused space in the queue.
[0044] Figure 19 Shows a non - restrictive example of resource release when the consumer releases the remaining space.
[0045] Figure 20 Shows an example of non - restrictive resource release to allow a new producer to start production earlier.
[0046] Figure 21 Shows a non - restrictive example of resource release to allow variable - sized output in a related partial release.
[0047] Figure 22 Shows an example queue memory allocation scheme to avoid fragmentation.
[0048] Figure 23 Shows a non - restrictive example of work amplification where each entry starts multiple consumers.
[0049] Figure 24 Shows an example non - restrictive multi - bucket (1 extra) startup scheme to start multiple different consumers on each entry.
[0050] Figure 25 Shows an example of non - restrictive work aggregation where a group of work items is aggregated into a single consumer startup.
[0051] Figure 26 Shows an example of non - restrictive work aggregation including partial startup timeout results.
[0052] Figure 27 Shows a connection allocation example that includes a set of separately computed fields that are aggregated into a single work item.
[0053] Figure 28Shows an example of using aggregation to sort work into multiple queues for re - convergent execution.
[0054] Figure 29 Shows exemplary non - restrictive work sorted into multiple queues for shader specialization.
[0055] Figure 30 Shows exemplary non - restrictive work sorted into multiple queues for coherent material shading in ray tracing.
[0056] Figure 31 Shows an exemplary non - restrictive dependency scheduling.
[0057] Figure 32 Shows an exemplary non - restrictive computational pipeline to graphics.
[0058] Figure 33 Shows exemplary non - restrictive computations that feed graphics to achieve a fully flexible compute front - end.
[0059] Detailed Description of Exemplary Non - Restrictive Embodiments
[0060] This technology extends the compute programming model to provide data marshalling features available in some systems of the graphics pipeline, thereby improving efficiency and reducing overhead. In particular, the techniques herein allow programmers / developers to create computational kernel pipelines in ways not achievable with currently available GPU compute APIs.
[0061] Exemplary non - limiting embodiments provide a simple scheduling model based on scalar counters (e.g., semaphores) that represent the availability of abstract hardware resources. In one embodiment, the system scheduler responsible for starting new work reserves / acquires resources ("idle" space or "ready" work items) by decrementing the corresponding semaphore, and user code (i.e., the application) releases resources (consumers: by adding back to the "idle" pool of the resource pool from which they read, producers: by incrementing back to the "ready" pool of the resource to which they are writing). In such an exemplary non - limiting arrangement, reserve / acquire can always be conservatively done by the scheduler decrementing the semaphore, and resource release is always done by the user (application) code since it determines when it is appropriate to perform the release and always does so by incrementing the semaphore. In such an embodiment, resource release can thus be done programmatically, and the system scheduler only needs to keep track of the state of such counters / semaphores to make work - start decisions (a fallback provision can be provided in case the user or application software fails to do this as expected so that the scheduler can increment the semaphore and thus release the resource). The release of resources does not have to be done by the user's application; it can be implemented by some system software, e.g., that can ensure all the correct accounting is done. In one embodiment, the semantics of the counter / semaphore are defined by the application, which can use the counter / semaphore to represent, e.g., the availability of free space in an available memory buffer, the cache pressure caused by data flows in a network, or the presence of work items to be processed. In this sense, our approach is a "resource - constrained" scheduling model.
[0062] A novel aspect of this approach is that the application can flexibly organize custom data marshalling among nodes in the network, using efficient lock-free algorithms, using queues or other data structures chosen by the application. By decoupling the management of hardware counters / semaphores to drive scheduling decisions for the management of such data structures, we have implemented a framework for expressing work amplification and aggregation, which is a useful aspect of graphics pipeline scheduling and helps in efficiently processing workloads with different amounts of data parallelism (such as compute workloads). A resource reservation system is provided that allocates resources at startup to ensure that thread blocks can run to completion, thus avoiding deadlocks and expensive context switches. The resource reservation system is relatively simple and does not need to scale or grow depending on the number of work items from the same task / node or the number of concurrently executing thread groups. For example, in one embodiment, during such scheduling of any particular task, the scheduler only looks at two counters SF and SR for each resource (when there are multiple resources / phases in the pipeline, there will be or may be multiple pairs of counters) – this pair has some implications for the type of execution node graph that can be constructed. In the example embodiment, the semaphore itself is not considered part of the scheduler, but is provided by the hardware platform to support the scheduler. Thus, the scheduler and the application can each manipulate the semaphore in certain ways, and the scheduler can monitor the application's manipulation of the semaphore.
[0063] Another aspect involves tuning the expansion / shrink support of the graphics pipeline to make these concepts useful for compute workloads – providing dynamic parallelism.
[0064] When previously attempting to schedule computational workloads that exhibit irregular parallelism, there has generally been no attempt to directly incorporate the concept of backpressure into the producer-consumer data flow, which aids in efficient scheduling and allows the data flow to stay on-chip. Backpressure - the ability of the consumer to slow down the producer to prevent the consumer from being overwhelmed by the data stream that the producer is generating - allows for the efficient transfer of large amounts of data through a fixed-size buffer. Since physical memory is a finite resource and is shared with many other parts of the application, it is highly desirable to be able to allocate a fixed-size buffer whose size is less than the potential number of work items being transferred through the buffer. Graphics pipelines typically support such backpressure scheduling, for example, Jonathan Ragan Kelley, "Keeping Many Cores Busy: Scheduling the Graphics Pipeline, Beyond Programmable Shading II" (Thursday, July 29, 2010, SIGGRAPH); Kubisch, GPU-Driven Rendering (NVIDIA, GTC Silicon Valley, April 4, 2016); Node.js, "Backpressure in Streams", https: / / nodejs.org / en / docs / guides / backpressuring-in-streams / . Another option is to allocate a buffer large enough between the producer and the consumer to accommodate the "worst case", but the parallelism of the machine only allows a small fraction of this "worst case" allocation to be active at the same time - the rest of the worst-case memory allocations will be wasted or idle. Backpressure scheduling allows for the efficient use of fixed-size memory allocations, minimizing waste.
[0065] Example non-limiting methods can also explicitly avoid any ongoing work that requires a context switch. For example, in modern GPUs, the amount of state that needs to be saved and restored for a context switch can be too large to be used in a high-performance environment. By making the scheduler directly aware of the resource constraints imposed in the chip, such as the amount of on-chip buffering available, our method enables the entire system to flow through the workload while minimizing the external memory bandwidth required to transfer transient data.
[0066] As we expand the range of applications processed by processors such as GPUs into new areas such as machine learning and ray tracing, the current methods will help to advance the core programming models for efficiently processing such workloads, including those that exhibit irregular parallelism. In particular, the example features of our model allow for the reduction and / or complete elimination of external memory traffic when scheduling complex computational workloads, thus providing a way to extend performance beyond the limits of external memory bandwidth.
[0067] Another aspect includes APIs (Application Programming Interfaces) that support the technology. Such APIs can exist on many different hardware platforms, and developers can use such APIs to develop applications.
[0068] Another aspect includes the advantageous use of a memory cache to capture the data stream between computing levels more quickly with memory, thereby using on-chip memory for data communication between graphics pipeline levels. This allows for optimization by leveraging on-chip memory, thus avoiding expensive global (e.g., frame buffer) off-chip memory operations by pipelining the computational data stream, increasing bandwidth, and computational power.
[0069] Yet another aspect relates to data marshalling, which arranges work at a finer granularity. Instead of starting all the threads in a grid or other thread sets simultaneously and waiting for them all to complete, we can start selected threads or thread groups in a pipeline fashion, thereby achieving fine-grained synchronization by not synchronizing in batches. This means we can run producers and consumers simultaneously - just like in a pipeline model.
[0070] Exemplary non - restrictive scheduling model
[0071] Figure 4 An exemplary non - restrictive scheduling model 100 is shown schematically, showing the basic unit of scheduling. In the shown scheduling model 100, an input payload 102 is provided to a thread block 104, and the thread block 104 generates an output payload 106. One of the features of the scheduling model 100 is a system - defined programming model that provides explicit definitions for the input payload of the thread block 104 and the output payload from the thread block. In one embodiment, the explicit definitions of the payloads 102, 106 define the size of the data grouping.
[0072] For example, the definition can conform to a general program model that includes an explicit system statement that thread block 104 will consume an N-byte input payload 102 and produce an M-byte output payload 106. Thread block 104 (as opposed to the system in one embodiment) will be concerned with the meaning, interpretation, and / or structure of the input and output data payloads 102, 104, and for each invocation of the thread block, the amount of data that the thread block actually reads from its input payload 102 and writes to its output payload 106. On the other hand, the system in one embodiment does know that, according to the scheduling model, the input payload 102 is for the input of thread block 104 and the size of this (opaque to the system) input data payload, and similarly, the output payload 106 is for the output of the thread block and the size of this (opaque to the system) output data payload. For each thread block 104 and its respective input payload 102 and output payload 106, there is this explicit statement of input and output dependencies.
[0073] Figure 5 Shows the use of such a definition in one embodiment to enable the system to schedule thread blocks 104 in a very efficient pipeline. Well-defined inputs and outputs allow for synchronization within the object scope and obviate the need for synchronization of unrelated work. Such well-defined input and output declarations increase occupancy and enable resources to be acquired in a conservative manner (as shown by the "acquire" arrow) before kernel launch and released programmatically (as shown by the "release" arrow) before execution completion. Figure 5 Shows that, in one embodiment, the scheduler performs resource acquisition before thread block launch (i.e., subtracts a certain known amount pre-declared by the thread block from the SF semaphore to reflect the size of the output payload, thereby reserving the amount of resources required for the thread block), while the thread block (requesting) itself can release resources programmatically after resource usage and before execution completion, and the scheduler will record and track the appropriate release amount, which can be calculated programmatically - to provide a deadlock-free guarantee, avoid the need for context switching, and achieve additional flexibility. In one embodiment, the release can be performed at any point in the program flow and can be performed at different time increments for a given thread block to enable the thread block to reduce its resource reservation when it no longer requires as many resources. In such an embodiment, like a group of diners leaving the table after finishing their meals, it is up to the thread block to decide when to release resources; the system-based scheduler will identify and utilize the release by retaining the release amount. The aggregate operations of how much resource is acquired and how much is released should be consistent to ensure that the resources acquired each time are ultimately released, avoiding cumulative errors and related deadlocks and the halting of the forward process. However, in one embodiment, when and how to acquire and release resources can be determined programmatically by the thread block but monitored, enforced, and utilized by the system-based scheduler.
[0074] Figure 5 Examples show some abstract input resources 107 and output resources 108 (e.g., on-chip memory, main memory allocations for the system to materialize data marshalling for thread blocks 104, hardware resources, computing resources, network bandwidth, etc.). In some embodiments, the size / range of such input and output resource 107, 108 allocations is fixed, i.e., a finite number of bytes in the case of memory allocation. Since the system knows the sizes of the inputs and outputs 102, 106 for each individual thread block 104, it can calculate the number of thread blocks that can run in parallel and schedule the start of the thread blocks based on whether the thread block is a producer or a consumer. For example, if two thread blocks 104a, 104b participate in a certain producer / consumer relationship, the system can schedule the producer and know which output resource 108 the producer will use, and then can schedule the consumer thread block from that output resource. By knowing the sizes of the input payload 120 and the output payload 106, the system can use optimal resource constraint scheduling to correctly schedule the pipeline. See, for example, Somasundaram et al., "Node Allocation in Grid Computing Using Optimal Resource Constraint (ORC) Scheduling," IJCSNS International Journal of Computer Science and Network Security, Vol. 8, No. 6 (June 2008).
[0075] Although the most common parameter related to the amount of input and output resources 107, 108 will be the memory size, the embodiments are not limited thereto. The input and output resources 107, 108 can be, for example, network bandwidth, communication bus bandwidth, computing cycles, hardware allocation, execution priority, access to input or output devices such as sensors or displays, or any other shared system resource. As an example, if multiple thread blocks 104 need to share limited network bandwidth, a similar resource-constrained scheduling can be applied to the network bandwidth as a resource constraint. Each thread set or block 104 will declare the bandwidth it uses per invocation, enabling the system to schedule within the bandwidth constraint. In this case, a value or other metric will be used to specify the amount of network bandwidth required (i.e., how much network bandwidth resources the thread block will consume upon invocation). The amount of parallelism that the GPU can support will be a function of the resource size required to support the thread blocks 104. Once the system knows the size / amount of resources required, the system can optimally schedule the computational pipeline, including which thread blocks 104 will run concurrently, and allocate input / output resources before the producers and consumers start. In one embodiment, before starting a thread block, the system enables it to acquire space in the output resource 108 and track utilization accordingly. In an exemplary non-limiting embodiment, for speed and efficiency, such system operations are performed by programmable hardware circuits (e.g., hardware counters and associated logic gates implementing a set or pool of free and release counters / semaphores pairs). Such programmable hardware circuits can be programmed using traditional API calls by system software and applications running on the system.
[0076] Figure 6 Shows an example view of resource-constrained scheduling, using semaphores S F and S R to indicate the free and ready capacity of the constrained resources. S F can indicate the total capacity of the resource, and S R can indicate the capacity of the resource that has been allocated or otherwise claimed. Knowing what these values are, the system can calculate or otherwise determine how many (and which) thread blocks to start to avoid resource overload and, in some embodiments, optimize resource utilization.
[0077] Producer 110 and consumer 112 are connected through resource 114. Resource 114 can be a software construct, such as a queue. Each core defines a fixed-size input and output - synchronization within the scope of the object. More specifically, each core technically defines the maximum input or output; the producer can choose to produce up to this amount of product. The consumer may (in the merged startup case) use fewer work items at startup than the maximum amount. The system can assume that producer 110 and consumer 112 will always run to completion, and the system's hardware and / or software scheduling mechanism can monitor system operations and, based on the monitoring, maintain, update, and view S F and S R to determine how many thread blocks to start. In one embodiment, the system uses semaphores to ensure that input / output resources will be available when concurrent processes require these resources, thus ensuring that the processes can run to completion rather than deadlock, have to be preempted or paused, and wait for resources to become available, or require expensive context switches and associated overhead.
[0078] In one implementation, the system's hardware and / or software scheduling agent, mechanism, or process runs in the background independently of the consumer and producer. The scheduling agent, mechanism, or process views semaphores S F and S R , to determine how many producers and consumers the system can start at a given time step. For example, the scheduling agent, mechanism, or process can determine to start a producer only when there is available space and determine how many producers can be started based on the value of the available space semaphore S F . The scheduling agent, mechanism, or process similarly determines the number of valid consumer entries for the constrained resource based on the ready semaphore S R . Whenever a new thread block is started, the scheduling agent, mechanism, or process updates the semaphore, and whenever the thread block updates the semaphore before completion and termination, the scheduling agent, mechanism, or process also monitors the state change.
[0079] In many implementations, multiple producers and consumers will run concurrently. In this case, since the input / output resources are obtained when the producers and consumers start, the input / output resources are not clearly divided into occupied and unoccupied spaces. Instead, there will be a small portion of resources in the intermediate space. For example, a producer that has just started running has not yet produced valid output, so the ready semaphore S R will not register the producer's output as valid data, and the value of the semaphore will be less than the amount of memory space that the producer will ultimately need when the output is valid data. In one embodiment, the producer will only account for this actual data usage by incrementing the ready semaphore S R when the data produced by the producer becomes valid. The exemplary embodiment uses two semaphores SF and S R to illustrate the time sequence of these events connected with concurrent processing and resource acquisition at startup. For example, before starting a thread block, the system reserves the output space that the thread block will need in the I / O resource 114, thus ensuring that the space is available when needed and avoiding deadlock situations that may result from resource acquisition during later execution. It does this by borrowing the free semaphore S F to track this resource reservation. Once the thread block starts and begins to fill or otherwise use the resource, the amount of the resource being used moves to the ready state – tracked by updating the semaphore S R Allocating resources before startup not only avoids potential deadlocks, it also has the advantage of being able to use algorithms that do not need to handle the possibility of resource unavailability. This potential large simplification of the allocation phase of the data marshalling algorithm is an advantage.
[0080] A similar analysis applies to the consumer side of resource allocation. Before the system starts a consumer thread block, it checks the ready semaphore S R to ensure that it is non-zero. When the system starts a consumer thread block, it tracks that the consumer will consume resources in the future but has not consumed them yet by decrementing the ready semaphore S R Only after the consumer consumes the reserved resources can the system state convert the resources to available space, release the resources associated with the work item (in the example embodiment, this is done by the consumer itself), and increment the free space semaphore S F . Between reservation and release, the work item is in an indeterminate state.
[0081] In one embodiment, the scheduling agent, mechanism, or process can be simple because, as described above, it uses two semaphores to track resources. In the example implementation, the semaphores S F and S RIt can be implemented by a collection of hardware circuits such as registers and / or counters, which are inexpensive to maintain, update, and use. In each cycle or at other periodic or aperiodic frequencies, the system updates the registers / counters through simple and inexpensive hardware-based updates (e.g., incrementing and decrementing, adding or subtracting integer values to / from registers / counters, etc.) to account for resource consumption. Additionally, this implementation can be directly mapped to the old GPU scheduling implementation (due to the simplicity of the current scheduling technique, not requiring all the complexities of the old scheduler), so no additional hardware design is needed. In other implementations, additional simplified hardware circuits dedicated to the current scheduler can be provided, thus reducing chip area and power requirements. Since the current scheduler can be implemented in an abstract manner around a simple counter that tracks resource utilization, a simple scheduler can be achieved. The semaphore values can be updated through hardware or software mechanisms. If software instructions are used, the instructions can be embodied in the system code or application code. In one embodiment, the signals for requesting and releasing resource reservations are provided by the producer and consumer thread blocks themselves. Additionally, S F and S R increases / decreases can be performed programmatically in response to requests and release signals from producers and consumers, thus taking advantage of early resource release. Although S F and S R can be considered part of the scheduler, which ultimately controls these values, higher-level system components can also ensure that misbehaving producers and consumers do not tie up resources or stop progressing.
[0082] A sufficient number of such registers are provided to support the maximum number of tasks required for concurrent processing, i.e., representing the number of states needed to support all possible concurrently launched consumers and producers. In one embodiment, this means providing a pair of semaphores for each concurrently executing thread block or application. The scheduler can start any number of concurrent consumer and producer nodes in a streamlined manner, and the amount of state that the scheduler needs to track for resource reservation and usage does not depend on the number of concurrently executing thread groups from the same task / node. Instead, in one embodiment, the amount of state that the scheduler maintains for a given node does not grow accordingly and thus does not depend on the number of queued work items or tasks (e.g., threads) waiting to execute. The number of states maintained by the scheduler will increase with the number of resources to be allocated and the number of active shader programs (tasks) in the system, which in the pipeline scheduling model depends on the number of stages in the pipeline because the output resources of one stage typically include the input resources of subsequent stages in the pipeline. In one embodiment, the scheduler maintains one register for each "resource" (e.g., memory, cache, etc.) to represent S F , and one register for each "task" to represent S R .
[0083] However, the techniques herein are not limited to creating pipelines, but can support more complex relationships and dependencies. In such cases, the amount of state that the scheduler needs to maintain will depend on the topology of the node graph for which it is responsible for scheduling concurrent execution. As described above, since in some implementations the scheduler is designed to be simple and efficient, based on the two values S for each resource as described above F 、S R , it may not support arbitrary complexity / topology graphs, and in cases where such flexibility is required, another scheduling algorithm / arrangement can be invoked at the cost of increased scheduling complexity.
[0084] Furthermore, in the case where the allocated resource is memory, the scheduler is not limited to any memory type, but can be used with many different memory structures, such as linear memory buffers, stacks, associative caches, etc.
[0085] Because in order to keep the scheduler simple, the number of states maintained by the exemplary embodiment scheduler is of a fixed size, this results in a phenomenon we call backpressure. First, in the exemplary embodiment, the fixed-size constraint may mean that each resource uses only a pair of semaphores S F 、S R to control the allocation and release of resources, thereby controlling system concurrency. In one embodiment, each resource has only two values, so the amount of state maintained by the scheduler does not grow with the nature of the execution node topology graph that the scheduler is asked to schedule. In the exemplary implementation, each resource has a pair of such semaphores, and in a graph with multiple nodes and queues connecting them, each queue will have a pair of semaphores. When initializing these values S F 、S R at the start of system execution, they are initialized to the amount of resources available at those particular stages of the pipeline or graph topology. Then, as work is scheduled, the semaphores S F 、S R change (increase and decrease).
[0086] When a resource runs out, the particular node that needs that resource will stop being scheduled, while some other nodes will start executing because it has available resources and inputs.
[0087] In one embodiment, when work is started, the idle and ready semaphores S F 、S R are atomically modified along with the work log for that day. For example, starting work will atomically decrement the idle semaphore S F, thus preventing other work from starting under the false assumption that additional resources are available when in fact they have already been reserved. This is somewhat analogous to a maitre d' keeping track of the number of restaurant customers already drinking at the bar before seating, to ensure that they are allocated rather than new arrivals. This avoids the possibility of contention. Similarly, the dispatcher atomically decrements semaphore S F , in order to allocate resources to the producer. Additionally, in some embodiments, an application may release resources using a semaphore before completing the work. Thus, in an exemplary non - limiting embodiment, resource reservation is done by the scheduler, but resource release is done by the application / user code.
[0088] Figure 7 and Figure 8 shows an example of a one - to - one mapping between resource acquisition events and resource release events. Time progresses from left to right. Assume the resource is a buffer in memory. As Figure 7 shown, upon acquisition, S F is subtracted by the amount explicitly defined and declared by the producer thread block to indicate that that portion or amount of memory is being reserved for use. Figure 8 shows that the producer thread block will increment S R at some point in its life cycle to indicate that the producer thread block is now actually using the resource to store valid memory data. Since the data required by the consumer thread block is now valid, the scheduler can start the consumer thread block. The scheduler will also decrement S R when starting the consumer – completing the symmetry between the two semaphores and the producer / consumer tasks. At some point during execution, once the consumer consumes the resource, it can indicate to the system that it no longer needs the resource. Thus, the system converts the resource back to available space by adding the S F value in response to a release event executed by the application / consumer code, meaning that the scheduler can now start other producers. As Figure 13 shown, the scheduler will not start another producer until the previously started producer releases the memory reserved for it; thus, the next acquisition can only occur after the scheduler increments the S F value to indicate the amount of resources now available for re - acquisition. Thus, concurrency in the system is limited by the resource utilization life cycle.
[0089] Figure 9 is a flowchart of an example scheduler view that starts thread blocks based on whether the input is ready and the output is available, as shown by S R >0 and S F >0. In one embodiment, this functionality can be implemented in hardware. Multiple sets of hardware circuit registers can be provided, each register set corresponding to a task. AsFigure 9 As shown on the right, each register set (which may be referred to as a "task descriptor") contains, for example:
[0090] A reference to a task (e.g., a name or address that can be used to invoke a task instance);
[0091] S R A semaphore (uniquely associated with the task);
[0092] A reference to at least one S F semaphore (*S F ); in some embodiments, multiple taskless semaphore references are allowed to reference a common taskless semaphore S F ; other embodiments may also use a task identifier instead of a reference value to store the taskless semaphore).
[0093] Multiple tasks can reference the same semaphore, which is how implicit dependencies between tasks are implemented.
[0094] At each time unit or clock, the scheduler examines these semaphores and compares them to zero. If a task's semaphores are all greater than zero, the scheduler determines that it can start the task and does so, thereby decrementing the task's S F semaphore. When the task finishes using a resource, the task itself (in one embodiment) causes the hardware to increment S F and S R semaphores by using a mechanism that identifies which semaphores to increment. For ready semaphores, the task can send an increment signal to the task entry associated with the task. For free semaphores, the software will send an increment signal with an index to the free semaphore table. The software can reference semaphores by name and command them to increment. As will be understood, "increment" and "decrement" signals or commands can actually command an increase (addition) or decrease (subtraction) of any integer within a range.
[0095] At the software API level, objects that can be defined or declared by a developer or application programmer and serviced by the system software include:
[0096] At least one queue object, and
[0097] A task object that references the code to be run and the queue object to be used.
[0098] The system then maps the queues to semaphores and the tasks to hardware task descriptors.
[0099] The application will first specify the graph, i.e., the topology of the task nodes to be executed and their dependencies. In response to such a specification, the system software will create queues and initialize the hardware and the system. When the application triggers the execution of the graph, the GPU system software calls the root task of the graph and starts executing the graph in an autonomous manner separate from the CPU. See Figure 9A . When all queues are emptied and the graph execution is idle, the system can determine that the graph has completed. Another approach is to use tokens interpreted by the system to explicitly indicate the end of work. Figure 9A Shows three alternative ways to release resources: Option 1) Propagate an end-of-work token through the pipeline to specify the completed work; Option 2) Wait for the queue to become empty and any actively executing nodes to stop running; Option 3) The user code in the task determines when it is complete and signals (the most common scenario in the example embodiments).
[0100] Figure 10 Shows an example of resource-constrained scheduling using queues. As described above, the resource can be any type of memory structure, but it is often advantageous to use queues to transfer data from producers to consumers, which is similar to the model often used in a graphics pipeline. In some embodiments, semaphores and test logic are implemented in the hardware circuit, while the queues are software constructs and managed by the system software.
[0101] Figure 11 Shows how a producer can write data to a queue while a consumer reads data from the same queue. Thus, a queue can include a production section, a ready section, and a consumption section. The production section contains the valid data that the producer is currently writing, the ready section contains the valid data that the producer previously wrote to the queue and the consumer has not yet started consuming, and the consumption section can contain the valid data that the producer previously wrote and the consumer is currently reading.
[0102] Figure 12 Shows how an example scheduler uses queues to implement data marshalling. In this example, S F semaphore tracks how many free slots are in the queue, and S R semaphore tracks how many ready slots are in the queue (i.e., slots containing valid data that the consumer has not yet released). The scheduler uses these two semaphores to determine whether (when) to start new producers and consumers. Figure 12Four pointers are shown at the bottom: outer put, inner put, inner get, and outer get. The system software can use these four pointers to manage the queue. Although the semaphore at the top and the pointers at the bottom may seem to be tracking the same thing, in some embodiments, the semaphore is maintained and updated by the scheduling hardware circuitry, while the pointers at the bottom are maintained and updated by the application software. Decoupling the semaphore scheduling function from the queue management function in this way provides certain advantages, including enabling a variety of use cases. For example, looking at Figure 23 , the S R semaphores of two different producers can be incremented by different amounts. In one embodiment, the scheduler only looks at these semaphores to decide how many consumers to start. So, in this example, there is no one-to-one correspondence between producers and consumers, and the result is work amplification, i.e., multiple consumers are started for each entry (multiple instances of the same task, or different tasks, or both). This can be used to handle larger output sizes. Any number of consumers can be started for each entry, and all consumers "see" the same payload at their input. The consumers coordinate with each other to advance the queue pointer that manages the queue itself. Such data parallel operations can be used to enable different consumers to perform different work on the input data, or the same work on different parts of the input data, or more generally, the same or different work on the same or different parts of the input data. This functionality can be enabled by decoupling the management of the queue state from the new work scheduling function provided by the semaphore.
[0103] Figure 13 and 14 show that in some embodiments, if a producer does not need a resource at all, the producer can release the resource in advance. For example, assume a producer is started that conditionally generates an output. Before starting, the producer will declare the amount of resources it will need if it generates the output, and the scheduler handles this by subtracting from the S F the amount of resources the producer declares it needs, thus ensuring that the producer can run to completion and preventing another producer from competing for that resource. However, if the producer determines during execution and before completion that it no longer needs some or all of that resource (e.g., because it determines it will generate less or no output), the producer can indicate this to the system scheduler by releasing its reservation of the resource. The resource lifecycle is now much shorter, and the concurrency of the system can increase as the latency between releasing the resource and starting the next producer is reduced. Compared to the position in Figure 13 , Figure 14 moves the second producer to the left side of the figure, which shows this. Specifically, the first producer will increment the S F, which enables the scheduler to immediately allow the second producer to acquire resources so that the scheduler will again reduce S F to record the reservation for the second producer.
[0104] In addition, the release can be issued by the producer, the consumer, or even subsequent stages in the pipeline. Figure 15 Shows a pipeline of three stages where the final release operation is performed by the grandchild of the producer (i.e., the next consecutive stage after the producer in the pipeline sequence, a pipeline stage downstream in the pipeline). Thus, as long as the scheduler enforces consistent acquisition and release (i.e., the amount of resource reservation is always offset by the corresponding amount of resource release), resource reservation can be maintained across multiple pipeline stages.
[0105] Figures 16 - 20 Shows another example where the producer releases resources gradually. In Figure 16 , the producer acquires N units of resources. When producing partial output as shown in Figure 17 (e.g., the system reserves N time slots in the memory resource and the producer, but the producer only outputs data with a small number of time slot values shown by the cross-hatched portion, such as M - N time slot values), the producer can release the unused space M ( Figure 17 ) for use by other producers ( Figure 18 ). When the producer releases space M < N ( Figure 20 ), this means that it has stored some valid data in the resource but does not need the M time slots that the scheduler can reserve for another producer, which can now start to use the released space M ( Figure 19 ). In this example, the producer knows that N time slots are reserved for it and only uses M time slots (like a diner in a restaurant reserves a table for 8 people but only 6 people show up). In the case of partial release, the consumer is informed that it can only read N valid time slots ( Figure 20 ), so as shown in Figure 19 , the consumer reads and then releases a portion of the output (S Figure 20 += N - M), and the other portion M of the allocation is acquired by another producer. Thus, the semaphore S F , S F , S R can be used to indicate fractional resource usage, and the producer can return the unused fraction. Another resource release from an unrelated producer / consumer can come in a timely manner, so the scheduler can start another producer. This mechanism allows for handling variable-sized outputs up to a maximum size of N.
[0106] Figure 21 Shows that the usage of the resources being tracked by the scheduler can be of variable size.
[0107] Exemplary cache occupancy management
[0108] In one example embodiment, the resources acquired and released by the system scheduler can include caches, semaphores S F , S R Track the number of cache lines for the data flow between the producer and the consumer. Since many caches are designed to freely associate cache lines with memory locations, the memory locations are usually not contiguous, and the data fragmentation in the cache is quite unpredictable. However, the disclosed embodiment scheduler can use a single value S F to track and manage the use of the cache (and main memory), as this value only indicates how many cache lines the thread block can still acquire.
[0109] As Figure 22 shown, the non-limiting embodiment scheduler tracks the cache working set, i.e., the cross-hatched cache lines that are being used. The hardware cache is an excellent allocator, which means that the cache lines can be arbitrarily segmented. When the producer writes to the external memory through a fully associative cache, cache lines are allocated (similarly, if the producer does not write to part of the external memory, no cache lines are allocated). In a practical implementation, the cache is usually set to set associative, so they do not have perfect behavior in this regard, but are still quite close. What the system scheduler attempts to do is to schedule in an approximate way in which cache lines can be cheaply and dynamically allocated.
[0110] The disclosed example scheduler can ignore the way cache lines are allocated to fixed-size blocks of main memory within the cache because scheduling decisions are not based on actual or worst-case usage, but rather on the total number of cache lines that have been reserved / acquired and released, regardless of how those cache lines are allocated and potentially mapped to physical memory addresses. This means that allocating cache lines to fixed-size main memory blocks (an advantage in software development is the use of fixed-size memory block allocation) can be managed by the cache, and the system scheduler can track physical memory usage based only on the total number of cache lines used for physical memory allocation, without having to care which specific fixed-size memory blocks have been allocated or not. For example, a producer may only partially write to each block of a fixed-block memory allocation, and a consumer may only partially read. However, the scheduler does not need to track which blocks are read or written because the cache automatically keeps track of these blocks. Using an associative cache to track memory mapping can avoid the scheduler needing to use linked lists or any other mechanism that some other schedulers may need to use to track fragmentation of memory updates through physical external memory. The example scheduler can manage fragmentation by ignoring it using existing cache allocation mechanisms while still allowing data streams on the GPU to stay in the on-chip cache / through the on-chip cache. By managing the cache footprint in fixed-size blocks, the system benefits from on-chip data streams by leveraging the cache's ability to capture data streams that are sparse relative to larger memory allocations because the cache captures what is actually read and written. If the scheduler can manage its working set as a finite size, the cache provides a finite working size that the scheduler can take advantage of. Example non-limiting scheduling models enable this. Exemplary embodiments can layer on additional scheduling constraints that can make the cache more efficient.
[0111] In addition, when a consumer finishes reading (but not updating) data from a cache line, it may be necessary to destroy (invalidate the cache line without pushing it out to external memory) in order to free the cache line and return it to the scheduler without writing the cache line back to external memory. This means that data produced by a producer can be passed to a consumer through the on-chip cache without writing it to external memory, thus saving memory cycles and memory bandwidth and reducing latency. At the same time, the scheduler actively manages the cache through resource-constrained scheduling of cache lines, improving cache efficiency.
[0112] Generalized functions such as work amplification and work expansion
[0113] Figure 23 and the following figure shows an example non-limiting execution node graph topology that can be constructed to use the disclosed scheduler embodiments. As described above in connection with Figure 12 described Figure 23 .
[0114] Figure 24 By using shadow display, the consumers that are launched do not have to be the same program. Multiple instances of the same consumer or multiple different consumers (i.e., different instances of the same task or different tasks) can be launched from the same queue simultaneously.
[0115] Figure 25 Shows an example of work aggregation, which is a mechanism complementary to work expansion. Here, the producer places some work into a special queue that does not immediately launch a consumer. Instead, the queue is programmed with an aggregation window that allows it to group and process the work produced by the producer within a specific time window and / or based on a specific number of elements produced by the producer. In one example, the queue will attempt to collect a maximum of the user-specified number of work items before launching a consumer task instance. The queue will also report to the launched consumer how many items it was able to collect (minimum = 1, maximum = user-specified value). To use this feature, Figure 9 will be modified to compare with this programmed value, and S R will increment this programmed value. This use case is very useful. For example, hardware support can be used to ensure that all threads in a thread block have data to process when the thread block is launched.
[0116] Figure 26 Shows the situation that occurs during timeout for restoring consistency. If Figure 25 the aggregation window in is not fully filled, a timeout occurs and a partial launch is performed to reduce latency. In this partial launch, S R the semaphore decrements the number of actually launched elements instead of the user-defined programmable value. The timeout sets an upper limit on the time the queue waits before launching a consumer.
[0117] Figure 27 Shows an example of a join operation. Different producers 502 provide different works that attempt to fill a common structure with different fields corresponding to different producers. In the example shown, each of the three producers writes a field to each of the output data records 508(1), 508(n), where each output data record contains three fields 510(a), 150(b), 510(c) corresponding to the three producers 502(a), 502(b), 502(c) respectively. That is, the data producer 502(a) writes to the field 510(a) of each output record 508, the data producer 502(b) writes to the field 502(b), and the data producer 502(c) writes to the field 502(c). Thus, each producer 502 has its own independent concept of space in the output buffer. Each producer 502 has its respective free semaphore S F, which allows the producer to fill in its own fields 510 wherever there is space. Thus, there are three redundant free semaphores S F A, S F B, and S F C, all of which reference the same output record, but each keeps track of a different set of fields 510 in these output records. Thus, in this exemplary embodiment, these semaphores map one-to-one to resource occupancy, but other arrangements are possible.
[0118] In the example shown, the last producer 506 to fill its respective last record field is responsible for advancing the ready semaphore S R , to indicate that the output record 508 is readable by the consumer. Since the scheduling is very simple, a relatively small amount of scheduling hardware can be used to manage a relatively complex output data structure.
[0119] Figure 28 shows using aggregation to sort work into multiple queues for re-convergent execution. Here, the producer outputs to three different queues, which marshal the data to three different consumers. The scheduling mechanism described herein applies to each of the three queues.
[0120] Figure 29 shows an example of sorting work into multiple queues for shader specialization, Figure 30 shows work sorting to reduce convergence in real-time ray tracing. The rays can produce work with random convergence. By arranging multiple queues in these topologies and merging (as shown by the elements growing larger at the end of the queue), the consistency of processing can be increased to reduce random convergence. This aggregation can be used to divide work into consistent batches. This arrangement can be used when the system diverges.
[0121] Figure 31 shows a basic dependency scheduling model that generalizes the above, selecting multiple tasks to start based on dependencies on completed tasks.
[0122] Figure 32 and 33 shows examples of how to use these general concepts to implement graphics tasks and a graphics pipeline that provides data to a graphics rasterizer.
[0123] All patents and publications cited herein are incorporated by reference for all purposes as if expressly set forth.
Claims
1. A scheduler, comprising: A set of idle semaphores, A set of ready semaphores, and A hardware circuit operably coupled to the set of idle semaphores and the set of ready semaphores, wherein, The hardware circuit is configured to change the state of an idle semaphore in the set of idle semaphores in response to reserving a resource and when the resource is released, and is further configured to change the state of a ready semaphore in the set of ready semaphores in response to valid data output by a producer and in response to consumption of the valid data by a consumer, and Wherein, the scheduler starts a task in response to the state of the ready semaphores in the set of idle semaphores and the set of ready semaphores.
2. The scheduler according to claim 1, wherein the idle semaphore and the ready semaphore include counters.
3. The scheduler according to claim 1, wherein the scheduler starts the task when each of the idle semaphore and the ready semaphore is non - zero.
4. The scheduler according to claim 1, wherein the scheduler collaborates with a software - defined data marshalling that manages the resources.
5. The scheduler according to claim 1, wherein the hardware circuit reduces the idle semaphore in response to reserving the resources for a task, and the task increases the idle semaphore when releasing the reserved resources.
6. The scheduler according to claim 1, wherein the resources are used to transfer data between consecutive tasks in a graph.
7. The scheduler according to claim 1, wherein the semaphore represents the availability of available space in a memory buffer or the amount of cache pressure caused by the presence of data streams or work items to be processed in a network.
8. The scheduler according to claim 1, wherein the hardware circuit allocates resources at startup to ensure that a thread block can run to completion.
9. The scheduler according to claim 1, wherein the hardware circuit does not scale or grow with the number of work items from the same task / node or the number of concurrently executing thread groups.
10. The scheduler according to claim 1, wherein the hardware circuit monitors the resource constraints imposed in the chip to minimize the amount of off - chip memory bandwidth required for communicating transient data.
11. The scheduler according to claim 11, wherein the cache is configured to capture the data stream between computational phases without writing the data stream to external memory.
12. The scheduler according to claim 11, wherein the cache is configured to capture the data stream between computational phases without writing the data stream to external memory.
13. The scheduler according to claim 1, wherein the scheduler achieves fine - grained synchronization by not performing batch synchronization.
14. A GPU scheduling method, comprising: (a) Subtract from the idle semaphore to reserve a resource for the task to be started; (b) Change the state of the ready semaphore in response to valid data output by a producer and in response to consumption of the valid data by a consumer; (c) When the task is completed using the reserved resource, programmatically add to the idle semaphore and / or the ready semaphore, and (d) In response to testing the idle semaphore and the ready semaphore, reserve the resource for another task.
15. The GPU scheduling method according to claim 14, comprising performing (a)-(d) in hardware.
Citation Information
Patent Citations
High performance synchronization mechanisms for coordinating operations on a computer system
US11803380B2
System and method for hardware scheduling of indexed barriers
US20140282566A1
Techniques for representing and processing geometry within a graphics processing pipeline
US20190236827A1
Method and apparatus for dynamic allocation and management of semaphores for accessing shared resources
US7353515B1