Hierarchical work scheduling
By adopting a hierarchical work scheduling method in the graphics processing application, the system is subdivided into multiple scheduling domains, which solves the problem of large memory overhead of scheduler circuits when processing complex graphics and cannot scale across multiple engines in the prior art, and achieves the scaling and performance improvement of complex graphics scheduling capabilities.
Patent Information
- Application Number
- CN202380066996.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-28
- Publication Date
- 2025-05-16
AI Technical Summary
Scheduler circuits in existing graphics processing applications have problems such as high memory overhead, high execution overhead and inability to scale across multiple engines when processing complex graphics.
Using a hierarchical work scheduling method, the system is subdivided into multiple scheduling domains, each scheduling domain contains local scheduler circuits and work group processing elements, and the work items are allocated and scheduled through local scheduler circuits, and when necessary, return the work items to a higher level scheduler circuit for reallocation.
Scaling of complex graphics scheduling capabilities is achieved, reducing data movement and cross-domain bandwidth, improving performance and reducing power usage.
Smart Images

Figure CN120019362A_ABST
Abstract
Description
Background Art
[0001] Graphics processing applications typically include a workflow of vertex and texture information and instructions to process this information. The various work items (also called "commands") may be prioritized according to some order and queued into a system memory buffer for subsequent retrieval and processing. Scheduler circuitry receives instructions to be executed and generates one or more commands to be scheduled and executed at a processing resource such as a graphics processing unit (GPU) or other single instruction multiple data (SIMD) processing unit. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] The present disclosure may be better understood, and its numerous features and advantages made apparent to those skilled in the art by referencing the accompanying drawings.The use of the same reference numbers in different drawings indicates similar or identical items.
[0003] Figure 1 A block diagram of a computing system implementing hierarchical work scheduling is shown according to a particular implementation.
[0004] Figure 2 is a block diagram of a portion of a processor that implements a hierarchical scheduler circuit according to a particular implementation.
[0005] Figure 3 is a block diagram illustrating a graphics processing stacked die chiplet implementing a hierarchical scheduler circuit according to a particular implementation.
[0006] Figure 4 is a block diagram of a method for performing hierarchical work scheduling according to a specific implementation. DETAILED DESCRIPTION
[0007] The performance of GPU architectures and other parallel processing architectures continues to improve as applications perform large operations involving many iterations (or time steps) and multiple operations within each step. In order to avoid the overhead and performance degradation caused by launching these operations individually to the GPU, multiple work items (often collectively referred to as "graphs" or "work graphs") are launched via a single CPU operation rather than via separate CPU operations. Graph-based software architectures (often referred to as dataflow architectures) are common for software applications that process continuous data streams or events. However, centralized scheduling systems such as command processors may incur significant memory overhead, execution overhead, and cannot be expanded across multiple engines.
[0008] To address these issues and enable improved scheduling of complex graphs (particularly for multi-chiplet GPU architectures, or more generally, multi-chiplet single instruction multiple data (SIMD) processor architectures), Figures 1 to 4Systems and methods for hierarchical work scheduling for a scheduler circuit including a local execution graph are described. In a specific implementation, a method for hierarchical work scheduling includes consuming a first work item at a first scheduling domain having a local scheduler circuit and one or more work group processing elements. The first scheduling domain that consumes the first work item generates a set of new work items. Subsequently, the local scheduler circuit allocates at least one new work item from the set of new work items to be executed locally at the first scheduling domain. However, if the local scheduler circuit of the first scheduling domain determines that the set of new work items includes one or more work items that will overload the first scheduling domain with work if scheduled for local execution, these work items are returned to the next higher level scheduler circuit in the scheduling domain hierarchy for reallocation. Similarly, if a scheduling domain runs out of work, it can request new work from the next scheduler circuit in the hierarchy. In this case, the scheduler circuit requests any scheduling domain below to make work available to the first scheduling domain. In this way, the scheduling capability of complex graphs scales with chip size while keeping data movement to a minimum. In addition, cross-domain bandwidth is reduced, which becomes increasingly important as off-chip memory latency and power usage increase.
[0009] Note that while reference is made below to graphics processing and GPUs, these references are for illustrative purposes and are not intended to limit the following description. Rather, the systems and techniques described herein may be implemented for various forms of parallel processing of work items in addition to or in addition to graphics-related work items. Thus, references to graphics work scheduling and graphics work items apply equally to other types of work scheduling and work items, and similarly, references to GPUs apply equally to other types of SIMD processing units or other parallel processing hardware accelerators, such as accelerated processing units (APUs), discrete GPUs (dGPUs), artificial intelligence (AI) accelerators, and the like.
[0010] Figure 1 A block diagram of a computing system 100 with hierarchical work scheduling according to a specific implementation is shown. In a specific implementation, the computing system 100 includes at least one or more processors 102A-N, a structure 104, an input / output (I / O) interface 106, a memory controller 108, a display controller 110, and other devices 112. In a specific implementation, in order to support the execution of instructions for graphics and other types of workloads, the computing system 100 also includes a host processor 114, such as a central processing unit (CPU). In a specific implementation, the computing system 100 includes a computer, a laptop, a mobile device, a server, or any of various other types of computing systems or devices. It is noted that the number of components of the computing system 100 can vary. It is also noted that in a specific implementation, the computing system 100 includes Figure 1Other components not shown in FIG. In addition, the computing system 100 may be different from Figure 1 Constructed in the manner shown in .
[0011] The structure 104 represents any communication interconnect that conforms to any of the various types of protocols used to communicate between components of the computing system 100. The structure 104 provides data paths, switches, routers, and other logical components that connect the processor 102, the I / O interface 106, the memory controller 108, the display controller 110, and other devices 112 to each other. The structure 104 handles request, response, and data services and probe services to promote consistency. The structure 104 also handles interrupt request routing and configuration access paths to various components of the computing system 100. In addition, the structure 104 handles configuration requests, responses, and configuration data services. In a specific implementation, the structure 104 is bus-based, including a shared bus configuration, a crossbar configuration, and a hierarchical bus with a bridge. In other specific implementations, the structure 104 is packet-based and is a hierarchical structure with bridges, crossbars, point-to-point, or other interconnects. From the perspective of the structure 104, the other components of the computing system 100 are referred to as "user ends". The structure 104 is configured to process requests generated by various user ends and pass these requests to other user ends.
[0012] The memory controller 108 represents any number and type of memory controllers coupled to any number and type of memory devices. For example, the types of memory devices coupled to the memory controller 108 include dynamic random access memory (DRAM), static random access memory (SRAM), NAND flash memory, NOR flash memory, ferroelectric random access memory (FeRAM), etc. The memory controller 108 can be accessed by the processor 102, the I / O interface 106, the display controller 110, and other devices 112 via the structure 104. The I / O interface 106 represents any number and type of I / O interface (e.g., peripheral component interconnect (PCI) bus, PCI-Extended (PCI-X), PCIE (PCI Express) bus, Gigabit Ethernet (GBE) bus, Universal Serial Bus (USB)). Various types of peripheral devices are coupled to the I / O interface 106. Such peripheral devices include (but are not limited to) a display, a keyboard, a mouse, a printer, a scanner, a joystick or other type of game controller, a media recording device, an external storage device, a network interface card, etc. Other devices 112 represent any number and types of devices (eg, multimedia devices, video codecs).
[0013] In a specific implementation, each of the processors 102 is a parallel processor (e.g., a vector processor, a graphics processing unit (GPU), a general-purpose GPU (GPGPU), a non-scalar processor, a highly parallel processor, an artificial intelligence (AI) processor, an inference engine, a machine learning processor, other multi-threaded processing units, etc.). In a specific implementation, each parallel processor 102 is constructed as a multi-chip module (e.g., a semiconductor die package) that includes two or more basic integrated circuit dies that are communicatively coupled together with a bridge chip or other coupling circuits or connectors, so that the parallel processors are available (e.g., addressable) like a single semiconductor integrated circuit. As used in this disclosure, the terms "die" and "chip" are used interchangeably. Those skilled in the art will recognize that conventional (e.g., non-multi-chip) semiconductor integrated circuits are manufactured as wafers, or are manufactured as dies (e.g., single-chip ICs) that are formed in a wafer and later separated from the wafer (e.g., when the wafer is cut); multiple ICs are typically manufactured simultaneously in a wafer. The ICs and possibly discrete circuits and possibly other components (such as non-semiconductor packaging substrates, including printed circuit boards, interposers, and possibly other components) are assembled in a multi-die parallel processor.
[0014] In a specific implementation, each of the processors 102 includes one or more base IC dies that employ processing chiplets according to the specific implementation. The base die is formed as a single semiconductor chip that includes N communicatively coupled graphics processing stacked die chiplets. In a specific implementation, the base IC die includes two or more direct memory access (DMA) engines that coordinate DMA data transfers between a device and a memory (or between different locations in a memory).
[0015] As will be appreciated by those skilled in the art, in a specific implementation, parallel processors and other multithreaded processors 102 implement multiple processing elements (not shown) (also interchangeably referred to as processor cores or computing units) that are configured to execute multiple instances (threads or waves) of a single program on multiple data sets simultaneously or in parallel. Several waves are formed (or generated) and then dispatched to each processor in a multithreaded processing unit. In a specific implementation, a processing unit includes hundreds of processing elements, so that thousands of waves execute programs in the processor simultaneously. The processing elements in a GPU typically use a graphics pipeline formed by a series of programmable shaders and fixed-function hardware blocks to process three-dimensional (3-D) graphics.
[0016] The host processor 114 prepares and distributes one or more operations to one or more processors 102A-N (or other computing resources), and then retrieves the results of the one or more operations from the one or more processors 102A-N. Conventionally, the host processor 114 sends work to be performed by the one or more processors 102A-N by enqueuing various work items (also referred to as "threads") in a command buffer (not shown). Computer applications such as graphics processing applications perform a large number of operations (e.g., kernel launches or memory copies) that involve many iterations (or time steps) and multiple work items within each step. In a specific implementation, the computing system 100 utilizes a graph-based model to submit work to be performed by one or more processors 102A-N (or other parallel computing architectures) as an integrated whole by using a work graph rather than a single GPU operation.
[0017] In at least one specific implementation, the host processor 114 executes one or more work graphs. Specifically, a workload including multiple work items is organized as a work graph (or simply referred to as a "graph"), wherein each node in the graph represents a corresponding work item to be executed, and each edge (or link) between two nodes corresponds to a dependency (such as data dependency, execution dependency, or some other dependency) between two work items represented by two linked nodes. For illustration, the work graph 116 includes work items of nodes (AD) forming the work graph 116, wherein the edges are dependencies between the work items. In a specific implementation, the dependency indicates when a work item of a node must be completed before a work item of another node can be started. In a specific implementation, the dependency indicates when a node needs to wait for data from another node before it can start and / or continue its work item. In a specific implementation, one or more processors 102A-N execute the work graph 116 by executing the work item started at node A after being called by the host processor 114. As shown, the edges (shown as arrows) between node A and nodes B and C indicate that the work items of node A must complete execution before the work items of nodes B and C can begin. In a specific implementation, the nodes of work graph 116 include work items, such as kernel launches, memory copies, CPU function calls, or other work graphs (e.g., each of nodes AD may correspond to a subgraph including two or more other nodes [not shown]).
[0018] As follows about Figures 2 to 4Described in more detail, the computing system implements a scalable scheduling mechanism that can scale to arbitrarily large GPU configurations (across chiplets and packages), has minimal bandwidth requirements, and can achieve near-optimal memory usage. This scheduling mechanism includes a parallel, scalable scheduler circuit that subdivides the system into multiple scheduling domains and can also schedule simultaneously across multiple scheduling domains and within scheduling domains. The scheduling mechanism also includes memory / control hardware attached to each scheduling domain. Each scheduling domain includes its own scheduler circuit and can create work for itself by consuming work items in nodes of the corresponding work graph while keeping most of the memory traffic local. In a specific implementation, each scheduling domain communicates only with its parent scheduler circuit (e.g., one level up in the scheduling domain hierarchy), but does not communicate horizontally with other scheduler circuits within the same hierarchical level to avoid N-squared communication patterns.
[0019] View now Figure 2 , a block diagram of a portion of a processor 200 implementing a hierarchical scheduler circuit according to a specific implementation is shown. In a specific implementation, the processor 200 includes a fetch / decode logic component 202 that fetches and decodes instructions in a wave of a work group that is scheduled for execution by the processor 200. A specific implementation of the processor 200 executes the waves in the work group. For example, in a specific implementation, the fetch / decode logic component 202 fetches kernels of instructions executed by all waves in the work group. The fetch / decode logic component 202 then decodes the instructions in the kernels. The processor 200 also includes a cache memory 204, such as a level 1 (L1) cache memory, which is used to store local copies of data and instructions used during execution of the wave.
[0020] In a specific implementation, the processor 200 is used to implement the following Figure 1 1. As will be appreciated, in a specific implementation, each processor 102 includes multiple processing elements that operate according to the SIMD protocol to use multiple processor cores to simultaneously execute the same instructions on multiple data sets. Therefore, the smallest processing element is called a SIMD unit. In a specific implementation, the SIMD unit is divided into processor cores 208, which together form a workgroup processing element 210.
[0021] In a specific implementation, processor 200 includes one or more scheduling domains 206 (sometimes referred to as "node processors" due to their processing of work items in nodes of a work graph (e.g., work graph 116 as previously described)) that include local scheduler circuits 212 (also interchangeably referred to as work graph scheduler circuits [WGS]) associated with a set of work group processing elements 210. The various scheduler circuits and command processors described herein handle queue level allocation. As shown in FIG. Figure 2 As shown in , an example implementation of the processor 200 includes a number N=2 scheduling domains 206. In a specific implementation, each scheduling domain 206 includes a plurality of workgroup processing elements 210, a local scheduler circuit 212, a local queue 214, and a local cache 216. During the execution of a job, the local scheduler circuit 212 executes the job locally and does not interact with other scheduling domains. Instead, the local scheduler circuit 212 uses a dedicated memory area for scheduling and as a scratch space.
[0022] Although shown as including two scheduling domains 206, those skilled in the art will recognize that any number of scheduling domains across any number of domain hierarchies may be included at the processor 200. Furthermore, those skilled in the art will recognize that in a specific implementation, the processor 200 includes any number of nested scheduling domains. For example, such as the following regarding Figure 3 As described, the higher-level scheduling domain is generally a hierarchical structure with a smaller granularity than the scheduling domain 206 (e.g., including multiple scheduling domains 206). In other specific implementations, the lower-level scheduling domain is generally a hierarchical structure with a finer granularity (e.g., Figure 2 Each workgroup processing element 210 is its own lower-level scheduling domain and includes its own local scheduler circuitry. Nested scheduling domains include any number and combination of these various hierarchical scheduling domains.
[0023] In a specific implementation, the workgroup processing element 210 that enqueues the work item is attached to a single local queue 214 (or a small set of local queues), which sends a signal when the queue has a sufficient number of work items to be processed. In order to allocate efficiency, the queue memory is managed into blocks 218 (including empty blocks, blocks in the queue state, blocks in use, etc.). As used herein, a "block" (e.g., block 218) is a memory block (e.g., in a cache memory) and is used to group work items together to avoid processing each work item at the scheduling level. In a specific implementation, a block 218 includes multiple work items, each of which is potentially targeted at different nodes in the work graph. In this way, the local scheduler circuit 212 can provide local execution of work (e.g., minimum bandwidth use) with reduced memory access and use. In a specific implementation, the queue described herein (e.g., local queue) allows the addition and / or removal of work items. In a specific implementation, such a queue is implemented as a ring buffer and can be full.
[0024] In a specific implementation, the processor 200 also includes a global command processor (CP) 220 (also interchangeably referred to as a "global scheduler circuit" or "central CP"), which is a higher-level scheduler circuit for the scheduling domains 206 and communicates with all of these scheduling domains. In a specific implementation, when there are work items that need to be distributed for execution, the global CP 220 distributes (e.g., evenly) the work items across two or more scheduling domains 206 for execution at their corresponding workgroup processing elements 210. However, execution of the work items typically continues to generate additional new work items ("new work items"). In a specific implementation, the new work items 224 are generated by workers at the lowest level of the scheduling domain hierarchy. The lowest level in the scheduling domain hierarchy includes hardware (e.g., Figure 2 The lowest level includes hardware consumption and the generation of new work items due to consumption, and sending those new work items up / down through the scheduling domain hierarchy if the local scheduling domain overflows or runs empty with work items. For example, eight new work items 224 are generated by the work group processing element 210 due to the previous execution of work items. The new work items 224 are initially stored at the local cache 216.
[0025] Conventionally, such new work items 224 are passed from the local cache memory (e.g., local cache 216) of the scheduling domain 206 to the cache memory 204 for allocation by the global CP 220 in the next round of scheduling. As will be appreciated by those skilled in the art, if work needs to be sent to the global CP 220 for scheduling and allocation, this scheduling scheme results in work items being moved around inefficiently. Additionally, this is inefficient because work generated at a particular scheduling domain 206 cannot be consumed immediately locally, but needs to be passed to the global CP 220 for rescheduling. In order to improve the scheduling of work items, the present disclosure describes a rescheduling mechanism, namely, a scheduler circuit subdivides the system into multiple scheduling domains, and the scheduler circuit can also schedule simultaneously across multiple scheduling domains and within a scheduling domain. Despite the description of Figure 2 Although described in the context of scheduling domain 206 within processor 200, those skilled in the art will recognize that the concepts described are not limited to Figure 2A two-level scheduling domain hierarchy is shown, where a top-level global scheduling domain of the scheduling domain hierarchy includes the global CP 220 and a lower level of the scheduling domain hierarchy includes each individual scheduling domain 206 and its corresponding workgroup processing element 210. In a specific implementation, the scheduling domain hierarchy may be expanded in scope to include, for example, further lower levels within the scheduling domain hierarchy (e.g., scheduling domains within each individual workgroup processing element 210) and / or further higher levels within the scheduling domain hierarchy (e.g., a parent scheduling domain for each package in a multi-GPU system including multiple processors 200).
[0026] Utilizing the local scheduler circuit 212 of each scheduling domain 206 as described herein to distribute work to locally available workgroup processing elements 210 generally results in reduced memory traffic, latency associated with transferring data to cache memory 204, and also waiting for workgroup processing elements 210 to complete before reallocating new work items. In a specific implementation, each scheduling domain 206 communicates with the global CP 220, such as via the local scheduler circuit 212, only when it is out of work or has so much work that it needs to offload some work to other scheduling domains. Figure 2 As illustrated in FIG. 1 , the local scheduler circuit 212 allocates eight new work items 224 (eg, Figure 2 The eight new work items 224 generated by the scheduling domain 206 are shown being redistributed equally among the four workgroup processing elements 210).
[0027] In this way, the scheduling of work from the local cache 216 by the local scheduler circuitry 212 (rather than transferring work items to the global CP 220) reduces the amount of round-trip communication with the cache memory 204. Latency is also reduced due to the communication paths 226 between the local scheduler circuitry 212 and the local workgroup processing elements 210 within each scheduling domain 206 (as opposed to the longer communication paths (not shown) from the scheduling domain 206 to the cache memory 204 and then back to the workgroup processing elements for execution of work items).
[0028] In a specific implementation, each scheduling domain 206 maintains local queues (e.g., local queues 214) that are visible only internally (i.e., to components of the same scheduling domain 206) and shared queues (not shown) that are shared with other scheduling domains 206. Work flows in both directions, allowing a scheduling domain 206 to take work from a shared queue to place in its own local queue 214, or to migrate local work from a local queue 214 to a shared queue. In a specific implementation, only one shared queue above one level is always visible to a scheduling domain 206 - if the hierarchy is deeper, the scheduling domain 206 cannot skip levels, but another scheduling domain at a higher hierarchy level needs help. The difference with shared queues is that the scheduling domain 206 knows that it needs synchronized access to them. This means that all operations on shared queues require synchronized access, such as via atomics, to allow other scheduling domains to access them simultaneously. On the other hand, if other scheduling domains need to request access to do work stealing, local queues are safe to access without synchronization.
[0029] This architecture and method of distributing work without communicating with a centralized instance (e.g., global CP 220) allows the system to self-balance and reduce the amount of resource overload and underutilization. For example, in a specific implementation, local scheduler circuit 212 communicates with global CP 220 only when its associated scheduling domain 206 is overloaded and / or does not have enough work to keep busy (as opposed to the conventional method of communicating with global CP 220 each time a unit of work is completed). Instead, as described herein, the responsibility and computational overhead for work scheduling are distributed across multiple scheduler circuits at different levels of the hierarchy to gain the benefits of memory locality and reduced latency.
[0030] Reference now Figure 3, shows an example of a hierarchical work scheduling technique in a system implemented as a graphics processing stacked die chiplet according to a specific implementation. As shown, the graphics processing stacked die chiplet 302 includes one or more basic active interposer dies (AIDs) 304 (individually represented as AID 304A and AID 304B and collectively referred to as AID 304). It should be recognized that although the graphics processing stacked die chiplet 302 is described below in the specific context of GPU terminology for ease of illustration and description, in specific implementations, the described architecture is applicable to any of various types of parallel processors without departing from the scope of the present disclosure. In specific implementations, the graphics processing stacked die chiplet 302 is used, for example, as a multi-chip module. Additionally, in specific implementations, and as used herein, the term "chiplet" refers to any device that includes, but is not limited to, the following characteristics: 1) the chiplet includes an active silicon die that includes at least a portion of the computational logic components used to solve a problem (i.e., the computational workload is distributed across multiple of these active silicon dies); 2) the chiplets are packaged together as a monolithic unit on the same substrate; and 3) the programming model retains the concept of these separate computational dies as a combination of a single monolithic unit (i.e., each chiplet is not exposed as a separate device to an application using the chiplet to process the computational workload).
[0031] The basic active interposer die 304 of the graphics processing stacked die chiplet 302 includes N scheduling domains 306 (similar to Figure 2 In a specific implementation, each AID 304 includes one or more engine dies (EDs) 308, such as a shader engine die or other computing die or any suitable parallel processing unit. Figure 3 In a specific implementation, each discrete scheduling domain 306 includes a single ED 308. In other specific implementations (not shown), multiple EDs 308 may constitute a single scheduling domain 306. Although shown as including two scheduling domains 306, those skilled in the art will recognize that any number of scheduling domains across any number of domain hierarchies may be included at AID 304. For example, in a specific implementation, AID 304 is a component of a graphics processing stacked die chiplet 302 or other processor semiconductor package. The graphics processing stacked die chiplet 302 also includes a cache memory die 312 such as a last level cache and a global command processor (CP) die 314 (also interchangeably referred to as a "global CP die" or "global scheduler circuit"). As shown in FIG. Figure 3As shown in FIG, the graphics processing stacked die chiplet 302 is formed as a single semiconductor chip package including a first AID 304A and a second AID 304B that are two different scheduling domains 316 that are different from each other. As will be understood, each scheduling domain 316 is a higher-level scheduling domain (relative to the scheduling domain 306) that includes one or more lower-level scheduling domains 306.
[0032] As will be appreciated by those skilled in the art, the global CP die 314 acts as the top-level scheduler circuit in the scheduling domain and communicates with all of these scheduler circuits. However, this limits the scalability of current parallel processor architectures because the system relies on the one-way communication mode of the global CP to send data to each scheduling domain for work to be executed. As described in more detail with respect to specific implementations herein, the processor 200 and the graphics processing stacked die chiplet 302 include the global CP die 314 and one or more local scheduler circuits (e.g., Figure 2 The local scheduler circuit 212 and Figure 3 For example, in a specific implementation, each scheduling domain 316 also includes a working scheduler circuit 310 that is local to each AID 304.
[0033] exist Figure 3 In the example implementation shown in , the first AID 304A includes a scheduler circuit 318A and a work queue 320A that is local to the first AID 304A and a work scheduler circuit 310 that is one level higher (i.e., a “parent scheduler circuit”) in the scheduling domain hierarchy than the local scheduler circuits 318A and 318B (collectively referred to as local scheduler circuits 318) of each ED 308 in the nested scheduling domain configuration. Although previously described in Figure 2 304A) relinquishes work to other scheduling domains within the same hierarchy (e.g., scheduling domain 306 of ED 308B) or to one level up the hierarchy (e.g., scheduling domain 316 of AID 304A). Local scheduler circuits 318A and 318B are collectively referred to as local scheduler circuits.
[0034] In a specific implementation, the local queue 320A of the first ED 308A includes a queue of work items to be processed. Figure 3As illustrated in , the local queue 320A includes eight work items 322, which are known to have an amount that will cause the scheduling domain 306 of the first ED 308A to be overloaded with work (e.g., based on the number of work items to be scheduled for each available work group processing element). In other specific implementations, the first ED 308A determines the type of work queued at the local queue 320A. For example, some work items have a low amplification factor (as determined, for example, when the work items are packaged into a work graph via one or more API or library function calls), so that consuming one work item produces a single new work item (e.g., a linear 1 to 1 relationship) or several additional work items (e.g., a 1 to 2 or 1 to 3 relationship). The theoretical amplification factor is known from the work graph definition, but the actual amplification is determined at runtime. The application defines during graph definition time that the consumption of a particular work item can produce 1 to N number of new work items, and when the graph executes, it is instructed to stay within those bounds. Other work items have large amplification factors and produce a large number of new work items when consumed. In such implementations, the first ED 308A preemptively returns work items to the global CP die 314 for allocation to other scheduling domains to prevent itself from becoming overloaded.
[0035] In a specific implementation, a scheduling domain may also steal work from other scheduling domains when it is idle or underutilized. For example, the second ED 308B is idle, and an underutilization notification 324 is sent from the local scheduler circuit 318B to other scheduler circuits within the same scheduling domain hierarchy (e.g., the local scheduler circuit 318A of the first ED 308A) or to the next hierarchical higher level scheduler in the system (e.g., the job scheduler circuit 310 of the scheduling domain 316). Similarly, if there is no job at the queue 320A, an idle signal is sent by the job scheduler circuit 310 at the scheduling domain 316 to the global CP die 314, requesting to "steal" a job from another domain.
[0036] like Figure 3, the second ED 308B steals work directly from the first ED 308A to transfer to the local queue 320B of the second ED 308B (the local queues 320A and 320B are collectively referred to as local queues 320). As will be appreciated, this direct communication path between the local scheduler circuits 318 is not overly complex when there are few scheduling domains per hierarchy level, but grows exponentially as the number of scheduling domains increases. Therefore, in other implementations, instead of stealing work directly, the second ED 308B communicates to the global CP die 314 (e.g., moving up and down between scheduling domain hierarchy levels, rather than horizontally traversing peer-adjacent scheduling domains within the same hierarchy level) to request that the global CP die 314 find some work, such as by pinging other scheduling domains to return some work to the global queue at the cache memory die 312. In a specific implementation, a single work item is shared among many projects, thereby amortizing the cost of migrating a work item from one scheduling domain to another (eg, via a cache flush so that the input work item is visible and available to the new scheduling domain).
[0037] Reference now Figure 4 , a block diagram of a method 400 for performing hierarchical work scheduling according to a specific implementation is shown. For ease of illustration and description, the following reference Figures 1 to 3 Method 400 is described with reference to systems and devices and in their exemplary contexts. However, method 400 is not limited to these exemplary contexts, but rather is used in different implementations for any of a variety of possible system configurations using the guidance provided herein.
[0038] Method 400 begins at block 402, where a first scheduling domain receives a first work item from a global command processor, which is a higher-level scheduler circuit of the first scheduling domain. Figure 2 As described above, when there are work items that need to be allocated for execution, the global CP 220 distributes the work items across two or more scheduling domains 206 (e.g., evenly). At block 404, the first scheduling domain (including the local scheduler circuit and one or more work group processing elements) consumes the first work items to generate a set of new work items. As previously described with respect to Figure 2 As described in more detail, work items, when consumed / executed, typically continue to generate additional new work items (eg, new work items 224).
[0039] At block 406, the local scheduler circuit of the first scheduling domain determines whether the set of new work items includes one or more work items that would overload the first scheduling domain with work if scheduled for local execution. In a specific implementation, determining that the set of new work items includes one or more work items that would overload the first scheduling domain includes determining that a total number of the set of new work items exceeds a predetermined threshold. Figure 3 As illustrated in FIG. 1 , the local queue 320A includes eight work items, which are known to be of an amount that would cause the scheduling domain 306 of the first ED 308A to be overloaded with work (eg, based on the number of work items to be scheduled per available workgroup processing element).
[0040] In other specific implementations, determining the set of new work items includes one or more work items that will overload the first scheduling domain includes identifying one or more work items having a magnification factor greater than a predetermined threshold. For example, some work items have a low magnification factor, such that consuming one work item produces a single new work item (e.g., a linear 1-to-1 relationship) or several additional work items (e.g., a 1-to-2 or 1-to-3 relationship). Other work items have a large magnification factor and produce a large number of new work items when consumed, and at least one set of new work items is allocated by the local scheduler circuit to be executed in the first scheduling domain.
[0041] If the local scheduler circuit determines that one or more work items within the set of new work items will indeed overload the first scheduling domain, then method 400 proceeds to block 408, where those identified work items are assigned by the local scheduler circuit to the next higher level scheduler circuit in the hierarchy (e.g., the global command processor) for rescheduling. Figure 3 As discussed, in a specific implementation, the first ED 308A preemptively returns work items to the global CP die 314 for allocation to prevent itself from becoming overloaded with work. However, if the local scheduler circuit determines that it is currently able to execute the set of new work items, then at block 410, the local scheduler circuit allocates the work locally to the workgroup processing element. Figure 2 As shown in , the local scheduler circuit 212 allocates the new work item 224 by scheduling the work item for execution at the work group processing element 210 of the same scheduling domain 206 that originated the work item (as indicated by the * symbol). In this way, the scheduling of work from the local cache 216 by the local scheduler circuit 212 (rather than transferring the work item to the global CP 220) reduces the amount of round-trip communication with the cache memory 204.
[0042] Although described above in the context of moving work items to the next higher level scheduling domain in response to an overload of the local scheduler circuit, those skilled in the art will recognize that the operation of block 408 may also be triggered by other conditions. In a specific implementation, work items may also be returned to the next higher level scheduler circuit in the scheduling hierarchy when there is not enough work for effective SIMD unit execution. For example, if there is a collection of work items generated from the workgroup processing element 210 consisting of a large amount of work for program A and only a little bit of program B, so that the workgroup processing element 210 is not fully utilized, then for the case where program B is a very long running program to make the migration work worthwhile, the local scheduler circuit will push the work of program B to the next higher scheduling domain. At the next higher level scheduling domain, the work of program B is combined with other work items of program B to form a larger set of work of program B that will be sent back down to one of the lower scheduling domains for execution.
[0043] Additionally, in a specific implementation, the local scheduler circuitry also determines at block 412 whether one or more workgroup processing elements of its scheduling domain are underutilized and should request additional work for the next round of scheduling. In a specific implementation, the local scheduler circuitry 212 communicates with the global CP 220 when its associated scheduling domain 206 does not have enough work to stay busy. The local scheduler circuitry generates an underutilized signal requesting that additional work items be assigned to a first scheduling domain. In a specific implementation, the underutilized signal is transmitted to a second local scheduler circuitry within the same level of the scheduling domain hierarchy. Figure 3 , the second ED 308B is idle and sends an underutilization notification 324 to other scheduler circuits within the same scheduling domain hierarchy (e.g., the local scheduler circuit 318A of the first ED 308A) or to scheduler circuits at a higher level from the next hierarchy in the system (e.g., the scheduler circuit 318A of the scheduling domain 316). Figure 3 As shown in , the second ED 308B steals work directly from the first ED 308 A. If no work exists at queue 320A, an idle signal is sent to the next higher level scheduling domain (global CP die 314 in this implementation) requesting work to be stolen from another domain.
[0044] In other implementations, the underutilized signal from the local scheduler circuit is transmitted to the scheduler circuit at the next higher level of the scheduling domain hierarchy. For example, instead of stealing work directly, the second ED 308B communicates to the global CP die 314 (e.g., moving up and down between scheduling domain hierarchy levels, rather than horizontally traversing peer-adjacent scheduling domains within the same hierarchy level) to request that the global CP die 314 find some work, such as by pinging other scheduling domains to return some work to the global queue at the cache memory die 312.
[0045] Thus, as discussed herein, hierarchical work scheduling by local scheduler circuitry provides increased data locality, where work items are consumed closer to producers by implementing parallel, scalable scheduler circuitry that subdivides the system into (nested) scheduling domains. Each scheduling domain includes its own scheduler circuitry and can create work for itself while keeping most of the memory traffic local. This improves the efficiency of scheduling large, complex graphs (which include nodes that generate work and reduce) with minimal memory overhead (e.g., a constant number of items processed simultaneously rather than a constant total amount of work). In addition, by scheduling multiple domains simultaneously and independently of each other, the number of cross-domain connections is reduced and cross-domain communication is limited. Using this combined hardware / software approach, hierarchical work scheduling allows multi-level scheduling to be implemented from individual WGPs in a scheduling domain on a chiplet to multi-GPU scheduling.
[0046] Computer-readable storage media may include any non-transitory storage media or combination of non-transitory storage media that can be accessed by a computer system during use to provide instructions and / or data to the computer system. Such storage media may include, but are not limited to, optical media (e.g., compact disks (CDs), digital versatile disks (DVDs), Blu-ray disks), magnetic media (e.g., floppy disks, tapes, or magnetic hard drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or micro-electromechanical system (MEMS)-based storage media. Computer-readable storage media may be embedded in a computing system (e.g., system RAM or ROM), fixedly attached to a computing system (e.g., a magnetic hard drive), removably attached to a computing system (e.g., an optical disk or flash memory based on a universal serial bus (USB)), or coupled to a computer system via a wired or wireless network (e.g., a network accessible storage device (NAS)).
[0047] In a specific implementation, certain aspects of the above-mentioned technology can be implemented by one or more processors of a processing system that executes software. The software includes one or more sets of executable instructions that are stored or otherwise tangibly embodied on a non-transitory computer-readable storage medium. The software may include instructions and certain data that manipulate one or more processors to perform one or more aspects of the above-mentioned technology when executed by one or more processors. Non-transitory computer-readable storage media may include, for example, magnetic or optical disk storage devices, solid-state storage devices such as flash memory, cache, random access memory (RAM), or other one or more non-volatile memory devices, etc. The executable instructions stored on the non-transitory computer-readable storage medium may be source code, assembly language code, object code, or other instruction formats that are interpreted or otherwise executed by one or more processors.
[0048] It should be noted that not all activities or elements described above in the general description are required, a portion of a particular activity or device may not be required, and one or more additional activities may be performed, or elements may be included in addition to those described. Furthermore, the order in which the activities are listed is not necessarily the order in which they are performed. In addition, these concepts have been described with reference to specific implementations. However, it is understood by those of ordinary skill in the art that various modifications and changes may be made without departing from the scope of the present disclosure as set forth in the following claims. Therefore, the specification and drawings are to be regarded as illustrative rather than restrictive, and all such modifications are intended to be included within the scope of the present disclosure.
[0049] Benefits, other advantages, and solutions to problems have been described above with respect to specific implementations. However, benefits, advantages, solutions to problems, and any features that may cause any benefit, advantage, or solution to appear or become more significant should not be interpreted as key, essential, or basic features of any or all claims. In addition, the specific implementations disclosed above are merely illustrative, as the disclosed subject matter may be modified and practiced in different but equivalent ways that are obvious to those skilled in the art who benefit from the teachings herein. Except as described in the following claims, it is not intended to limit the details of the construction or design shown herein. Therefore, it is apparent that the specific implementations disclosed above may be changed or modified, and all such changes are considered to be within the scope of the disclosed subject matter. Therefore, the protection sought herein is as set forth in the following claims.
Claims
1. A method for execution at a processor, the method comprising: scheduling, by a first scheduler circuit of the processor, a first work item for consumption by a first work group processing element associated with the first scheduler circuit; a set of new work items generated as a result of the first work item being consumed by the first work group processing element; as well as At least one new work item from the set of new work items is allocated, by the first scheduler circuit, to a second scheduler circuit of the processor that is different from the first scheduler circuit.
2. The method according to claim 1, wherein: The first scheduler circuit is part of a first scheduling domain of a scheduling hierarchy, the first scheduling domain including a set of workgroup processing elements, the set of workgroup processing elements including the first workgroup processing element; and The second scheduler circuit is part of a second scheduling domain of the scheduling hierarchy.
3. The method of claim 2, wherein the second scheduling domain is at the same level in the scheduling hierarchy as the first scheduling domain and comprises a set of workgroup processing elements separate from the set of workgroup processing elements of the first scheduling domain.
4. The method of claim 2, wherein the second scheduling domain is at a higher level in the scheduling hierarchy than the first scheduling domain and includes a set of workgroup processing elements, the set of workgroup processing elements including the set of workgroup processing elements of the first scheduling domain and a set of workgroup processing elements of a third scheduling domain at a lower level in the scheduling hierarchy than the second scheduling domain.
5. The method of claim 4, wherein the second scheduling domain is a global scheduling domain for the scheduling hierarchy and the second scheduler circuit is a global command processor.
6. The method according to claim 5, further comprising: The first work item is received at the first scheduler circuit from the global command processor.
7. The method according to any one of claims 2 to 6, wherein: Allocating the at least one new work item is in response to the first scheduler circuit determining that the set of new work items includes one or more work items that would overload the first scheduling domain if consumed by at least one workgroup processing element of the first scheduling domain.
8. The method of claim 7, wherein determining that the set of new work items includes one or more work items that will overload the first scheduling domain includes identifying one or more work items having a magnification factor greater than a predetermined threshold.
9. The method of claim 7, wherein determining that the set of new work items includes one or more work items that will overload the first scheduling domain comprises determining that a total number of new work items in the set of new work items exceeds a predetermined threshold.
10. The method according to any one of claims 2 to 9, further comprising: The first work item is received at the first scheduler circuit in response to an indication from the first scheduler circuit that one or more workgroup processing elements of the first scheduling domain are idle.
11. The method according to any one of claims 2 to 10, further comprising: The at least one new work item is allocated to the second scheduler circuit in response to an indication from the second scheduler circuit that one or more workgroup processing elements of the second scheduling domain are idle.
12. A system, comprising: a first processor, the first processor comprising a first scheduling domain and a second scheduling domain, the first scheduling domain comprising a first scheduler circuit and a set of workgroup processing elements, and the second scheduling domain comprising a second scheduler circuit separate from the first scheduling circuit; and The first scheduler circuit is configured to: scheduling a first work item for consumption by a workgroup processing element of the set; as well as At least one new work item from a set of new work items generated by the work group processing elements consuming the first work item is assigned to the second scheduler circuit.
13. The system of claim 12, wherein the first scheduling domain and the second scheduling domain are part of a scheduling hierarchy of the first processor.
14. The system of claim 13, wherein the second scheduling domain is at the same level in the scheduling hierarchy as the first scheduling domain and comprises a set of workgroup processing elements separate from the set of workgroup processing elements of the first scheduling domain.
15. A system according to claim 13, wherein the second scheduling domain is at a higher level in the scheduling hierarchy than the first scheduling domain and includes a set of workgroup processing elements, the set of workgroup processing elements including the set of workgroup processing elements of the first scheduling domain and a set of workgroup processing elements of a third scheduling domain at a lower level in the scheduling hierarchy than the second scheduling domain.
16. The system of claim 15, wherein the second scheduling domain is a global scheduling domain for the scheduling hierarchy and the second scheduler circuit is a global command processor.
17. The system of claim 16, wherein the first scheduler circuit receives the first work item from the global command processor.
18. A system according to any one of claims 12 to 17, wherein the first scheduler circuit is configured to allocate the at least one new work item in response to determining that the set of new work items includes one or more work items that would overload the first scheduling domain if consumed by at least one workgroup processing element of the first scheduling domain.
19. A system according to claim 18, wherein the first scheduler circuit is configured to determine that the set of new work items includes one or more work items that will overload the first scheduling domain if consumed by at least one workgroup processing element of the first scheduling domain in response to at least one of the following: identifying one or more work items having a magnification factor greater than a predetermined threshold; or determining that the total number of new work items in the set of new work items exceeds a predetermined threshold.
20. The system according to any one of claims 12 to 19, wherein the first scheduling domain comprises: a local queue, the local queue being accessible only by components in the first scheduling domain; and A shared queue is accessible to both the first scheduling domain and the second scheduling domain and is configured to receive work from the local queue.
21. The system according to any one of claims 12 to 20, further comprising: A host processor is coupled to the first processor and is configured to issue a set of work items as a work graph to the first processor, the set of work items including the first work item.
Citation Information
Patent Citations
Method for dynamically adjusting scheduling interval based on time
CN113434280A
Multi-processor scheduling
EP2256632B1
Feedback guided split workgroup dispatch for gpus
US20190332420A1
Method of dispatching tasks in multi-processor computing environment with dispatching rules and monitoring of system status
US8015564B1
Computing session workload scheduling and management of parent-child tasks
US9417918B2